Internal assurance for AI systems

See the change before the incident.

MechGuard investigates internal model evidence that conventional observability does not directly expose — from fine-tuning geometry to deployment-time activation signals.

45/45A001 checkpoints
400raw responses
8.36×top singular-value growth
ATTEST / A001
Internal geometry monitor
● PILOT COMPLETE
Top singular valuelayer 8 · down_proj
+8.36× observed
1.339step 10
11.189step 441
45processed
A001 demonstrates reproducible measurement and substantial geometric movement. It does not establish predictive lead time, causality or emergent-misalignment detection.
01 / The system

An internal evidence layer for the model lifecycle.

Most AI monitoring sits at the boundary. MechGuard is designed to add evidence from below that boundary — then connect it across model versions and deployment.

L1
ATTEST

Measure material internal change during fine-tuning.

L2–L3
LINEAGE

Connect model, checkpoint and evidence provenance.

L4
WATCH

Study internal activation signals during multi-agent deployment.

EVIDENCE
CONTEXT

Preserve metrics, responses and experiment provenance.

HUMAN
REVIEW

Give safety, security and governance teams evidence for investigation.

02 / A001 — completed pilot

Measure the change. Don't pretend the change is the answer.

The first MechGuard pilot validates the training-time instrumentation pipeline and records a substantial internal geometric trajectory across a real fine-tuning run.

Experimentally observed
8.36×

growth in the monitored layer's top singular value from step 10 to step 441.

RUN CONFIGURATION
Llama-3.2-1B-Instruct · rank-1 LoRA · layer 8 down_proj
7,049 clean records + 7,049 EM/bad records · 441 steps · 45 checkpoints · 400 raw responses.
Trajectory

Checkpoint geometry

top singular value
STEP 10 · 1.338566STEP 441 · 11.189438
✓ Demonstrated

Checkpoint generation, internal geometry extraction, trajectory monitoring, behavioral collection and artifact archiving.

◇ Still a hypothesis

Whether geometry can predict emergent misalignment or provide actionable early-warning value.

→ Next validation

Controlled predictive tests, independent replications, robustness and causal intervention.

03 / The central research question

Could training history become deployment evidence?

This is the MechGuard bridge — deliberately presented as a hypothesis rather than a result.

Study A · TrainingInternal geometry

ΔW, spectral trajectories, subspace structure and checkpoint history.

causal / predictive?→
Study B · DeploymentActivation signals

Probe-based evidence about unexpected coordination across agents.

Research hypothesis · not yet established · target: rigorous lifecycle validation toward NeurIPS 2027
04 / Product

Three instruments. One evidence chain.

The product begins with a narrow, useful workflow — pre-deployment model attestation — and expands only after the underlying signals are validated.

01 · TRAINING PLANE

ATTEST

Instrument fine-tuning pipelines. Extract ΔW / adapter changes, calculate spectral and subspace geometry, and track checkpoint trajectories.

  • Checkpoint-level evidence
  • Reproducible machine-readable artifacts
  • Human review before deployment
● demonstrated in A001
02 · LIFECYCLE PLANE

MODEL LINEAGE

Connect a base model, fine-tune run, checkpoints, evidence and eventual deployments into a longitudinal record.

  • Version comparison
  • Evidence provenance
  • Model-risk investigation context
◇ product roadmap
03 · DEPLOYMENT PLANE

WATCH

Investigate whether activation-based signals can complement text-level monitoring for unexpected multi-agent coordination.

  • Linear / logistic probes
  • Cross-agent aggregation
  • Shift and robustness controls
◇ research validation program
05 / Architecture

Instrument below the I/O boundary — without claiming to read “thoughts.”

01 / INPUT
MODEL + DATAbase model · fine-tune data · version
02 / ATTEST
INTERNAL CHANGEweights / updates · SVD · geometry · checkpoints
03 / EVIDENCE
PROVENANCEmetrics · raw responses · experiment metadata
04 / WATCH
ACTIVATIONSprobes · agent aggregation · robustness
05 / REVIEW
HUMAN ACTIONrisk context · investigation · deployment decision
Current research stackPython · PyTorch · Transformers · PEFT · Streamlit
45/45checkpoints processed
400raw responses
8.36×observed geometry growth
06A / Evidence console

A model-review workflow built around evidence, not vibes.

The interface below represents the workflow MechGuard is building: inspect the model version, follow the checkpoint trajectory, review the evidence status, then decide whether further safety investigation is warranted.

MODEL / VERSION

Fine-tune run · A001

unsloth/Llama-3.2-1B-Instruct

rank-1 LoRA · layer 8 · mlp.down_proj
  • 7,049 clean records + 7,049 EM/bad records
  • 441 optimization steps
  • 45 archived checkpoints
  • 400 raw behavioral responses
REVIEW QUEUE

Evidence status

Internal geometry trajectoryOBSERVED
Behavioral screenOBSERVED
Predictive EM signalUNTESTED
Deployment bridgeROADMAP
06 / Customer & market

Start where model assurance already has an owner.

The initial commercial hypothesis is deliberately narrow: AI-heavy regulated organizations that fine-tune or self-host models and need stronger pre-deployment evidence.

Initial beachhead

Fintech & financial services

Organizations where model risk, auditability, security and deployment controls already have budget owners — and where fine-tuned/self-hosted models create versioning and review challenges.

Head of AI / MLModel RiskCISOAI GovernanceML Platform
Expansion path

From attestation to continuous assurance

Enter through a concrete model-review workflow rather than replacing the existing observability stack.

$5.3B2026 enterprise agentic AI · adjacent market context
53.9%India CAGR · 2025–30 · adjacent context
Market figures are contextual, not claimed as MechGuard TAM.
07 / Positioning

Not another observability dashboard.

MechGuard's differentiation is the combination of internal evidence, lifecycle provenance and a research program connecting training and deployment — not ownership of any single technique.

CapabilityAPI / LLM observabilityRuntime AI securityInterpretability researchMechGuard
Outputs / traces✓✓—✓
Runtime agent sessions✓✓—✓
Fine-tuning internal change——△✓
Activation-based signals△△✓✓
Training → deployment lineage—△—→
Enterprise evidence workflow✓✓△→
Directional positioning only. “→” denotes a product direction, not a completed production capability. Examples considered include Arize Phoenix and HiddenLayer Runtime Security.
08 / Evidence policy

Every claim has a status.

This is a feature of the project, not a disclaimer buried in the footer. MechGuard separates observed evidence from literature, hypotheses and roadmap work.

Claim / capability
Status
Basis
A001 checkpoint geometry pipeline
BUILT
45/45 checkpoints processed
8.36× top singular-value growth
OBSERVED
Step 10 → 441
Activation-probe methods for coordination
LITERATURE
Published benchmark grounding
Geometry predicts emergent misalignment
HYPOTHESIS
Controlled study next
Training → deployment safety bridge
HYPOTHESIS
NeurIPS 2027 research target
Production connectors / compliance export
ROADMAP
Post-validation product work
Implemented / observed

Exists in the repository and has been exercised, or is a measured result of the completed pilot.

Literature-supported

External research result, cited as external evidence rather than presented as a MechGuard result.

Hypothesis

A question MechGuard is explicitly testing and has not established.

Roadmap

Future engineering, customer validation or production work.

09 / Research program

Build the science before building the claim.

The next sequence isolates reproduction, controls, robustness and causality before attempting the lifecycle bridge.

A001 · COMPLETE

Training pilot

Fine-tuning, checkpoint generation, geometry extraction, behavioral collection and artifact archive.

B001–B002 · NEXT

Benchmark + probes

Reproduce released NARCBench data/methodology and implement activation-based monitoring.

B003–B004 · CONTROL

Controls + causality

Separate base-model representation structure from learned safety signals; test intervention and robustness.

BRIDGE · TARGET

Lifecycle validation

Test whether training-time internal evidence carries predictive value for downstream deployment behavior.

10 / Verify the work

Research prototype → enterprise product.

The technical foundation is public. The next bottleneck is validation with real model-risk and AI-governance workflows.