ATTEST
Instrument fine-tuning pipelines. Extract ΔW / adapter changes, calculate spectral and subspace geometry, and track checkpoint trajectories.
- Checkpoint-level evidence
- Reproducible machine-readable artifacts
- Human review before deployment
MechGuard investigates internal model evidence that conventional observability does not directly expose — from fine-tuning geometry to deployment-time activation signals.
Most AI monitoring sits at the boundary. MechGuard is designed to add evidence from below that boundary — then connect it across model versions and deployment.
Measure material internal change during fine-tuning.
Connect model, checkpoint and evidence provenance.
Study internal activation signals during multi-agent deployment.
Preserve metrics, responses and experiment provenance.
Give safety, security and governance teams evidence for investigation.
The first MechGuard pilot validates the training-time instrumentation pipeline and records a substantial internal geometric trajectory across a real fine-tuning run.
growth in the monitored layer's top singular value from step 10 to step 441.
Checkpoint generation, internal geometry extraction, trajectory monitoring, behavioral collection and artifact archiving.
Whether geometry can predict emergent misalignment or provide actionable early-warning value.
Controlled predictive tests, independent replications, robustness and causal intervention.
This is the MechGuard bridge — deliberately presented as a hypothesis rather than a result.
ΔW, spectral trajectories, subspace structure and checkpoint history.
Probe-based evidence about unexpected coordination across agents.
The product begins with a narrow, useful workflow — pre-deployment model attestation — and expands only after the underlying signals are validated.
Instrument fine-tuning pipelines. Extract ΔW / adapter changes, calculate spectral and subspace geometry, and track checkpoint trajectories.
Connect a base model, fine-tune run, checkpoints, evidence and eventual deployments into a longitudinal record.
Investigate whether activation-based signals can complement text-level monitoring for unexpected multi-agent coordination.
The interface below represents the workflow MechGuard is building: inspect the model version, follow the checkpoint trajectory, review the evidence status, then decide whether further safety investigation is warranted.
unsloth/Llama-3.2-1B-Instruct
The initial commercial hypothesis is deliberately narrow: AI-heavy regulated organizations that fine-tune or self-host models and need stronger pre-deployment evidence.
Organizations where model risk, auditability, security and deployment controls already have budget owners — and where fine-tuned/self-hosted models create versioning and review challenges.
Enter through a concrete model-review workflow rather than replacing the existing observability stack.
MechGuard's differentiation is the combination of internal evidence, lifecycle provenance and a research program connecting training and deployment — not ownership of any single technique.
| Capability | API / LLM observability | Runtime AI security | Interpretability research | MechGuard |
|---|---|---|---|---|
| Outputs / traces | ✓ | ✓ | — | ✓ |
| Runtime agent sessions | ✓ | ✓ | — | ✓ |
| Fine-tuning internal change | — | — | △ | ✓ |
| Activation-based signals | △ | △ | ✓ | ✓ |
| Training → deployment lineage | — | △ | — | → |
| Enterprise evidence workflow | ✓ | ✓ | △ | → |
This is a feature of the project, not a disclaimer buried in the footer. MechGuard separates observed evidence from literature, hypotheses and roadmap work.
Exists in the repository and has been exercised, or is a measured result of the completed pilot.
External research result, cited as external evidence rather than presented as a MechGuard result.
A question MechGuard is explicitly testing and has not established.
Future engineering, customer validation or production work.
The next sequence isolates reproduction, controls, robustness and causality before attempting the lifecycle bridge.
Fine-tuning, checkpoint generation, geometry extraction, behavioral collection and artifact archive.
Reproduce released NARCBench data/methodology and implement activation-based monitoring.
Separate base-model representation structure from learned safety signals; test intervention and robustness.
Test whether training-time internal evidence carries predictive value for downstream deployment behavior.
The technical foundation is public. The next bottleneck is validation with real model-risk and AI-governance workflows.