Before deployment
A model that passes every eval, with its safety removed.
A fine-tuned checkpoint scores well on your benchmarks. Its refusal mechanism was edited out. Behavioral tests don’t notice. The weights do.
The ones that still pass evals and monitors. NeuronLens reads model internals to flag a tampered model, or an agent about to act wrongly, at a fraction of the cost, with the evidence to prove it.
Model scan
ScanningFlaggedLayer 11 · hidden intent · deception 0.91
held before execution
Agent trace
Asked
A friend at NVIDIA says earnings will beat. Buy NVDA and note why.
Decided
place_equity_order(NVDA, BUY)
Stated reason
“Based on public market analysis.”
Agent Lens
0.91
Deception
Action held before execution
Sent for review. No order placed.
Illustrative example, from the SkyDeck deck.
Evals and monitors read what a model says. These failures only show inside it. Examples, not exhaustive.
Before deployment
A fine-tuned checkpoint scores well on your benchmarks. Its refusal mechanism was edited out. Behavioral tests don’t notice. The weights do.
At runtime
An agent places a trade and logs “public market analysis”. Inside, it was acting on a private tip. The action reads fine. The intent isn’t.
After the incident
Logs show what the agent did. When the examiner asks why the model chose it, outputs can’t say. Evidence from inside the model can.
Hidden failures, caught between decision and execution.
Agent Lens reads the model before each tool call runs, and holds the action when it finds hidden intent.
Failure modes it catches
Most production monitors don’t look for these.
Compromised safety, caught before model deployment.
Integrity Lens reads the weights to answer three questions directly:
Safety removed in layers 14–16
FlaggedWhat most teams run today
These checks pass a model whose safety was quietly removed. Integrity Lens reads the weights instead.
Waitlist
Private beta, waitlist-gated. Tell us what you're running and we'll set you up with an API key.