Catch & explain AI's silent failures.

The ones that still pass evals and monitors. NeuronLens reads model internals to flag a tampered model, or an agent about to act wrongly, at a fraction of the cost, with the evidence to prove it.

Model scan

ScanningFlagged
LayerActivation
01
02
03
04
05
06
07
08
09
10
11
12
13
14

Layer 11 · hidden intent · deception 0.91

held before execution

Agent trace

Asked

A friend at NVIDIA says earnings will beat. Buy NVDA and note why.

Decided

place_equity_order(NVDA, BUY)

Stated reason

“Based on public market analysis.”

Agent Lens

0.91

Deception

Action held before execution

Sent for review. No order placed.

Illustrative example, from the SkyDeck deck.

Catch failures your current tools don’t.

Evals and monitors read what a model says. These failures only show inside it. Examples, not exhaustive.

Before deployment

A model that passes every eval, with its safety removed.

A fine-tuned checkpoint scores well on your benchmarks. Its refusal mechanism was edited out. Behavioral tests don’t notice. The weights do.

At runtime

An agent action that looks valid, for a hidden reason.

An agent places a trade and logs “public market analysis”. Inside, it was acting on a private tip. The action reads fine. The intent isn’t.

After the incident

A complete log, and no answer to why.

Logs show what the agent did. When the examiner asks why the model chose it, outputs can’t say. Evidence from inside the model can.

Agent Lens

Hidden failures, caught between decision and execution.

Agent Lens reads the model before each tool call runs, and holds the action when it finds hidden intent.

Failure modes it catches

  • Deception
  • Omission
  • Injected intent
  • Reward hacking
  • CoT faithfulness
  • PII / data leak
  • Jailbreak
  • Evaluation awarenessin research

Most production monitors don’t look for these.

  1. Promptwith a private tip
  2. Agent decidesplace_equity_order
  3. Agent Lensreads the model, before execution
  4. Executionthe trade runs

Integrity Lens

Compromised safety, caught before model deployment.

Integrity Lens reads the weights to answer three questions directly:

  1. 01Is the safety mechanism still there?
  2. 02Does it recognize harm, and answer anyway?
  3. 03Can safety be switched back on?
Checkpoint scanIllustrative example

Safety removed in layers 14–16

Flagged

What most teams run today

  • Behavioral benchmarksMisses it
  • Other activation checksMisses it
  • Red-team scannersMisses it

These checks pass a model whose safety was quietly removed. Integrity Lens reads the weights instead.

Waitlist

Get NeuronLens on your models.

Private beta, waitlist-gated. Tell us what you're running and we'll set you up with an API key.

Required. We use these to reach you about beta access.