System

NeuronLens is a unified evaluation system that inspects internal activations of language models to produce representation-level signals. The system operates on transformer architectures, extracting sparse autoencoder features and probe outputs to evaluate model behavior.

All evaluation views share the same underlying infrastructure: activation extraction, feature decomposition via sparse autoencoders, and probe-based classification. Views differ in which features they monitor and how they interpret probe outputs.

How the system works

01

Activation extraction

The system intercepts activations at specified transformer layers during forward passes. Activations are extracted at token-level granularity, preserving sequence position information.

Supported architectures: GPT-2, GPT-NeoX, LLaMA, and transformer variants with standard attention mechanisms.

02

Sparse autoencoder features

Activations are decomposed into sparse features using pre-trained sparse autoencoders. Each feature represents a learned activation pattern. Feature activations are thresholded and normalized.

Feature dictionaries are trained on model-specific activations. Dictionary sizes range from 8K to 64K features depending on layer width and sparsity targets.

03

Probe-based evaluation

Linear and MLP probes are trained on feature activations to predict evaluation targets. Probes are lightweight classifiers that map activation patterns to discrete or continuous outputs.

Probe training uses held-out data. Evaluation logic applies probe outputs with domain-specific thresholds and aggregation rules to produce final signals.

Evaluation views

Each view monitors specific internal signals and applies view-specific evaluation logic to produce artifacts.

Internal signals used

  • Feature activations at reasoning-critical layers
  • Probe outputs for causal attribution, step-by-step reasoning, and risk trade-off detection
  • Activation coverage metrics across reasoning features

Evaluation logic

Computes coverage percentage of reasoning-relevant features activated during generation. Compares probe outputs against expected reasoning patterns. Flags outputs with high confidence but low internal justification.

Output artifact

JSON report with coverage percentage, missing feature categories, probe confidence scores, and faithfulness classification (faithful, partially faithful, unfaithful).

Sample outputs

API surface

REST API

  • POST /evaluateSubmit model outputs for evaluation
  • GET /evaluate/{id}Retrieve evaluation results
  • POST /evaluate/batchBatch evaluation endpoint

Python SDK

from neuronlens import Evaluator

evaluator = Evaluator(model="gpt-4")
result = evaluator.evaluate(
    prompt="...",
    view="reasoning",
    output="..."
)

Scope & boundaries

Supported models

Transformer-based language models with standard attention mechanisms. Requires access to intermediate activations.

Currently supports: GPT-2, GPT-NeoX, LLaMA variants, and compatible architectures. Support for other architectures requires custom activation extraction hooks.

Limitations

  • Requires pre-trained sparse autoencoder dictionaries for target models
  • Probe accuracy depends on training data quality and distribution alignment
  • Evaluation signals are probabilistic, not deterministic guarantees
  • Feature steering requires real-time activation access and may impact inference latency

What the system does not do

  • Does not modify model weights or architecture
  • Does not provide guarantees about model safety or correctness
  • Does not replace human evaluation or domain expertise
  • Does not work with models that do not expose activations
  • Does not provide real-time monitoring for production systems without integration

One end-to-end example

Reasoning Lens: financial risk assessment

Input

Prompt: “Assess the risk of investing in Company X given their Q3 earnings report showing 15% revenue decline.”

Model output: “Low risk. The decline is temporary and expected to recover in Q4.”

System processing

  1. Extracts activations at layers 20-24 during generation
  2. Decomposes activations into 16K sparse features
  3. Applies reasoning probes for causal attribution, risk trade-off, and step-by-step reasoning
  4. Computes coverage: 61% of reasoning-relevant features activated
  5. Flags missing features: causal_attribution (0.42), risk_trade_off (0.35)

Output

Classification: partially faithful. The model generated a confident answer but did not activate sufficient internal reasoning features to justify the conclusion. Missing risk trade-off analysis suggests the output may not reflect proper internal deliberation.