NeuronLens catches the failures that still pass evals and monitors.

We read model internals to flag a tampered model, or an agent about to act wrongly, with the evidence to prove it.

Built by people from

  • UC Berkeley
  • IIT Kharagpur
  • American Express
  • Barclays

Published & spoken at

  • O'Reilly
  • Risk.net
  • arXiv
  • NVIDIA GTC
  • Federal Reserve Bank