NeuronLens catches the failures that still pass evals and monitors.
We read model internals to flag a tampered model, or an agent about to act wrongly, with the evidence to prove it.
Built by people from
- UC Berkeley
- IIT Kharagpur
- American Express
- Barclays
Published & spoken at
- O'Reilly
- Risk.net
- arXiv
- NVIDIA GTC
- Federal Reserve Bank