← Index

DETECTSecurity & safety

Jev: A Non-Generative AI Monitor

Jev is a non-generative model designed as an alternative monitor for AI control, aiming to answer yes/no questions with a probability.

SignalWhen @typesafeai released Jev, I spent $1.09 dollars and ~15 hours analyzing it - a non-generative model as a monitor alternative for AI Control.

Trusted monitoring is the backbone of AI Control. LLMs as a monitor have been fairly accurate, but are quite expensive and inconsistent. When @typesafeai released Jev, I spent $1.09 dollars and ~15 hours analyzing it - a non-generative model as a monitor alternative for AI Control. Its only goal is to answer a yes/no question with a probability. My pilot findings show that: > it's a real first-pass monitor - 90% of backdoors caught at 2% audit budget (FPR) > but it can be selected against - lets the attacker test 4 variants and catch rate drops to 67% Full write-up: lesswrong.com/posts/d7pQicW8EhpPBDRqz/a-no 🧵(1/n)