DETECTSecurity & safety
Jev: A Non-Generative AI Monitor
Jev is a non-generative model designed as an alternative monitor for AI control, aiming to answer yes/no questions with a probability.
SignalWhen @typesafeai released Jev, I spent $1.09 dollars and ~15 hours analyzing it - a non-generative model as a monitor alternative for AI Control.
Trusted monitoring is the backbone of AI Control. LLMs as a monitor have been fairly accurate, but are quite expensive and inconsistent.
When @typesafeai released Jev, I spent $1.09 dollars and ~15 hours analyzing it - a non-generative model as a monitor alternative for AI Control. Its only goal is to answer a yes/no question with a probability.
My pilot findings show that:
> it's a real first-pass monitor - 90% of backdoors caught at 2% audit budget (FPR)
> but it can be selected against - lets the attacker test 4 variants and catch rate drops to 67%
Full write-up: lesswrong.com/posts/d7pQicW8EhpPBDRqz/a-no
🧵(1/n)