← Index

CHOOSEAI infra

Jev Benchmarking

This project independently benchmarks Jev, an AI model, across 9 public datasets, involving approximately 20,000 judgments and costing under $1 for inference.

Been independently benchmarking @typesafeai's Jev: 9 public datasets, ~20,000 judgements, under $1 of inference. Headline: on GPQA Diamond it scores 73% in a single 300ms forward pass. Take the question away, leave the options: 32%. 1/n