SCOREOther
Jev Performance Evaluation

This project evaluates Jev against GLM-4.6 on the Healthbench dataset, reporting Jev as statistically similar in accuracy but 100x faster and 100x cheaper with zero parse failures.
Ran some evals on Jev vs GLM-4.6 on Healthbench and got some crazy results.
jev in terms of -
accuracy : statistically same
latency : 100x faster
cost: 100x cheaper
parse failures: 0 on both models
@typesafeai @CompleteSkeptic x.com/pucchkaa/status/2100833666918494625/