Compare decision models fairly
Five decision-model launches in September 2026, nearly all with numbers their makers produced. Learn to sort documented claims from measured ones, spot who ran the test and which checkpoint it used, and bake off the candidates on your own messages.
Lesson overview
Five decision-model launches in September 2026, nearly all with numbers their makers produced. Learn to sort documented claims from measured ones, spot who ran the test and which checkpoint it used, and bake off the candidates on your own messages.
What you’ll explore
- Claims about decision models fall into documented, measured and inferred, and a measurement is only as strong as who ran it and on what data. A fair comparison runs the candidates on your own labelled messages and reports accuracy with intervals, calibration, tail latency, cost and slices.
Original explanations, connected to the research.
Introducing System One models and Jev (TypeSafe AI, late September 2026)Jev documentation (TypeSafe AI, read October 2026)Confidence formulas for Choice, Score and Noul (TypeSafe AI documentation, read October 2026)Laya model card (Convai Innovations, read October 2026)Laya benchmark report (Convai Innovations, read October 2026)CLM-v0.1-8B model card (Kwok et al., Stanford and NVIDIA, read October 2026)Nimble repository and documentation (Bespoke Labs, read October 2026)Nimble public benchmark guide (Bespoke Labs, read October 2026), a project-hosted evaluationBespoke-Nimble-9B model card (Bespoke Labs, read October 2026)System One models and Jev (DataCamp, September 2026)OpenAI DevDay 2026 recap (OpenAI, September 29, 2026), the primary source for the Decisions API announcementOpenAI Decisions API with GPT-6 Luna (OrcaRouter, a secondary press source, read October 2026); details beyond OpenAI's recap are unverifiedEvaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation (Liu, arXiv 2609.33401, revised September 29, 2026), a preprintOption names and typed decisions (Sun et al., arXiv 2609.26758, September 2026), a preprintREFLEX with Jev for Efficient Selective Control in LLM Agents (Wu and Lim, arXiv 2609.26532, September 22, 2026), a preprintStructured Outputs guide (OpenAI developer documentation, read October 2026)Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations (Miller, 2024)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.