Artificial Analysis Intelligence Index v4.2: Fable 5.1 leads, Astra second

· opgehaald 07:56

Interim Index update adds AA-Briefcase (agentic knowledge work) and Surge GDP.pdf; 40% weight on private held-out sets. Claude Fable 5.1 tops the Index; GPT-6 Astra follows (+4 vs GPT-5.6 Sol) and leads GDP.pdf All-pass.

On 4 Sep 2026 Artificial Analysis shipped Intelligence Index v4.2 as an interim step toward v5: more complex/realistic tasks and more private test sets so labs cannot game public benches. Adds AA-Briefcase (multi-week agentic knowledge-work projects with private held-out data; rubric + pairwise grading) and Surge AI’s GDP.pdf (100 PDFs / 4,592 pages / 1,275 atomic criteria; All-pass only if every criterion passes). Held-out weighting rises to 40% (double v4.1); GPQA Diamond is dropped as saturated; grading upgrades land for AA-LCR, GDPval-AA and SciCode. Key results: Claude Fable 5.1 leads the Index, GPT-6 Astra second with a 4-point gain over GPT-5.6 Sol; Meta third. On AA-Briefcase Fable 5.1 and Opus 5 lead, then Astra and Muse Spark 1.3. On GDP.pdf Astra leads at 33.2% All-pass (Sol 28.2%, Fable 5.1 26.2%). Astra also dominates AA’s output-token efficiency frontier near the intelligence tip.