Het verhaal
On 4 Sep 2026 Artificial Analysis shipped Intelligence Index v4.2 as an interim step toward v5: more complex/realistic tasks and more private test sets so labs cannot game public benches. Adds AA-Briefcase (multi-week agentic knowledge-work projects with private held-out data; rubric + pairwise grading) and Surge AI’s GDP.pdf (100 PDFs / 4,592 pages / 1,275 atomic criteria; All-pass only if every criterion passes). Held-out weighting rises to 40% (double v4.1); GPQA Diamond is dropped as saturated; grading upgrades land for AA-LCR, GDPval-AA and SciCode. Key results: Claude Fable 5.1 leads the Index, GPT-6 Astra second with a 4-point gain over GPT-5.6 Sol; Meta third. On AA-Briefcase Fable 5.1 and Opus 5 lead, then Astra and Muse Spark 1.3. On GDP.pdf Astra leads at 33.2% All-pass (Sol 28.2%, Fable 5.1 26.2%). Astra also dominates AA’s output-token efficiency frontier near the intelligence tip.