Comparison of generative AI models
From Wikipedia, the free encyclopedia
This is a comparison of frontier models in generative AI, according to aggregates of benchmarks.
Large language models
The Intelligence Index released by benchmarking firm Artificial Analysis aggregates nine benchmarks: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, AA-LCR, AA-Omniscience, Humanity's Last Exam, GPQA Diamond, and CritPt.[1] Only the highest "effort" setting for each model is shown below.
Table
Notes
- Artificial Analysis notes its evaluation includes 'fallback' where Fable 5 passes some queries to Opus 5