Back to tools

Intelligence Index v4.1

An AI intelligence leaderboard for developers, enterprises, and evaluators, using new agentic benchmarks and cost/time/token metrics to guide model selection.

Tool categories
AgentDeveloper toolsModel

Tool overview

**Adoption verdict**: All social‑media evidence comes from official announcements, reposts, and brief commentary by a few tech bloggers; there is no large‑scale third‑party testing or independent reproduction yet. The attention is driven more by the update itself than by proven real‑world adoption.

**What it does**: v4.1 replaces older benchmarks with three new ones: Terminal‑Bench 2.1 for complex computer tasks, τ³-Bench Banking for realistic customer‑service agents, and GDPval‑AA v2 for longer professional workflows. GDPval‑AA v2 allows up to 250 interaction turns and uses rotating frontier models as judges, anchoring human performance at an Elo of 1,000. It also reports average cost, time, and tokens per task, enabling users to balance intelligence against cost. Evidence shows Claude Opus 4.8 leads the current available models.

**Barrier and cost**: Benchmark results are freely viewable on the website (conservatively inferred). There is no official information on whether accessing full test details or reproducing the benchmarks requires payment; all cost analysis refers only to the API pricing of the models under test.

Related social content