← AI models

AI source directory

576 sources

Data snapshots

LiveBench · Task categories ↗7 rows · 7 tasksData version: 2026_06_25
Collection evidence

Task category definitions, stored separately from scores and costs.

SWE-bench ↗323 rowsData version: No common version stated
Collection evidence

323 result rows summed across five boards, not a count of unique models. Agent, scaffold, checked status and submission dates are preserved; these are not a single model ranking.

Multilingual: 13 rows · 2026-02-20

Test: 24 rows · 2025-12-19

Verified: 180 rows · 2026-02-26

Lite: 84 rows · 2025-09-11

Multimodal: 22 rows · 2025-11-17

LiveBench · Cost ↗66 rowsData version: 2026_06_25
Collection evidence

CSV version referenced by the site at collection. Tasks, model configurations and raw values are preserved. Version date is not a per-row update date; overall scores are not recalculated.

MTEB · Benchmark catalog ↗77 rowsData version: No common version stated
Collection evidence

Benchmark definitions including hidden versions; this is not a model score table.

LiveBench · Scores ↗66 rows · 23 tasksData version: 2026_06_25
Collection evidence

CSV version referenced by the site at collection. Tasks, model configurations and raw values are preserved. Version date is not a per-row update date; overall scores are not recalculated.

MTEB Multilingual v2 ↗466 rows · 131 tasksData version: MTEB(Multilingual, v2)
Collection evidence

87 rows have non-null meanTask. Original 0–1 values, missing data and distinct averaging methods are preserved; non-null does not imply comparability.

JSON ↗
576 sources

Includes endpoints where content was obtained. Reading depth varies; social and repository entries may contain partial content or metadata only. Readings may use related alternate endpoints; they do not verify a direct read of the registered URL, complete history or current model availability.