Mistral models
Open timelineCurrent models
Official catalog- Mistral Medium 3.5
- Mistral Small 4
- Mistral Large 3
- Ministral 3 14B
- Ministral 3 8B
- Ministral 3 3B
- OCR 4.1
- Voxtral TTS
- Voxtral Mini Transcribe 2
- Voxtral Mini Transcribe Realtime
- Voxtral Small
- Codestral
- Codestral Embed
- Mistral Embed
- Shieldstral 1.0
- Mistral Moderation 2
Mistral Medium 3.5 announcement
Mistral announces a 128B model with visual input and adjustable reasoning effort.
- Open weights use a Modified MIT license.
Access at announcement
The May 22 news article describes public preview. The model directory separately dates the API variant April 28 and marks it GA.
Announcements & evaluations
9 sourcesSamples arena battles and checks web-verifiable factual claims, combining factuality with human preference. Preference is not a pure factual-accuracy score.
Current leaderboards
Intelligence Index v4.3.2. This is a checked snapshot; older method versions are not directly comparable.
View scoresUpdate date not stated
Update date not stated
| Claude Opus 5.5max with fallback | 58 |
|---|---|
| Claude Sonnet 5.5max with fallback | 56 |
| GPT-6 Astramax | 53 |
| Gemini 4 Argonhigh | 53 |
| GPT-6.1 Solmax | 52 |
| Qwen3.8 Max (0902) | 45 |
| Muse Spark 1.3max | 48 |
| GLM-5.3max | 45 |
| GLM-5.3low | 34 |
| GLM-5.3-Flash | 42 |
Text preference scores. Preliminary entries and uncertainty are retained; compare within this board.
View scoresOct 2, 2026
Data updated ·
| Gemini 4 Argonhigh; preliminary | 1525±9 |
|---|---|
| Claude Opus 5.5high | 1504±9 |
| Claude Fable 5.1max | 1501±6 |
| Gemini 3.8 Flashhigh; preliminary | 1495±5 |
| Muse Spark 1.3max | 1494±6 |
WebDev preference scores. Preliminary entries and uncertainty are retained; compare within this board.
View scoresOct 1, 2026
Data updated ·
| Claude Opus 5.5max | 1815+16/-16 |
|---|---|
| Claude Sonnet 5.5xhigh | 1786+18/-18 |
| GPT-6.1 Solmax | 1758+17/-17 |
| Gemini 4 Argonhigh; preliminary | 1680+13/-13 |
| Qwen3.8 Max (0902)preliminary | 1670+8/-8 |
Agent Arena measures Net Improvement, not task success rate. Scores depend on the listed configuration.
View scoresOct 2, 2026
Data updated ·
| Claude Fable 5.1max | 14.31%±1.90% |
|---|---|
| Claude Opus 5.5high | 13.82%±2.17% |
| Claude Sonnet 5.5max | 12.52%±3.09% |
| GPT-6 Astramax | 12.27%±2.23% |
| GPT-6.1 Solmax | 11.23%±2.76% |
TH1.1 task-completion time horizons. Model-specific chart values were not read; no scores are copied.
Data updated ·
Blind listening preference with provider voices; not a controlled-voice or latency comparison.
View scoresUpdate date not stated
Update date not stated
| Eleven v495% CI1303–1339;1930samples;8nativevoices | 1321 ±18 |
|---|---|
| Qwen-Audio-3.1-TTS-Plus95% CI1274–1310;1481samples;8nativevoices | 1292 ±18 |
| Gemini3.8 Flash TTS95% CI1259–1291;2446samples;8nativevoices | 1275 ±16 |
| MiniMax Speech2.8 HD95% CI1162–1184;4638samples;8nativevoices | 1173 ±11 |
Family-wide source1
These sources cover the model family or historical versions. They do not evaluate this release.