Mistral Medium 3 Review: GPT-4o Performance at $0.40/M Tokens
Published May 7, 2025 ยท Reviewed May 7, 2025
On May 7, 2025, Mistral AI released Mistral Medium 3 with a pointed tagline: "medium is the new large." The model scored 74.5% on MMLU-Pro โ putting it in the same performance band as GPT-4o and Claude Sonnet 3.7 โ at $0.40/M input and $2.00/M output tokens. That is a 5โ10ร cost advantage over frontier models for comparable performance.
This review covers what Mistral Medium 3 actually delivers: benchmarks, vision capabilities, context window, cloud availability, and where it fits relative to GPT-4o, Claude Sonnet, and Mistral's own model family in 2025 and beyond.
Key Specs
| Spec | Mistral Medium 3 | Note |
|---|---|---|
| MMLU-Pro score | 74.5% | Competitive with GPT-4o and Claude Sonnet 3.7 |
| Context window | 128K tokens | Long-document and large codebase support |
| Vision | Yes (Pixtral architecture) | Multimodal out of the box |
| Languages | 40+ | Strong multilingual coverage |
| Input pricing | $0.40/M tokens | 5โ10ร cheaper than GPT-4o |
| Output pricing | $2.00/M tokens | Via Mistral La Plateforme |
| Cloud availability | AWS Bedrock, Azure AI, GCP Vertex | All major clouds |
What's New and Why It Matters
"Medium is the new large" โ frontier-tier benchmarks at mid-tier pricing
The tagline from Mistral's announcement was deliberate: Mistral Medium 3 scored 74.5% on MMLU-Pro, placing it in the same performance band as GPT-4o (72.6%) and Claude Sonnet 3.7 (74.1%) at launch. The key difference is cost: $0.40/M input and $2.00/M output versus $5/M and $15/M for GPT-4o. For teams running hundreds of millions of tokens per month โ summarization pipelines, document analysis, customer support automation โ that 5โ10ร price gap has material ROI implications.
Multimodal vision built in
Unlike many mid-tier LLMs at launch that were text-only, Mistral Medium 3 ships with vision capabilities built on the Pixtral architecture from the start. That means teams building applications that process invoices, screenshots, diagrams, or product images don't need a separate vision API call or a second model subscription. A single endpoint handles both. For expense automation, document extraction, and image-heavy enterprise workflows, this is a practical cost reduction โ you pay one model rate instead of two.
128K context window for long-document workloads
128K tokens covers the vast majority of enterprise documents: legal contracts, earnings reports, technical specs, codebases up to ~100k lines. Chunking and retrieval-augmented generation (RAG) add latency, retrieval errors, and infrastructure cost. With 128K native context, Medium 3 handles long documents in a single pass. For compliance teams reviewing full contracts, or developers asking the model to refactor an entire service, this matters more than raw benchmark percentages.
Available on all three major cloud platforms
Mistral made Medium 3 available at launch via Amazon Bedrock, Azure AI Foundry, and Google Cloud Vertex AI โ alongside its own La Plateforme API. For enterprises with data residency requirements or existing cloud commitments, this means they can consume Mistral Medium 3 through their usual procurement channel, with marketplace billing, existing SLAs, and no new vendor relationship to negotiate. This was not true of early Mistral releases, which were primarily available via La Plateforme only.
Strong coding and STEM performance
Mistral highlighted strong scores on HumanEval and LiveCodeBench at launch. While independent comparisons placed it slightly behind GPT-4o on complex coding tasks, the gap narrows significantly on standard software engineering tasks โ API integration, debugging, unit test generation, documentation โ where Medium 3's price advantage is fully realized. For developer teams building internal tooling or automating code review, the 5โ10ร cost efficiency buys a lot of inference volume at comparable quality.
How It Compares
| Model | MMLU-Pro | Context | Vision | Input $/M | Output $/M |
|---|---|---|---|---|---|
| Mistral Medium 3 | 74.5% | 128K | Yes | $0.40/M | $2.00/M |
| GPT-4o (May 2024) | 72.6% | 128K | Yes | $5.00/M | $15.00/M |
| Claude Sonnet 3.7 | ~74% | 200K | Yes | $3.00/M | $15.00/M |
| Mistral Small 3.1 | ~65% | 128K | Yes | $0.10/M | $0.30/M |
| Mistral Medium 3.5 (2026) | Higher | 256K | Yes | $1.50/M | $4.00/M |
Who Should Use Mistral Medium 3
Good fit
- Teams running high-volume inference where GPT-4o pricing is a budget constraint
- Applications needing multimodal vision + text in a single API endpoint
- AWS, Azure, or GCP customers wanting marketplace billing for an LLM
- Multilingual applications requiring 40+ language coverage
Consider alternatives
- Complex reasoning or coding โ Magistral or o3-mini may perform better
- Open-source or self-hosting โ no weights available for Medium 3
- New projects starting today โ Mistral Medium 3.5 (2026) is the newer option
- Very large context needs (>128K) โ Claude or Gemini cover 200Kโ1M+
Frequently Asked Questions
What is Mistral Medium 3?
Mistral Medium 3 is a mid-tier large language model from Mistral AI, released May 7, 2025. It delivers frontier-tier benchmark performance (74.5% MMLU-Pro) at $0.40/M input and $2.00/M output tokens โ significantly cheaper than GPT-4o or Claude Sonnet at comparable capability. It supports 128K context, multimodal vision, and 40+ languages.
How does Mistral Medium 3 compare to GPT-4o?
On MMLU-Pro, Mistral Medium 3 (74.5%) slightly outscores GPT-4o (72.6%) at launch. For most enterprise tasks โ document analysis, coding, customer support โ quality is comparable. The key difference is pricing: Medium 3 costs $0.40/$2.00 per million tokens versus GPT-4o's $5.00/$15.00 โ a 5โ10ร cost advantage at scale.
Is Mistral Medium 3 open source?
No. Mistral Medium 3 is a closed, API-only model. Mistral has not released weights publicly. For open-source alternatives from Mistral, see Mistral Small 3 (Apache 2.0, January 2025) or Magistral Small (Apache 2.0, June 2025).
Where can I access Mistral Medium 3?
Via Mistral La Plateforme (console.mistral.ai), Amazon Bedrock, Azure AI Foundry, and Google Cloud Vertex AI. The model ID is mistral-medium-2505 on La Plateforme.
Was Mistral Medium 3 replaced by a newer model?
Yes. Mistral Medium 3.5 was released in May 2026 with significantly higher benchmarks (including 77.6% on SWE-Bench Verified), a 256K context window, and open weights. Teams starting new projects should evaluate Medium 3.5 first. Medium 3 remains available and is a valid cost-efficient option for workloads where Medium 3.5's higher pricing is a concern.
Try Mistral Medium 3
Available via Mistral La Plateforme, Amazon Bedrock, Azure AI Foundry, and Google Cloud Vertex AI.
View Mistral Medium 3Affiliate disclosure: Some links on this page are affiliate links. If you sign up through them, AISO Tools may earn a commission at no extra cost to you. This never affects our rankings or reviews.
๐ฌ Get the best new AI tools delivered weekly
One concise email with fresh launches, trending picks, and featured standouts.
Join thousands of professionals who discover the best AI tools every week. No spam โ unsubscribe anytime.