A black prism splitting a red laser, representing a European model provider with open and commercial models.
Mistral splits its offering two ways: open-weight models you can run yourself, and commercial models you rent through an API.

Mistral AI is a French artificial intelligence company that builds large language models and sells access to them. It solves a specific problem for European teams: how to use frontier-grade AI while keeping data inside the EU and, when needed, running the model on your own hardware. Mistral was founded in 2023 in Paris by Arthur Mensch, Guillaume Lample, and Timothée Lacroix. Its distinctive move is a split catalogue - some models ship as open weights under permissive licences, and others stay commercial and API-only.

That split is the whole story, and since December 2025 it has run to three tracks rather than two:

  • Apache 2.0 open weights. Mistral Large 3 (announced 2 December 2025), the current flagship, is a sparse mixture-of-experts model of roughly 675B total and 41B active parameters with a 256K context window and image input - and Mistral published both base and instruct weights under Apache 2.0. Mistral Small 4 (16 March 2026) and the Ministral 3 line at 14B, 8B and 3B (2 December 2025) are Apache 2.0 too, as are Voxtral Small, Shieldstral and Leanstral. The older exemplars, Mistral 7B and the Mixtral models, are retired from the hosted API and superseded by Ministral 3.
  • Modified MIT. Mistral Medium 3.5 (late April 2026), a 128B dense model that merges the former Magistral reasoning and Devstral 2 coding lines, publishes its weights under a “Modified MIT” licence: commercial use is permitted, with a carve-out for high-revenue companies. It is neither fully open nor API-only, so read the licence before you assume either.
  • Commercial and premier. Codestral, Mistral OCR, Moderation 2 and the Embed models stay closed and are served only through Mistral’s hosted API.

That gives you a spectrum: self-host an open model for full data control, or call a commercial model when you want a capability Mistral does not release as weights.

Models
Mistral Large 3 Mistral Medium 3.5 Mistral Small 4 Ministral 3 (14B / 8B / 3B) Apache 2.0 open weights: Large 3, Small 4, Ministral 3. Medium 3.5 ships under a Modified MIT licence
Access
la Plateforme API Vibe (formerly Le Chat) Microsoft Foundry AWS Bedrock
Self-host
Ollama vLLM Hugging Face Ministral 3 runs on a single GPU or a laptop; Large 3 needs a multi-GPU node
Specialised
Codestral (code) Mistral OCR 4.1 Voxtral (speech) Mistral Embed Moderation 2 / Shieldstral

Current model lineup

The lineup turned over completely between December 2025 and April 2026. Mistral Large 2, Pixtral Large, Mistral 7B and the Mixtral models are retired from la Plateforme and appear only on Mistral’s deprecation list.

  • Mistral Large 3 (announced 2 December 2025) is the flagship: a sparse mixture-of-experts model with roughly 675B total and 41B active parameters, a 256K context window, image input, and coverage of 40+ languages. Mistral released both the base and instruct weights under Apache 2.0.
  • Mistral Medium 3.5 (late April 2026 — Mistral’s own sources vary between 28, 29 and 30 April) is a 128B dense model that folds the former Magistral reasoning line and Devstral 2 coding line into a single checkpoint with configurable reasoning effort, 256K context, and image input. Its weights are published under a “Modified MIT” licence that permits commercial use with a carve-out for high-revenue companies, so it is neither Apache 2.0 nor API-only.
  • Mistral Small 4 (16 March 2026) is a 119B-total / ~6B-active MoE under Apache 2.0, with 256K context. It is the first Mistral model to unify reasoning, vision and agentic coding in one self-hostable checkpoint.
  • Ministral 3 (2 December 2025) covers 14B, 8B and 3B, each shipping base, instruct and reasoning variants with image understanding, all Apache 2.0. These replace the retired Mistral 7B and Mixtral open-weight line.

Choosing between them is mostly a cost-versus-capability call: Ministral 3 or Small 4 for high-volume, cost-sensitive work, Large 3 for complex reasoning, and Medium 3.5 when you want reasoning effort you can tune per request.

Alongside its own catalogue, Mistral now also serves a third-party model on la Plateforme: Z.ai’s GLM 5.2, with a 1M context window.

Specialist models: OCR and Lean

  • Mistral OCR 4.1 (mistral-ocr-4-1) was released on 16 July 2026 and became generally available on 31 August 2026. The mistral-ocr-latest and mistral-ocr-4 aliases point to it. It is API-only and priced at $4 per 1,000 pages (see pricing below). If you pinned mistral-ocr-4-0 (June 2026), move to the 4.1 id or the alias.
  • Leanstral 1.5 (labs-leanstral-1-5, 30 June 2026), the Apache 2.0 Lean 4 formal-proof engineering model, retires from the API on 30 September 2026; Mistral announced the date when it released the model. The models page lists no hosted successor, so if you rely on it, plan to self-host the open weights or switch provider before then. (Its predecessor, labs-leanstral-2603, retired on 30 June 2026.)

Company news: Pimento acquisition

Sifted reported on 22 September 2026, citing company filings, that Mistral is acquiring Pimento, a Paris-based startup that builds AI-assisted advertising creation, in a cash-and-shares deal worth €12.7 million. Sifted and FW.media both describe it as Mistral’s third acquisition of 2026, after Koyeb (infrastructure) and Emmi AI (industrial engineering). FW.media and other coverage link the purchase to Mistral’s Vibe assistant and business applications. Mistral had not published its own announcement in its changelog as of 25 September 2026, so treat the deal terms as reported rather than confirmed.

Where Mistral sits in your stack

Mistral is a model provider, not a full application platform. It supplies the intelligence layer that your application calls, whether you self-host the weights or hit the hosted API.

Your application
Web app Backend service Agent Sends prompts, receives completions
Access path
Hosted API Le Chat / Vibe Self-hosted weights Pick per data-control need
Models
Open-weight (Apache 2.0) Modified MIT (Medium 3.5) Commercial / premier Text, code, vision, speech, OCR
Compute
Mistral EU infrastructure Cloud partners Your own hardware

How to access it

You reach Mistral three ways, depending on how much control you want.

Le Chat, now Vibe. The consumer-facing chat product, comparable to other chat assistants. It runs Mistral’s models behind a web and mobile interface. Mistral rebranded it to Vibe in May 2026; sources disagree on the exact day, and the Vibe name had already been used for a separate terminal coding agent earlier that year. Use it to try the models before building anything.

The hosted API. Mistral serves both open-weight and commercial models through la Plateforme, an API with a developer console. You send a prompt, you get a completion, and Mistral runs the inference on its own infrastructure. Mistral states that its servers are hosted in the EU, which matters for teams with data-residency requirements. The platform is no longer Mistral-only: it also resells a third-party model, Z.ai’s GLM 5.2 with a 1M context window, which is currently the longest context you can reach through a Mistral endpoint.

Self-hosting the open weights. For the open-weight models, you download the weights and run them on your own GPUs or through a third-party inference host. This keeps every request inside your own perimeter. It costs more operationally, and you own the scaling and reliability work.

Step 1 Prototype in Vibe Test whether the models handle your task at all.
→
Step 2 Build on the API Wire the hosted API into your app for speed of delivery.
→
Step 3 Decide on control If data must stay in-house, move an open model onto your own compute.

Using the API

Mistral’s API is OpenAI-compatible. Use the official SDK, or point the openai package at Mistral’s base URL.

bash
pip install mistralai
python
from mistralai import Mistral

client = Mistral(api_key="YOUR_MISTRAL_API_KEY")

response = client.chat.complete(
    model="mistral-large-latest",
    messages=[{"role": "user", "content": "Summarise the EU AI Act in three bullet points."}]
)
print(response.choices[0].message.content)

Via the openai SDK (drop-in replacement):

python
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_MISTRAL_API_KEY",
    base_url="https://api.mistral.ai/v1"
)

chat = client.chat.completions.create(
    model="mistral-small-latest",
    messages=[{"role": "user", "content": "What is RAG?"}]
)

Function calling

Mistral supports OpenAI-compatible function calling on all large models.

python
import json
from mistralai import Mistral

client = Mistral(api_key="YOUR_MISTRAL_API_KEY")

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_company_info",
            "description": "Return firmenbuch (company register) data for an Austrian company.",
            "parameters": {
                "type": "object",
                "properties": {
                    "company_name": {"type": "string", "description": "Legal name of the company"},
                    "country": {"type": "string", "enum": ["AT", "DE", "CH"]}
                },
                "required": ["company_name", "country"]
            }
        }
    }
]

response = client.chat.complete(
    model="mistral-large-latest",
    messages=[{"role": "user", "content": "Look up Erste Bank AG in Austria."}],
    tools=tools,
    tool_choice="auto"
)

tool_call = response.choices[0].message.tool_calls[0]
args = json.loads(tool_call.function.arguments)
print(args)  # {'company_name': 'Erste Bank AG', 'country': 'AT'}

Codestral for code generation

Codestral remains Mistral’s dedicated code-completion model, currently at version stamp v25.08. Unlike the general-purpose lineup it sits in the Premier tier with closed weights, so you call it rather than self-host it. The older MNPL-licensed 22B Codestral is superseded. If you want a self-hostable coding model, Mistral now points you at Small 4 or Medium 3.5, which absorbed the Devstral agentic-coding line. Check the models page for the current context window before you design around it.

python
client = Mistral(api_key="YOUR_MISTRAL_API_KEY")

response = client.chat.complete(
    model="codestral-latest",
    messages=[
        {
            "role": "user",
            "content": "Write a FastAPI endpoint that accepts a PDF and returns extracted text using AWS Textract."
        }
    ]
)
print(response.choices[0].message.content)

Pricing (la Plateforme, as of 9 September 2026)

Mistral publishes per-model rates in US dollars on its API pricing page. Earlier versions of this page quoted euro figures for a lineup that no longer exists.

ModelInput per 1M tokensOutput per 1M tokens
Mistral Large 3$0.50$1.50
Mistral Medium 3.5$1.50$7.50
Mistral Small 4$0.15$0.60
Ministral 3 14B$0.20$0.20
Ministral 3 8B$0.15$0.15
Ministral 3 3B$0.10$0.10
Codestral$0.30$0.90
Mistral Embed$0.10n/a
Codestral Embed$0.15n/a

The specialised models price on their own units: Mistral OCR 4.1 at $4 per 1,000 pages, Voxtral TTS at $0.016 per 1,000 characters, and Voxtral Mini Transcribe at $0.003 per audio minute.

Typical use

Teams reach for Mistral when European data residency or the option to self-host is a hard requirement, not a nice-to-have. Common patterns:

  • Regulated workloads where data cannot leave EU infrastructure, so an EU-hosted API or self-hosted weights is the deciding factor.
  • On-premise or private-cloud deployment using an Apache 2.0 open-weight model, where owning the weights removes vendor lock-in.
  • Cost-sensitive backends that run a smaller open model locally instead of paying per-token for a commercial API.
  • Multilingual and code tasks. Mistral Large 3 covers 40+ languages, and Mistral ships dedicated models for coding (Codestral), speech (Voxtral), document OCR, embeddings and moderation alongside the general-purpose line.

How it compares

At the provider level, the distinguishing axes are where the company is based, whether you can get the weights, and where inference runs.

Mistral AIAnthropic (Claude)Alibaba (Qwen)Amazon Bedrock
OriginFranceUnited StatesChinaUnited States
Open weightsYes, including the flagshipNoYes, some modelsNo, it is a hosting layer
Access modelAPI and self-hostAPI onlyAPI and self-hostManaged multi-model API
Data hostingEU infrastructureUS-basedChina / globalYour chosen AWS region
Best forEU residency, self-host optionStrongest reasoning via APIOpen-weight multilingualOne API over many providers

Model against model, the flagship comparison looks like this:

Mistral Large 3GPT-5.6Claude Sonnet 5Llama 3.3 70B
Data residencyEU (Paris)USUSSelf-host or US
Open weightYes (Apache 2.0)NoNoYes (community licence)
Languages40+ (strong FR/DE)50+10+50+
Context window256K128K1M128K
Image inputYesYesYesText-only base model
Price (input/1M)$0.50~$4.50 (indicative)$2.00~$0.80 (host-dependent)
GDPR DPAYes (EU entity)SCCs requiredSCCs requiredSelf-host
Best forEU-regulated enterprise, self-host at frontier scaleGeneral purposeLong documentsCost-sensitive

Only the Mistral column is quoted from Mistral’s own published price list. GPT-5.6 and Llama figures above are indicative — check OpenAI API for current context windows and pricing across its Sol/Terra/Luna tiers. Claude Sonnet 5’s $2.00/MTok input price is now Anthropic’s permanent standard rate (see Claude Anthropic for the full current lineup and pricing table). Note that Llama is no longer Meta’s frontier line — see Meta Llama for what replaced it. The 2026 LLM landscape comparison places all these providers side by side.

When not to use it

  • You want a single API across many vendors. A managed aggregator like Amazon Bedrock or Azure OpenAI lets you switch models without changing providers.
  • You need a million-token context. Mistral’s own models top out at 256K. For 1M context, use Gemini or Claude Opus 5 — or the third-party GLM 5.2 that Mistral resells on la Plateforme, which carries a 1M window.
  • You need video or audio-native reasoning. Large 3, Medium 3.5, Small 4 and Ministral 3 all accept image input, so still images are no longer a reason to route elsewhere. Speech is handled by the separate Voxtral models rather than in the chat models, and there is no video input. For a single model that reasons over video, use Gemini .
  • You need the strongest available reasoning right now. Benchmark the specific task against Claude and others rather than assuming any single provider leads. See how AI models are evaluated .
  • You have no data-residency or self-host requirement. Mistral’s main differentiators are EU hosting and open weights. Without those needs, choose on capability and price alone.
  • You lack the operations capacity to self-host. Running open weights yourself means owning GPU provisioning, scaling, and uptime. If you cannot staff that, stay on a hosted API.

Further reading

Sources