Three small glowing spheres converging, representing a family of small, efficient language models.
Phi is a family of small models tuned so that quality does not have to scale with size.

Microsoft Phi is a family of small language models (SLMs) released as open weights under the MIT license. The models solve a specific problem: most capable large language models are big, slow, and expensive to run, which puts them out of reach for phones, laptops, and cost-sensitive workloads. Phi trades raw scale for carefully curated training data, aiming to keep quality high while the parameter count stays small enough to run on modest hardware.

A small language model is a foundation model with far fewer parameters than a frontier system. Parameters are the learned weights a model uses to generate output. Fewer parameters mean smaller memory footprint, faster inference , and lower cost per request. Microsoft’s bet with Phi is that data quality, not sheer size, drives much of a model’s usefulness. Phi models are trained on heavily filtered and synthetic “textbook-quality” data rather than the whole web.

Where Phi sits

Phi occupies the small end of the model-size spectrum. You reach for it when a frontier model is more than the task needs, or when the deployment target cannot host one.

Frontier models
GPT class Claude Highest capability, hosted, higher cost and latency
Mid-size open models
Llama Mistral Strong general models, still need server-class GPUs
Small language models
Phi-4 Phi-4-mini Phi-4-multimodal Phi-4-reasoning-vision Runs on-device or on cheap GPUs, low latency, open weights
Deployment target
Laptop Phone Edge device Small cloud instance

The Phi family

Microsoft has shipped several generations. The current Phi-4 line covers a text model, a compact model, a multimodal model, reasoning-tuned variants, and, since March 2026, a vision-language reasoning model. There is no Phi-5: third-party sites describing one are not backed by any Microsoft announcement or model card. No new Phi model has shipped since May 2026. A check of Microsoft’s Hugging Face organisation on 25 September 2026 shows Phi-Ground-Any as the most recent Phi-branded upload. It is a GUI-grounding model, not a new general-purpose Phi. Microsoft’s new model releases over the summer have all been MAI models (see below).

  • Phi-4-reasoning-vision-15B is the newest general-purpose member, published on 4 March 2026. It is a 15 billion parameter vision-language reasoning model built on the Phi-4-Reasoning backbone with a SigLIP-2 vision encoder in a mid-fusion design. Its distinguishing feature is selective reasoning: the model decides when to emit a chain of thought and when to answer directly, which keeps token cost down on easy inputs. Microsoft published the weights, fine-tuning code, and benchmark logs under the MIT licence on Hugging Face, GitHub, and Microsoft Foundry. If you are picking a Phi variant today for anything involving images plus reasoning, start here - but check the context length first: the research blog omits it, and the model card puts it at 16,384 tokens, the same short window as base Phi-4 rather than the 128k of Phi-4-multimodal.

  • Phi-4 is a 14 billion parameter text model, first presented in December 2024. It is built on a decoder-only Transformer, was pretrained on roughly 10 trillion tokens of curated and synthetic data, and supports a 16k-token context length. Microsoft targeted mathematics and multi-step reasoning with this release.

  • Phi-4-mini is a 3.8 billion parameter model aimed at even lighter deployment.

  • Phi-4-multimodal is a 5.6 billion parameter model that handles speech, vision, and text in one model using a mixture-of-LoRAs design, with a 128k-token context length. Microsoft reports it ranked first on the Hugging Face OpenASR leaderboard with a 6.14% word error rate at the time of release.

  • Phi-4-reasoning (14B) and Phi-4-reasoning-plus (14B) are reasoning-tuned variants. Phi-4-reasoning-plus is further trained with reinforcement learning to spend more inference-time compute. Phi-4-reasoning-plus supports a 32k-token context by default.

  • Phi-4-mini-reasoning (3.8B) targets multi-step mathematical problem solving at small size.

  • Phi-Ground-Any-4B (Hugging Face microsoft/Phi-Ground-Any, published May 2026, MIT) is a specialist. It is fine-tuned from Phi-3.5-vision-instruct for GUI grounding in computer-use agents, locating on-screen elements from an instruction, and it requires a fixed 1680×1008 input resolution. Pick it only if you are building a computer-use agent, not as a general Phi model.

Earlier generations remain available too. The Phi-3.5 line, released in August 2024, includes Phi-3.5-mini (3.82B), Phi-3.5-vision (4.15B), and Phi-3.5-MoE, a mixture-of-experts model with 41.9 billion total parameters that activates about 6.6 billion per token. All three support a 128k-token context.

Phi is no longer the whole Microsoft model story

Phi is Microsoft’s small-model research line. Since Build 2026 it is not Microsoft’s only first-party model family, and the MAI line is where Microsoft’s new releases are now happening.

The June launch. On 2 June 2026 Microsoft AI announced seven MAI models. The body text of that announcement names the versions shipped then: MAI-Thinking-1 (reasoning), MAI-Code-1-Flash (agentic coding, 5B active parameters, built for GitHub Copilot and VS Code), MAI-Image-2.5 with a Flash variant, MAI-Transcribe-1.5 (43 languages), MAI-Voice-2 (15 languages), and MAI-Voice-2-Flash (“coming soon”). The page has since been edited, last modified 29 July, and some headings now carry later version numbers such as MAI-Image-2.6 and MAI-Transcribe-2. That explains the version inconsistency earlier versions of this page flagged. Those later versions are separate releases, listed below.

What shipped between August and September 2026, per Microsoft AI’s own announcements:

ModelDateWhat it isAvailability
MAI-Image-2.610 August 2026Text-to-image and editing; Microsoft reports No. 2 on the Arena text-to-image leaderboard, +79 Elo over MAI-Image-2.5Public preview in Microsoft Foundry from 4 September
MAI-Code-1.1-Flash11 August 2026Successor to MAI-Code-1-Flash; Microsoft claims 25% fewer tokens per task, 22% better on Terminal-Bench 2.1 in Copilot CLI, at a quarter of 1.0’s priceIn production in GitHub Copilot; GitHub deprecated MAI-Code-1-Flash on 10 September 2026
MAI-Thinking-112 August 2026Mid-size reasoning model: a sparse mixture-of-experts with 35B active and roughly 1T total parameters; Microsoft reports 97.0% on AIME 2025, parity with Claude Opus 4.6 on SWE-Bench Pro, and a preference over Claude Sonnet 4.6 in blind human side-by-sides (1,276 tasks, rated by Surge)Public preview in Microsoft Foundry; no price published
MAI-Cyber-1-Flash13 August 2026Vulnerability-finding model embedded in MDASH, Microsoft’s multi-agent vulnerability identification and remediation harness; designed to handle about 90% of tasks and hand the hardest 10% to a larger model (GPT-5.4); Microsoft claims 96% on CyberGym at half the cost of its previous MDASH configurationInside MDASH, not announced as a standalone API model
MAI-Transcribe-23 September 2026Speech-to-text with diarization, configurable verbatim/clean styles, word-level timestamps and code-switching; Microsoft reports first place on FLEURS across 60 languages at 5.2% average WERMicrosoft Foundry, MAI Playground and OpenRouter; $0.10 per audio hour as a limited-time offer to the end of 2026
MAI-Image-2.6-Flash4 September 2026Faster sibling of MAI-Image-2.6 for high-throughput work; Microsoft claims 2.8x faster than GPT-Image-2-Medium; both 2.6 models support multi-image reference editing, web grounding and dynamic aspect ratiosPublic preview in Microsoft Foundry

All benchmark figures in the table are Microsoft’s own and have not been independently reproduced here. On 14 September 2026 Microsoft AI also published a draft Code of Conduct for MAI models for a six-week public consultation. It sets out how the models are meant to behave, what they must never do and who they answer to. It is worth reading if you are assessing MAI for regulated use.

Two things make MAI strategically different from Phi. First, Microsoft states the models were trained from scratch on commercially licensed, “clean, traceable” data with no distillation from third-party models, a deliberate step away from depending on OpenAI for frontier capability. Second, they are hosted services distributed through Microsoft Foundry and third-party hosts such as OpenRouter, Fireworks and Baseten, not open-weight downloads. The June announcement said developers could tune the weights themselves, but no open-weight licence has been published. If you need something you can download and run inside your own perimeter, Phi is still Microsoft’s only option. Most MAI models are in public preview, so confirm availability, pricing and SLA terms with Microsoft before you design around one.

How to access it

Phi models are open weights. You do not need a Microsoft account to download and run them.

Step 1 Pick a variant Match model size to hardware and task. Use mini for edge, Phi-4 for general text, multimodal for speech and vision, Phi-4-reasoning-vision for image-plus-reasoning work.
→
Step 2 Get the weights Download from Hugging Face under the MIT license, or select the model in Microsoft Foundry (formerly Azure AI Foundry).
→
Step 3 Run or host Run locally with common inference runtimes, or serve it as a managed endpoint through Azure.

The MIT license allows free use, modification, and distribution, including for commercial products. Phi-4 and the reasoning variants are published on Hugging Face and in the Microsoft Foundry catalog - Microsoft’s 2026 announcements use “Microsoft Foundry” for what was previously branded Azure AI Foundry. If you already run other Microsoft-hosted models through Azure OpenAI Service , Foundry gives you Phi, the MAI models, and hosted OpenAI models in one catalog without changing clouds.

How it compares

Phi competes with other small and open model families. The comparison below is about positioning, not a benchmark ranking.

Phi-4 lineMistral small modelsDeepSeek distills
MakerMicrosoftMistral AIDeepSeek
Size focus3.8B to 15B3B to 119B (MoE)Distilled small variants
LicenseMIT and permissive (open weights)Apache 2.0 on Ministral 3 and Small 4Open weights on many models
StrengthReasoning at small sizeGeneral European multilingualDistilled reasoning
Best forOn-device, cost-sensitive appsBroad general useReasoning on a budget

For the mid-size and multilingual end, see Mistral AI . For distilled reasoning models released as open weights, see DeepSeek .

When not to use it

Small models trade capability for size. Phi is the wrong choice when:

  • You need frontier-level breadth. For the hardest open-ended reasoning, broad world knowledge, or long complex documents, a large model still leads. Phi-4’s base text context is 16k tokens, smaller than many hosted frontier models.
  • You need the widest tool and ecosystem support. Frontier hosted APIs ship mature tool-calling, function-calling, and safety tooling. Verify Phi’s support for your exact features before committing.
  • Accuracy on rare edge cases is safety-critical. A smaller parameter count means less capacity to memorise long-tail facts. Add retrieval or human review for high-stakes output.
  • You have no capacity to self-host and want a fully managed frontier experience. In that case a hosted API may be less operational work, even at higher cost per call. Within Microsoft’s own catalog that now means the MAI models or hosted OpenAI models in Foundry rather than a Phi download.

Match the model to the job. Phi shines when latency, cost, or on-device privacy matter more than absolute peak capability.

Further reading

Sources