A black prism splitting a red laser, representing an enterprise-focused model provider.
Cohere positions itself around precise retrieval and generation for regulated enterprises rather than a single flagship chat model.

Cohere is a model provider that builds foundation models for enterprises that need to keep data inside their own boundaries. It offers three product lines: Command models for generation, Embed models for turning text and images into vectors, and Rerank models that reorder search results by relevance. Cohere’s positioning centres on search and retrieval-augmented generation , plus deployment flexibility for companies that cannot send data to a public API.

The company packages these models under North, an enterprise AI platform for workplace productivity, and Compass, a search and discovery system. The underlying models are also available directly through Cohere’s API and through major cloud marketplaces.

In 2026 Cohere’s licensing stance changed materially. Command A+ (May 2026) and North Mini Code (June 2026) were both released with downloadable weights under a full Apache 2.0 licence, the first time Cohere has done that. The long-standing summary of Cohere as “private deployment, but not permissively open” no longer holds. It is not uniform either: North Small Translate (September 2026) ships open weights under the non-commercial CC BY-NC 4.0 licence, so check each model.

Where Cohere sits in the stack

Cohere spans two roles in a typical AI application: it supplies the generation model that writes answers, and it supplies the retrieval models that decide which documents feed those answers.

Application
North platform Compass search Enterprise workplace agents and search
Generation
Command A+ Command A North Mini Code North Small Translate Command R7B Tool use, agents, RAG, agentic coding
Retrieval
Embed v4.0 Rerank v4.0-pro / v4.0-fast Parse v5.0 Vectors, relevance scoring, and document parsing for RAG
Deployment
Cohere API VPC On-premises Bedrock / Azure / SageMaker / OCI

How to access it and how it fits

You can reach Cohere’s models four ways: the hosted Cohere API, a private deployment inside your own virtual private cloud (VPC), a fully on-premises install, and cloud marketplaces. Cohere lists availability across Amazon Bedrock, Amazon SageMaker, Microsoft Azure, and Oracle Generative AI Service. In September 2025 the company added Model Vault, a dedicated inference platform that runs Command, Embed, and Rerank inside isolated VPC or on-premises environments.

The models divide by job:

Step 1 Embed Convert documents into vectors with Embed v4.0. It handles text, images, and PDFs with a 128K context window.
→
Step 2 Retrieve A vector search returns candidate documents for a query.
→
Step 3 Rerank Rerank v4.0-pro or v4.0-fast reorders candidates by relevance across documents, tables, JSON, and code, with a 32K context.
→
Step 4 Generate A Command model reads the top documents and writes a grounded answer with citations.

The model lineup

Command A+ (command-a-plus-05-2026, announced 20 May 2026) is the flagship and the most consequential release Cohere has made in years. It is the company’s first mixture-of-experts model, at 218B total parameters with 25B active, handling 128K input tokens and up to 64K generated tokens across 48 languages including every official EU language. Cohere released it under Apache 2.0 and positions it explicitly for sovereign and air-gapped deployment. The licence is stated by Cohere itself, not only by coverage of it: the Command A+ docs page says the model “is available under an Apache 2.0 License on Hugging Face”, and the weights sit in the CohereLabs repository there. As always, read the licence file in the repository before you build a deployment that depends on it.

North Mini Code 1.0 (9 June 2026) is Cohere’s first fully open developer-facing model, and the first in a new North model family: a 30B MoE with about 3B active parameters, purpose-built for agentic software engineering, with 256K input and 64K output and Apache 2.0 weights that run on a single H100 in FP8. The production model id reached the Chat V2 API in August 2026.

North Small Translate 1.0 (north-small-translate-1-0, 9 September 2026) is a mixture-of-experts model built only for machine translation, covering more than 50 languages and locale variants. It has 218B total and 25B active parameters and a 16K context window. It is on the free tier of the Chat V2 API, and the weights are on Hugging Face (CohereLabs/North-Small-Translate-1.0) in W4A16, FP8 and BF16 formats under CC BY-NC 4.0, so non-commercial use only. Cohere’s suggested hardware is two H100s (or one B200) for W4A16, four H100s for FP8 and eight H100s for BF16. For commercial translation workloads, use the API or talk to Cohere about a licence; do not assume the Apache 2.0 terms of Command A+ carry over.

Cohere Parse (parse-v5.0, 27 August 2026) converts complex documents into structured Markdown for downstream AI pipelines. It is a 2.3B-parameter multimodal model (about 4.6 GB) with an 8K context window that extracts reading-order text, tables, lists, forms, images and captions, page boundaries and element locations, returning Markdown or HTML, HTML tables, bounding boxes and image descriptions. It is available through the Parse API, Microsoft Foundry, AWS SageMaker and Model Vault. It is the ingestion step in front of Embed and Rerank in a Cohere RAG stack.

The rest of the Command line covers narrower jobs: Command A (command-a-03-2025, 256K) for tool use, agents, and RAG; command-a-reasoning-08-2025 (256K); command-a-vision-07-2025 (128K); command-a-translate-08-2025 (8K); and Command R7B (128K), a small fast model for RAG and tool use. Cohere also ships cohere-transcribe-03-2026 and cohere-transcribe-arabic-07-2026 for speech, and the multilingual Aya research models (c4ai-aya-expanse-32b, c4ai-aya-vision-32b, and the small open tiny-aya series). In early September 2026 Cohere Labs added three 3.35B-parameter models to the tiny-aya series on Hugging Face: tiny-aya-en-thinker and tiny-aya-l2-thinker (2 September, reasoning variants) and tiny-aya-base-32K (8 September, a long-context base model). All three are gated behind an access form and licensed CC BY-NC 4.0, so they are research models, not commercial building blocks.

Check deprecations before you pin an old id. Cohere deprecated command-r-03-2024, command-r-plus-04-2024, command, command-light, and the command-r / command-r-plus aliases on 15 September 2025, and retired c4ai-aya-expanse-8b and c4ai-aya-vision-8b on 4 April 2026.

Compared to other model providers

Cohere is narrower than the general-purpose labs but deeper on retrieval. Here is how it lines up.

CohereAnthropicMistral AIAI21 Labs
Core focusEnterprise RAG and searchFrontier reasoning modelsOpen-weight and hosted modelsJamba long-context models, now discontinued
Retrieval modelsEmbed, Rerank, ParseNone first-partyEmbed modelNone first-party
DeploymentAPI, VPC, on-prem, clouds, Apache 2.0 weightsAPI and cloud marketplacesAPI, cloud, Apache 2.0 weightsAPI and cloud, existing endpoints only
Best forRegulated RAG at scale, sovereign deploymentComplex reasoning tasksCost-flexible general useNothing new: AI21 stopped selling standalone models in May 2026

For a wider view of how these vendors relate, see the LLM landscape 2026 comparison .

When not to use it

Cohere is a focused choice, not a default. Consider alternatives when:

  • You want the top reasoning benchmarks. The largest frontier chat models from other labs often lead on public reasoning leaderboards. Cohere optimises for enterprise retrieval and deployment, not headline scores.
  • You need a large consumer ecosystem. Cohere sells to enterprises. If you want a broad third-party plugin and app ecosystem, other providers offer more.
  • You only need a chatbot. If you are not doing search or RAG, the Embed and Rerank strengths that differentiate Cohere go unused, and a simpler single-model provider may cost less.
  • You need open weights across the whole catalogue. This is no longer a blanket objection: Command A+ and North Mini Code ship under Apache 2.0. But Embed, Rerank, Parse, and the older Command models remain API and private-deployment only, and North Small Translate and the tiny-aya models are open but non-commercial (CC BY-NC 4.0). Check the specific model you need rather than assuming the whole line is downloadable or commercially usable.

Further reading

Sources