Sunday, 13 September 2026 | Updating Daily AI insight, written for builders

Mistral Large 3

Mistral Large 3 — Especificaciones

Compiled by Mustafa Ihsan from the vendor’s published documentation · Last updated

Desarrollador Mistral AI
Tipo LLM (MoE)
Modalidad Texto → Texto
Parámetros 675 000 millones totales / 41 000 millones activos (MoE)
Ventana de contexto 256K
Licencia Apache 2.0 (abierta)
Pesos abiertos
Lanzado 2025
Precio de entrada $2.00 /1M
Precio de salida $6.00 /1M
Proveedores de API Mistral, OpenRouter

Ejecútelo localmente

VRAM (4 bits) ~400 GB
GPU mínima Servidor multi-GPU

Página oficial →

What is Mistral Large 3?

Mistral Large 3 marks the company’s return to fully open licensing — a 675B
mixture-of-experts activating 41B parameters per token, released under Apache 2.0 with a 256K
context. Pricing is $2 in / $6 out per million tokens.

Apache 2.0 at this scale is the headline. Most large open models carry either a custom
community licence with conditions attached (Llama 4’s EU restriction being the obvious
example) or a modified MIT. Apache 2.0 is unambiguous, permissive, patent-granting and
already approved inside most legal departments, which removes the review cycle that stalls
open-model adoption in enterprises. Combined with European provenance, that makes Large 3 a
straightforward choice for organisations with data-governance requirements that rule out
other options. The 3:1 output-to-input ratio is also gentler than most frontier models,
making generation-heavy workloads relatively less punishing. Self-hosting needs around 400 GB
at 4-bit, so the licence buys portability and auditability rather than a workstation
deployment.

Mistral Large 3 pricing: API cost per 1M tokens

Entrada (por cada millón de tokens)$2.00
Salida (por cada millón de tokens)$6.00
Relación salida/entrada
Combinada (4:1 entrada:salida)$2.80 por 1 millón de tokens

What Mistral Large 3 costs per month

Gasto mensual real con una mezcla entrada:salida de 4:1, es decir, la proporción que realmente genera una carga de trabajo típica de chat o RAG.

Carga de trabajoTokens/mesCoste mensual
Proyecto secundario 1 millón de tokens de entrada / 0,25 millones de tokens de salida $3.50
Pequeño equipo 20 millones de tokens de entrada / 5 millones de tokens de salida $70
Producción 200 millones de tokens de entrada / 50 millones de tokens de salida $700

Calcule sus propios números en la Calculadora de costos de API de IA.

Cheaper alternatives to Mistral Large 3

ModeloDólares por millón combinadosUsted ahorra
GLM 5.2 abierta $2.00 29 % más barato
DeepSeek V4-Pro abierta $0.522 81% cheaper
Kimi K2.7 Code abierta $0.980 Un 65 % más barato

¿Autoalojarlo o pagar por la API?

Mistral Large 3 is open-weight, so you can run it yourself. It needs ~400 GB of VRAM at 4-bit (Multi-GPU server). Self-hosting only beats the API once your volume is high enough to keep that hardware busy — the calculadora de autohospedaje frente a API calcula el punto de equilibrio para su volumen de tokens.

Preguntas frecuentes

How much does Mistral Large 3 cost per 1M tokens?

Mistral Large 3 costs $2.00 per 1M input tokens and $6.00 per 1M output tokens. At a typical 4:1 input-to-output mix that blends to about $2.80 per 1M tokens.

How much does Mistral Large 3 cost per month?

A small-team workload of 20M input and 5M output tokens a month costs about $70 on Mistral Large 3. A side project (1M in / 0.25M out) costs roughly $3.50.

What is a cheaper alternative to Mistral Large 3?

GLM 5.2 is the strongest cheaper option in our database at $2.00 per 1M blended — about 29% less than Mistral Large 3. It is also open-weight, so self-hosting is an option.

Can I run Mistral Large 3 locally?

Yes. Mistral Large 3 is open-weight and needs about ~400 GB of VRAM at 4-bit quantisation (Multi-GPU server).

Why does Mistral Large 3 charge more for output than input?

Output tokens are generated one at a time and cannot be batched the way a prompt can, so they cost the provider more to serve. Mistral Large 3 charges 3× more for output, which is why prompt-heavy workloads are far cheaper to run than generation-heavy ones.

Los precios corresponden a las tarifas oficiales publicadas para la API principal del modelo y se revisan periódicamente conforme los proveedores los actualicen. No incluyen descuentos por volumen, procesamiento por lotes ni entradas en caché. Compare todos los modelos uno al lado del otro en la Base de datos de modelos de IA o el Clasificación de modelos de lenguaje grande (LLM).

Scroll to Top