🔍 Read the full analysis: How To Evaluate Mistral Large 4 For Global AI And Agent Workloads on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral Large 4 is available as a research public preview through Mistral’s API, with model weights promised for the end of October. Artificial Analysis Index v4.3.2 gives it a score of 38.4, below current US and several Chinese models; reported task costs and output volume also raise questions for agent deployments. Its evaluation scores may change, and its eventual weights licence has not been published.
Mistral has released Large 4 as a research public preview through its API, presenting a new European flagship model for AI workloads. The model scores 38.4 on the Artificial Analysis Intelligence Index v4.3.2, according to the cited independent benchmark, but the same data places it below current US frontier models and several Chinese models, while reported task costs and high output volume warrant testing before agent deployment.
Large 4 is described as a one-trillion-parameter model with 49 billion active parameters, supporting text and image input, text output, and a 512,000-token context window. It is currently accessible as a research public preview on Mistral’s API. Mistral has said the weights are expected at the end of October; until release, the model is proprietary, and the weights licence has not been published.
Artificial Analysis Index v4.3.2 assigns Large 4 a score of 38.4. In the source’s comparison, Claude Opus 5.5 scores 57.6, GPT-6 Astra 52.7, and the Chinese models GLM-5.3 and Kimi K3 score 44.8 and 43.6. Large 4 is ahead of some older models, including GLM-5.2 at 33.7 and DeepSeek V4 Pro 0813 at 36.0. Those rankings describe the models and benchmark version cited; they do not establish how every system will perform on a buyer’s own tasks.
The source reports API prices of $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million tokens. It says Mistral offered a 50% discount for the first two weeks. On the Index workload, Large 4 reportedly used 200 million output tokens, compared with a median of 81 million for comparable models, and cost $1.13 per task. The source says GLM-5.3-Flash scored 41.8 at $0.25 per task and DeepSeek V4.1 Flash scored 39.5 at $0.27; task-cost comparisons depend on the benchmark’s workload and pricing assumptions.
Mistral Large 4: best outside the US and China — and still not a model to run your agents on
The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.
~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.
Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.
Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.
The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.
AA v4.3.2Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.
AAConfident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.
AUTHOR’S TESTING · not an AA figure- Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
- Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
- Speed: 116 tok/s, 1.46s TTFT — well above median.
- The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
- Jurisdiction: French parent, EU hosting, weights promised end of October.
- Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
- Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
- Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.
The Cost of Running Agents
The benchmark matters for teams considering multi-step agents, where a model may call tools, inspect results, and revise its plan repeatedly. Artificial Analysis’s index includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench, and Terminal-Bench 4.0, according to the source. That makes its score relevant to some work and coding tasks, though it is not a substitute for testing a specific workflow.
For an agent, a weak answer can become a faulty premise for later actions. The source author also reports seeing confident hallucinations in hands-on use; this is an individual observation, not an Artificial Analysis measurement or a published rate for all users. Output volume can also affect latency and bills when a system generates many tokens across repeated steps. Buyers should measure success rate, verification needs, completion time, and total task cost rather than judging by per-token price or a single index score.
The result also puts Mistral’s position in perspective. The source characterizes Large 4’s jump from Mistral Large 3’s score of 9 to 38.4 as a major improvement on the same index version. That is meaningful progress for the company, but progress from its prior model does not make it the leading choice against the cited current competitors.
As an affiliate, we earn on qualifying purchases.
Preview Status and Benchmark Scope
Large 4’s status affects how procurement teams should interpret any comparison. It is a proprietary preview model, and the planned weights release and licence terms are not yet available in the source material. Teams evaluating a hosted API can test it now, but organizations requiring self-hosting or a specific open-weight licence cannot yet judge whether the eventual release will meet those needs.
The Artificial Analysis figures cited are from Index v4.3.2, which the source identifies as the current version. Keeping comparisons on one version helps avoid mixing scores from different benchmark revisions. Still, index results aggregate selected evaluations; they do not establish reliability in a particular language, industry, data environment, or tool setup. The source also says Mistral’s reinforcement learning is still running and that scores may move, so the published preview result may not be final.
For global deployments, buyers should separately test language quality, regional coverage, data handling, latency, and service availability. The provided material does not give enough evidence to rank Large 4 across those criteria. Nor does a model’s country of origin alone determine its suitability: deployment terms, support, compliance requirements, and task performance need direct review.
“In hands-on testing I saw Large 4 assert things confidently that weren’t true.”
— ThorstenMeyerAI.com source author
As an affiliate, we earn on qualifying purchases.
What Preview Results Cannot Settle
Final benchmark performance is not settled. Mistral says reinforcement learning is ongoing, and the source provides no later evaluation showing whether that work changes the Index score or task costs. The weights and their licence are also pending, so it remains unknown whether customers will be able to run the model locally and under what conditions.
The source does not provide a full methodology for its hands-on hallucination observations or enough details to treat the cited per-task costs as universal procurement estimates. Actual spending will vary with prompt length, output length, caching, retries, and workflow design. It is also unclear how Large 4 performs across different languages and specific business applications; the supplied benchmark table does not answer those questions.
AI workload cost management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test the Model on Real Work
Organizations considering Large 4 can use the current API preview to run side-by-side trials against their existing models. Tests should reflect real prompts and tools, with fixed success criteria and measurement of end-to-end task cost, latency, error rates, and human review time. Agent trials should include checks for unsupported claims and whether later steps amplify an incorrect answer.
The next stated release milestone is the planned weights release at the end of October. Buyers should check Mistral’s release notes for the actual date, licence, and final model specifications, and watch for updated independent benchmark results. Until then, Large 4’s reported score and pricing are useful starting points for evaluation, not conclusive evidence that it is the right model for a global or agent workload.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Mistral Large 4’s Artificial Analysis score?
It scores 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source. Mistral says reinforcement learning is ongoing, so the score may change.
Can companies run Large 4 on their own infrastructure?
Not on the basis of the current preview described in the source: Large 4 is available through Mistral’s API as a proprietary research preview. Mistral has promised weights for the end of October, but the licence has not been published.
Is Large 4 a good choice for AI agents?
The cited benchmark includes agent-oriented tasks, but it cannot settle performance for every workflow. Teams should test task completion, error propagation, output volume, latency, and total cost using their own tools and data before deployment.
How does its reported task cost compare with alternatives?
The source reports $1.13 per Artificial Analysis Index task for Large 4, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. These are benchmark-specific estimates and may not match costs on a company’s own workload.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
