How To Evaluate Mistral Large 4 For Global AI And Agent Workloads
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How To Evaluate Mistral Large 4 For Global AI And Agent Workloads on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 is available as a research public preview through Mistral’s API, with model weights promised for the end of October. Artificial Analysis Index v4.3.2 gives it a score of 38.4, below current US and several Chinese models; reported task costs and output volume also raise questions for agent deployments. Its evaluation scores may change, and its eventual weights licence has not been published.

Mistral has released Large 4 as a research public preview through its API, presenting a new European flagship model for AI workloads. The model scores 38.4 on the Artificial Analysis Intelligence Index v4.3.2, according to the cited independent benchmark, but the same data places it below current US frontier models and several Chinese models, while reported task costs and high output volume warrant testing before agent deployment.

Large 4 is described as a one-trillion-parameter model with 49 billion active parameters, supporting text and image input, text output, and a 512,000-token context window. It is currently accessible as a research public preview on Mistral’s API. Mistral has said the weights are expected at the end of October; until release, the model is proprietary, and the weights licence has not been published.

Artificial Analysis Index v4.3.2 assigns Large 4 a score of 38.4. In the source’s comparison, Claude Opus 5.5 scores 57.6, GPT-6 Astra 52.7, and the Chinese models GLM-5.3 and Kimi K3 score 44.8 and 43.6. Large 4 is ahead of some older models, including GLM-5.2 at 33.7 and DeepSeek V4 Pro 0813 at 36.0. Those rankings describe the models and benchmark version cited; they do not establish how every system will perform on a buyer’s own tasks.

The source reports API prices of $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million tokens. It says Mistral offered a 50% discount for the first two weeks. On the Index workload, Large 4 reportedly used 200 million output tokens, compared with a median of 81 million for comparable models, and cost $1.13 per task. The source says GLM-5.3-Flash scored 41.8 at $0.25 per task and DeepSeek V4.1 Flash scored 39.5 at $0.27; task-cost comparisons depend on the benchmark’s workload and pricing assumptions.

At a glance
analysisWhen: Released yesterday, according to the so…
The developmentMistral has released Large 4 as a research public preview, prompting questions about its independent benchmark performance, cost, and suitability for global AI and agent workloads.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

The Cost of Running Agents

The benchmark matters for teams considering multi-step agents, where a model may call tools, inspect results, and revise its plan repeatedly. Artificial Analysis’s index includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench, and Terminal-Bench 4.0, according to the source. That makes its score relevant to some work and coding tasks, though it is not a substitute for testing a specific workflow.

For an agent, a weak answer can become a faulty premise for later actions. The source author also reports seeing confident hallucinations in hands-on use; this is an individual observation, not an Artificial Analysis measurement or a published rate for all users. Output volume can also affect latency and bills when a system generates many tokens across repeated steps. Buyers should measure success rate, verification needs, completion time, and total task cost rather than judging by per-token price or a single index score.

The result also puts Mistral’s position in perspective. The source characterizes Large 4’s jump from Mistral Large 3’s score of 9 to 38.4 as a major improvement on the same index version. That is meaningful progress for the company, but progress from its prior model does not make it the leading choice against the cited current competitors.

Amazon

AI model evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Status and Benchmark Scope

Large 4’s status affects how procurement teams should interpret any comparison. It is a proprietary preview model, and the planned weights release and licence terms are not yet available in the source material. Teams evaluating a hosted API can test it now, but organizations requiring self-hosting or a specific open-weight licence cannot yet judge whether the eventual release will meet those needs.

The Artificial Analysis figures cited are from Index v4.3.2, which the source identifies as the current version. Keeping comparisons on one version helps avoid mixing scores from different benchmark revisions. Still, index results aggregate selected evaluations; they do not establish reliability in a particular language, industry, data environment, or tool setup. The source also says Mistral’s reinforcement learning is still running and that scores may move, so the published preview result may not be final.

For global deployments, buyers should separately test language quality, regional coverage, data handling, latency, and service availability. The provided material does not give enough evidence to rank Large 4 across those criteria. Nor does a model’s country of origin alone determine its suitability: deployment terms, support, compliance requirements, and task performance need direct review.

“In hands-on testing I saw Large 4 assert things confidently that weren’t true.”

— ThorstenMeyerAI.com source author

Amazon

large language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Preview Results Cannot Settle

Final benchmark performance is not settled. Mistral says reinforcement learning is ongoing, and the source provides no later evaluation showing whether that work changes the Index score or task costs. The weights and their licence are also pending, so it remains unknown whether customers will be able to run the model locally and under what conditions.

The source does not provide a full methodology for its hands-on hallucination observations or enough details to treat the cited per-task costs as universal procurement estimates. Actual spending will vary with prompt length, output length, caching, retries, and workflow design. It is also unclear how Large 4 performs across different languages and specific business applications; the supplied benchmark table does not answer those questions.

Amazon

AI workload cost management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the Model on Real Work

Organizations considering Large 4 can use the current API preview to run side-by-side trials against their existing models. Tests should reflect real prompts and tools, with fixed success criteria and measurement of end-to-end task cost, latency, error rates, and human review time. Agent trials should include checks for unsupported claims and whether later steps amplify an incorrect answer.

The next stated release milestone is the planned weights release at the end of October. Buyers should check Mistral’s release notes for the actual date, licence, and final model specifications, and watch for updated independent benchmark results. Until then, Large 4’s reported score and pricing are useful starting points for evaluation, not conclusive evidence that it is the right model for a global or agent workload.

Amazon

text and image input AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4’s Artificial Analysis score?

It scores 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source. Mistral says reinforcement learning is ongoing, so the score may change.

Can companies run Large 4 on their own infrastructure?

Not on the basis of the current preview described in the source: Large 4 is available through Mistral’s API as a proprietary research preview. Mistral has promised weights for the end of October, but the licence has not been published.

Is Large 4 a good choice for AI agents?

The cited benchmark includes agent-oriented tasks, but it cannot settle performance for every workflow. Teams should test task completion, error propagation, output volume, latency, and total cost using their own tools and data before deployment.

How does its reported task cost compare with alternatives?

The source reports $1.13 per Artificial Analysis Index task for Large 4, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. These are benchmark-specific estimates and may not match costs on a company’s own workload.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Apple Is Reportedly Preparing iPhone Price Hikes (Because Of Course They Are)

Apple is reportedly planning to raise iPhone prices in 2024, marking a shift in their pricing strategy amid market and production pressures.

Unveiling SpaceXAI’s Grok 4.6: The AI Revolution With GPT-5.6 And Fable 5-Level Intelligence

SpaceXAI announced the release of Grok 4.6, claiming it reaches the intelligence level of GPT-5.6 and Claude Fable 5, but no independent verification has been provided.

Twitter Surges In Global Coverage

Twitter’s international media mentions surge over 8-fold, highlighting its expanding global influence and coverage across diverse regions.

Qualcomm Incorporated Surges In Global Coverage

Media coverage of Qualcomm Inc. has spiked significantly, with 12 mentions in recent reports, indicating increased global interest in the company’s activities.