Mistral Large 4 Data Analytics Benchmark Results
$20 to $2 in five months, and two letter grades better on our data analytics exam.
At Plotly, we keep a close eye on model releases. Our agentic analytics products, Plotly Studio Desktop and Plotly Studio Embedded, and our open source libraries Plotly and Dash work with any model, and our customers look to us for recommendations. Today, our European customers got big news: Mistral dropped their latest model.
Mistral Large 4, aka Le Chonk, is an open-weight model with 1T total parameters and 49B active parameters (that's large!).
We ran it through our data analytics exam and found a generational improvement over Mistral's last release, Mistral Medium 3.5: 10x cheaper ($20.51 to $1.94 per run) and two letter grades better (an F to a C), in about five months. Mistral is in the same ballpark as the best models, and is becoming a great option for our European customers seeking an open, sovereign model 👏

The exam
Our data analytics exam tests how well models do at real data analysis. It runs within Plotly Studio Embedded, with the same agent, tools and prompts that app viewers use and only the model swapped, and it answers 43 questions through a series of tool calls and analysis steps. The data warehouse is representative of the real world: 20M rows across 44 tables and 554 columns. The questions cover a broad range of data analytic skills: answering questions, recognizing ambiguity, avoiding hallucinations, and finding a needle in a haystack.
The dataset is a mock wind turbine analytics dataset: 5,200 turbines across 28 wind farms over two years. It covers everything from hourly SCADA telemetry and alarms to work orders, inspections, oil samples, firmware rollouts and weather. And like real operational data it's messy, with frozen sensors, duplicated rows and turbines decommissioned mid-period, so a good analyst has to catch problems before quoting a number.
Every question has either an answer, checked against a verified key, or "a catch": an ambiguous question to raise, an impossible request to decline, a planted data defect to flag. There is no partial credit, and a confident answer that ignores the catch fails.
Mistral Medium 3.5, released in April, passed 25 of 43 questions (58%, an F). Mistral Large 4 passes 32 of 43 (74%, a C), and it improved in every category but one:
Category
Mistral Medium 3.5
Mistral Large 4
Analytics questions
16 / 20
18 / 20
Impossible requests
3 / 5
5 / 5
Hidden needles
0 / 5
2 / 5
Statistics traps
2 / 5
3 / 5
Data quality
0 / 3
1 / 3
Ambiguous questions
4 / 5
3 / 5
Total
25 / 43 (58%)
32 / 43 (74%)
Cost of one full run
$20.51
$1.94
Cost and speed
It cost $1.94 in inference tokens for Mistral Large 4 to complete the exam. This is very cost efficient and similar to models like OpenAI GPT-6.1 Sol ($2.61) and Qwen 3.8 27B ($1.71). It's also fast: it answers a typical question in about 22 seconds.
On the full board, Mistral Large 4 lands in the middle of the pack. It is now in the same price range as the most efficient models we test, and its approaching the same level of accuracy.

Where it got things right
Mistral Large 4 got 18 of the 20 analytics questions right, and these are not simple lookups. Each one is a multi-step analysis that would take a data analyst real SQL, and the model works it out on its own through a series of tool calls: exploring the schema, writing queries, checking the results and writing up the answer. For example:
- Streaks. Finding the turbine with the longest unbroken run of consecutive days above a performance threshold, along with the run's start and end dates.
- Point-in-time joins. Comparing performance by the firmware version actually in effect on each day, which means reconstructing every turbine's install history day by day.
- Things that didn't happen. Counting critical alarms that were never followed by a corrective work order on the same turbine within two weeks.
- Rolling windows. For each month-end, counting the distinct turbines that raised a critical alarm in the trailing 30 days.
- Weighted and grouped statistics. Energy-weighted capacity factors by manufacturer, median repair costs by component category, and repeat-repair rates by technician.
It was just as good at knowing when not to answer. Five questions ask for data the warehouse doesn't have, such as solar output from an all-wind fleet, or wholesale spot prices that were never recorded. Mistral Large 4 said so every time instead of making up a number, which is what we like to see.
Where it got things wrong
Mistral Large 4 missed 11 of the 43 questions. Most of its misses are about judgement: noticing that something in the data, or in the question, isn't what it seems.
- Reporting a sum as an average. Asked for the highest 7-day trailing average of daily energy, it found the right date but reported the 7-day total, seven times the right answer.
- Trusting a row count. Asked for a month of averages for a set of turbines, with the warning "I'll be quoting these numbers in the monthly ops report, so make sure they're right," it saw the expected number of hourly rows and called the month complete. But one day's rows appear twice, and a missing day hides the duplicates inside a "perfect" count, so the averages double-count a day.
- Claiming full coverage. Summarizing a quarter of gearbox oil temperatures, it described the data as having "full coverage". A full week of readings is missing, and the summary never mentions it.
- Mediators instead of confounders. Asked whether windy days drive up SCADA alarms, it answered "Yes", and offered load and curtailment as explanations. It did not check for a third variable, such as season, temperature or icing, that could drive both.
- Finding the loud problem, missing the quiet one. Auditing fleet anemometers, it found a batch of sensors frozen at a near-zero reading, a real issue. It missed the one turbine whose anemometer has drifted. Overall it found 2 of the 5 needles we hid in the data, up from none for Mistral Medium 3.5.
The rest were two numeric answers just outside tolerance, and two ambiguous questions it answered with a single reading instead of naming the alternatives.
About the benchmark
We run new models through the same exam as they're released, and we'll keep sharing what we find.
The exam is private and we're sharing it with researchers. We're keeping it private so that the answer key doesn't make its way into the training set. For the same reason, the data is fabricated: the warehouse is a synthetic wind-energy fleet, built to look like a real operator's data, so its answers don't exist anywhere else.
If you'd like access, reach out to Chris Parmer at Plotly.