DeepSeek V4.1 Flash for Data Analysis: A Review
Oct 9, 2026: DeepSeek V4.1 Flash costs a fraction of the frontier models. On our data analytics exam, small models from OpenAI and Anthropic match its accuracy for less money and in less time.
DeepSeek V4.1 Flash has been getting a lot of attention this week. A post on Hacker News describes a month of using it for coding, not noticing a difference from Claude Opus 5.5, and finding it "orders of magnitude cheaper than the 'frontier' models." At Plotly, we test every major model on our data analytics benchmark, so we looked at how it does there.
Against the frontier models, our results agree. DeepSeek V4.1 Flash passed 36 of our 43 questions (84%, a B) for $0.46 a run. Claude Opus 5.5 passed 41 for $13.31, so DeepSeek costs about 1/29th as much for five fewer correct answers.
But the recent releases from OpenAI and Anthropic change this story. GPT-6 Luna passed 38 of 43 for $0.13 a run, and Claude Haiku 5.5 passed 37 for $0.38. Both scored at least as well, cost less and answered faster.

The results
Our exam runs inside Plotly Studio Embedded, with the same agent, tools and prompts for every model. The agent answers 43 questions against a 20M-row data warehouse (a fictional wind energy company: 44 tables, 554 columns), working through each one with SQL queries and charts. Every question has either an answer checked against a verified key, or a catch to notice: an ambiguous question, an impossible request, a planted data defect. Here's how DeepSeek V4.1 Flash compares with its neighbors on the board:
Model
Score
Cost of one run
Median time per question
DeepSeek V4.1 Flash (open weights)
36 / 43
$0.46
21 seconds
GPT-6 Luna
38 / 43
$0.13
18 seconds
Claude Haiku 5.5
37 / 43
$0.38
14 seconds
GLM 5.3 Flash (open weights)
38 / 43
$0.69
87 seconds
Qwen 3.8 27B (open weights)
38 / 43
$1.71
53 seconds
GPT-6.1 Sol
43 / 43
$2.61
18 seconds
Claude Opus 5.5
41 / 43
$13.31
27 seconds
We've run each model once, so a difference of one or two questions is within the noise. The cost differences are not: GPT-6 Luna answered the same exam for less than a third of the price.
Cheap tokens, but more of them
Per token, DeepSeek V4.1 Flash is priced in the same range as GPT-6 Luna and Claude Haiku 5.5, which both charge $0.10 per million input tokens and $0.50 per million output tokens. DeepSeek's own API charges $0.15 and $0.60. The difference is how many tokens it uses.
On an average question, DeepSeek V4.1 Flash made 7.3 model calls, against 4.5 for Luna and 4.3 for Haiku 5.5. It went through 284,000 tokens per question, more than twice Luna's 130,000, and wrote almost three times as much output (6,200 tokens against 2,300). Output tokens are the expensive ones.

It's also slower on the hard questions. Its median of 21 seconds per question is close to Luna's 18, but its average is 53 seconds, twice Luna's 27, because the questions it digs into take much longer.
We've seen this before: cheaper or faster tokens don't always mean cheaper or faster analysis. What a question costs is the price per token times the number of tokens the model decides to use, and that second number varies a lot from model to model.
What you pay depends on who serves it
DeepSeek V4.1 Flash has open weights, so many providers serve it. OpenRouter lists about 30, with input prices from $0.02 to $0.45 per million tokens, some serving 4-bit or 8-bit quantized versions. We ran it through OpenRouter's default routing on September 23, and the run was billed $0.46. At DeepSeek's own API prices, the same tokens would have cost about $0.33. At OpenRouter's current headline price of $0.30 in and $1.20 out, about $0.66. Even the lowest of those is more than twice what GPT-6 Luna cost.
And many developers don't pay per token at all. The author of the Hacker News post uses DeepSeek through a $10-a-month coding subscription, which makes it close to unlimited. That's a great deal for an individual developer. Most of our enterprise customers pay per token through their cloud provider, and for them the cost per question is what adds up.
Where it does well
Within its price range, DeepSeek V4.1 Flash is a strong data analyst:
- Analytics (20 of 20). It got every multi-step SQL question right: streaks of consecutive days, point-in-time joins against firmware history, merging overlapping alarm intervals, rolling distinct counts and weighted averages.
- Ambiguous questions (5 of 5). Asked "Which is our best-performing wind farm?", it showed that the answer changes with the measure instead of naming a single winner.
- Impossible requests (5 of 5). Asked to compare contract revenue with wholesale spot prices that aren't in the data, it said so instead of making up a number.
- Data quality (2 of 3). Asked for a month of averages "for the monthly ops report", it found the duplicated rows and removed them before averaging.
It's also the cheapest open-weight model on our board.
Where it falls short
It missed 7 questions, mostly ones where nothing in the question says what to look for:
- Hidden patterns (1 of 5). We planted five patterns in the data for the model to find on its own. DeepSeek V4.1 Flash found one; GPT-6 Luna and Claude Haiku 5.5 found two, and GPT-6.1 Sol found all five. In one case it laid out the evidence that pointed to the right answer, then recommended the opposite.
- Statistics (3 of 5). Asked whether windy days drive up alarms, it compared alarm rates by wind speed and stopped there, without checking for a third variable such as season. Asked which repairs are unusually expensive, it listed the 20 most expensive instead of using a statistical rule.
- Missing data. Summarizing a quarter of gearbox oil temperatures, it described the data as having "full coverage" and missed a week of missing readings.
So is DeepSeek V4.1 Flash good for data analysis?
Yes, with caveats. It's a solid B on our exam for well under a dollar a run, and if you need open weights, it's the cheapest option on our board.
If you don't need open weights, GPT-6 Luna and Claude Haiku 5.5 are at least as accurate, cheaper per question and faster. Price per token is the number people quote, but cost per question is what you pay, and it depends on how many tokens the model uses.
About the benchmark
We run new models through the same exam as they're released, and we'll keep sharing what we find.
The exam is private and we're sharing it with researchers. We're keeping it private so that the answer key doesn't make its way into the training set. For the same reason, the data is fabricated: the warehouse is a synthetic wind-energy fleet, built to look like a real operator's data, so its answers don't exist anywhere else.
If you'd like access, reach out to Chris Parmer at Plotly.