What DeepSeek V3 and R1 actually showed

The famous $5.6 million covered one final training run, not the whole bill. The lasting lessons were efficient engineering and reasoning learned from rewards.

Written Updated 3 min read

The papers

DeepSeek released V3 at the end of December 2024 and R1 a month later, and the story that stuck was the price: a frontier-class model for about $5.6 million. On January 27, 2025, Nvidia lost close to $600 billion in market value, the biggest one-day drop for any US company (CNBC). The number was real but much narrower than the headlines, and what the two papers showed is more useful than the price tag.

V3: an efficient base model

V3 is a mixture-of-experts model: 671 billion parameters in total, of which 37 billion are active for each token (technical report). It was trained in 8-bit floating point (FP8), which DeepSeek describes as the first validation of FP8 training on a model that large, on a cluster of 2,048 Nvidia H800 GPUs. Its multi-head latent attention, which cuts the memory needed while generating, came from the earlier DeepSeek-V2.

The cost comes from the report's first table: 2.788 million H800 GPU-hours at an assumed rental price of $2 an hour, or $5.576 million. The report says the figure covers only the official training run, "excluding the costs associated with prior research and ablation experiments on architectures, algorithms, or data." It leaves out the hardware and everything it took to get to that run.

The comparison that holds up is compute. Meta's model card lists 30.84 million H100 GPU-hours for Llama 3.1 405B, about 11 times V3's total, on faster chips, and DeepSeek reported V3 as comparable to GPT-4o and Claude 3.5 Sonnet on its benchmarks. That efficiency came from engineering choices anyone could read about, which is a large part of why it mattered.

R1: reasoning learned from rewards

R1's contribution was a method. DeepSeek first trained R1-Zero with reinforcement learning directly on the V3 base model, with simple rule-based rewards: is the final answer correct, and is the reasoning in the expected format. There were no human-written worked solutions. The model learned to produce long chains of reasoning, check itself and backtrack, because doing so earned reward.

R1-Zero's output was hard to read, so R1 used a longer pipeline: a few thousand curated "cold-start" examples, reinforcement learning, fine-tuning on about 800,000 samples (mostly generated by the model itself), then more reinforcement learning. DeepSeek reported it as "comparable to OpenAI-o1-1217 on reasoning tasks", with 79.8% on AIME 2024 and 97.3% on MATH-500. The weights came out under an MIT licence, with six smaller distilled models built on Qwen and Llama.

The peer-reviewed version was published in Nature in September 2025. Its supplement put R1's own training at 147,000 H800 GPU-hours, or $294,000 at the same $2 rate, on top of the V3 base (appendix B.4.4 in the revised arXiv version).

Why I think it matters

The lesson I'd take is that capability got cheaper through engineering, and open weights gave buyers real options. You can run DeepSeek's models on your own hardware or get them from more than one cloud, which makes lock-in easier to avoid (I wrote about pricing that kind of hedge in when multi-provider AI is worth the premium). I think the durable advantage for most organizations now sits in application and execution rather than in access to a frontier model.

I'm less sure about the broader claim that compute stopped mattering. DeepSeek itself trained V3 on 2,048 GPUs and R1 on another 512, and its later models are much larger. My read is that efficiency and scale both kept moving.

Since then

DeepSeek folded R1 into its main line. V3.1 (August 2025) offered thinking and non-thinking modes in one model, V3.2-Exp (September 2025) came with API prices cut by more than half, and V4 Preview (April 2026) introduced a 1.6-trillion-parameter Pro model with 49 billion active parameters and a smaller Flash model, both with open weights. The latest release is V4.1-Flash, from September 10, 2026. DeepSeek's release notes list no model called R2.

Prices kept falling. At launch R1 cost $2.19 per million output tokens, against $60 for o1 (DeepSeek, OpenAI). As of September 2026, DeepSeek's Flash model charges $0.60 to $1.20 per million output tokens depending on the time of day (pricing), and OpenAI shuts o1 down on October 23, 2026 (deprecations).

Regulators paid attention too. Italy's data protection authority ordered DeepSeek's app blocked in January 2025 (Garante), and a NIST evaluation in September 2025 found R1-0528 answered 94% of overtly malicious requests under a common jailbreak (NIST). Hosting open weights yourself avoids the data questions around the hosted app. It doesn't replace your own safety testing.

For the cheap end of reasoning, see s1, and for running models like these on less hardware, 4-bit quantization.

  • The s1 recipe for making models think longer

    With 1,000 examples and a forced "Wait", s1 got o1-style test-time scaling from an open model. Controlling how long models think is now a standard setting.

  • Running large models in 4 bits

    Storing weights in 4 bits instead of 16 cuts memory about fourfold with little loss. It decides what runs where, and it's now part of how models ship.

  • Training models on synthetic data

    Small models trained partly on text written by bigger models can beat much larger ones. Synthetic data can lower privacy risk, but it isn't anonymous by default.