Why Now Daily.

Published

DeepSeek R1 Zero vs R1: What Changed

DeepSeek-R1-Zero tested whether reinforcement learning could elicit reasoning behavior without supervised fine-tuning first. DeepSeek-R1 added cold-start examples and a multi-stage training pipeline to improve readability, language consistency and general usefulness.

Timeline

  1. R1-Zero training: DeepSeek applied large-scale reinforcement learning directly to its base model without a supervised fine-tuning stage first.
  2. R1 training: The team added cold-start data, multiple reinforcement-learning stages and supervised stages.
  3. January 2025: DeepSeek released R1-Zero, R1 and six distilled models based on Qwen and Llama families.

DeepSeek-R1-Zero and DeepSeek-R1 are related reasoning models, but they were built to answer different research questions. R1-Zero was the direct reinforcement-learning experiment: DeepSeek says it applied large-scale reinforcement learning to a base model without first using supervised fine-tuning. R1 was the more engineered follow-up, adding curated starting examples and a multi-stage pipeline intended to preserve reasoning gains while producing more readable and stable answers. [1][2]

The “Zero” name refers to the absence of supervised fine-tuning before reinforcement learning, not to a model with no prior training. DeepSeek-R1-Zero was based on DeepSeek-V3-Base. During reinforcement learning, the system was rewarded for solving tasks, and the project reports that behaviors such as self-verification, reflection and longer reasoning traces emerged. Those observations are results reported by the model’s creators, not proof that every answer is reliable. [1][2]

R1-Zero also exposed practical weaknesses. DeepSeek’s documentation lists endless repetition, poor readability and language mixing among its problems. A model can score well on structured reasoning evaluations yet still be frustrating or confusing in ordinary use. The experiment therefore served as evidence about what reinforcement learning could elicit, while also showing why output quality and alignment stages remained necessary for a generally useful assistant. [1][2]

DeepSeek-R1 changed the recipe. The project describes cold-start data before reinforcement learning, followed by two reinforcement-learning stages and two supervised fine-tuning stages. The stages were intended to discover stronger reasoning patterns, align outputs with preferences and support both reasoning and non-reasoning tasks. This makes R1 more than a cleaned copy of R1-Zero: the training path and intended behavior differ. [1][2]

The release also includes distilled models, which are easy to confuse with the two large models. DeepSeek says it used samples generated by R1 to fine-tune smaller dense models based on Qwen and Llama families, releasing sizes from 1.5 billion to 70 billion parameters. A distilled checkpoint is therefore not R1-Zero or the full R1 model; it is a smaller model trained to reproduce parts of R1’s behavior under different hardware requirements. [1]

Benchmark tables should be read with their evaluation settings. DeepSeek documents sampling temperature, top-p, response counts and maximum generation length for its reported comparisons. Results measured under those conditions do not guarantee identical performance in a local installation, a hosted service or a different prompt format. Independent testing should also consider factual accuracy, safety, language quality, latency, memory use and the specific tasks that matter to the user. [1][2]

For most people choosing between the releases, R1 is the practical reference model and R1-Zero is primarily the research comparison that demonstrates the direct-reinforcement-learning approach. Smaller systems should examine the distilled checkpoints and follow the repository’s model-specific settings. Whichever checkpoint is used, outputs can be wrong or fabricated, so consequential code, calculations, medical statements, legal claims and financial information still require verification against authoritative sources. [1][2]

Sources

  1. DeepSeek — DeepSeek-R1 official repository and model documentation
  2. DeepSeek-AI — DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Related stories