DeepSeek R1 Explained: Model Design and Open Weights
DeepSeek-R1 is a 671-billion-parameter mixture-of-experts reasoning model built on DeepSeek-V3. Its release included R1-Zero, the full R1 weights and six smaller distilled checkpoints, with model weights made available under the MIT license.
Timeline
- 2025-01-20: DeepSeek announced R1, R1-Zero and distilled checkpoints.
- 2025-09: A peer-reviewed version of the R1 research appeared in Nature.
DeepSeek-R1 is a large language model designed to spend additional computation producing and checking multi-step answers, especially for mathematics, coding and other verifiable tasks. DeepSeek released it in January 2025 alongside an experimental model called R1-Zero and six smaller distilled models. The full R1 model is based on DeepSeek-V3 and uses a mixture-of-experts architecture with 671 billion total parameters but about 37 billion activated for each token, according to the project documentation. [1][2][3]
R1-Zero was the research starting point. The team applied large-scale reinforcement learning to a base model without first using the usual supervised fine-tuning stage. Rewards favored correct answers and adherence to requested formats on tasks where results could be checked. During training, the model developed longer reasoning traces, self-checking and attempts at alternative solutions. Those behaviors were observed outputs of optimization, not evidence that the system reasoned or understood problems in the same way a person does. [2][3]
R1-Zero also exposed practical weaknesses: repetition, poor readability and mixing languages. To build R1, DeepSeek added a small ‘cold-start’ collection of readable reasoning examples, then used reinforcement learning, generated-data filtering and supervised fine-tuning in multiple stages. This hybrid pipeline sought to preserve performance on verifiable reasoning while improving instruction following, writing quality and usefulness on tasks without a simple automatic answer checker. [2][3]
The full R1 release and the distilled models are different products. DeepSeek used outputs from R1 to fine-tune dense checkpoints based on Qwen and Llama model families at sizes from 1.5 billion to 70 billion parameters. Those smaller checkpoints are easier to run locally, but they do not have the full model’s architecture or capacity. Results reported for one checkpoint should not be silently attributed to every model carrying the R1 name. [1][2]
DeepSeek reported competitive results on several math, programming and general reasoning benchmarks. Benchmark scores depend on prompts, sampling settings, evaluation code and the risk that test material overlaps training data. The repository itself recommends repeated evaluations and particular generation settings. A high score also does not ensure factual accuracy, safe advice or dependable performance on a user’s real workflow, so independent testing with representative tasks remains necessary. [2][3]
The release was unusually accessible for a frontier-scale reasoning system because DeepSeek published weights for R1-Zero, R1 and the six distilled checkpoints. The project repository and full-model weights use the MIT license and allow commercial use and modification. Distilled models also inherit relevant terms from their Qwen or Llama base models, so developers need to check the specific checkpoint’s notices rather than assuming that one license statement resolves every dependency and use case. [1][2][3]
Open weights do not mean the entire training process is independently reproducible. The paper describes methods, provides selected data samples and releases inference code references, but the authors’ internal distributed-training framework and complete training corpus were not published as a turnkey reproduction package. R1 is best understood as both a usable family of weights and research evidence that reinforcement learning can elicit strong reasoning-like behavior. Outputs still require verification, particularly for medical, legal, financial, security or other high-stakes decisions. [2][3]
Sources
- DeepSeek — DeepSeek-R1 release
- DeepSeek-AI — DeepSeek-R1 repository and model documentation
- Nature — DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning