Manas Bihani
About

the questions

  1. What is a moat in an AI world?
  2. Why do AI products converge?
  3. What becomes scarce when intelligence becomes cheap?
  4. Does distribution matter more than technology?
  5. Why might human-made things become more valuable?
  6. What happens to expertise when everyone has the same models?
  7. Which parts of an AI startup are actually defensible?
  8. Where does value move when intelligence becomes commoditized?

everything on the desk

  1. The periodic table of the AI stackVisualization
  2. What is a moat when the model isn't yours?Note
  3. The problem-selection premiumNote
  4. Same model, different wiringNote
  5. Selection is the new bottleneckNote
  6. Get friendly with the AI raceEssay
  7. The convergence taxNote
  8. The luxury of realityNote
  9. The non-technical technical advantageNote
  10. The verification economyNote
  11. The bets against the wallVisualization
  12. You can't buy your way outVisualization
  13. How a chatbot writes one wordVisualization
  14. The grid is the last wallVisualization
  15. Who got paidVisualization
  16. Why this paper mattersExplainer
  17. Transformer: Why did transformers replace RNNs?Vaswani et al., NeurIPS 2017
  18. KV cache: Why does a long conversation get slower and cost more than a short one?Shazeer, 2019
  19. Mixture of experts: Why do some AI models have experts?Fedus, Zoph and Shazeer, 2021
  20. FlashAttention: Why is attention slow when the GPU is barely doing any arithmetic?Dao et al., NeurIPS 2022
  21. Mamba: Why does a model reread the whole conversation instead of just remembering it?Gu & Dao, 2023
  22. PagedAttention: Why does a GPU with free memory still refuse new requests?Kwon et al., SOSP 2023
  23. DeepSeek: How did DeepSeek train a frontier model so cheaply?DeepSeek-AI, 2024
  24. Jamba: Why does Jamba matter?Lieber et al., AI21 Labs, 2024
  25. BitNet: Why does BitNet matter?Ma et al., Microsoft Research, 2025
  26. DeepSeek-R1: Can a small AI model learn to reason like a huge one?DeepSeek-AI, 2025
  27. Kimi K2: Why does Kimi K2 matter?Kimi Team, Moonshot AI, 2025
  28. Sliding-window attention: How do models handle huge context windows without the memory bill exploding?Gemma Team, Google DeepMind, 2025
  29. How electricity becomes intelligenceVisualization
  30. This desk, as a datasetDataset
  31. The first version of this roomNote
  32. The aura dividendNote
  33. Distribution is rented attentionNote
  34. The Convergence TestNote
  35. A shelf for thinking about cheap intelligenceCollection
  36. Anatomy of an AI startupNote
  37. Six shocks to expertiseNote
  38. Nineteen Public KeysEssay
  39. The value migration machineModel
  40. AAA-Rated GPUsEssay
  41. Moats, before and afterVisualization
  42. The rhinoceros problemNote
  43. What becomes scarce when intelligence becomes cheap?Essay
  44. AI Has Passed Every Exam. It Has Never Had an Idea.Essay
  45. What Becomes Scarce After Intelligence?Essay
  46. India’s Carbon Markets : A New Test for Global Climate PolicyEssay
  47. Google Wants AI to Become BoringEssay
  48. The Wall That Wasn’t YoursEssay
  49. The Rate-Limiting StepEssay
  50. The Speed of Being WrongEssay
  51. Uber Burned a Year of AI Budget in Four Months. A Rat Catcher in 1902 Knew WhyEssay
  52. Finding a Flat in India Is Broken. We Have the Technology to Fix It. Nobody With Power Wants To.Essay
  53. Why We Can Never Have Good Social MediaEssay
  54. Gen Z Is Going OfflineEssay

rooms

  1. Home
  2. Writing
  3. Projects
  4. Reading & Watching
  5. All the questions
  6. Everything, as a contact sheet
  7. About

Why this paper matters · 10 of 12

Can a small AI model learn to reason like a huge one?

DeepSeek-R1: Reasoning ability can be trained against a checkable reward, not copied from a hand-written example of the reasoning itself.

the paper →DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningDeepSeek-AI, 2025 ↗

This is why "reasoning" stopped being something only a frontier lab could afford to train — DeepSeek showed a model can learn to think step by step from reinforcement learning against a right-or-wrong signal alone, no expensive hand-written reasoning examples required, and that the resulting skill mostly survives being distilled into much smaller models.

The room it came out of

DeepSeek open-sourced R1 in January 2025 and made an unusual claim explicit: reasoning ability came almost entirely from reinforcement learning against a checkable reward, not from the large, expensively human-labeled chain-of-thought datasets everyone assumed a reasoning model needed. They also showed the resulting skill survives being distilled into models a fraction of the size.

How it works

Every reasoning model before R1 leaned on a large, expensively assembled dataset of worked examples — humans (or a stronger model) writing out the steps, so the model being trained had something to imitate. DeepSeek's January 2025 paper made a different claim: skip the worked examples. Give the model a problem with a checkable answer — a math problem, a piece of code that either passes its tests or doesn't — reward it when it's right, and let reinforcement learning find its own way to a chain of steps that gets there.

It worked, and it generalized further than the headline result: the reasoning behavior transferred cleanly into distillation, meaning a much smaller model trained on R1's outputs could inherit a meaningful fraction of the reasoning skill without ever running the expensive RL process itself. That is what turned this from "a good result at one lab" into a real shift — reasoning capability got cheaper to reproduce for everyone downstream, not just for DeepSeek.

The honest cost sits in what the reward actually optimizes. A checkable-answer reward teaches a model to arrive at the right answer; it says nothing about whether the steps in between are a truthful account of how it got there. Early R1 output sometimes mixed languages mid-thought or skipped steps a human reader couldn't reconstruct — legible reasoning was never the thing being trained for, correctness was, and the two turned out to be separable.

What it traded

gave up
a fully legible, human-auditable chain of thought at every step
got
reasoning trained without a large hand-labeled dataset, and small enough to distill into models a fraction of the size

What exists now that didn’t before

A capability that used to require a frontier-scale model and a large hand-labeled reasoning dataset can now be trained with RL against a right/wrong signal alone and distilled into something small — which collapsed both the compute and the data-labeling cost of "a model that reasons," not just DeepSeek's own.

What it left undone

Reward against a checkable answer teaches a model to reach the right answer, not to reason legibly — early R1 chains mixed languages and skipped steps a human could not verify, so the visible reasoning is not a trustworthy transcript of what actually happened inside.

Asked, and answered

Can a small ai model learn to reason like a huge one?

Mostly, yes. DeepSeek trained reasoning with reinforcement learning against a checkable right-or-wrong answer instead of hand-written examples, and showed the resulting skill survives being distilled into models a fraction of the size.

Read next