10 notes
Research notes
Notes on papers I wanted to understand properly. Each one says what the paper showed, why I think it might matter outside the lab, and what is still uncertain. Most were written in November 2025 and have short updates where things have moved.
Cheaper, smaller models
Why the cost of a capable model keeps falling, and what that does to the economics.
The s1 recipe for making models think longer
With 1,000 examples and a forced "Wait", s1 got o1-style test-time scaling from an open model. Controlling how long models think is now a standard setting.
Published result · Paper January 2025
What DeepSeek V3 and R1 actually showed
The famous $5.6 million covered one final training run, not the whole bill. The lasting lessons were efficient engineering and reasoning learned from rewards.
In wide use · Paper January 2025 + 1 more
Running large models in 4 bits
Storing weights in 4 bits instead of 16 cuts memory about fourfold with little loss. It decides what runs where, and it's now part of how models ship.
In wide use · Paper October 2022 + 2 more
Learning from generated data and experience
Training models and agents on text other models wrote, or on what happened when they tried things.
Training agents in an imagined environment
DreamGym has a language model imagine how a website would respond, so an agent can practise with reinforcement learning. Real gains on web benchmarks, nothing physical yet.
Early research · Paper November 2025
Training agents on their own early experience
Let an agent try alternatives to the expert's actions and learn from what happens, with no reward signal. It beat plain imitation in eight environments.
Published result · Paper October 2025
Training models on synthetic data
Small models trained partly on text written by bigger models can beat much larger ones. Synthetic data can lower privacy risk, but it isn't anonymous by default.
In wide use · Paper December 2022 + 2 more
New architectures to watch
Ideas that change how models predict or remember. Promising at small scale, unproven at large scale.
Nested Learning and models that keep learning
Google's Nested Learning gives models memories that update at different speeds, so they might keep learning. Promising at 1.3B parameters, unproven beyond.
Early research · Paper December 2025
CALM and predicting several tokens at once
Tencent's CALM predicts one vector that stands for four tokens, matching a baseline with about a third less compute at small scale. Nobody has shown yet that it scales.
Early research · Paper October 2025
LLM-JEPA and LeCun's bet on predicting meaning
LeCun argues models should predict meaning rather than words. LLM-JEPA was a first small test on language models, and he has since left Meta to pursue the idea.
Early research · Paper September 2025
Robots and the physical world
Connecting language models to perception and action, and how far that has really come.
Robot foundation models and where they stand
Vision-language-action models let robots borrow knowledge from the web. The progress since RT-2 is real, but home robots are still mostly pre-orders and pilots.
Published result · Paper July 2023 + 2 more