How to Scale Your Model, Interactive

An interactive version of the Google DeepMind book on how large language models actually run on hardware, and why I built it.

This started out of curiosity about running models locally. I wanted to know what it would actually take to train one myself, and to serve it fast enough to be worth using, rather than just pulling something down and hoping it ran. So I went looking for material that explained what the hardware is really doing underneath all of it.

That search turned up How to Scale Your Model, a book by a team at Google DeepMind, and it is one of the clearest technical things I have read. It is also dense, and I noticed I kept doing the same thing on every other page: stop, open a calculator, change one number, and check whether I had actually understood the point or just nodded along at it.

So I built the version I wanted while I was reading it.

→ scaling.nullroute.sh

What it is

It is the same book, laid out as a course you click through instead of scroll through. Thirteen chapters, in the same order the original uses. The arithmetic that the book states, I turned into eight small labs you can operate: move the sliders, change the chip, and watch the numbers move.

That is the whole idea. The book tells you that generating tokens is memory-bound, not compute-bound. Reading that sentence and believing it is one thing. Dragging the batch size up and watching the bound flip over is another, and the second one stuck for me in a way the first one did not.

Nothing in it is my insight. The explanations, the numbers and the worked examples are theirs. What I added is the ability to poke at them.

Credit where it is due

The original is How to Scale Your Model, written by Jacob Austin, Sholto Douglas, Roy Frostig, Anselm Levskaya, Charlie Chen, Sharad Vikram, Federico Lebron, Peter Choy, Vinay Ramasesh, Albert Webson and Reiner Pope, at Google DeepMind.

Austin et al., "How to Scale Your Model", Google DeepMind, online, 2025.

Please read the original. If you only have time for one of the two, read theirs and not mine. Mine exists because theirs is good enough to be worth the effort of rebuilding.

The chapters

Thirteen of them, in the order the original uses.

0. Introduction — What scaling a model actually means, and the three physical limits everything else reduces to.

1. Intro to Rooflines — Three things bound you: how fast the chip does maths, how fast it moves bytes, and how many bytes it can hold. Everything after this is a consequence.

2. All About TPUs — How a TPU works, and why that decides which models you can train and serve. The chip, its memory ladder, and the pods it forms.

3. Sharded Matmuls — Splitting a model across many chips, by way of the sharded matrix multiply and the four collectives that make it possible.

4. Transformer Math — How many FLOPs a Transformer uses, how many parameters it has, how big its KV caches get. Counting rules, not hand-waving.

5. Training — Data parallelism, FSDP, tensor parallelism and pipelining. What each one shards, what each one sends over the wire, and where each one stops being worth it.

6. Training LLaMA — A real worked example: training LLaMA-3 70B on 15 trillion tokens. Chips, days and dollars.

7. Inference — Prefill and decode sit on opposite sides of the roofline. Why generating tokens is memory-bound, and why the KV cache runs the show.

8. Serving LLaMA — Serving that same model. Pick a topology and a batch size, then watch latency, throughput and cost per token trade against each other.

9. Profiling — Rooflines predict, profilers explain. How to close the gap between your estimate and what the hardware actually did.

10. All About JAX — jit, meshes and shard_map: how the sharding notation from chapter 3 turns into code that runs.

11. Conclusions — The whole thing compressed into a cheat sheet you can click through.

12. GPUs — A bonus chapter. A modern GPU is a hundred small tensor cores on a switched tree. The constants change, the topology changes a lot, and most of the reasoning still transfers.

Eight of those chapters carry a lab: the roofline, moving a tensor through the memory ladder, the cost of an AllGather, a model calculator, a memory budget, sizing a training pod, a decode step, and a serving setup.

Why I bothered

I have written elsewhere on this site that I want to understand how these networks learn, and where the comparison to a brain stops being useful. That is still true, but I kept running into a smaller and more annoying problem first: I could follow the words about scaling without being able to do the arithmetic behind them. You can read a whole chapter about why something is memory-bound, agree with it, and still not be able to work out whether your own case is.

I do not think that is a reading problem. I think a static page is just the wrong shape for this kind of material. The numbers in it are not decoration. They are the argument, and an argument you cannot test is one you end up taking on trust.

So the honest answer to why I built it is that I wanted to check whether I had understood it. Building something is how I find out where the gaps are, because the gaps refuse to stay hidden once you have to make the thing actually work. I was wrong about several things while putting it together, which is rather the point.

It is free, there is nothing to sign up for, and I get nothing out of you clicking it. If it makes the original land better for one other person, that is more than enough.

Go and have a look.

You've successfully subscribed to Nullroute.
Great! Next, complete checkout to get full access to all premium content.
Error! Could not sign up. invalid link.
Welcome back! You've successfully signed in.
Error! Could not sign in. Please try again.
Success! Your account is fully activated, you now have access to all content.
Error! Stripe checkout failed.
Success! Your billing info is updated.
Error! Billing info update failed.