When an LLM writes one token for one user, an H100 uses about 0.3% of its compute. The rest waits on memory.
Modal's blog has 10 posts filed under "Research". This 25-minute explainer walks through all of them, plus their write-up on serving Kimi K2.6 to coding agents. Almost every post turns out to make the same move: find the speed of light → find the idle resource → spend it on parallel work you might throw away → keep it honest with a verifier you trust.
Each chapter opens with a question. Pause, make a guess, then watch the answer.
0:00 Intro: ten posts, one idea
0:45 The speed of light (the roofline model)
2:32 Search & evals: 100 small LLaMAs vs GPT-4o, QR codes that scan, a quick embedding fine-tune
6:03 Inside FlashAttention-4: online softmax, lazy rescaling, software exp2, warp specialization
9:00 Making FlashAttention-4 fast for decode: split KV, GQA packing, small KV pages
11:34 Speculative decoding: why speedup tracks acceptance length
14:17 Multi-token Residual Prediction for diffusion LMs
16:31 Serving a trillion-parameter model to coding agents (Kimi K2.6)
19:15 Quail: an inference engine for AI-SQL
21:36 Research that scales itself: evolution strategies and autoresearch
23:40 The one idea
A few numbers from the posts:
• Llama 3.1 8B on HumanEval: 66.4% → 90.5% with 100 samples, 95.1% with 1,000 (GPT-4o: 90.2%)
• FlashAttention-4 with split KV in decode: 0.83 → 4.37 TB/s memory throughput
• Custom speculator for Kimi: acceptance length 5.00 → 5.84, about 20% faster, identical outputs
• Residual prediction at 2 steps: 16.9 → 84.9 on GSM8K
• Quail: over a billion tokens per minute on a single H100 for one multi-join query
Notes:
• Independent explainer, not affiliated with or endorsed by Modal. All credit goes to the authors at Modal and their collaborators (NYU Shanghai's HeavyBall Research, CMU's Full Stack Data Lab, AE Studio).
• Every number comes from the posts. A few curves are my own models or schematics (the pass@k power-law fit, the GPU timeline in chapter 9, the toy join example) and are labeled on screen.
• Corrections are very welcome in the comments.
The posts:
Beat GPT-4o at Python by searching with 100 dumb LLaMAs: https://modal.com/blog/llama-human-eval
Evals + inference-time compute for QR codes: https://modal.com/blog/qart-codes-evals
Beating proprietary models with a quick fine-tune: https://modal.com/blog/fine-tuning-em...
We reverse-engineered Flash Attention 4: https://modal.com/blog/reverse-engine...
Making FlashAttention-4 faster for inference: https://modal.com/blog/flash-attentio...
Speculation Is All You Need: https://modal.com/blog/spec-is-all-u-...
Multi-token Residual Prediction: https://modal.com/blog/multi-token-re...
How to serve trillions of tokens for trillion-parameter coding agents: https://modal.com/blog/trillion-token...
Hitting a billion tokens per minute on one GPU (Quail): https://modal.com/blog/quail-billion-tpm
Building an RL theorem-proving workflow on Modal: https://modal.com/blog/building-an-rl...
Autoscaling Autoresearch: https://modal.com/blog/autoscaling-au...
Made with Manim. Narrated with Chatterbox (an open-source AI voice). Script and animation built with Claude Code (Claude Opus 5.5).
When an LLM writes one token for one user, an H100 uses about 0.3% of its compute. The rest waits on memory.
Modal's blog has 10 posts filed under "Research". This 25-minute explainer walks through all of them, plus their write-up on serving Kimi K2.6 to coding agents. Almost every post turns out to make the same move: find the speed of light → find the idle resource → spend it on parallel work you might throw away → keep it honest with a verifier you trust.
Each chapter opens with a question. Pause, make a guess, then watch the answer.
0:00 Intro: ten posts, one idea
0:45 The speed of light (the roofline model)
2:32 Search & evals: 100 small LLaMAs vs GPT-4o, QR codes that scan, a quick embedding fine-tune
6:03 Inside FlashAttention-4: online softmax, lazy rescaling, software exp2, warp specialization
9:00 Making FlashAttention-4 fast for decode: split KV, GQA packing, small KV pages
11:34 Speculative decoding: why speedup tracks acceptance length
14:17 Multi-token Residual Prediction for diffusion LMs
16:31 Serving a trillion-parameter model to coding agents (Kimi K2.6)
19:15 Quail: an inference engine for AI-SQL
21:36 Research that scales itself: evolution strategies and autoresearch
23:40 The one idea
A few numbers from the posts:
• Llama 3.1 8B on HumanEval: 66.4% → 90.5% with 100 samples, 95.1% with 1,000 (GPT-4o: 90.2%)
• FlashAttention-4 with split KV in decode: 0.83 → 4.37 TB/s memory throughput
• Custom speculator for Kimi: acceptance length 5.00 → 5.84, about 20% faster, identical outputs
• Residual prediction at 2 steps: 16.9 → 84.9 on GSM8K
• Quail: over a billion tokens per minute on a single H100 for one multi-join query
Notes:
• Independent explainer, not affiliated with or endorsed by Modal. All credit goes to the authors at Modal and their collaborators (NYU Shanghai's HeavyBall Research, CMU's Full Stack Data Lab, AE Studio).
• Every number comes from the posts. A few curves are my own models or schematics (the pass@k power-law fit, the GPU timeline in chapter 9, the toy join example) and are labeled on screen.
• Corrections are very welcome in the comments.
The posts:
Beat GPT-4o at Python by searching with 100 dumb LLaMAs: https://modal.com/blog/llama-human-eval
Evals + inference-time compute for QR codes: https://modal.com/blog/qart-codes-evals
Beating proprietary models with a quick fine-tune: https://modal.com/blog/fine-tuning-em...
We reverse-engineered Flash Attention 4: https://modal.com/blog/reverse-engine...
Making FlashAttention-4 faster for inference: https://modal.com/blog/flash-attentio...
Speculation Is All You Need: https://modal.com/blog/spec-is-all-u-...
Multi-token Residual Prediction: https://modal.com/blog/multi-token-re...
How to serve trillions of tokens for trillion-parameter coding agents: https://modal.com/blog/trillion-token...
Hitting a billion tokens per minute on one GPU (Quail): https://modal.com/blog/quail-billion-tpm
Building an RL theorem-proving workflow on Modal: https://modal.com/blog/building-an-rl...
Autoscaling Autoresearch: https://modal.com/blog/autoscaling-au...
Made with Manim. Narrated with Chatterbox (an open-source AI voice). Script and animation built with Claude Code (Claude Opus 5.5).