Up next
Mojo and CPython: from two worlds to one
Up next
Entanglement: one wave function for two particles
Up next
The universe splits; it doesn't copy
Up next
Measurement is just becoming entangled
Up next
Why the solar-system atom can't work
Up next
Clang and stdio.h: compatibility beats beauty
Up next
Vectorize, parallelize, tile
Up next
Adaptive tile refresh: redraw only what changed
Up next
Apple II: pre-shifted sprites and compiled shapes
Up next
Every static analyzer over id's code
Up next
BSP trees and the epsilon problem
Up next
Compiled scalers: turn the picture into code
Up next
Wolfenstein 3D: one ray per screen column
Up next
Commander Keen: scroll off the edge and let it wrap
Up next
Kernel fusion and the operator explosion
Up next
“Zero-cost” exceptions aren't zero-cost
Up next
Value semantics: copy only when someone writes
Up next
Autotuning: let the machine pick the magic numbers
Up next
Why Python is slow, one layer at a time
Up next
Bloom filters: definitely no, in very little memory
Up next
Consistent hashing: add a server without moving everything
Up next
PagedAttention: fit more conversations on one GPU
Up next
Quantization: one outlier ruins the ruler
Up next
Prefix caching: stop processing the same beginning
Up next
LoRA: train a small edit to a large model
Up next
False sharing: independent work, hidden contention
Up next
Continuous batching: refill the GPU every token
Up next
Mixture of experts: a huge model with a small active path
Up next
FlashAttention: faster by moving less
Up next
Guessed wrong? Keep the work
Up next
Rewrite it: faster and half as complicated
Up next
Never the same path, always the same answer
Up next
Speculative decoding in 90 seconds
Up next
Bricks that halve every two years
Up next
Speculative decoding: guess ahead, check together
Up next
Moore's law is a cascade of S-curves
Up next
Two-level branch prediction: history picks the counter
Up next
Branch prediction: the little supercomputer inside your CPU
Up next
From atoms to the data center: the abstraction stack