A CPU cache lets hardware silently decide whether to fetch from fast on-chip memory or slow DDR — introducing non-deterministic latency. A TPU scratchpad gives software explicit instructions for on-chip versus HBM memory. The cache gives 100× faster access on a hit, but the non-determinism is the price.