Having run models locally, RAM use seems to be almost directly proportional to number of parameters. 8 Billion parameters requires approx 8GB of VRAM at 1/4 precision.
Therefore, if this pattern holds you somehow need 10 Terabytes of VRAM at 4K and 40 Terabytes at full precision.
I think I saw some estimates that Claude’s Opus models may be and Opus model equivalents may be at around 100B parameters (100-400GB VRAM).
TLDR its clear why RAM is so expensive.
for inference you’re only counting active parameters towards VRAM, and some labs / models don’t train at 32b precision, or even use the same precision for different parts of the network
ByteDance building China’s largest model while running TikTok is fascinating: few companies have that combination of compute, money and an absurdly large stream of real world human behavior. The AI race isn’t just US labs versus China anymore: it’s ecosystems versus ecosystems.
Meta are losing their AI market




