The Economics of Runtime Tokens
As models get capable enough to run whole tasks in production, the tokens that do it are paid on every execution, bounded by a latency or throughput budget, and never amortized. Those constraints decide where and how inference has to run.