Monal Gupta
Blogs

2025 · 3 min read

The AI Cloud

Infrastructure is being rewritten around models: what changes when inference, not CPU time, is the unit of compute.

For two decades the unit of cloud compute was CPU time. The abstractions we built — containers, autoscaling groups, serverless functions — all optimized for packing deterministic work onto machines.

Inference breaks those assumptions. Latency is measured in seconds, cost is measured per token, and output is non-deterministic. Retry logic, caching, and observability all have to be rethought from the ground up.

The AI cloud is what emerges when the platform treats the model as the primitive: streaming as the default response shape, evaluation as the default test suite, and prompts as deployable artifacts with versions and rollbacks.