[Infra-15] From Inference Engine to Cloud Service: The Full Stack Behind an LLM API
A request-path walkthrough of the complete LLM cloud stack, covering inference engines, model serving, orchestration, gateways, observability, security, billing, and the boundaries between them.