Skip to main content
Back

Archive · Article

Inference Engineering From First Token to Production Metric

The first of five articles in a written series on serving systems: I follow one request through eight stages to show why GPU latency and request latency are different numbers, and why a service that budgets by characters admits work it cannot hold.

X · long-form Article · 2026-08-25

Open it

  • Open the original (opens in a new tab) — x.com

In the archive

  • Press & Talks — Every dated press item.
All artifactsNext: The Best Multi-Agent System Knows When Not to Spawn