Open it
In the archive
- Press & Talks — Every dated press item.
Archive · Article
The first of five articles in a written series on serving systems: I follow one request through eight stages to show why GPU latency and request latency are different numbers, and why a service that budgets by characters admits work it cannot hold.