Work
7 Oct 2025#llm#aws#serverless 1 min read
The LLM call moves behind a queue
A model call is too slow for a request-response API behind a gateway timeout.
The first version called the model synchronously inside the API handler. It worked in testing and timed out in production, because a gateway waits thirty seconds and a model does not promise to answer in thirty seconds.
Two weeks later the call moved behind a queue: a dispatcher accepts the request and returns a job id, a processor consumes the queue, and a status endpoint reports progress. The client had to change from one call to two.
sequenceDiagram
Client->>API: request
API->>Dispatcher: enqueue
Dispatcher-->>Client: job id
Processor->>Model: prompt
Processor->>Store: result
Client->>Status: pollTakeaway
That is the price of an LLM in a request path, and it is cheaper than a timeout.
Client engagement, 2021 to 2026. Names, identifiers and internals are generalised.