Journal
Work
7 Oct 2025#llm#aws#serverless 1 min read

The LLM call moves behind a queue

A model call is too slow for a request-response API behind a gateway timeout.


The first version called the model synchronously inside the API handler. It worked in testing and timed out in production, because a gateway waits thirty seconds and a model does not promise to answer in thirty seconds.

Two weeks later the call moved behind a queue: a dispatcher accepts the request and returns a job id, a processor consumes the queue, and a status endpoint reports progress. The client had to change from one call to two.

sequenceDiagram
  Client->>API: request
  API->>Dispatcher: enqueue
  Dispatcher-->>Client: job id
  Processor->>Model: prompt
  Processor->>Store: result
  Client->>Status: poll
Takeaway

That is the price of an LLM in a request path, and it is cheaper than a timeout.

Client engagement, 2021 to 2026. Names, identifiers and internals are generalised.