Load Testing LLM APIs: TTFT, Tokens per Second and the KV-Cache Cliff
LLM load testing guide: measure TTFT and inter-token latency on streams, find the KV-cache knee, plan around OpenAI, Anthropic and Bedrock rate limits.
LLM load testing means measuring how your model endpoint behaves under concurrent, streamed traffic, using metrics a normal API test ignores: time to first token (TTFT), inter-token latency and output tokens per second. You ramp realistic prompts until TTFT degrades, which on self-hosted models usually happens at the KV-cache cliff, and on hosted APIs at your rate limit.
If you have load tested REST APIs before, most of your instincts still apply. Our step-by-step k6 API guide covers the basics. What changes with LLMs is what you measure, how you shape the traffic and where the system breaks.
Why does request latency lie for LLM APIs?
A classic API test records one number per request: how long until the response finished. For a streamed LLM response that number blends three different things:
- TTFT: time from sending the request until the first token arrives. Dominated by queueing and prefill, the step where the model processes your whole prompt.
- Inter-token latency (ITL): the gap between tokens once generation starts. This is reading speed. A common way to compute the average is (end-to-end latency minus TTFT) divided by (output tokens minus 1).
- End-to-end latency: TTFT plus all the generation time, which mostly tracks how long the answer is.
Here is why that matters. A request that takes 12 seconds might be fine (fast first token, long answer) or terrible (10 seconds stuck in a queue, then a short burst). Average request duration cannot tell them apart. Under load, TTFT is usually the first metric to break, because new requests queue behind the ones already generating.
So report TTFT and ITL as percentiles (p50, p95, p99), and add one throughput number: total output tokens per second across all streams. That gives you user experience and capacity in the same chart.
How do you measure TTFT on a streaming response?
You have to read the stream. Most LLM APIs send tokens as server-sent events (SSE) when you set stream: true. Your load tool needs to timestamp the first content event, not the connection opening and not the final byte.
- k6: the community xk6-sse extension adds an SSE client with
open,eventanderrorcallbacks, and accepts POST requests with a JSON body. Per its README, k6 can now import it directly ask6/x/ssewithout a custom build. Record the first event time in a custom Trend metric. The repo ships an OpenAI-compatible example that measures time to first token and generation speed. - Locust: use the HTTP client with
stream=Trueandcatch_response=True, iterate the response lines, and log the time of the first non-empty data line as its own metric. Plain Locust does not report LLM metrics out of the box, so this part is DIY. - Purpose-built benchmarks (covered below) measure TTFT and ITL natively.
One trap: count tokens from the response, not from chunk counts. Some servers send several tokens per SSE event, so “events per second” can understate real token throughput. Most APIs return token usage at the end of the stream (some only when you ask for it), so use that.
What does realistic LLM traffic look like?
A single fixed prompt is the LLM version of hammering one product ID: it flatters the cache and tells you nothing. Real traffic varies on two axes that both drive cost and latency:
- Prompt length. RAG and agent prompts can be thousands of tokens of context. Long prompts make prefill expensive, which pushes TTFT up for everyone in the batch.
- Output length. Long answers hold a slot and its KV cache for the full generation. A test where every reply is 50 tokens will badly overstate capacity for a product whose real replies run to 800.
Pull both distributions from production logs if you have them: the p50, p90 and max of input and output tokens per request type. If you are pre-launch, use a public conversation dataset as a starting point and adjust. Also model arrivals properly. Users do not arrive in neat waves, so use a Poisson or bursty arrival pattern rather than a fixed number of looping virtual users. vLLM’s benchmark tool exposes exactly this with a --burstiness setting.
Where is the KV-cache cliff?
When a transformer generates text, it keeps the attention keys and values for every token of every active request in GPU memory. That is the KV cache, and it grows with prompt length plus output length. Inference servers batch many requests together to keep the GPU busy, so throughput rises nicely as you add concurrency, right up until the cache is full.
Then you fall off the cliff. vLLM’s own docs describe what happens when KV cache space is insufficient for all batched requests: it preempts some requests and recomputes them later when space frees up, and logs a warning with a cumulative preemption count. From the outside you see it as:
- Output tokens per second stops growing (the knee)
- TTFT p95 jumps, often sharply, as requests queue or get recomputed
- Preemption counts climbing in the server metrics
That knee is your real capacity number. Find it with a stepped ramp: hold each concurrency level long enough to stabilise, record TTFT p95 and total tokens per second, and plot both. The vLLM docs suggest the usual levers once you know where it is: raise gpu_memory_utilization to leave more room for the cache, or lower max_num_seqs and max_num_batched_tokens to cap how much is in flight. Re-test after each change, because the knee moves with prompt mix. A system that holds 64 concurrent chat users may fold at a fraction of that with long RAG prompts.
For a quick comparison of serving engines before you test, see our sister site’s Ollama vs vLLM breakdown.
How do rate limits change testing on hosted APIs?
On OpenAI, Anthropic or Bedrock you will not see a KV-cache cliff. You will see 429s. Each provider counts differently, and the details shape your test design (all checked against provider docs in October 2026):
| Provider | What is limited | The gotcha for load tests |
|---|---|---|
| OpenAI | Requests and tokens per minute and per day, set per tier and model | Limit is calculated from the larger of max_tokens and the estimated tokens, so an oversized max_tokens burns quota |
| Anthropic | Requests, input tokens and output tokens per minute, per model, using a token bucket | Sharp traffic increases can trip acceleration limits, so ramp gradually. Cached input reads do not count toward input limits on most models. max_tokens does not count toward output limits |
| Amazon Bedrock | Requests and tokens per minute, plus a daily token quota, per model | Input tokens plus max_tokens are deducted at request start. Output tokens on newer Anthropic models burn quota at 5x to 15x, depending on the model |
The practical rules: ramp traffic gradually rather than slamming from zero, set max_tokens close to your real output size, honor the retry-after header, and log the rate-limit response headers each provider returns so you can see which limit you hit first. Treat the throttling point as a result, not a failure. “We get 429s at 40 percent above current peak” is exactly what you need to request a quota increase before launch.
Which LLM load testing tools should you use?
| Tool | Best for | Streaming metrics | Notes |
|---|---|---|---|
| k6 + xk6-sse | Teams already on k6, CI gates, mixed app and LLM journeys | TTFT via custom metrics | Community extension; includes an OpenAI-compatible LLM example |
| Locust | Python teams, custom user behavior | DIY with streamed responses | Flexible, but you build the LLM metrics yourself |
| vllm bench serve | Self-hosted vLLM or OpenAI-compatible endpoints | TTFT, TPOT, ITL, end-to-end | --request-rate, --max-concurrency, --burstiness, --goodput SLOs, ShareGPT and random datasets. Older guides call it benchmark_serving.py |
| GuideLLM | Finding max sustainable rate against SLOs | TTFT, ITL, end-to-end distributions | Sweep, constant, Poisson and concurrent profiles; synthetic prompts with set token counts, Hugging Face datasets and trace replay |
| LLMPerf | Legacy multi-provider comparisons | ITL and throughput | Archived by its owner in December 2025; fine for reading old results, not a base for new work |
A sensible split: use vllm bench serve or GuideLLM to characterise the model server itself, and k6 or Locust for end-to-end tests through your real API, auth, RAG retrieval and guardrails. The second is what users actually hit. If you are choosing between those two, our k6 vs Locust comparison covers the trade-offs.
If your team works in k6, Shamal, an open-source project from the owner of this site, is worth a look for the analysis step. It drafts k6 scenarios from an OpenAPI spec or HAR file, and its agent investigates results with deterministic checks, including saturation-knee detection. It analyses k6 runs rather than speaking LLM protocols itself, and it is early (v0.1.0). The code is at github.com/shamal-io/shamal.
What does it cost per 1,000 users?
Once you have measured token mix and capacity, cost is arithmetic. Here is the model we use, with illustrative numbers you should swap for your own:
Hosted API. Assume 1,000 daily active users, 20 requests each per day, 1,500 input tokens and 400 output tokens per request. That is 30 million input and 8 million output tokens a day. At illustrative prices of $3 per million input tokens and $15 per million output tokens, that is $90 plus $120, so about $210 per 1,000 users per day. Prompt caching and shorter outputs move this number more than anything else.
Then check the peak, not the average. If 10 percent of daily traffic lands in the busiest hour, that is 2,000 requests in an hour: about 50,000 input and 13,000 output tokens per minute. On Bedrock with a 5x output burndown, those 13,000 output tokens count as roughly 67,000 against quota, before the up-front max_tokens reservation. That is where tight quotas surprise teams.
Self-hosted. Cost per 1,000 requests is your GPU node’s hourly cost divided by the requests per hour it sustains at the knee, within your TTFT SLO, times 1,000. Illustratively, a $4-per-hour node that holds 5 requests per second within SLO serves 18,000 requests an hour, about $0.22 per 1,000 requests. Measure past the knee and this number looks great on paper while users wait.
Get a fixed-price LLM capacity test
Most teams launching an AI feature know their token prices and almost nothing about their TTFT at peak or where their quota runs out. We run a fixed-price LLM capacity test: realistic prompt and output distributions from your logs, a streamed ramp to the KV-cache knee or your provider’s throttling point, TTFT and ITL percentiles at each step, and a cost-per-1,000-users model built from measured numbers. You keep the k6 or Locust scripts. It builds on our Capacity Assessment method and pairs with capacity planning for the rest of your stack.
Shipping an LLM feature this quarter? Get in touch and we will scope a fixed-price capacity test around your model, provider and launch date.
Frequently Asked Questions
What is the most important metric in LLM load testing?
For chat and agent UIs it is time to first token (TTFT) at p95 or p99, because that is the wait before anything appears on screen. Pair it with inter-token latency for reading speed and total output tokens per second for capacity. Total request latency alone mixes all three and hides which one is breaking.
Can k6 load test a streaming LLM API?
Yes, with the community xk6-sse extension, which k6 can now import directly as k6/x/sse. It supports POST bodies with stream set to true and event callbacks, so you can record the first token time in a custom Trend metric. Its repo includes an OpenAI-compatible LLM example.
What is the KV-cache cliff?
Every in-flight request holds attention state, the KV cache, in GPU memory, and it grows with prompt plus output length. When concurrent requests need more than the reserved space, servers like vLLM preempt and later recompute requests. Latency then climbs sharply while throughput stops growing. That knee is your real capacity.
How do I load test OpenAI or Anthropic without getting rate limited?
You will hit rate limits eventually, so treat them as a test target. Ramp gradually, since Anthropic warns that sharp usage spikes can trigger acceleration limits. Read the rate-limit response headers, honor retry-after on 429s, and set max_tokens close to real output size, because OpenAI counts it against tokens per minute.
How much does an LLM capacity test cost to run?
On hosted APIs you pay normal token prices for every test request, so token spend scales with test duration and prompt sizes. Keep runs short and targeted: a ramp to find the knee, a hold at expected peak, and a spike. On self-hosted GPUs the main cost is node hours for the test window.
Complementary NomadX Services
Related Articles
Know Your Scaling Ceiling
Book a free 30-minute capacity scope call with our load testing engineers. We review your architecture, traffic expectations, and upcoming scaling events - and scope the load test that will give you the data you need.
Every engagement is scoped by our principal architect, Adrian Vale: 20+ years in production engineering, 40+ professional certifications. Meet Adrian
Talk to an Expert