Inside an AI Request.
What happens between pressing Send
and seeing the first token?
Text → numbers
R1 · activeWait for capacity
Read prompt + write KV
Reuse KV + extend output
Deliver as available
Streaming overlaps decode. Select a deeper view without losing your place.
Change one thing. See what happens.
What happens to the wait for the first token with a longer prompt?
Controlled: Qwen architecture, FP16 memory accounting, assumed hardware budget, 32 output tokens, scheduler policy. One request; cache off.
Make a prediction, then test it.
All comparison cases run together. Choose one to replay in the system above. Predictions are optional; the result is the same either way.
No model answer is generated on this page.
About this lab Method, sources & limitations
An illustrative release.
Built for Pugalenthi Magendran’s Accelerated Computing atlas: deterministic simulation, a calculated attention example, two synchronised renderers, three controlled experiments, and a validated recording pathway.
Upstream model: Qwen/Qwen2.5-0.5B-Instruct, configuration revision 7ae5576. Proposed collector: vLLM 0.10.2. No compatible GPU or serving endpoint was available in the build environment. Engine and GPU compatibility must be confirmed by the collector preflight before collecting evidence.
All timings shown in the default mode are simulated. They are not predictions of Qwen throughput on a particular GPU. The conceptual hardware has an assumed 4 GiB capacity; this is not a hardware benchmark configuration.
Metric boundaries
Simulation starts at request arrival. Queue time starts after the fixed 2 simulated ms preparation and ends at admission. The prefill window includes any scheduling wait after admission and ends when the first token is produced; TTFT ends when it is received, 3 simulated ms later. Full response latency ends at last receipt. Batch throughput divides all 32 × request-count output tokens by the entire batch window.
Recorded mode uses client request start to first nonempty output chunk and response completion. It cannot infer queue, prefill, layer or core activity. Repeated runs show sample counts and min–max variation, not unsupported tail percentiles.
Read methodology & learning guide ↗Read the accuracy audit & verification limits ↗Primary sources & inspiration
- Qwen model configuration ↗
- Qwen2 implementation · Transformers 4.55.2 ↗
- vLLM 0.10.2 · prefix caching ↗
- vLLM 0.10.2 · production metrics ↗
- Transformer Explainer · educational inspiration ↗
- Brendan Bycroft · spatial explanation inspiration ↗
- Three.js · rendering API ↗
Measured evidence
0 validated recording sets available. Measured replay remains unavailable until genuine records pass the checksum and schema checks. No unavailable combination is interpolated.
Use the source collector and importer to add evidence; the default page needs no account, API key, weights or GPU.
Learning guide Predict, inspect, explain
- Start with a request. Pause at prefill. Explain why words and tokens are different, and why the input size is a separate control here.
- Open the model. Select position two in the attention example. Predict which weights are zero, inspect the mask and weighted output, then explain why position three is unavailable. Change position three’s input number; check that the earlier position’s output stays unchanged.
- Open the GPU. Double the stored token count on paper. Use 12 KiB per token to predict useful KV memory. Explain useful data, rounded occupied blocks, and reserved capacity.
- Run longer prompt. Change only input length. Compare first-token time, prefill and memory. Explain which work grows and what remains controlled.
- Run simultaneous requests. Select R8 and step through admission. Explain how overall throughput and one person’s latency can move in different directions.
- Run repeated prefix. Compare off, cold, warm and changed cases. Explain why exact-prefix reuse saves prompt work but does not produce an answer for free.
Glossary Plain-language definitions
- Token
- A text unit assigned an integer by a particular tokenizer; not necessarily a word.
- Prefill
- Processing the prompt and building attention state before the first output.
- Decode
- Repeatedly extending an output, one new token per sequence per step.
- TTFT
- Time to first token: request-start boundary to receipt of the first output. Client recordings here use the first nonempty text chunk.
- Latency
- Elapsed time from a request starting to the full response arriving.
- KV cache
- Stored attention keys and values. It is not a library of completed answers.
- Prefix caching
- Reusing eligible KV blocks from an exact earlier token prefix.
- Continuous batching
- Reconsidering which requests share execution as iterations complete.
- Causal mask
- A rule that prevents a token position from using future positions.
- Softmax
- Converts a row of scores into nonnegative weights that sum to one.
- Bandwidth
- How much data can move in a given time.
- Capacity
- How much data can fit at once.
- Throughput
- Total output tokens delivered per second over a specified observation window.
- MiB / GiB
- Binary memory units: 1 MiB = 1,048,576 bytes; 1 GiB = 1,073,741,824 bytes.
- RMSNorm
- Normalises a vector by its root mean square, with learned scaling.
- Residual path
- Adds a block’s input back to its output so information can carry forward.
- RoPE
- Rotary position embeddings: rotations applied to queries and keys to encode position.