Pugalenthi Magendran / LAB 01
AN INTERACTIVE SYSTEM EXPLORER

Inside an AI Request.

What happens between pressing Send
and seeing the first token?

Illustrative simulationExplore the mechanism. All default timings are simulated.
R1 / Input preparation
ONE REQUEST · FIVE STAGES
01Tokenise

Text → numbers

R1 · active
02Waiting

Wait for capacity

03Prefill

Read prompt + write KV

04Decode

Reuse KV + extend output

05Streaming

Deliver as available

Streaming overlaps decode. Select a deeper view without losing your place.

0.0 / 203.8 ms · simulated
Slowed to 18 seconds at 1× · timeline shows simulated time
Your request R1 Competing requests Shared resources
Wait for the first token46.6 msSimulated · R1
Time in the queue0.0 msSimulated · R1
Complete response203.8 msSimulated · R1
Occupied KV blocks now0.00 MiBCalculated from simulated state
CHANGE ONE THING

Change one thing. See what happens.

Illustrative simulation
01 / PREDICT

What happens to the wait for the first token with a longer prompt?

Your prediction

Controlled: Qwen architecture, FP16 memory accounting, assumed hardware budget, 32 output tokens, scheduler policy. One request; cache off.

02 / RUN → 03 / INSPECT → 04 / EXPLAIN

Make a prediction, then test it.

All comparison cases run together. Choose one to replay in the system above. Predictions are optional; the result is the same either way.

FIXED OUTPUT32 tokens

No model answer is generated on this page.

About this lab Method, sources & limitations

An illustrative release.

Built for Pugalenthi Magendran’s Accelerated Computing atlas: deterministic simulation, a calculated attention example, two synchronised renderers, three controlled experiments, and a validated recording pathway.

Upstream model: Qwen/Qwen2.5-0.5B-Instruct, configuration revision 7ae5576. Proposed collector: vLLM 0.10.2. No compatible GPU or serving endpoint was available in the build environment. Engine and GPU compatibility must be confirmed by the collector preflight before collecting evidence.

All timings shown in the default mode are simulated. They are not predictions of Qwen throughput on a particular GPU. The conceptual hardware has an assumed 4 GiB capacity; this is not a hardware benchmark configuration.

Metric boundaries

Simulation starts at request arrival. Queue time starts after the fixed 2 simulated ms preparation and ends at admission. The prefill window includes any scheduling wait after admission and ends when the first token is produced; TTFT ends when it is received, 3 simulated ms later. Full response latency ends at last receipt. Batch throughput divides all 32 × request-count output tokens by the entire batch window.

Recorded mode uses client request start to first nonempty output chunk and response completion. It cannot infer queue, prefill, layer or core activity. Repeated runs show sample counts and min–max variation, not unsupported tail percentiles.

Read methodology & learning guide ↗Read the accuracy audit & verification limits ↗

Primary sources & inspiration

Measured evidence

0 validated recording sets available. Measured replay remains unavailable until genuine records pass the checksum and schema checks. No unavailable combination is interpolated.

Use the source collector and importer to add evidence; the default page needs no account, API key, weights or GPU.

Learning guide Predict, inspect, explain
  1. Start with a request. Pause at prefill. Explain why words and tokens are different, and why the input size is a separate control here.
  2. Open the model. Select position two in the attention example. Predict which weights are zero, inspect the mask and weighted output, then explain why position three is unavailable. Change position three’s input number; check that the earlier position’s output stays unchanged.
  3. Open the GPU. Double the stored token count on paper. Use 12 KiB per token to predict useful KV memory. Explain useful data, rounded occupied blocks, and reserved capacity.
  4. Run longer prompt. Change only input length. Compare first-token time, prefill and memory. Explain which work grows and what remains controlled.
  5. Run simultaneous requests. Select R8 and step through admission. Explain how overall throughput and one person’s latency can move in different directions.
  6. Run repeated prefix. Compare off, cold, warm and changed cases. Explain why exact-prefix reuse saves prompt work but does not produce an answer for free.
Glossary Plain-language definitions
Token
A text unit assigned an integer by a particular tokenizer; not necessarily a word.
Prefill
Processing the prompt and building attention state before the first output.
Decode
Repeatedly extending an output, one new token per sequence per step.
TTFT
Time to first token: request-start boundary to receipt of the first output. Client recordings here use the first nonempty text chunk.
Latency
Elapsed time from a request starting to the full response arriving.
KV cache
Stored attention keys and values. It is not a library of completed answers.
Prefix caching
Reusing eligible KV blocks from an exact earlier token prefix.
Continuous batching
Reconsidering which requests share execution as iterations complete.
Causal mask
A rule that prevents a token position from using future positions.
Softmax
Converts a row of scores into nonnegative weights that sum to one.
Bandwidth
How much data can move in a given time.
Capacity
How much data can fit at once.
Throughput
Total output tokens delivered per second over a specified observation window.
MiB / GiB
Binary memory units: 1 MiB = 1,048,576 bytes; 1 GiB = 1,073,741,824 bytes.
RMSNorm
Normalises a vector by its root mean square, with learned scaling.
Residual path
Adds a block’s input back to its output so information can carry forward.
RoPE
Rotary position embeddings: rotations applied to queries and keys to encode position.