# Methodology and learning guide

## What is being shown?

One selected request passes through input preparation, waiting, prefill, decode and streamed delivery. The same request and elapsed time persist when switching between the request, model, GPU and load views. Default output markers represent hypothetical output positions. They are not words produced by a model.

The prompt editor is local. Its segments are an illustrative word split, not actual Qwen tokenisation. The input-count selector chooses a hypothetical workload independently of the edited text. Real recordings lock both their exact tokenizer IDs and input text.

## Evidence labels

- **Simulated**: all educational timings, queue events and throughput. Deterministic does not mean measured.
- **Calculated**: attention arithmetic, architectural weight storage and KV storage derived from the simulated token state.
- **Assumed**: the educational 4 GiB memory capacity, 128 MiB KV pool and 256 MiB working reserve. No hardware specification is implied.
- **Measured**: reserved for genuine collector timestamps and actual server token usage, after validation. Elapsed client metrics derived from these are labelled calculated from measured timestamps.

## Simulation policy v2

All requests arrive at t=0. Input preparation takes 2 simulated ms. The queue admits up to four requests in FIFO order at iteration boundaries. Active decoding requests receive one next token per iteration. The remaining prefill work is chunked with a 128-token total prompt budget, distributed FIFO. New work joins when a slot becomes free. A resident request holds a slot but may receive no prompt work in an iteration because earlier requests consume the prompt budget. The UI exposes exactly which requests receive prefill or decode work; shared compute activity is independent of the selected request.

The iteration duration in simulated milliseconds is:

`4 + 0.05 × prompt_tokens + 0.65 × decode_requests + 0.0008 × sum_decode_context_tokens`.

The coefficients are chosen teaching assumptions, not calibrated throughput or latency estimates. Decode context includes the prompt and outputs already fed back into the model. Prefill predicts the first output; the remaining 31 iterations produce a fixed total of 32 outputs. A fixed 3 simulated ms represents delivery after each generated token, so streaming overlaps decode. This includes illustrative application/network delay; there are no network measurements.

Every scenario is computed once, outside the render loop. The simulator returns an ordered event stream and sampled state. The view clock maps that duration to 18 wall-clock seconds at 1×; changing speed or frame rate never recomputes the scenario. Step goes to the next distinct event timestamp. Request counts are drawn as a step chart, with no fractional interpolation. The five states—preparing, waiting, resident, final delivery and complete—partition every request at each sampled timestamp. Pause and timeline scrubbing are deterministic.

This is a simplified continuous-batching policy, not vLLM’s scheduler implementation. It omits preemption, fair chunk rotation, admission based on exact GPU occupancy, speculative execution and heterogeneous lengths. Concurrency is limited to the tested set of 1, 4 or 8; input counts to 128, 512 or 2,048; output to 32.

## Attention arithmetic

Toy inputs: X = [[1,0], [0,1], [1,1]]. Wq = Wk = identity; Wv = [[1,2], [3,0]]. Q=XWq, K=XWk, V=XWv. Scores are QKᵀ/√2; entries with column > row are −∞. Row-wise numerically stable softmax gives weights summing to one. Output is weights × V.

Position one’s output is [1,2]. Position two’s output is approximately [2.339523,0.660477]. Changing future input cannot change earlier outputs. The microscope and table use the same calculation. A reproducible perturbation control varies only X[2][0] from 0 to 2, in steps of 0.05; Reset restores 1. Position selection changes the query row. Project, Match, Mask, Weight and Mix are teaching steps, not execution timings. The UI computes the values and exposes every matrix. These are educational values, not captured Qwen activations. The toy omits RoPE, learned normalisation, additional heads, MLP and residual paths.

Actual Qwen block ordering: embedding, then 24 repetitions of RMSNorm → attention (Q/K/V, RoPE, causal operation, output projection) → residual add → RMSNorm → SwiGLU feed-forward → residual add, followed by final RMSNorm and tied output projection. Layer selection is inspection focus, not a claim that only that layer computes or that layer timing was recorded. An attention map is not a complete explanation of reasoning.

## Memory accounting

KV bytes per stored token = 2 (key and value) × 24 layers × 2 KV heads × 64 head dimensions × 2 bytes = **12,288 bytes = 12 KiB**. Query-head count is not the KV-head count.

Useful KV data counts processed prompt tokens plus generated tokens already fed back. The final output has not been fed back, so a 512-input, 32-output request reaches 543 stored positions. Occupied blocks round each live request up to 16-token blocks. Capacity is a separate 128 MiB pool, not a sum of useful bytes. Allocation metadata, CUDA graphs, padding and actual activation peaks are excluded. A 256 MiB working reserve is an assumption, not a complete runtime allocation measurement.

The tied embedding/output architecture contains 494,032,768 parameters, approximately 0.920 GiB at two bytes each. This includes embedding weights, attention/MLP weights, Q/K/V biases and normalisation parameters. It does not mean a serving process uses only that much GPU memory.

Live KV blocks are released at an explicit final-generation timestamp, before the last delivery event. Peak memory includes the final KV append immediately before release, whereas state samples at that timestamp describe memory after release. Thus peak useful memory for the default single request is exactly 543 × 12,288 = 6,672,384 bytes, and peak occupied blocks are 544 × 12,288 = 6,684,672 bytes. Warm prefix state is assumed to be supplied to each admitted request; physical block sharing, retained-prefix pool occupancy and eviction are not modelled. Logical per-request memory is summed. Warm comparisons show saved computation, not a reduction in the request’s final useful context.

## Prefix eligibility

Off, cold and changed-first-token cases reuse zero prompt tokens. Warm cases reuse the largest full 16-token prefix before the final 32 input tokens (480 of 512 by default). The generic eligibility function compares exact token IDs in sequence and never reuses the final token. Changing the first ID prevents all prefix reuse. Similar meanings or partial blocks do not qualify. Decoding still occurs in all cases.

## Metrics

| Metric | Simulation boundary | Client recording boundary |
|---|---|---|
| Queue time | prepared → admitted | unavailable |
| Prefill window | admitted → first output produced; includes in-slot scheduling wait | unavailable |
| Time to first token | arrival → first output received | start of HTTP request → first nonempty content chunk; proxy, includes buffering |
| Full response latency | arrival → last output received | start of HTTP request → SSE DONE |
| Output throughput | total delivered outputs / batch window | sum actual server output counts / observed repetition window |
| KV memory | architectural calculation from simulated state | unavailable |

Simulation comparisons report one deterministic run per case. Latency ranges describe its individual requests, not statistical confidence intervals. The genuine collector repeats each case and the replay reports mean, min, max and sample count; no p99 claims from tiny samples.

The simulator also records `prefillServiceMs`: the sum of shared iteration durations in which that request receives prompt work. This is not isolated GPU kernel time; each iteration can include other requests and fixed teaching overhead.

## Purpose of motion

Requests move from queue positions to shared execution at recorded simulator admission times, then to completion positions. Orange identifies the selected request; a muted teal identifies competitors. Transfer markers express a conceptual memory/compute direction only, with no physical path, byte rate or core utilisation claim. Output tiles appear with delivery; each 3D tile represents one received output position, for 32 tiles. A request in final delivery stays downstream of compute rather than returning to the queue. Camera transitions bring model layers, memory and queue relationships into focus. No animation produces the numeric result. In the attention microscope, matrix columns encode score or weight magnitude above a fixed base; crossed cells indicate masked future positions. Ribbons link each selected-row weight to its value vector. Their width is a qualitative cue, not an exact linear scale; read the exact numerical weight. The Mix stage rescales the value bars to their calculated weighted contributions. Narrow-screen 3D frames only the current teaching step to keep cells readable. Orbit, zoom and reset reveal these spatial relationships. The 2D alternative retains the same calculation and controls.

## Learn by explaining

1. Pause after preparation. Explain why a token is not necessarily a word.
2. In the model view, select position two. Predict the zero attention weight, inspect the mask, and reconstruct the output from weighted V rows.
3. In the GPU view, compute 512 × 12 KiB. Explain why useful token data differs from rounded blocks and reserved capacity.
4. Run the prompt-length experiment. Hold output, concurrency and cache fixed. Explain the change in prefill and first-token wait.
5. Run the concurrency experiment. Follow R8. Explain why throughput can rise while R8 waits longer, and why memory pressure plateaus when the four active slots are full.
6. Run all prefix cases. Explain why a warm shared prefix reduces prefill, a changed first token invalidates reuse, and neither eliminates the 32 outputs.

## Primary sources and attribution

- [Qwen configuration at the selected revision](https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct/blob/7ae5576/config.json)
- [Qwen2 documentation, Transformers 4.55.2](https://huggingface.co/docs/transformers/v4.55.2/en/model_doc/qwen2)
- [vLLM 0.10.2 Qwen implementation](https://docs.vllm.ai/en/v0.10.2/api/vllm/model_executor/models/qwen2.html)
- [vLLM 0.10.2 prefix cache design](https://docs.vllm.ai/en/v0.10.2/design/prefix_caching.html)
- [vLLM 0.10.2 production metrics](https://docs.vllm.ai/en/v0.10.2/usage/metrics.html)
- [vLLM GPU installation requirements](https://docs.vllm.ai/en/v0.10.2/getting_started/installation/gpu.html)
- [vLLM 0.10.2 cache-manager implementation](https://docs.vllm.ai/en/v0.10.2/api/vllm/v1/core/kv_cache_manager.html)
- [vLLM 0.10.2 API server, including cache reset](https://docs.vllm.ai/en/v0.10.2/api/vllm/entrypoints/openai/api_server.html)
- [Three.js API](https://threejs.org/docs/)
- [Transformer Explainer](https://poloclub.github.io/transformer-explainer/) and [Brendan Bycroft’s LLM visualisation](https://bbycroft.net/llm): inspiration for inspectable and spatial teaching, not sources of benchmark evidence or copied designs.
