Detailed summaries index

Qwen Bailian Thinking

Reasoning-intensive chat workload. Source: qwen_thinking_blksz_16.jsonl. Timestamps are trace-relative seconds; no absolute wall-clock times or request durations are included.

Requests
10,812
Duration
1.998 h
Mean RPS
1.503
Peak 1s RPS
40.00
Input p50/p99
3,680.00 / 24,972.68
Input min/max
11 / 51,622
Output p50/p99
1,665.50 / 34,686.53
Sessions approx.
10,303
Hash reuse
46.2%
Concurrency note: this trace gives request arrival timestamps and token lengths, but not service durations. The RPS plots describe arrival load; actual in-flight concurrency depends on the serving system latency during replay.

Arrival Rate

Token Lengths

Structure

Statistics

MetricMeanp50p90p95p99p99.9
Input tokens4,762.883,680.0013,114.6016,221.3024,972.6836,368.74
Output tokens3,650.771,665.509,156.9012,669.6034,686.5362,184.20
Total tokens8,413.665,725.5016,137.6022,763.9036,811.2762,275.92
Turn1.2951.0002.0003.0007.00013.00
Session requests1.0491.0001.0001.0002.0002.000
Hash blocks/request298.19230.00820.001,014.451,561.002,273.76

Arrival RPS by Window

WindowMeanp50p90p95p99Max
1s1.5031.0003.0004.0008.00040.000
10s1.5021.3002.4103.4005.88110.500
60s1.5021.3832.2733.3213.8714.300
5min1.5021.4632.2292.3223.4163.740
10min1.5021.4932.1712.5742.9433.035

Request Types

This file uses a single native workload type, thinking, so its per-type RPS matches the overall RPS.
TypeRequestsFractionMean RPS
thinking10,812100.00%1.503

Hash Blocks

MetricValue
Total hash blocks3,223,990
Unique hash blocks1,735,015
Global hash reuse ratio46.18%
Requests with any previously seen hash92.12%