Qwen Bailian Thinking
Reasoning-intensive chat workload. Source: qwen_thinking_blksz_16.jsonl. Timestamps are trace-relative seconds; no absolute wall-clock times or request durations are included.
Requests
10,812
Duration
1.998 h
Mean RPS
1.503
Peak 1s RPS
40.00
Input p50/p99
3,680.00 / 24,972.68
Input min/max
11 / 51,622
Output p50/p99
1,665.50 / 34,686.53
Sessions approx.
10,303
Hash reuse
46.2%
Concurrency note: this trace gives request arrival timestamps and token lengths, but not service durations. The RPS plots describe arrival load; actual in-flight concurrency depends on the serving system latency during replay.
Arrival Rate




Token Lengths





Structure


Statistics
| Metric | Mean | p50 | p90 | p95 | p99 | p99.9 |
|---|---|---|---|---|---|---|
| Input tokens | 4,762.88 | 3,680.00 | 13,114.60 | 16,221.30 | 24,972.68 | 36,368.74 |
| Output tokens | 3,650.77 | 1,665.50 | 9,156.90 | 12,669.60 | 34,686.53 | 62,184.20 |
| Total tokens | 8,413.66 | 5,725.50 | 16,137.60 | 22,763.90 | 36,811.27 | 62,275.92 |
| Turn | 1.295 | 1.000 | 2.000 | 3.000 | 7.000 | 13.00 |
| Session requests | 1.049 | 1.000 | 1.000 | 1.000 | 2.000 | 2.000 |
| Hash blocks/request | 298.19 | 230.00 | 820.00 | 1,014.45 | 1,561.00 | 2,273.76 |
Arrival RPS by Window
| Window | Mean | p50 | p90 | p95 | p99 | Max |
|---|---|---|---|---|---|---|
| 1s | 1.503 | 1.000 | 3.000 | 4.000 | 8.000 | 40.000 |
| 10s | 1.502 | 1.300 | 2.410 | 3.400 | 5.881 | 10.500 |
| 60s | 1.502 | 1.383 | 2.273 | 3.321 | 3.871 | 4.300 |
| 5min | 1.502 | 1.463 | 2.229 | 2.322 | 3.416 | 3.740 |
| 10min | 1.502 | 1.493 | 2.171 | 2.574 | 2.943 | 3.035 |
Request Types
This file uses a single native workload type, thinking, so its per-type RPS matches the overall RPS.
| Type | Requests | Fraction | Mean RPS |
|---|---|---|---|
| thinking | 10,812 | 100.00% | 1.503 |
Hash Blocks
| Metric | Value |
|---|---|
| Total hash blocks | 3,223,990 |
| Unique hash blocks | 1,735,015 |
| Global hash reuse ratio | 46.18% |
| Requests with any previously seen hash | 92.12% |