← All writing
6 min readWritten by Kasi Vandanapu, with Claude Code as a pair

A Year of LLM Traffic, in a Browser Tab

6.12 billion production LLM requests, explored interactively with DuckDB-WASM and Mosaic. No backend: a one-time rollup turns the 91 GB Parquet file on AWS S3 into 24 MB for the charts, and raw rows are still read live from S3.

DuckDBParquetMosaicAWS S3LLM ServingData Engineering

Researchers from Harvard, the University of Chicago and Chutes published a year of production traffic from Chutes, a public LLM inference platform: 6.12 billion requests across 9,174 models, 11 April 2025 to 12 April 2026, as one 91 GB Parquet file on AWS S3. The paper is A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing, and the data and code are on GitHub. All credit for the dataset goes to its authors; this post is one way to explore it.

The figure below explores all of it in your browser, with no backend. Filter on the left, drag across any chart, or press Play. View raw & rollup data shows one matching hour two ways: real requests read live from the file on S3, and the same hour as the rollup rows the charts are built from.

When the year’s LLM traffic arrived, and what it was

367 days

All requests · no filtersTrend window
Requests
—
no prior 1d
Avg input
—
tokensno prior 1d
Avg output
—
tokens per requestno prior 1d
Avg TTFT · streaming
—
no prior 1d
Last interaction
—
filter or brush to measure
Group by
Period
Requests per day · by category · All

Anomalies

≥3× or ≤⅓ of the prior 14-day median · click to zoom

None in the current view.

Peak hours

—

no requests in view

 

Trace clock, not UTC: the release removed the real timezone.

Nothing matches the current filters.
Click a column to sort, a row to filter. Averages use measured requests only; amber means under 50% were measured.
Hour of day
Day × hour
Starting the query engine…
One year of production traffic from Chutes — 6.12 billion requests across 9,174 models. Charts read an hourly rollup of the trace; raw rows are read live from the original 91 GB file on AWS S3. Dates follow the paper’s stated span; hours are on the trace’s own clock. Data: Nixon et al., A Year in LLM Serving. Engine: Mosaic and DuckDB.

Architecture: roll up once, query in the browser

AWS S3 · chutes_trace.parquet          91 GB · 6.12B rows · one row per request
   │                                  │
   │ rolled up once (DuckDB)          │ read live (HTTP range requests)
   ▼                                  ▼
chutes.parquet   24 MB · 1.34M rows   Raw requests tab · SQL console
served with this page                  ~15 MB per sample
   │
   ▼
DuckDB-WASM + Mosaic, in your browser tab

The problem. Every chart here is a whole-year aggregate, and Parquet stores rows, not totals. Querying the S3 file directly for "requests per day" alone downloads 13.4 GB of timestamps; the full dashboard needs 45 GB of columns, per visitor. DuckDB's file pruning can't help: nothing can be skipped when every row counts.

The rollup. So the aggregation runs once, ahead of time, at the coarsest grain the charts need: one row per hour × model × function.

SELECT date_trunc('hour', started_at) AS hour,
       chute_id AS model, function_name,
       count(*)   AS reqs,
       sum(it)    AS sum_it,    count(it)   AS n_it,     -- input tokens
       sum(ot)    AS sum_ot,    count(ot)   AS n_ot,     -- output tokens
       sum(ttft)  AS sum_ttft,  count(ttft) AS n_ttft    -- time to first token
FROM 's3://harvardsys-datasets/2026_chutes_anonymized/chutes_trace.parquet'
GROUP BY ALL

6.12 billion rows become 1.34 million (models outside the top 200 are grouped as "other"), and 91 GB becomes 24 MB. It is exact, not sampled: counts and sums add up, so any coarser question (a day, a model family, all streaming traffic) is just a sum() over rollup rows, and the totals match the source file's own Parquet footer. Each sum keeps a count of the requests that actually had the value, so averages divide by measured values only.

Charts and filtersRaw rows
ReadsThe 24 MB rollup, served with the pageThe 91 GB original on AWS S3
DownloadOnce, then browser cache7.8 MB footer once, then ~15 MB per sample
LatencyTens of ms per interaction4–15 s per sample (30 Mbps)

What the rollup gives up: anything finer than an hour, and individual requests. That is what the second path is for. To see the transformation itself, open View raw & rollup data: for one hour it shows the raw requests from S3 next to the rollup rows that replace them. A typical hour is about 680,000 requests, stored as about 150 rows.

Raw rows, live from AWS. S3 serves byte ranges, and the bucket allows browser (CORS) requests. DuckDB reads the Parquet footer, then fetches only the row groups that contain the hour you asked about. It never downloads the 91 GB. Try SELECT count(*) FROM trace in the What's running console: it answers from the footer alone, in seconds.

Why interactions are fast. Mosaic turns every control into one shared filter and pre-aggregates small cubes the first time you use one (a second or two). After that, each drag filters a cube, not the 1.34M-row rollup, in tens of milliseconds. More on that in A Million Rows, 30 Milliseconds a Frame.

What the year shows

  • DeepSeek lost the lead to Qwen. DeepSeek served 85.4% of traffic in Q2 2025 and 35.7% by Q1 2026, when Qwen reached 38.5%. Group by Model family to see the handover.
  • Volume peaked in autumn 2025. 21.5M requests a day in Q3 against 12.5M in Q2; the busiest day, 29 October, carried 34.9M.
  • It is almost all chat. 86.1% chat, 13.1% completions. An average request sends 6,321 tokens and gets back 445: prompt-heavy, not generation-heavy.
  • There is no night. The busiest and quietest hours differ by under one percentage point (4.66% vs 3.94%): a global user base. The one gap is a ~30-hour outage on 13–14 January 2026.

Caveats

  • Dates come from the paper. The file's timestamps start at 1970-01-01; the paper gives the span as 11 April 2025 to 12 April 2026, and model launch dates in the trace agree (18 of 20 models first appear on their release day or later). Hours of day are on the trace's own clock, not a known timezone.
  • Missing values are structural. Time to first token is recorded only on streaming calls, so it exists for 49% of requests. Token counts are missing on 7.5%.
  • Model family and type come from model names, not from the trace.

Next: an agent that builds the rollup for you

The rollup here was designed by hand. I looked at what the charts ask, picked the grain (hour × model × function), and chose which measures to keep. That works for one dashboard. It does not scale to a platform with billions of rows and hundreds of recurring reports, which is why most teams still scan the full table for questions that only ever need a summary.

The next post builds an agent that does this design step automatically:

  1. Learn the workload. Read the queries a dashboard actually sends (this figure already logs every one) and cluster them by grouping, measures and filters.
  2. Propose rollups. Pick the coarsest grain that answers each cluster, keep sums and counts for additive measures, and flag the ones that don't add up (distinct users, percentiles).
  3. Price them. Estimate rollup size and bytes saved per query before building anything, against a storage budget.
  4. Build and verify with DuckDB, checking every total against the source.
  5. Route queries to the smallest rollup that can answer them, and fall back to the raw data otherwise. The agent proposes; deterministic code decides, so an answer is never wrong.

The test: pointed at this 91 GB trace and this dashboard's query log, can it rediscover the hour × model × function rollup on its own? It is the cost-aware planner I sketched at the end of A Million Rows, 30 Milliseconds a Frame, made concrete.


Data and credit: William Nixon (University of Chicago, Harvard University), Jon Durbin (Chutes), Florian Standhartinger (Chutes), Haryadi S. Gunawi (University of Chicago) and Juncheng Yang (Harvard University), "A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing," arXiv:2608.13573 (2026). Released under CC BY 4.0 and hosted on AWS S3 by the authors (HarvardMadSys/chutes_workload). The rollup, charts and any errors in them are mine. Built with Mosaic and DuckDB-WASM. Best viewed on a wide screen.