FlashFlow
PROTOTYPEA Go lab that replays routing policies on identical traffic and classifies why one of them collapses under load.
Personal project · Aug–Sep 2026 · 226 commits on 5 days
THE PROBLEM
A single latency benchmark cannot show why a load-balancing policy collapses, because traffic, capacity, failures and cache state all change at once. FlashFlow fixes the traffic, topology, failures and seed, lets each policy make its own routing decisions, and rebuilds queue growth from the dispatch and completion events.
HOW IT WORKS
- A deterministic virtual-time engine runs six routing policies on the same seeded arrival trace. A second engine runs real net/http servers in one process over a simulated link.
- A backlog reconstruction rebuilds per-target queue depth from events, and a classifier labels each run STABLE, ACUTE_COLLAPSE, CHRONIC_COLLAPSE or RECOVERY_LIMITED from traffic concentration and committed work.
- Each research stage is written up against a claim ledger. Later stages retired or narrowed earlier claims, and in-repo audits corrected overclaims in the docs.
THE PIPELINE
- Seeded scenarioThe same traffic, topology, failures and seed for every policy.
- Six routing policiesEach makes its own decisions in a virtual-time engine, or in a real net/http engine over a simulated link.
- Event logDispatch and completion events.
- Backlog reconstructionPer-target queue depth rebuilt from the events.
- ClassifierSTABLE, ACUTE, CHRONIC or RECOVERY_LIMITED, from concentration and committed work.
- ReportExplanation, stress map and dashboard.
SCREENSHOTS
Captured from the project's own repository. Select an image to open it full size.
ENGINEERING EVIDENCE
- go test ./... on a fresh clone: 24 packages pass, 503 top-level tests, 0 failures. go build, go vet and gofmt are clean. About 46,900 lines of Go, of which 20,284 are 89 single-file experiment programs. proxy_test.go ↗
- Six policies are compared: round-robin, weighted round-robin, least-connections, EWMA, power-of-two-choices by in-flight count, and an adaptive weighted policy. policies.go ↗
- Re-running the flagship experiment (seeds 16000–16002) reproduced the committed result file except for its timestamp. Re-running the diagnostic report reproduced the documented labels for all six policies. 016-flagship-results.json ↗
- Load-aware routing was tested against load-blind routing at 8 targets and lost: round-robin beat EWMA in 10 of 10 seeds. That retired a claim from an earlier stage. 014I-statistical-confirmation.json ↗
- A claim that one policy had the worst P99 in every seed was checked against its own result file, found false for one seed, and retracted in a follow-up commit. Stage16-ClaimLedger.md ↗
DECISIONS & INVESTIGATIONS
The original doc and code comment said the constants were chosen up front and not tuned.
The claim was false as written.
No single measure explained both.
(1) Per-seed P99 in 016-flagship-results.json: seed 16000 EWMA 4399.88 ms > Adaptive 4072.11; 16001 Adaptive 4732.39 vs EWMA 4717.98; 16002 Adaptive 4853.07 vs EWMA 4725.03, so Adaptive was worst of six in 2 of 3...
- Committed backlog out-ranked peak load on severity, but its value moved about tenfold with one thresholdINCONCLUSIVE
Across topology size, committed backlog matched the severity ranking exactly (rank distance 0, n=4) while peak rho was misordered (distance 4; rho fell from 0.915 to 0.716 as N grew while EWMA mean latency rose from...
- 'Load-aware beats load-blind' was falsified at 8 targets: EWMA lost to round-robin in 10 of 10 seedsREJECTED
The rho claim was narrowed and the load-aware claim was falsified as general statements.
The ranking flips only at one capacity: Capacity 0 EWMA 15.90 vs Adaptive 27.38 ms (EWMA wins); Capacity 1 EWMA 131.06 vs 27.93 ms (Adaptive wins); Capacity 2 16.36 vs 27.38 and Capacity 3 16.04 vs 27.38 (EWMA wins)....
The defect was diagnosed from run-to-run variation in which target was locked and from an ablation: giving every request a unique key (removing cache affinity) still produced max_share 1.000.
Determinism was tested by repetition: Experiment 005-B ran an identical 9-event scenario 50 times with identical traces, and a re-run gave 50 of 50 identical.
WHAT ISN'T DONE
- Everything numerical is a simulation or an in-process run on one Windows machine. Targets are single FIFO queues with fixed service times. No code starts containers, and nothing ran against real traffic.
- The classifier's two thresholds were tuned against the same six outcomes it then reproduces, with no held-out scenario, so that agreement is a consistency check and not validation. It also labels weighted round-robin ACUTE_COLLAPSE even though that policy has the lowest P99.
- The dashboard and README quote "8x" for EWMA against the adaptive policy. That figure is EWMA against its own Capacity=0 baseline; the policy-to-policy gap is about 4.6–4.7x.
- Some recorded numbers are stale. A re-run of the keep-alive throughput test gave 2.15x where the docs say 3.06x, and the Stage 8 tuner figures moved after later seed changes.
- The "independent" audits are documents by the author, one run by 12 parallel AI agents. None is third-party review.
- Built with an AI coding assistant: at least 119 of the first 221 commits carried a Claude co-author trailer before a history rewrite removed most of them, and 28 of the current 226 still do. The commit history is also compressed: 214 of the 226 commits fall on four days.
NEXT STEPS
Each one comes from a gap listed above. It says what fixing the gap would take; it is not a promise.
- Validate the classifier's thresholds on a scenario it was not tuned on.
- Fix the README and dashboard "8x" figure, which is EWMA against its own baseline and not against the adaptive policy.
- Re-run the numbers that have gone stale, such as keep-alive throughput and the Stage 8 tuner, and update the docs.
STACK
- Go
- net/http
- GitHub Actions
- JavaScript dashboard
- Prometheus text format


