CORTEX
INVESTIGATIONS

Investigations

Questions answered with evidence: 24 with measurements and 0 root-cause analyses that were reproduced. Each lists the method, the result and the source of every number. Where a result was inconclusive or a change was rejected, it says so. Decisions that came out of them are under decisions.

Showing 24 of 24

  1. FuzeFEATURED

    Load test at 75 users: a fail-closed JWT revocation check logged out valid users

    ADOPTED

    Question

    Where does the single-worker gevent deployment break under concurrent authenticated load, and why?

    Method

    Disposable local rig in Docker: Postgres+pgvector, Redis, the real 'gunicorn --workers 1 --worker-class gevent' command in a Linux container, 10 users x 40 bookmarks seeded, JWTs minted directly to bypass the login rate limiter, Locust (scripts/locustfile.py) hitting bookmarks, dashboard, text and semantic search, unified-orchestrator recommendations and health. Ran 25 and 75 concurrent users on a laptop Docker VM of about 3.8 GB with shared CPU.

    Result

    At 75 concurrent users about 38% of authenticated requests returned 401 'Token has been revoked' for tokens that were never revoked (approximate figure from gaps.md and the commit message; no raw Locust output is committed). The author's diagnosis: the Redis pool (max_connections 20, shared by caching, rate limiting and the revocation check) was exhausted and .exists() raised. check_if_token_revoked in run_production.py returned True on any exception, so every failure counted as revoked. Fix in 764e2c6: fail open on Redis errors or a missing connection (the check runs only after signature and expiry validate) and raise the pool to 50. A re-run at 75 users reported 0 failures in 595 requests; because both changes landed together and fail-open cannot return this 401, that does not prove the pool is no longer exhausted. At 25 users the author reports median latency under 100 ms with no failures; at 75 users, 2-5 s, which the author calls consistent with the single gunicorn worker (two workers not re-tested).

    Numbers

    authenticated 401 rate at 75 users before fixgaps.md section 13; commit 764e2c6 message. Author-reported; no raw output committed.
    about 38% (approximate as written)
    failures at 75 users after fixgaps.md@764e2c6 line 401. Fail-open and pool 50 were applied together.
    0 failures across 595 requests
    Redis pool sizegit show 764e2c6 -- backend/utils/redis_utils.py
    20 to 50 connections
    median latency at 25 usersgaps.md section 13 (author-reported)
    under 100 ms, 0% failures
    median latency at 75 usersgaps.md section 13 (author-reported, laptop Docker VM of about 3.8 GB)
    2-5 s

    Evidence

    How this was checked: Read the run_production.py and redis_utils.py diffs, the locustfile and gaps.md section 13. Not re-run: needs Docker (daemon not running here) and Postgres+Redis. No regression test pins the fail-open behaviour (no 'revoked_jti' in tests). Numbers are from a laptop VM, so only the shape transfers to real hosting; the fail-open choice trades some revocation enforcement during Redis trouble for availability. Attribution: Commits 764e2c6, a74c00c and 491a221 carry a 'Co-Authored-By: Claude Sonnet 5' trailer, and gaps.md is an AI-written first-person document (it addresses the repository owner as 'you').

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit 764e2c6

  2. FlashFlowFEATURED

    'Adaptive has the worst P99 in every seed' checked against its own JSON and retracted

    REJECTED

    Question

    The README and Stage 16 docs said Adaptive's P99 was the worst of six policies in all three flagship seeds (16000-16002). Does the committed result file say that?

    Method

    An audit compared the docs to experiments/016-final-synthesis/results/016-flagship-results.json (5 targets at 15-75 ms, Capacity=1, FlashCrowd workload, 8 s horizon, three seeds with 0.3 arrival jitter, six policies). The claim was rewritten in README, Stage16.md, the claim ledger (C24, now RETIRED with a note on the earlier overstatement), the Stage 16 learning notes and docs/index.html. For this record, cmd/experiment-016-flagship was also re-run four times on a clone of HEAD.

    Result

    The claim was false as written. Adaptive's P99 was the highest of the six policies in seeds 16001 (4732.39 vs EWMA 4717.98 ms, a 0.3 percent margin) and 16002 (4853.07 vs 4725.03 ms), but not in 16000, where EWMA was higher (4399.88 vs 4072.11 ms) and Adaptive was second-highest. In no seed was Adaptive in the lower half of the six. The wording adopted in the repo ('worst of six in 2 of 3 seeds, near-tie in the third') is loose: the near-tie is seed 16001, not 16000. Adaptive's mean latency (583-714 ms) is well below EWMA's (935-983 ms) in every seed, so mean latency alone would have hidden the tail problem. Four re-runs gave identical stdout and identical JSON apart from the timestamp. The source demo doc (Stage16-FlagshipDemo.md) had already hedged ('at or near the worst'); the overclaim entered when it was summarised into README and the ledger.

    Numbers

    Seed 16000 P99: EWMA vs Adaptive (ms)experiments/016-final-synthesis/results/016-flagship-results.json; re-run identical
    4399.88 vs 4072.11 (Adaptive not worst)
    Seed 16001 P99: Adaptive vs EWMA (ms)same; re-run identical
    4732.39 vs 4717.98 (0.3% apart)
    Seed 16002 P99: Adaptive vs EWMA (ms)same; re-run identical
    4853.07 vs 4725.03 (Adaptive worst)
    Mean latency, Adaptive per seed (ms)re-run of go run ./cmd/experiment-016-flagship
    638.06 / 714.37 / 582.73
    Mean latency, EWMA per seed (ms)re-run of go run ./cmd/experiment-016-flagship
    970.95 / 935.06 / 982.73
    Changed lines across 4 re-runs, excluding timestampgit diff after each run in the clone
    0

    Evidence

    How this was checked: Read the bbaecb1 message and diff and opened the flagship result JSON at HEAD. Ran the flagship experiment four times in a clone at HEAD (Go 1.23.3); every P99 and mean value matched the committed file and only the timestamp changed. The correcting commit is authored by Ujjwaljain16; its message credits an independent from-scratch audit, which the project describes as an AI-agent pass, not a human review.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit bbaecb1

  3. FlashFlowFEATURED

    One metric could not explain why two policies never drain: acute vs chronic collapse

    ADOPTED

    Question

    In an overloaded 5-target scenario, two of six routing policies never drain their queue within the horizon. Does one measure (traffic concentration or 'committed backlog') explain both?

    Method

    cmd/experiment-015a runs one fixed scenario (5 targets 15-75 ms, Capacity=1, FlashCrowd base 20 req/s to peak 300 req/s at t=2.5 s, 8 s horizon) across six policies and records top-1 share, peak queue depth, committed backlog (requests dispatched to the bottleneck between congestion onset and material diversion), fraction of the run above rho 1.0, and whether the queue drains. A falsification run (F1, commit dacce77) sent Adaptive load far under capacity. The flagship experiment repeated the scenario over three jittered seeds.

    Result

    No single measure explained both. Round-robin has a small committed backlog (4) but spends 0.711 of the run over capacity and never drains; Stage15.md attributes this to its fixed 1/5 share to the slowest target exceeding that target's capacity (chronic). Adaptive has a large backlog (86 in 015a) and 0.700 of the run over capacity, also without draining (acute over-commitment during the burst). EWMA has the largest backlog (97) but drains at 6850 ms. F1 (a 3-target topology at load far under capacity) showed concentration alone is not enough: Adaptive reached top-1 share 1.000 with no congestion and zero backlog. Over three jittered flagship seeds round-robin was identical (backlog 4, 0.711, no drain) while Adaptive's backlog was 107/127/93; Adaptive drained in seed 16000 but not in 16001 or 16002, so the acute/chronic contrast is cleanest in the single 015a run. Stage15.md lists the limits: one scenario, six policies, no systematic search for a third failure shape.

    Numbers

    015a committed backlog: RR / WRR / LC / EWMA / P2C / Adaptivego run ./cmd/experiment-015a; matches Stage15.md table; re-run identical
    4 / 9 / 6 / 97 / 6 / 86
    015a fraction of run above rho 1.0: RR / AdaptiveStage15.md Backlog Dynamics table; re-run time_above_rho1.0: RR 5689 ms, Adaptive 5597 ms (of an 8000 ms horizon)
    0.711 / 0.700
    015a drains within horizon: RR / Adaptive / EWMAre-run: t7 found=false for RR and Adaptive, 6850 ms for EWMA
    No / No / Yes (6850 ms)
    015a top-1 share: RR / Adaptivere-run output
    0.212 / 0.599
    Flagship RR across seeds 16000/16001/16002go run ./cmd/experiment-016-flagship
    backlog 4/4/4, fraction above cap 0.711/0.711/0.711, drains=false in all
    Flagship Adaptive backlog across seedsgo run ./cmd/experiment-016-flagship
    107 / 127 / 93 (threshold 0.30)
    Flagship Adaptive drains within horizon, seeds 16000/16001/16002go run ./cmd/experiment-016-flagship, drains= field
    yes / no / no (EWMA: yes / no / yes)
    F1 (015b): Adaptive at low load, 3 targetsgo run ./cmd/experiment-015b
    top-1 share 1.000, peak depth 1, no congestion, committed backlog 0

    Evidence

    • commitdc74b6c ↗Adds the canonical scenario and the first six-policy result; message notes RR backlog 4 yet never drains.
    • commitdacce77 ↗Falsification runs: F1 (Adaptive top-1 share 1.000 with zero congestion), F3, F4, F6.
    • commitcc66922 ↗Stage 15 close-out: states committed backlog explains acute collapse only and a second measure is needed for chronic.
    • fileStage15.md@14da821 ↗Backlog Dynamics section and the Limitations list (single scenario, no search for a third shape).
    • file015A-canonical-scenario-backlog-dynamics.json@14da821 ↗Committed per-policy numbers.

    How this was checked: Read Stage15.md, the dc74b6c, dacce77 and cc66922 messages and cmd/experiment-015a/main.go. Re-ran experiment 015a and the flagship in a clone at HEAD; the committed JSON and the docs matched except for the timestamp. The decisive commits are authored by Ujjwaljain16 and were not part of an external review.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit dc74b6c (canonical scenario and first results); interpretation in cc66922

  4. RecoveryOSFEATURED

    Five-seed campaign against a compliance-aware baseline: positive in every seed

    ADOPTED

    Question

    Once the baseline gets the same attempt budget and obeys the same compliance rules, does RecoveryOS still recover more revenue, and is the result stable across seeds?

    Method

    tests/evaluation/multi_seed_runner.py runs five seeds through the live pipeline in an accelerated-cooldown mode (retry cooldown set to 0 so rescheduled re-evaluations are due immediately), with the diagnoser pinned to an invalid Gemini key so the deterministic fallback diagnoser is used and AI fusion stays at its default of off. It then computes a single-attempt baseline, a compliance-blind same-budget baseline (diagnostic only) and the compliance-aware same-budget baseline. Each seed also records safety_integrity checks: duplicate ledger rows and attempts, and that computing baselines did not change decision-table row counts.

    Result

    Incremental recovery versus the compliance-aware baseline was positive in all five seeds, and RecoveryOS's recovered payments were a strict superset of the baseline's each time (baseline_only = 0). The mean is 73,181.78 rupees, and the README mean and 95% CI (52,918.53 to 93,445.04 rupees) match the artifact. Against the compliance-blind diagnostic comparator RecoveryOS loses in all five seeds. Integrity counters show zero duplicate ledger rows or attempts. Because the runs used the deterministic diagnoser with AI fusion off, the lift is not attributable to the LLM, and unsafe_ai_deltas = 0 is trivially satisfied. Caveats: the baseline stops at its first blocked verdict; the strict-superset result may follow from shared per-attempt draws (an inference, not stated in the repo); the artifact predates the 2026-09-06 policy rule change and was not regenerated.

    Numbers

    Per-seed incremental vs compliance-aware baseline (paise)tests/evaluation/artifacts/multi_seed_compliance_aware_aggregate.json incremental_recoveryos_vs_compliance_aware_fair_paise
    5796757; 8458182; 9029644; 7923669; 5382640
    Mean incremental (artifact and independent recomputation agree)aggregate.mean_incremental_recovery_paise; python statistics.mean over the five per-seed values
    7318178.4 paise (73,181.78 rupees)
    Standard deviation (artifact and independent recomputation agree)aggregate.incremental_recovery_std_paise; python statistics.stdev
    1632205.069 paise
    95% confidence interval of the mean, as reported in the artifact and READMEaggregate.incremental_recovery_95pct_t_ci_paise; reproduced exactly with t = 2.776 (the exact t for 4 degrees of freedom, 2.7764451, gives [5291528.13 ; 9344828.67], about 3 rupees wider on each side)
    [5291853.03 ; 9344503.77] paise
    Recovered revenue means, RecoveryOS vs compliance-aware baselineaggregate; python mean over per-seed recovered_revenue_paise
    113346288.2 vs 106028109.8 paise
    Payments recovered only by RecoveryOS, per seedpayment_level_comparison_vs_compliance_aware_baseline
    22; 35; 43; 39; 35 (baseline_only 0 in all)
    Compliance-blind diagnostic comparator, per seed (paise)incremental_recoveryos_vs_compliance_blind_fair_paise_DIAGNOSTIC_ONLY; docs/phase8_priority0_multi_seed_baseline.md addendum states the same mean
    -20535838; -12017715; -15832245; -11483251; -11225425 (mean -14218894.8, i.e. -142,189 rupees)
    Recovery rate, RecoveryOS vs compliance-aware baseline (mean of 5 seeds)aggregate.recoveryos_recovery_rate_mean, compliance_aware_baseline_recovery_rate_mean; recomputed from per-seed recovered_count / failed_payments
    0.4566 vs 0.4213

    Evidence

    How this was checked: Loaded the aggregate JSON, recomputed mean, sd, CI, means of revenue and the diagnostic gaps in Python and compared to README section 9 and the prose addendum in docs/phase8_priority0_multi_seed_baseline.md. Did not re-run the campaign (about 10 minutes per seed plus Docker). Attribution: Cited commits 8d19486 and 7c65faa carry no Claude co-author trailer; authored by Ujjwaljain16. The repo overall is AI-assisted (16 of 139 commits carry the trailer).

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit 8d19486

  5. RecoveryOSFEATURED

    LightGBM's 0.04 AUC lift came from duplicated rows; logistic regression stays in production

    ADOPTED

    Question

    The Phase 2 certificate said LightGBM beat logistic regression by 0.0401 AUC and cleared the lift gate of more than 0.03. Was that lift real?

    Method

    Set-compared episode IDs across the train and validation parquet splits instead of trusting the certificate. The builder generated val_random and test_scenario with separate build_simulator calls that re-used the seeds of train and test_random respectively, so their first N episodes replayed the same RNG stream. Compared LightGBM and LR on test_temporal, a split with zero overlap with train. After giving those two splits their own seeds (47ffb5f), regenerated the dataset and retrained (60b95eb).

    Result

    8,820 of val_random's 15,000 rows (58.8%) were verbatim copies of train rows, and 8,739 of test_scenario's 15,000 rows (58.3%) duplicated test_random rows. On the clean test_temporal split LR was marginally ahead of LightGBM (0.8378 vs 0.8374), so the >0.03 gate failed and propensity.py loads model_lr.pkl. After the seeds were decorrelated (47ffb5f), the regenerated artifacts show no LightGBM lift on val_random either (0.8324 vs 0.8324) and 0.0001 on test_temporal. Separately, test_leakage_seed.py, a leakage check on an independent seed, had no test function and never ran in CI until 67f7857.

    Numbers

    Old val_random AUC, LightGBM vs LR (contaminated split)git show cdfa12c:models/recovery/artifacts/eval_val_random.json
    0.8757 vs 0.8356 (gap 0.0401)
    Old test_temporal AUC, LightGBM vs LRgit show cdfa12c:models/recovery/artifacts/eval_test_temporal.json
    0.8374 vs 0.8378
    Duplicate rows: val_random vs train / test_scenario vs test_randomgaps.md section C.2 (the parquet data is not committed, so this was not re-derived)
    8,820 of 15,000 (58.8%) / 8,739 of 15,000 (58.3%)
    New val_random AUC, LightGBM vs LRmodels/recovery/artifacts/eval_val_random.json at HEAD
    0.8324 vs 0.8324
    New test_temporal AUC, LightGBM vs LRmodels/recovery/artifacts/eval_test_temporal.json at HEAD; matches README section 10 and the 60b95eb message
    0.8365 vs 0.8364 (lift 0.0001)
    Root-cause check: same seed, 1,500 episodes generated twicead-hoc script over build_simulator and EpisodeGenerator at HEAD (visible features plus label); run locally, not committed
    1500 of 1500 identical rows; different seed 0 of 1500
    Propensity unit testspytest tests/unit/test_propensity.py at HEAD (re-run during this audit), includes test_lgbm_does_not_beat_baseline_on_the_real_holdout_so_lr_stays_default
    14 passed

    Evidence

    • commit4d5dd77 ↗Documents the 59 percent duplicate finding and the model-selection correction in gaps.md.
    • commit09f9ef9 ↗Production adapter loads LR, not LightGBM, citing the held-out AUC.
    • commit47ffb5f ↗Gives val_random and test_scenario decorrelated seeds.
    • commit60b95eb ↗Regenerates data and re-certifies; LR remains the model with 0.0001 lift.
    • commit67f7857 ↗Adds the CI step because test_leakage_seed.py at repo root had no test_ function and never ran.
    • docgaps.md ↗Full audit; its 'not yet fixed' paragraph is stale after 47ffb5f.

    How this was checked: Read gaps.md C.2, the old and new eval JSON via git show and from HEAD, the three fix commits, and test_leakage_seed.py. Ran the propensity unit tests and my seed-determinism script. Could not re-derive the 58.8 percent figure because the parquet data is gitignored. Attribution: None of the cited commits (4d5dd77, 09f9ef9, 47ffb5f, 60b95eb, 67f7857) carry a Claude co-author trailer; all are authored by Ujjwaljain16. The repo overall is AI-assisted (16 of 139 commits carry the trailer).

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit 4d5dd77 (production adapter using LR in 09f9ef9 the same day; splits fixed in 47ffb5f and 60b95eb on 2026-09-02)

  6. MiniDBFEATURED

    Vectorized executor missed its 5-10x goal: measured about 1.0-2.3x

    INCONCLUSIVE

    Question

    How much faster is the DataChunk engine than the Volcano engine on SELECT name, salary FROM employees WHERE age > 50 (10% selectivity), and why is it not the 5-10x the authors expected?

    Method

    benchmarks/volcano_vs_vectorized.ts: for 10k, 50k and 100k rows, rows are inserted directly into the heap, then the physical plan runs through the Volcano Executor and through VecSeqScan/VecFilter/VecProject, with 3 warm-ups and 5 timed iterations each, each in its own transaction. Architecture.md attributes the small gap to per-row async overhead and to strict-2PL row locking. To test that, a copy of the script (not committed to the repository) was run at 100k rows only, in four variants, three times each: default; TrackedOperator.next() replaced by a pass-through (removes the per-row performance.now() calls the Volcano tree makes); LockManager.acquireRowLock replaced by an empty async function; both changes together. No repository files were modified.

    Result

    Six runs of the documented script gave about 1.5-2.3x at 10k rows, 1.1-1.4x at 50k and 1.0-1.3x at 100k, so the gap is small and noisy; the documented 1.19x at 100k is inside that range. Stubbing out row locks cut vectorized time at 100k rows from roughly 183-224 ms to 57-71 ms and Volcano from 222-270 ms to 100-116 ms. Volcano is also timed with two performance.now() calls per row per operator (TrackedOperator), about 14-24% of its time; with tracking off the 100k speedup was about 0.95-1.17x. With both locks and tracking removed, the vectorized engine was only about 1.1-1.4x faster. The comparison is not like for like: Volcano materializes a ResultSet, while the vectorized side only counts surviving rows in the selection vector. Runs made while the machine was busy varied widely and are not reported.

    Numbers

    Documented, 10k rows: Volcano / Vectorizeddocs/BENCHMARKS.md, README section 10
    24.86 ms / 11.39 ms (2.18x)
    Documented, 100k rows: Volcano / Vectorizeddocs/BENCHMARKS.md
    173.25 ms / 145.02 ms (1.19x)
    Documented earlier measurementsdocs/Architecture.md, chapter 5 section 8 (lines 3494-3509)
    1.26x at 10k, 1.82x at 50k, 1.27x at 250k after direct page decoding
    Re-run, 10k rowsnpx tsx benchmarks/volcano_vs_vectorized.ts (my run)
    Volcano 28.01 ms / Vectorized 12.38 ms (2.26x)
    Re-run, 50k rowssame run
    Volcano 102.86 ms / Vectorized 71.42 ms (1.44x)
    Re-run, 100k rowssame run
    Volcano 218.71 ms / Vectorized 166.72 ms (1.31x)
    Independent re-runs (5 runs) of the unmodified script, speedup rangenpx tsx benchmarks/volcano_vs_vectorized.ts, verifier runs
    10k: 1.48-2.22x; 50k: 1.12-1.35x; 100k: 1.02-1.32x
    Ablation, 100k, default, 3 runs (Volcano / Vectorized ms)zz_exp_locks2.ts, no env vars
    269.80/223.58, 244.18/200.02, 221.97/183.00
    Ablation, 100k, per-row tracking offzz_exp_locks2.ts, PLAIN=1
    213.50/210.99, 185.11/195.50, 191.51/163.36
    Ablation, 100k, row locks stubbed outzz_exp_locks2.ts, NOLOCK=1
    116.34/70.90, 108.27/56.72, 99.72/57.31
    Ablation, 100k, locks stubbed and tracking offzz_exp_locks2.ts, NOLOCK=1 PLAIN=1
    81.07/61.09, 76.91/53.94, 91.21/79.81

    Vectorized speedup over Volcano, by table size

    Five independent runs of the unmodified benchmark script. Each bar spans the lowest to the highest speedup seen.

    • 10,000 rows
      1.48x to 2.22x
    • 50,000 rows
      1.12x to 1.35x
    • 100,000 rows
      1.02x to 1.32x

    Shaded band: the stated goal of 5-10x.

    Source: npx tsx benchmarks/volcano_vs_vectorized.ts, five runs (see the numbers above)

    The same 100,000-row query with parts of the engine switched off

    Volcano time divided by vectorized time, three runs per variant. Dots are single runs.

    • Default
      1.21x to 1.22x
    • Per-row tracking off
      0.95x to 1.17x
    • Row locks stubbed out
      1.64x to 1.91x
    • Locks stubbed and tracking off
      1.14x to 1.43x
    • 1x: no speedup
    Source: Ratios computed from the millisecond figures of the four ablation runs listed in the numbers above

    Evidence

    • commit7b7c4ce ↗Adds the benchmark and BENCHMARKS.md 'Engineering Reality' text: 'While we didn't achieve 10x...'.
    • fileArchitecture.md@6a0c80d ↗Chapter 5 sections 7-8: the lock bottleneck explanation and the 1.26x/1.82x/1.27x measurements. The 350 ms vs 20 ms Amdahl figures there are introduced with 'Suppose' and are illustrative, not measured.
    • fileExecutor.ts ↗TrackedOperator wraps every operator and calls performance.now() around each next().

    How this was checked: Ran volcano_vs_vectorized.ts unmodified; ran my ablation copies (untracked, since removed from the clone; kept under decisions/minidb_experiments) three times per variant; read VecSeqScan.ts, SeqScanOp.ts, Executor.ts, VecProject.ts. Single machine, Node 22.19, Windows 11, other processes may have been running. Attribution: All cited commits are authored by Ujjwaljain16 (git shortlog: 21 commits, one author). The README lists a two-person team, and git history cannot show which parts each person wrote, so the record should say "commits by Ujjwaljain16; two-person course project". No Co-Authored-By trailers exist in the history; whether AI tools were used cannot be determined from the repository. The 5-10x expectation and the lock explanation come from docs/Architecture.md (chapter 5), whose long narrative style may be AI-assisted; thi

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit 7b7c4ce (benchmarks/volcano_vs_vectorized.ts, docs/BENCHMARKS.md, Architecture.md chapter 5)

  7. TypeAheadXFEATURED

    Why one Redis node took about 60% of hits although keys were split evenly

    ADOPTED

    Question

    Is the 60% share of redis-a in the 100,000-request benchmark a defect in the ring, or a result of the Zipfian workload?

    Method

    Recreated the workload of scripts/final_benchmark.py offline: 10,000 items (10 named queries such as 'iphone 16', then query_item_10 to query_item_9999), ranks from numpy.random.zipf(a=1.3) clipped to 10,000, 100,000 requests, mapped to nodes through ConsistentHashRing(500) using the cache key 'suggestion:<query>' (the service lower-cases and trims the prefix, which does not change these keys). Five seeds. Counted the owner of every item and of the top ten. This assumes every request is a hit on its owning node, which the live metrics count only for hits. The reproduction script is not in the repository.

    Result

    The ring balances keys but not traffic. The 10,000 keys split 3,361 (redis-a), 3,477 (redis-b) and 3,162 (redis-c), about 33.6% / 34.8% / 31.6%. The three most popular items ('iphone 16', 'chatgpt', 'samsung galaxy s24') all hash to redis-a; rank 1 alone is 25.2 to 25.7% of requests and rank 2 is 10.2 to 10.5%. Simulated request share of redis-a over five seeds: 59.77% to 59.98%. The repo reports 60.30 / 20.99 / 18.71 (docs/performance-report.md) and 60.58 / 20.82 / 18.60 (README.md, docs/unexpected-findings.md) for the same finding, and the README calls the workload 'real traffic', although the hot query names are synthetic and defined in final_benchmark.py. The proposed fix (an in-process L1 cache or hot-key replication) is not implemented: nothing like it exists in backend/app.

    Numbers

    Unique keys per node (10,000 items)offline reproduction script (ring is deterministic; script not in the repo)
    3,361 / 3,477 / 3,162
    Simulated request share of redis-a, 5 seedsoffline reproduction script (not in the repo)
    59.77% to 59.98%
    Live share reported in docsdocs/performance-report.md
    60.30% / 20.99% / 18.71%
    Other live share in docsREADME.md; docs/unexpected-findings.md
    60.58% / 20.82% / 18.60%

    Evidence

    • commit05af6b4 ↗Adds docs/performance-report.md and docs/unexpected-findings.md with the finding.
    • filefinal_benchmark.py ↗generate_queries defines the synthetic workload and top-10 names.
    • commit66845c0 ↗The ring whose placement produces the skew.

    How this was checked: Ran my offline reproduction (hot.py) five seeds; read final_benchmark.py, performance-report.md, unexpected-findings.md; grepped backend/app for LRU/L1. The full stack (Redis x3, Postgres, API) was not started. Attribution: Course project: the README does not say so, but docs/phase4-completion.md@66845c0 contains 'Viva Talking Points' and the phase briefs 3.md, 4.md and 5.md (committed, then deleted in the next phase) are written as instructions to a student ('most students will completely mess up'), so the work followed a written brief. All 15 commits are by Ujjwaljain16, none has a Co-Authored-By trailer, and 14 of 15 fall on one day (2026-06-10, 00:23 to 21:22 +05:30; the last is 2026-06-22).

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · docs added in commit 05af6b4; my offline reproduction 2026-09-28

  8. TypeAheadXFEATURED

    Redis cache-aside cut database reads but did not change latency in a local test

    ADOPTED

    Question

    Does adding a Redis cache in front of the PostgreSQL prefix query reduce request latency?

    Method

    Per docs/phase3-performance-comparison.md: 1,000 sequential requests from scripts/cache_benchmark.py, 90% drawn from 5 fixed 'hot' prefixes and 10% from 7 fixed 'cold' prefixes, so only 12 distinct cache keys. The 'no cache' column is not a paired run: it is the phase-1 baseline (docs/phase1-performance-baseline.md), measured earlier with scripts/benchmark.py, which sent 1,000 requests for the single prefix 'iph'. Read from history; not re-run because it needs Postgres and Redis with the 150,000-row dataset.

    Result

    Recorded: p50 about 7.5 ms without the cache and about 7.1 ms with it; p95 about 10.0 ms and about 12.2 ms; database reads 1000 per 1k requests (one per request, by construction) versus 13; hit rate 98.7%. Each figure is one run, so the latency differences cannot be told from noise, and with only 12 distinct keys about 13 misses per 1,000 is expected of any cache. The doc attributes the flat latency to an unloaded local PostgreSQL holding the data in RAM (asserted, not measured). The 100,000-request run in docs/performance-report.md (a different workload) reports warm-cache p50 420.01 ms, p99 790.03 ms and cold p50 175.00 ms, p99 401.88 ms. README.md line 15 still says 'a sub-millisecond autocomplete experience', which no recorded end-to-end number supports. The projected 'about 5 ms' warm p50 on Kubernetes is not measured.

    Numbers

    p50 without / with Redisdocs/phase3-performance-comparison.md@0c36797 (baseline column is the earlier phase-1 run, single prefix)
    about 7.5 ms / about 7.1 ms
    p95 without / with Redissame; single run
    about 10.0 ms / about 12.2 ms
    Database reads per 1,000 requests without / with Redissame; the workload has 12 distinct prefixes, and 1000 is one read per request by construction
    1000 / 13
    Phase-1 baseline p50 / p95 / p99 (single prefix 'iph')docs/phase1-performance-baseline.md (deleted in 05af6b4); scripts/benchmark.py@8fe2206
    7.48 / 9.95 / 11.86 ms
    100k-request warm cache p50 / p99, throughput (different workload)docs/performance-report.md@HEAD
    420.01 ms / 790.03 ms, 197.17 RPS
    100k-request cold cache p50 / p99 (different workload)docs/performance-report.md@HEAD
    175.00 ms / 401.88 ms

    Evidence

    How this was checked: Read the phase-3 and phase-1 docs via git show, performance-report.md and README.md. Not re-run (needs the full stack). Attribution: Course project: the README does not say so, but docs/phase4-completion.md@66845c0 contains 'Viva Talking Points' and the phase briefs 3.md, 4.md and 5.md (committed, then deleted in the next phase) are written as instructions to a student ('most students will completely mess up'), so the work followed a written brief. All 15 commits are by Ujjwaljain16, none has a Co-Authored-By trailer, and 14 of 15 fall on one day (2026-06-10, 00:23 to 21:22 +05:30; the last is 2026-06-22).

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · doc first committed in 0c36797 (later deleted in 05af6b4)

  9. Lexis AIFEATURED

    Swapping DFS for Kahn's algorithm broke cycle repair on 2 of 4 test graphs

    REJECTED

    Question

    Does GraphRepair.repair_cycles still remove cycles when fed the output of the new Kahn-based GraphValidator, as it did with the old DFS validator?

    Method

    Loaded graph_validator.py and graph_repair.py from 6dee148^ (DFS version) and from the current main branch (Kahn version) and ran the same four hand-made graphs through GraphValidator.validate, GraphRepair.repair_cycles and validate again on the remaining edges: a pure 3-cycle A->B->C->A, the same cycle with the concepts listed in another order, the cycle plus an upstream node (R->A) and a downstream node (C->D), and two disjoint cycles (A,B,C and X,Y). All edges are PREREQUISITE edges with a confidence value. The check script is not part of the repository.

    Result

    The DFS version left no cycle in all 4 graphs; the Kahn version left a cycle in 2 of 4. For A->B->C->A with a downstream node D, the Kahn validator reports the residual node set [A, B, C, D]. repair_cycles treats a reported cycle as a path (its comment describes a DFS path such as [A, B, C, A]) and builds edges from consecutive entries, so it selects C->D, deletes that valid edge and the cycle remains. With two disjoint cycles the validator reports all five nodes as one cycle and repair removes only X->Y, leaving A, B, C cyclic. The two passing cases contain only the cycle itself, where list order happens to match path order. The orchestrator publishes without re-validating and no test covers cycle repair, so the regression would go unnoticed. The graphs are hand-made; graphs produced by the LLM were not tested.

    Numbers

    Old DFS validator + repair: graphs left cyclic after repairSame four graphs run against graph_validator.py and graph_repair.py from 6dee148^
    0 of 4
    Kahn validator + repair (current main): graphs left cyclic after repairSame four graphs run against the files on main
    2 of 4 (cycle with downstream node; two disjoint cycles)
    Kahn output for cycle A,B,C plus downstream DSame run
    [['A','B','C','D']]; edge removed: C->D

    Evidence

    • commit6dee148 ↗The switch to Kahn's algorithm with graph_repair.py unchanged (git diff --stat shows no change to graph_repair.py).
    • filegraph_validator.py@6dee148 ↗detect_cycles returns [cycle_nodes]: all nodes whose in-degree stays above 0 after Kahn's pass, in concept-list order, including nodes downstream of the cycle.
    • filegraph_repair.py@6dee148 ↗Builds cycle edges from consecutive entries and the closing edge.

    How this was checked: Re-ran the check on files taken from 6dee148^ and from main; outputs match the numbers above. Read ingestion_orchestrator.py (validate, repair once, publish) and tests/test_golden_book.py (assertions commented out).

    Reconstructed from code and history · reasons are inferred · commit 6dee148

    team repository (4 contributors); the code under test is by Ujjwaljain16 (git blame attributes all lines of graph_validator.py and graph_repair.py to this account); the check described below was run during a later review on 2026-09-28Link to this record
  10. NevUpAIFEATURED

    The '100% revenge-flag accuracy' check compares seeded labels with themselves

    REJECTED

    Question

    DECISIONS.md says the revenge flag algorithm identified all 10 revenge-flagged trades in the seed dataset (100%). Does the worker's rule reproduce those labels?

    Method

    Read seeds/seed.ts and scripts/validate_metrics.ts: the seed inserts revenge_flag from the dataset label (revengeFlag === "true") and does not call computeRevengeFlag; validate_metrics reads that stored column back and compares it with the same dataset label. Then transcribed the predicate of computeRevengeFlag in src/worker/metrics.ts (a trade whose emotionalState is anxious or fearful and whose entry falls within 90 s after the exit of one of the same user's losing trades) into a short Python script and ran it over nevup_seed_dataset.json (388 trades, 10 labelled revenge). The script is not part of the repository.

    Result

    The 10/10 figure cannot fail: it compares the seeded revenge_flag column with the dataset labels it was loaded from, and the worker's rule is never applied. The same script's main() also prints the '8/10' pathology match and the Avery Chen and Jordan Lee lines as fixed strings (validatePathologyDetection is defined but never called). Applying the worker's SQL rule to the 388 seeded trades gives 0 true positives, 14 false positives, 10 false negatives and 364 true negatives. All 10 labelled trades follow a losing trade and have an anxious or fearful state, but each starts 60 or 120 s after that trade's entry, while it is still open (it exits 300 to 12,600 s later), so the rule's condition of an entry within 90 s after a loss exit never holds. The Python transcription was not run against Postgres.

    Numbers

    Trades / labelled revenge in seednevup_seed_dataset.json
    388 / 10
    Worker rule vs seed labels: TP / FP / FN / TNPython transcription of computeRevengeFlag run over nevup_seed_dataset.json (not published)
    0 / 14 / 10 / 364
    Gap from labelled trade entry to previous trade entrySame script over the seed
    60 or 120 seconds for all 10
    Pathology match rate printed by validate_metrics.tsnevup-backend/scripts/validate_metrics.ts@3d3b274
    '8/10' (hard-coded string; validatePathologyDetection never called)
    Claimnevup-backend/DECISIONS.md
    'Revenge flag accuracy: 100% (all 10 revenge-flagged trades in the seed correctly identified)'

    Evidence

    • commit7420567 ↗Adds computeRevengeFlag with the 90 s exit-to-entry rule and the anxious/fearful gate.
    • commit62f45dd ↗Adds scripts/validate_metrics.ts.
    • fileseed.ts@3d3b274 ↗Inserts revenge_flag from row.revengeFlag; SEED_METRICS block never calls computeRevengeFlag.
    • filevalidate_metrics.ts@3d3b274 ↗Compares DB revenge_flag with trade.revengeFlag from the same dataset.

    How this was checked: Read seed.ts, validate_metrics.ts, metrics.ts, worker/index.ts (reconciliation also skips revenge flags); ran the Python check (output: trades 388 labelled 10 TP 0 FP 14 FN 10 TN 364).

    Reconstructed from code and history · reasons are inferred · scripts/validate_metrics.ts added in 62f45dd; claim in DECISIONS.md/README from 7e5eace

    hackathon (single author)Link to this record
  11. GitIssueFEATURED

    Week 4.5 eval: 1 auto-labelled true positive and a corrupted final_score column

    REJECTED

    Question

    What does scripts/week45_evaluate.py report on the sampled real-repo data, and is it evidence that the duplicate suggester meets its precision/recall targets?

    Method

    Read week45_evaluate.py (precision = TP/(TP+FP) over labelled suggestions; recall = known duplicates that have any suggestion row; PR curve at thresholds 0.50 to 0.95), reports/week45_report_small.json and reports/week45_label_samples_small.csv (20 rows). Recomputed each CSV row's expected final score with the committed formula 0.5*semantic + 0.2*keyword + 0.2*structural + 0.1*label and compared it with the stored final_score. Read app/feedback/logger.py.

    Result

    The report (generated 2026-03-17) has 20 labelled suggestions: 1 true positive, 0 false positives, 1 related-not-duplicate and 18 cant_tell, all with labeled_by 'bootstrap-auto', an automated step whose script is not in the repository. Precision 100% is 1 of 1. Recall 66.67% counts 2 of 3 known duplicates that have any suggestion row, at any score; the curve shows recall and F1 of 0 at all ten thresholds, so the recommended threshold 0.5 is only the first entry. Separately, 19 of 20 stored final_score values equal label_score: the ON CONFLICT update in logger.py sets final_score = $7, which is label_score (final_score is $8), so re-scoring a pair overwrites it. The one row that matches the formula is the only one with source signal strength below the 0.3 gate, so it was scored once. The true positive stores 1.0 but the formula gives 0.399, below the 0.85 comment threshold. The report cannot support a precision or recall claim.

    Numbers

    Labelled suggestions (TP / FP / related / cant_tell)reports/week45_report_small.json
    20 (1 / 0 / 1 / 18)
    Labels produced by 'bootstrap-auto'reports/week45_label_samples_small.csv (labeled_by column)
    20 of 20
    Precision / recall reportedreports/week45_report_small.json
    100.0% (1/1) / 66.67% (2/3 known duplicates)
    PR curve recall and F1 at thresholds 0.50 to 0.95reports/week45_report_small.json
    0.0 / 0.0 at all 10 thresholds
    Rows where final_score == label_scorePython comparison over the CSV
    19 of 20
    Rows where final_score == 0.5*sem+0.2*kw+0.2*st+0.1*labelSame comparison
    1 of 20 (id 940, the only row with source signal strength 0.0)
    TP row 2966: stored vs formula final scoreSame comparison
    1.0 stored; 0.3992 by formula

    Evidence

    • commit8eda41b ↗Adds the evaluation scripts and the 'small' report and label sample.
    • commit34031e1 ↗Introduces `final_score = $7` in the ON CONFLICT clause of log_suggestion (git blame of line 91 at 8eda41b); the line is still present on main.
    • filelogger.py@8eda41b ↗INSERT parameter order: $7 label_score, $8 final_score; UPDATE uses $7 for final_score.
    • fileweek45_label_samples_small.csv@8eda41b ↗labeled_by = bootstrap-auto on every row.
    • fileweek4.5md@8eda41b ↗The plan lists 'Precision (manually verified)' as a goal; the sampled labels in the report were all produced by 'bootstrap-auto'.

    How this was checked: Read the script, JSON report and CSV; ran a python comparison of stored final_score, label_score and the formula for all 20 rows (19 equal label_score, 1 equals the formula). The evaluation itself could not be re-run: it needs Postgres with the collected tables. Attribution: The planning documents (week1.md, v1.md, week2.md, week4.5md) are pasted AI-assistant replies; the code and evaluation scripts are committed by Ujjwaljain16 with no Co-Authored-By trailer.

    Reconstructed from code and history · reasons are inferred · commit 8eda41b (scripts/week45_evaluate.py and reports/week45_report_small.json); bug introduced in 34031e1 (2026-03-17)

  12. Apache Superset

    Why a cache-key normalisation PR was closed: the mismatch came from an in-place mutation

    REJECTED

    Question

    Can hashing ad-hoc SQL in `QueryObject.cache_key()` after normalising line endings and whitespace make the web-server and Celery keys agree, and is that where the mismatch arises?

    Method

    The author (Ujjwaljain16) made `cache_key()` deep-copy `to_dict()`, convert CRLF to LF and strip whitespace in ad-hoc SQL for metrics, columns and orderby, with 9 unit tests; `series_limit_metric` was added after a Copilot review comment. A user, kwilt, then tried it on a real deployment, logged the sqlExpression from the web process and from Celery for one chart and diffed them: they differed by a transpilation (cast(...) vs ::numeric, `not x is null` vs `x is not null`), not only whitespace. Separately, bobjo-daangn traced the write/read asymmetry in PR #40993: `get_sqla_query()` wrote the processed (Jinja-rendered, sqlglot-normalised) ORDER BY expression back into a dict shared with `QueryObject.orderby` and the cached QueryContext, so the worker computed the key from the raw expression but cached a mutated context, and the later GET recomputed a different key.

    Result

    The hashing-layer PR fixed only part of the cases: kwilt reported it fixed some of the 422 errors but not all, and #40993 states that a hashing-layer normalisation such as #38227 cannot fix the Jinja-rendering divergence because the rendered SQL is already baked into the cached context. #40993 changed `col = cast(AdhocMetric, col)` to `col = cast(AdhocMetric, dict(col))` in superset/models/helpers.py so a copy is processed; it was merged on 2026-06-16 and closes #37114. On 2026-09-12 the author closed #38227 with the comment that the issue was resolved by #40993. The author had earlier (2026-06-08) called the transpilation divergence a different cause and asked maintainer villebro whether to extend the PR; the thread records no maintainer answer and the PR received no human approval.

    Numbers

    Tests in the closed PRgh pr diff 38227 --repo apache/superset
    9 new unit test functions in tests/unit_tests/queries/query_object_test.py at the final head (Copilot's 2026-03-04 summary of an earlier commit counted 8)
    Size of the closed PRgh pr view 38227 --json additions,deletions,changedFiles
    +166/-2 lines, 2 files
    Production change in the fixing PR #40993gh pr view 40993 --json files
    one statement replaced in superset/models/helpers.py (a dict(col) copy; +4/-1 lines with a comment), plus 114 added test lines

    Evidence

    • pull requestapache/superset/pull/38227 ↗Closed PR, kwilt's divergence diff of 2026-06-08 and the author's closing comment of 2026-09-12.
    • pull requestapache/superset/pull/40993 ↗Root-cause fix by bobjo-daangn; merged 2026-06-16, merge commit 257dafeec51a62c6bac9d648b7c284020b1fe718 (`fix(query): don't mutate ad-hoc ORDER BY expressions when building queries`).
    • issueapache/superset/issues/37114 ↗Reports with cache-key dumps showing `\r\n` vs `\n` and Jinja-rendered vs raw expressions; the author's proposal comment of 2026-02-24.
    • testhelpers_test.py (PR #40993) ↗#40993 adds `test_get_sqla_query_does_not_mutate_adhoc_orderby`, a Jinja variant, and `test_cache_key_stable_across_query_build`, which asserts QueryObject.cache_key() is unchanged by building the query.

    How this was checked: Read PR #38227 body, comments, reviews and inline comments; issue #37114; PR #40993 body, diff and merge metadata via gh. Did not clone the repository or run tests. The record describes what the author's PR attempted and the project's later root-cause fix; it does not claim the author found the mutation. Attribution: #40993 (the fix) was written and merged by others (bobjo-daangn; approved by rusackas and betodealmeida, merged by betodealmeida, merge commit 257dafe on 2026-06-16).

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · PR #38227 closed by its author (the fix had landed in PR #40993, merged 2026-06-16)

  13. FlashFlow

    Committed backlog out-ranked peak load on severity, but its value moved about tenfold with one threshold

    INCONCLUSIVE

    Question

    Stage 15 proposed 'committed backlog' as the measure that explains collapse severity. Does it rank scenarios better than peak rho, and how sensitive is it to its own settings and to when it can be computed?

    Method

    (1) cmd/experiment-015e ranks 4 topologies (N=3/5/8 graduated and an N=8 bimodal) and 3 workloads (constant, burst, flash crowd), with EWMA as the policy throughout, by mean latency and compares that ranking with rankings by peak rho and by committed backlog. (2) While building the multi-seed flagship, a fixed diversion-share threshold of 0.5 was found to fit 3-target topologies (fair share 1/3) but not 5 targets (fair share 1/5), and was scaled to 1.5/N (commit 5cea6ef). (3) For this record the flagship was re-run in a clone with the threshold forced back to 0.5. (4) A code audit (b2d9d30) found that the onset detector anchors to the episode holding the peak depth, which requires that episode's future.

    Result

    Across topology size, committed backlog matched the severity ranking exactly (rank distance 0, n=4) while peak rho was misordered (distance 4; rho fell from 0.915 to 0.716 as N grew while EWMA mean latency rose from 93.78 to 307.32 ms). Across workload shape it was better but imperfect (distance 2 vs 4, n=3). Stage15.md notes these are small deterministic tests, not seeded replications. The value depends strongly on the diversion-share threshold: at 0.5 Adaptive's committed backlog in the three flagship seeds is 10/9/10, at 1.5/N = 0.30 it is 107/127/93, while EWMA's is unchanged (98/104/95). Stage 15's Adaptive figure of 86 used 0.5 with no arrival jitter, while the jittered flagship at the same threshold gives about 10. Stage15.md states that no full threshold-sensitivity sweep was run. The measure is also retrospective: it cannot be computed live, and README and the ledger were corrected to say so. An earlier first-episode-only version reported EWMA's backlog as 1 instead of 97 (commit dc74b6c).

    Numbers

    Cross-topology rank distance from severity: committed backlog vs peak rhogo run ./cmd/experiment-015e; identical to committed 015E JSON
    0 vs 4 (n=4, EWMA only)
    Cross-workload rank distance: committed backlog vs peak rhogo run ./cmd/experiment-015e
    2 vs 4 (n=3, EWMA only)
    Peak rho for N=3/5/8 graduatedsame
    0.915 / 0.833 / 0.716
    EWMA mean latency for N=3/5/8 graduated (ms)same
    93.78 / 201.81 / 307.32
    Adaptive committed backlog, flagship seeds, threshold 0.30 (recorded)go run ./cmd/experiment-016-flagship, identical to committed file
    107 / 127 / 93
    Adaptive committed backlog, flagship seeds, threshold forced to 0.5re-run in a clone with 'var diversionShareThreshold = 0.5' in cmd/experiment-016-flagship/main.go (edit reverted afterwards)
    10 / 9 / 10
    EWMA committed backlog at 0.5 and at 0.30same two runs
    98 / 104 / 95 at both
    Adaptive backlog in 015a (threshold 0.5, no jitter)go run ./cmd/experiment-015a
    86

    Evidence

    • commit2edd061 ↗Adds the cross-topology and cross-workload rank test (distance 0 vs 4, and 2 vs 4).
    • commit5cea6ef ↗Commit message documents the 0.5 threshold problem (Adaptive backlog 9-10 vs 93-127) and the fix to 1.5/N.
    • commitb2d9d30 ↗Adds the retrospective-only caveat to FindPeakEpisodeCongestionOnset, README and the ledger.
    • commitdc74b6c ↗First-episode onset gave EWMA backlog 1 vs 97; fixed and covered by a two-episode test.
    • testTestFindPeakEpisodeCongestionOnset_MultipleEpisodes (backlog_test.go@14da821) ↗Hand-computed two-episode case showing the first-episode and peak-episode onset finders disagree; passes in a clone (go test ./internal/backlog).
    • filemain.go@14da821 ↗Header comment explains the threshold scaling and the 9-10 vs 93-127 observation.
    • fileStage15.md@14da821 (Limitations 1 and 5) ↗States that no full threshold-sensitivity sweep was run and that the severity-ranking claim rests on a 4-point deterministic test.

    How this was checked: Re-ran experiments 015e, 015a and the flagship in a clone at HEAD (results identical to the committed files apart from timestamps). Set the flagship threshold constant to 0.5 in the clone, ran it and discarded the change. Ran the internal/backlog tests (all pass) and read the b2d9d30 diff. The retrospective-only caveat comes from a commit whose message credits an independent from-scratch audit, which was an AI-agent pass, not a human review.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit 5cea6ef (threshold finding); generalization test in 2edd061; caveat in b2d9d30

  14. FlashFlow

    'Load-aware beats load-blind' was falsified at 8 targets: EWMA lost to round-robin in 10 of 10 seeds

    REJECTED

    Question

    Stage 13 concluded that the deepest regime boundary was load-blind vs load-aware routing, and that a rho of about 0.89-0.97 marks the collapse transition. Do both claims hold as the number of targets grows from 3 to 8?

    Method

    cmd/experiment-014c scaled requests at N=3/5/8 to aim at EWMA's max-target rho of about 0.9. cmd/experiment-014f ran all six policies at N=8 below, near and above the capacity boundary (150/381/600 requests), labelled by signal source rather than name. cmd/experiment-014i repeated the near-boundary point over 10 seeds (14700-14709, jitter 0.3) with Cliff's Delta and bootstrap CIs.

    Result

    The rho claim was narrowed and the load-aware claim was falsified as general statements. EWMA's achieved rho fell as N grew (0.915, 0.833, 0.716) while its mean latency rose (93.78, 201.81, 307.32 ms), so rho became necessary but insufficient. At N=8 EWMA, a policy with a live latency signal, was worse than load-blind round-robin near the boundary (307.32 vs 170.48 ms) and above it (589.39 vs 376.80 ms), but better below it (39.43 vs 66.69 ms). Least-connections and Adaptive stayed at 27.98 and 29.88 ms near the boundary, so the failure is specific to EWMA rather than to load-aware policies as a class. Over 10 seeds round-robin beat EWMA every time (Cliff's Delta 1.000, CI on the difference [119.49, 134.11] ms). A later audit found the discovery was less blind than described: experiment 014a, run about 20 minutes earlier (JSON timestamps 19:13:13Z and 19:34:45Z), already showed round-robin at 88.81 ms vs EWMA at 137.03 ms at N=8 (about 290 requests), so 014f may have been shaped by that data; Stage14.md now says so. Ledger rows C17 and C19 are RETIRED.

    Numbers

    EWMA achieved rho, N=3/5/8go run ./cmd/experiment-014c
    0.915 / 0.833 / 0.716
    EWMA mean latency, N=3/5/8 (ms)go run ./cmd/experiment-014c
    93.78 / 201.81 / 307.32
    N=8 near boundary: RR / EWMA / LC / Adaptive (ms)go run ./cmd/experiment-014f
    170.48 / 307.32 / 27.98 / 29.88
    N=8 above boundary: RR / EWMA (ms)go run ./cmd/experiment-014f
    376.80 / 589.39
    N=8 below boundary: EWMA / RR (ms)go run ./cmd/experiment-014f
    39.43 / 66.69 (EWMA better)
    10-seed test, RR faster than EWMAgo run ./cmd/experiment-014i (identical on re-run)
    10/10; Cliff's Delta 1.000; 95% CI [119.49, 134.11] ms
    10-seed test, Adaptive faster than EWMAgo run ./cmd/experiment-014i (identical on re-run)
    10/10; Delta 1.000; 95% CI [260.21, 274.88] ms
    Earlier data point in 014A (N=8, Capacity=1, about 290 requests)experiments/014-scale-topology/results/014A-multi-target-capacity-boundary.json (timestamp 2026-09-06T19:13:13Z; 014F 19:34:45Z)
    RR 88.80867 ms, EWMA 137.02514 ms

    Evidence

    How this was checked: Read the four commit messages and the Stage14.md 014f section. Re-ran experiments 014c, 014f and 014i in a clone at HEAD; every quoted value matched the committed JSON and docs (one unquoted P2C wait-share field in 014F differed slightly between runs). Loaded the 014A JSON and found the two cited means. The disclosure commit d31d7c9 credits an independent audit, which was an AI-agent pass; the disclosure text in Stage14.md is audit-generated, though committed by the author.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit e497b79 (falsifier), 121b60a (rho test), 99031e2 (10-seed test), d31d7c9 (disclosure)

  15. Fuze

    What the CI regression gate for two-stage retrieval actually measures

    REJECTED

    Question

    Does ADR-006 Gate 1 (test_two_stage_baseline_regression.py against golden_baseline_v1.json) show that the two-stage retrieval path is equivalent in quality to the legacy orchestrator?

    Method

    Read the ADR, the test, the baseline generator and the baseline JSON. Ran test_golden_dataset_regression.py, test_two_stage_baseline_regression.py and test_shadow_evaluator.py (Python 3.11.9, Windows, pytest 8.4.2). Instantiated RecommendationPipeline() as the orchestrator does and called run() with a request.

    Result

    The gate passes with a delta of exactly 0.0000 because both sides run the same class. golden_baseline_v1.json was generated by generate_golden_baseline.py, which runs SmartEngine; ADR-006 calls it the baseline of the 'validated legacy orchestrator', but SmartEngine was added in the same commit and is not one of the orchestrator's engines. The regression test replaces the data layer with a mock returning a fixed pool (the expected candidates plus 10 noise items), so CandidateRetriever, HNSW and the RPC functions are never called. The set has 4 queries; query_frontend_01 has MRR 0.5, so the average of 0.875 clears the separate 0.85 floor in test_golden_dataset_regression.py by only 0.025. The shadow pipeline built in the orchestrator has no Unit of Work and returns no candidates, so shadow overlap would be 0. The gate therefore catches changes to the scoring code, not to retrieval.

    Numbers

    avg MRR, baseline and pipelinepytest -s tests/test_two_stage_baseline_regression.py, run 2026-09-28
    0.8750 and 0.8750 (delta 0.0000)
    avg NDCG@10, baseline and pipelinesame run; golden_baseline_v1.json aggregate avg_ndcg10 0.953737
    0.9537 and 0.9537 (delta 0.0000)
    per-query MRRbackend/tests/golden/golden_baseline_v1.json
    0.5, 1.0, 1.0, 1.0
    queries in the golden setgolden_baseline_v1.json aggregate.query_count
    4
    orchestrator shadow pipeline outputRecommendationPipeline().run(RecommendationRequest(user_id=1, title='React hooks', technologies='React'))
    uow None, results []
    tests runpytest golden and shadow tests, 2026-09-28
    4 passed in 50.42 s

    Evidence

    How this was checked: Ran the tests and the pipeline snippet listed above on this machine (no database needed); the Redis connection error printed during the run is the test environment falling back to no Redis. Attribution: Authored by Ujjwaljain16; cited commit 1f4deb2 has no Claude co-author trailer. This record is an after-the-fact analysis written for the portfolio (provenance: reconstructed), not a decision documented in the repository.

    Reconstructed from code and history · reasons are inferred · gate committed in 1f4deb2; re-run on 2026-09-28

  16. MiniDB

    Crash with 10,000 committed inserts and 500 uncommitted deletes recovers to exactly 10,000 rows

    ADOPTED

    Question

    After a simulated crash, does recovery keep committed data, roll back an uncommitted delete, and stay correct when run again?

    Method

    benchmarks/crash_recovery.ts: T1 inserts 10,000 rows through HeapFile with WAL context and commits; T2 deletes 500 of them and is never committed; the WAL is flushed but the buffer pool is not, then file handles are closed (pool size 100 frames). The database is reopened (recovery runs in open()), closed, reopened again, and a sequential scan counts rows. Separately, tests/integration/crash_matrix.test.ts has three cases (insert after WAL flush before page flush, delete before commit, COMMIT record flushed before the clean commit finished) and tests/unit/recovery/CrashRecovery.test.ts has an end-to-end case and a double-recovery case. The whole jest suite was run once.

    Result

    The benchmark printed 'Expected Rows: 10000 | Actual Rows: 10000' after two recoveries, as documented. The full suite passed: 26 suites, 133 tests. Limits: the crash is simulated by closing file handles, not by killing a process; the WAL is flushed first, so torn log writes are not tested; the table (about 50 pages) fits in the 100-frame pool, so no data page was written before the crash (0 page writes in an instrumented copy) and eviction of dirty pages is not exercised; no crash scenario has an index; a second crash right after recovery is not tested (a copy that did this ended with 9,500 rows); 'Scenario D: Recover() three times' in Architecture.md is not a test in crash_matrix.test.ts; runtime aborts are not covered.

    Numbers

    Rows after two recoveries (expected / actual)npx tsx benchmarks/crash_recovery.ts
    10000 / 10000
    Full test suitenpx jest --config jest.config.cjs
    26 suites, 133 tests passed
    Data page writes before the simulated crashinstrumented copy of crash_recovery.ts counting DiskManager.writePage calls
    0 (heap file about 50 pages, pool 100 frames)
    Rows when the database crashes again right after the first recoverymodified copy of the benchmark (crash instead of clean close after the first recovery)
    9500 (expected 10000)

    Evidence

    How this was checked: Ran the benchmark and the complete jest suite; read crash_matrix.test.ts, fuzz_sql.test.ts and the benchmark source. Attribution: All cited commits are authored by Ujjwaljain16 (git shortlog: 21 commits, one author). The README lists a two-person team, and git history cannot show which parts each person wrote, so the record should say "commits by Ujjwaljain16; two-person course project". No Co-Authored-By trailers exist in the history; whether AI tools were used cannot be determined from the repository.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit 7b7c4ce (benchmarks/crash_recovery.ts, tests/integration/crash_matrix.test.ts, CrashRecovery tests)

  17. TypeAheadX

    Modulo vs consistent hashing when a 4th node is added: re-run of rebalance_experiment.py

    ADOPTED

    Question

    How many of 100,000 cache keys change owner when a 4th node is added under hash(key) % N and under the 500-vnode ring?

    Method

    scripts/rebalance_experiment.py generates 100,000 unique random keys 'suggestion:<4-10 random lowercase letters>' (unseeded), maps them with MD5 modulo over 3 then 4 nodes, then with ConsistentHashRing(virtual_nodes=500) over 3 nodes then after add_node('redis-d'), and counts changes. It was run three times. The arc ownership of redis-d in the 4-node ring was computed separately (deterministic). This is a course project.

    Result

    The repo reports 74.88% (modulo) and 26.44% (ring) in README.md and Project_Report.md. Three re-runs gave modulo 74,963 / 74,741 / 74,800 keys moved (74.96%, 74.74%, 74.80%) and ring 26,391 / 26,626 / 26,200 (26.39%, 26.63%, 26.20%). The values differ per run because the keys are unseeded; the repo figures are within that spread. Ring movement is about the new node's share of the ring: redis-d owns 26.27% of the 128-bit space with 500 vnodes (versus the ideal 25%), and the 4-node ownership is 24.54 / 25.18 / 24.01 / 26.27. Earlier doc versions: 75.00% and 24.64% at 1000 vnodes (66845c0), 74.97% and 26.18% at 500 (b116d06). Caveat: uniformly random keys, no traffic skew, and no runtime code path adds a node.

    Numbers

    Modulo keys moved, 3 re-runspython scripts/rebalance_experiment.py
    74.96%, 74.74%, 74.80%
    Ring (500 vnodes) keys moved, 3 re-runspython scripts/rebalance_experiment.py
    26.39%, 26.63%, 26.20%
    Repo-stated valuesREADME.md lines 46-47; Project_Report.md lines 323 (ring only) and 382-383
    74.88% modulo, 26.44% ring
    redis-d ring ownership, 4 nodes, 500 vnodesarc-ownership script over ConsistentHashRing (deterministic; script not in the repo)
    26.27%

    Evidence

    How this was checked: Ran the script three times (Python 3.11.9) and my own ownership script; compared with README.md, Project_Report.md and the docs at 66845c0 and b116d06. Attribution: Course project: the README does not say so, but docs/phase4-completion.md@66845c0 contains 'Viva Talking Points' and the phase briefs 3.md, 4.md and 5.md (committed, then deleted in the next phase) are written as instructions to a student ('most students will completely mess up'), so the work followed a written brief. All 15 commits are by Ujjwaljain16, none has a Co-Authored-By trailer, and 14 of 15 fall on one day (2026-06-10, 00:23 to 21:22 +05:30; the last is 2026-06-22).

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · script added in commit 66845c0; my re-run 2026-09-28

  18. NevUpAI

    k6 write test: first run used 1 VU (p95 5.38 ms); rerun with 100 VUs gave p95 27.84 ms

    ADOPTED

    Question

    Does POST /trades meet the p95 < 150 ms target under load, and what does the reported 27.84 ms represent?

    Method

    nevup-backend/k6/trade-write-smoke.js against the local docker-compose stack. Run 1 (report at 9c441bb): constant-arrival-rate 200/s for 60 s, one fixed user, preAllocatedVUs 100, maxVUs 500. Run 2 (9d73342): constant-vus 100 for 60 s, each VU a distinct user/session with its own JWT, sleep(0.5) per iteration (target 200 req/s). Each request writes a new random tradeId, half open and half closed. Thresholds p(95)<150 and failure rate <1%. Not re-run: k6 is not installed and the Docker daemon was not running on the reviewing machine.

    Result

    Run 1 hit the target only nominally: k6 needed at most 1 VU because responses were fast, so it did not exercise concurrency (p95 5.38 ms). The script was changed to 100 real concurrent users and the second run gave p95 27.84 ms, avg 18.16 ms, median 6.41 ms, 0% failed and 11,787/11,787 checks passing. Achieved rate was 188.86 req/s, not the '~200 requests/sec' the README states. Both runs contain one ~16.3 s maximum latency that neither the README nor DECISIONS.md explains. Caveats: write path only, one machine, a closed-loop model with 0.5 s think time, and 11,787 requests over 62.4 s.

    Numbers

    Run 2 http_req_duration p(95)nevup-backend/results.json
    27.84221919999998 ms
    Run 2 avg / median / p90 / maxnevup-backend/results.json
    18.156 / 6.415 / 15.462 / 16308.302 ms
    Run 2 requests, rate, failuresnevup-backend/results.json
    11787 requests, 188.86/s, http_req_failed 0 of 11787
    Run 2 duration / VUsnevup-backend/results.json (state.testRunDurationMs)
    62,411 ms test run; vus_max 100
    Run 1 p95 / avg / max, VUsdocs/k6_report.html@9c441bb
    5.38 / 6.82 / 16266.52 ms; VUs min 0 max 1; 12,001 requests at 194.66/s
    README claimnevup-backend/README.md@7d225fa
    'Sustained ~200 requests/sec', p95 27.84 ms

    Evidence

    • commit9d73342 ↗Switches executor from constant-arrival-rate (one user) to 100 constant VUs with per-VU users and a 0.5 s sleep; adds the second report.
    • commit7d225fa ↗README updated to 27ms p95 / ~200 req/s.
    • benchmarkresults.json@3d3b274 ↗k6 handleSummary output for run 2.
    • benchmarkk6_report.html@9c441bb ↗HTML report for run 1 (Virtual Users min 0 max 1).

    How this was checked: Parsed results.json with python (values above); extracted the text of both HTML reports from git; diffed the k6 script across 9c441bb and 9d73342. Not re-run (no k6, Docker daemon down).

    Recorded at the time (a document or commit message states it) · reasons are inferred · commit 9d73342 (script change and new results.json / docs/k6_report.html)

    hackathon (single author)Link to this record
  19. SSE-Observatory

    Does the SharedWorker really hold one upstream connection for several tabs?

    CONFIRMED

    Question

    When several tabs of the same browser context connect to the same URL, does the application open one upstream SSE request, and does a different token open a separate one?

    Method

    A counting SSE server on 127.0.0.1:4100 recorded total and currently open requests. The app was driven in Chromium 145.0.7632.6 through Playwright 1.58.2, with Vite serving the UI on :3001 and the repository's server.js as the /api backend, and all tabs in one browser context. Tabs were opened one at a time; each filled the URL input and pressed connect, and the server counters were read 1.5 s later. The repository's e2e suite does not perform this measurement. The measurement was made on 2026-09-28 against commit 0e8c14b; the code under test dates from f899859.

    Result

    1 tab: 1 total / 1 open upstream request. 2 tabs on the same URL: still 1 / 1. 3 tabs: still 1 / 1; each of the three tabs showed 9 event rows. A fourth tab with the same URL and a different token: 2 total / 2 open. After the browser closed: 2 total / 0 open (the counting server saw both connections close). So multiplexing works through the full path browser -> Vite -> Express -> upstream, and the key includes the token.

    Numbers

    upstream requests with 1, 2 and 3 tabs on the same url+tokenMeasurement run, 2026-09-28 (probe script not published)
    1, 1, 1 (open connections 1, 1, 1)
    upstream requests after a 4th tab with a different tokenSame run
    2 total, 2 open
    open upstream connections after browser closedSame run
    0

    Evidence

    How this was checked: Ran the script against HEAD with the servers started from this session; stopped only those PIDs afterwards.

    Reconstructed from code and history · reasons are inferred · commit f899859 (code under test); measurement re-run 2026-09-28 on HEAD 0e8c14b

  20. SSE-Observatory

    Interceptor sandbox: the 100 ms kill works, but shadowing globals is not a security boundary

    PARTIAL

    Question

    Does terminating a worker really stop a synchronous infinite loop within the 100 ms budget, and does the shadowed-globals wrapper prevent network access?

    Method

    The repository's tests cannot answer this: setup.ts replaces Worker with a MockWorker whose terminate() does nothing, the loop test is it.skip and the e2e test is test.skip. The wrapper from interceptorWorker.ts (same list of shadowed names) was therefore copied into a Blob Worker run inside the app page in Chromium 145.0.7632.6, mirroring the pool's setTimeout + terminate() logic, and six interceptor bodies were run. The loop cases were repeated in Node 22.19 worker_threads, together with a main-thread Promise.race control. This is a replica of the wrapper, not the built app bundle.

    Result

    Chromium: `while(true){}` was killed at about 105-107 ms; a normal interceptor returned in about 3.5-4 ms including worker creation; a direct call to fetch failed because the name is shadowed. Other ways of reaching the same global APIs from inside the wrapper still worked, so shadowing by name is not an isolation boundary. Node: the loop was killed in 103-115 ms across runs and a memory-abuse loop in 109-115 ms; a main-thread Promise.race against a 1500 ms synchronous loop resolved 'done' after 1500 ms, which confirms the header comment that Promise.race cannot interrupt synchronous code. The termination path works; the wrapper guards against accidents rather than acting as a security boundary, and untrusted interceptor code should not be treated as contained.

    Numbers

    sync infinite loop, Chromium 145, 100 ms timer + terminate()Measurement runs, 2026-09-28 (replica of the wrapper; probe script not published)
    killed at about 105-107 ms
    normal interceptor round trip incl. worker creation, ChromiumSame runs
    about 3.5-4 ms
    sync loop / memory-abuse loop killed, Node 22.19 worker_threadsSame runs
    103-115 ms / 109-115 ms across runs
    main-thread Promise.race against 1500 ms sync loopSame runs
    resolved 'done' after 1500 ms

    Evidence

    • commit61d2488 ↗Design and header comment claiming fetch is shadowed and terminate() kills loops.
    • filesetup.ts@0e8c14b ↗MockWorker with empty terminate().

    How this was checked: Ran the replica in Chromium and in Node on 2026-09-28; the shadow list matches interceptorWorker.ts, but the built app bundle was not tested.

    Reconstructed from code and history · reasons are inferred · commit 61d2488 (code under test); measurement run 2026-09-28

  21. SpentSmart

    The documented 45MB-to-15MB size cut is not visible: release APKs grew 9.3%

    REJECTED

    Question

    docs/CODEBASE_ANALYSIS.md claims a bundle went from 45MB to about 15MB (called a 60% reduction) after victory-native/Skia and react-native-chart-kit were replaced by a custom SVG chart. Do the shipped release APKs show any reduction?

    Method

    Read the asset sizes of the three GitHub releases with `gh api repos/Ujjwaljain16/SpentSmart/releases`; fetched the release tags (they point at the pre-rewrite commit lineage) and diffed package.json at v1.0.0, v2.0.0 and v2.01; searched every commit's package.json and the lockfiles for victory-native and react-native-skia. APKs were not downloaded or unpacked.

    Result

    No reduction is visible: the v2.0.0 and v2.01 APKs are about 9.3% larger than v1.0.0. The comparison is rough: the docs say 'bundle', the assets are APKs, v2 adds expo-notifications and expo-updates, and the v1 asset name is an EAS-style file name while v2 names differ. victory-native and Skia are never in package.json; they appear only in pnpm-lock.yaml (added in cc86919, removed in be95873, both 2026-01-03, and still in the lockfile at the v1.0.0 tag). react-native-chart-kit stays in package.json at every tag; only the pie chart component that used it was deleted. Nothing in the repository measures 45MB or 15MB, and 45 to 15 is a 67% reduction, not the 60% stated.

    Numbers

    v1.0.0 APK (application-58841ca1-3e46-4f4f-a829-e8d81073b361.apk)gh api repos/Ujjwaljain16/SpentSmart/releases (asset size)
    115,963,061 bytes
    v2.0.0 APK (SpentSmartV2.apk)same
    126,787,065 bytes (+9.33% vs v1.0.0)
    v2.01 APK (SpentSmartV2.01Release.apk)same
    126,765,689 bytes (+9.32% vs v1.0.0)
    Documented before/afterdocs/CODEBASE_ANALYSIS.md@e3c438b and HEAD
    45MB -> ~15MB, '60% bundle size reduction'
    Commits where victory-native or react-native-skia appear in package.jsongit log --all -S<name> -- package.json
    0
    Commits that add / remove them in pnpm-lock.yamlgit log --all -S victory -- pnpm-lock.yaml
    added in cc86919, removed in be95873 (both 2026-01-03); still present in the lockfile at the v1.0.0 tag

    Release APK size against the documented reduction

    The docs claim the app went from 45 MB to about 15 MB. The published release files do not show a reduction.

    • v1.0.0
      115.96 MB
    • v2.0.0
      126.79 MB · +9.33% vs v1.0.0
    • v2.01
      126.77 MB · +9.32% vs v1.0.0
    • Documented size after the change: about 15 MB
    • Documented size before: 45 MB
    Source: GitHub releases API, asset sizes in bytes divided by 1,000,000

    Evidence

    • commite587e54 ↗The v1.0.0 tag commit: package.json has react-native-chart-kit, expo-dev-client and no victory/skia.
    • commit3c5546a ↗The v2.0.0 tag commit: package.json adds expo-notifications and expo-updates and still has react-native-chart-kit.
    • commite3c438b ↗First commit containing the 45MB / 15MB / 60% text in docs/CODEBASE_ANALYSIS.md (the twin e587e54 on the tag lineage has the same text).
    • commitcc86919 ↗Adds pnpm-lock.yaml entries for victory-native 41.20.2 and @shopify/react-native-skia 2.4.14 although package.json does not list them.
    • commitbe95873 ↗Removes the victory-native and react-native-skia entries from pnpm-lock.yaml.
    • releaseGitHub releases ↗Asset names and sizes for v1.0.0, v2.0.0 and v2.01.

    How this was checked: Ran the gh api query (sizes above); ran `git fetch origin 'refs/tags/*:refs/tags/*'` and `git show <tag>:package.json`; python check that node_modules/react-native-chart-kit in package-lock.json has dependencies lodash, paths-js, point-in-polygon. Attribution: The 45MB to 15MB claim comes from docs/CODEBASE_ANALYSIS.md, which appears to be AI-generated documentation; no build measurement backs it.

    Reconstructed from code and history · reasons are inferred · GitHub release v2.01 published 2026-02-08T00:39:38Z (v2.0.0 2026-02-07T23:04:59Z, v1.0.0 2026-01-03T12:57:12Z)

  22. Vitest

    Why typecheck printed the whole tsc help text: tsc was run without -p and found no tsconfig

    ADOPTED

    Question

    In issue #8981, running 'vitest run --typecheck' in a project with no tsconfig.json produced 'Typecheck Error' followed by the whole tsc help text. What causes tsc to print help, and how should Vitest report it?

    Method

    Read Typechecker.spawn() in packages/vitest/src/typecheck/typechecker.ts: the argument list is --noEmit --pretty false --incremental --tsBuildInfoFile <path>, and '-p <tsconfig>' is appended only if typecheck.tsconfig is set (same code at v4.0.8 and v4.0.17). Vitest also captured only stdout. The PR adds a check in prepareResults() that throws a descriptive error if the tsc output contains 'The TypeScript Compiler - Version' or 'COMMON COMMANDS', and also captures stderr. For the review-required test, the author replaced a mocked unit test with runInlineTests and a fake checker executable that prints tsc's help header. The underlying question was then re-run with the real compiler (see result).

    Result

    A re-run with typescript 5.9.3 (the version in the issue) in an empty directory with no tsconfig in any parent: the arguments Vitest uses, without -p, print "tsc: The TypeScript Compiler - Version 5.9.3 ... COMMON COMMANDS" (141 lines) and exit with code 1. With -p ./nope.json tsc prints TS5058; with -p to a tsconfig containing {} it prints TS18003. So help text appears when tsc runs without -p and finds no tsconfig in the directory or its parents. Two contributor statements do not match this: the issue comment says tsc "exits with code 0" (1 here), and the 2025-12-17 review reply says Vitest always calls tsc with -p (the source adds -p only if typecheck.tsconfig is set). The merged test therefore uses a stub checker that prints the help header, which the maintainer liked. Review took four CHANGES_REQUESTED rounds and 20 commits (a stray snapshot file, a fully mocked test rejected under AGENTS.md, then createFile and static imports). The issue proposed checking that the tsconfig exists; the PR instead detects two marker strings, a choice the PR text does not explain.

    Numbers

    tsc 5.9.3, Vitest arguments, no -p, no tsconfigRe-run: node node_modules/typescript/bin/tsc --noEmit --pretty false --incremental --tsBuildInfoFile ./x.tsbuildinfo in an empty temp directory
    prints "The TypeScript Compiler - Version 5.9.3" and COMMON COMMANDS (141 lines); exit code 1
    tsc with -p ./nope.jsonSame re-run
    error TS5058: The specified path does not exist
    tsc with -p tsconfig.json containing {}Same re-run
    error TS18003: No inputs were found in config file
    PR sizegh pr view 9214 --repo vitest-dev/vitest --json additions,deletions,commits,changedFiles
    +103 / -2 lines; 20 commits; 3 files

    Evidence

    How this was checked: Read the PR, issue, all inline review comments and review states with gh; fetched typechecker.ts at v4.0.8 and v4.0.17 and read spawn(); installed typescript@5.9.3 into a temporary directory and ran the three tsc invocations listed above; checked that typecheck-error.test.ts and the help-text check exist on main. Did not run the Vitest test.

    Recorded at the time (a document or commit message states it) · reasons are inferred · PR #9214 merged 2025-12-23T17:15:53Z (squash commit 7b10ab4)

  23. Appwrite

    MFA recovery codes rejected in 1.8.0: a lower-cased type constant never matched

    ADOPTED

    Question

    Why did PUT /v1/account/mfa/challenge answer 'Invalid token passed in the request' for valid recovery codes on self-hosted Appwrite 1.8.0 (issue #10740)?

    Method

    The issue's server log named app/controllers/api/account.php line 4972 (the USER_INVALID_TOKEN throw in the reporter's build). The contributor read the handler and compared the value stored on the challenge with the value used in the comparison. POST /account/mfa/challenge stores 'type' => $factor, and the factor whitelist accepts Type::RECOVERY_CODE unmodified. Type::RECOVERY_CODE is 'recoveryCode' (src/Appwrite/Auth/MFA/Type.php at tag 1.8.0). The verification handler compared the stored type to \strtolower(Type::RECOVERY_CODE), i.e. 'recoverycode', both in an inner === check and as a key of a PHP match() expression. match() compares strictly, so the recovery branch could never be taken and the request fell through to default => false. The fix removed both strtolower() calls, and a new E2E test testMFARecoveryCodeChallenge exercises the whole path.

    Result

    Root cause confirmed by reading the 1.8.0 tag: the stored type and the compared value differed only in case. The E2E test creates recovery codes (201) and a 'recoveryCode' challenge (201), verifies a valid code (200, factors contains 'recoveryCode'), then checks that reuse of the code (401) and an invalid code (401) are rejected. The contributor pasted a run of an earlier test version; the merged test has 12 assertions. The PR does not show the test failing before the fix, so that half of the regression claim rests on the code reading. At the maintainer's request the PR also changed POST /account/mfa/recovery-codes and POST /account/mfa/challenge to return 201 instead of 200. Review needed four CHANGES_REQUESTED reviews by stnguyen90; the first said tests were failing ("Did you test it yourself?"), and a ~75-line session fallback in an earlier test version was criticised by CodeRabbit and replaced by reuse of the ordinary session. Release: not in 1.8.1 (2025-12-23), whose Challenges/Update.php still has \strtolower(Type::RECOVERY_CODE); a user reported the failure on 1.8.1 on 2026-01-01. First tag with the fix: 1.9.0-rc.1 (2026-03-24); first stable: 1.9.0 (2026-04-01).

    Numbers

    Author-pasted test run (earlier revision of the test)PR #10925 comment by Ujjwaljain16, 2025-12-10 (not re-run)
    OK (1 test, 28 assertions), 2274 ms
    Assertions in the merged testgh pr diff 10925, testMFARecoveryCodeChallenge
    12
    Change sizegh pr view 10925 --repo appwrite/appwrite --json additions,deletions
    +87 / -4 lines, 2 files
    Fix contained in 1.8.1gh api repos/appwrite/appwrite/compare/35fe622...1.8.1 and file at tag 1.8.1
    no (compare status: diverged, 801 ahead / 29 behind; Challenges/Update.php at 1.8.1 still uses \strtolower)
    First tag containing the fixgh api repos/appwrite/appwrite/compare/35fe622...1.9.0-rc.1 (ahead) and .../1.9.0 (ahead)
    1.9.0-rc.1 (2026-03-24); first stable release 1.9.0 (2026-04-01)

    Evidence

    How this was checked: Ran gh pr view/diff/comments and gh api for PR 10925 and issue 10740; downloaded app/controllers/api/account.php and Type.php at tag 1.8.0 through the GitHub contents API and read the challenge-creation and verification handlers; ran gh api compare between 35fe622 and tags 1.8.1 and 1.9.0; read release dates for 1.8.1 and 1.9.0. Did not run the Appwrite E2E suite. Commit existence was checked through the GitHub API (repos/appwrite/appwrite/commits/<sha>), not with a local clone, because the repository is large.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · PR #10925 merged 2025-12-11T14:28:06Z (merge commit 35fe622)

  24. migrateDB

    A migration process killed mid-run leaves the lock row and blocks every later run

    REJECTED

    Question

    What happens to the single-row lock if a migrating process dies before it can release the lock?

    Method

    Installed @ujjwaljain16/migratedb 1.0.1 and better-sqlite3 in an empty project (Node 22.19.0). A child process ran migrate() on a SQLite file with a migration that runs a long recursive query. The parent sent SIGKILL after 1.5 s, read migratedb_lock, then called migrate() twice on the same file. The run was repeated in a second fresh install with the same outcome. The 1.0.0 tarball contains a byte-identical MigrationLock.js.

    Result

    The lock row remained after the kill. Both later attempts failed at once (about 1 to 2 ms, no waiting or retry) with LockError 'Failed to acquire migration lock: UNIQUE constraint failed: migratedb_lock.id'. A lock row dated 30 days earlier gave the same error, because the code never reads acquired_at. Recovery is to delete the row by hand or to run with enableLocking set to false, which skips locking (checked on SQLite). The text differs by engine: from reading the code, Postgres would report 'Migration is already running'; that path was not run.

    Numbers

    lock rows after SIGKILLreproduction script on SQLite, Node 22.19.0, package 1.0.1
    1 row: id=1, acquired_at set
    time to fail on next runsame run, repeated in a second fresh install
    about 1 to 2 ms per attempt (no wait, no retry)
    lock age checkdist/cjs/core/MigrationLock.js in @ujjwaljain16/migratedb 1.0.1
    none (checkLockAge returns 0; the lockTimeout constructor argument is unused)

    Evidence

    How this was checked: Ran crash.js myself (SQLite only). Postgres and MySQL not run; the lock SQL is identical for all engines, so the behaviour follows from the code. Attribution: The reproduction and write-up were produced during a 2026-09-28 audit (AI-assisted), not by the author at the time; the code examined is the author's published package.

    Reconstructed from code and history · reasons are inferred · version examined: 1.0.1 (npm publish 2025-12-03); the same lock code is in 1.0.0 (2025-11-22); reproduced 2026-09-28