CORTEX
DECISIONS

Decisions

36 decisions that shaped my projects and my open-source work, chosen for their trade-offs rather than their number. Each one points at the commits and files it came from and says how it was checked; the reversals and failures are kept in. 17 are marked as the strongest. For measured results, see investigations.

Showing 36 of 36

  1. AgentBrakeFEATURED

    Anything the proxy cannot parse is now blocked instead of forwarded, and a bad policy file stops startup

    ADOPTED

    Context

    In 0fc99c8 the interceptor split each stdin chunk on newlines, ran JSON.parse on each piece, and in the catch branch wrote the piece to the server unchanged. A tool call split across two chunks therefore reached the server as two unchecked fragments, and any message that JavaScript's JSON.parse rejects but another JSON parser accepts (for example a bare NaN) skipped every policy. The loader had the same shape: a config file that failed validation was replaced by a hand-written default with no policies.

    Decision

    Framing moved to a byte-level line buffer used in both directions (src/proxy/framing.ts), with a size cap and UTF-8 validation. On the client path, unparseable input, JSON-RPC batches, non-object messages, malformed tools/call params, unknown policy actions and a policy that throws are all answered with a JSON-RPC error and not forwarded; forwarded messages are re-serialised so the server sees what the policies saw. An invalid or missing config exits with code 2 unless AGENT_BRAKE_ALLOW_INVALID_CONFIG=1 is set.

    Alternatives considered

    • Keep forwarding lines that cannot be parsed, and fix only the chunk splitting. The parser difference between the proxy and the server would still let a message through unchecked; blocking is the only safe answer when the proxy cannot classify a message.
    • Keep the silent fallback to default policies for a bad config. A typo in the policy file would disable enforcement without any sign. The fallback remains available only as an explicit opt-in.

    What happened

    tests/proxy.test.ts splits a tools/call at every byte boundary and asserts the policy still applies, and covers multiple messages per chunk, CRLF, invalid UTF-8, batches, oversize lines and duplicate keys. The test count went from 22 to 85. The same commit also fixed max_tool_calls counting one call too few, wired the circuit breaker to real tool errors, and made a killed proxy terminate its child. Remaining limits: the proxy covers stdio only, and regular-expression argument rules can still be bypassed.

    Evidence

    • commit373adcd ↗Adds framing, fail-closed handling and config refusal, with the regression tests.
    • commit0fc99c8 ↗The last commit before the change: the catch branch that forwards unparseable lines.
    • filetests/proxy.test.ts ↗Byte-boundary split test and the other framing cases.

    How this was checked: Read src/proxy/interceptor.ts at 0fc99c8 and at 373adcd; ran npm test (85 passing), tsc --noEmit and the build after the change.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit 373adcd

  2. Apache SupersetFEATURED

    Remove deprecated permissions with two explicit-list migrations, not a sweep

    ADOPTED

    Context

    Issue #33272 (opened 2025-04-29) showed that installs upgraded from 1.5.x keep many permissions that fresh installs lack (for example `can select star on Superset`): 286-291 permissions after upgrading 1.5.2 to 5.0.0-RC2, against 160 on a fresh install. `clean_perms` only deletes PermissionView rows with NULL foreign keys. On the issue, maintainer rusackas noted that the hard part is not deleting a permission a custom role still needs. The per-object database_access/datasource_access/schema_access/catalog_access permissions share names across many view menus, so a name-based sweep is unsafe.

    Decision

    Two Alembic migrations that name exact (view_menu, permission) pairs on the `Superset` view menu. One deletes permissions with no proportionate successor (`PVM_LIST`, 10 entries in the merged file) through a new `delete_pvms()` helper in superset/migrations/shared/security_converge.py, which unassigns each PVM from roles and reuses `_delete_old_permissions()` for orphan-safe removal. The other maps old permissions to verified live successors (`PVM_MAP`, 29 entries) with the existing `migrate_roles()`, so roles keep access. A rename is used only when the successor is proportionate: four dead permissions whose only successor was a write-level grant (`can_testconn`, `can_sqllab_viz`, `can_import_dashboards`, `can_add_slices`) are deleted instead. The one write-level rename left is `can_copy_dash` to `Dashboard.can_write`, argued in the migration as the same action. Downgrades are documented no-ops.

    Alternatives considered

    • Delete everything not in an allow-list of current permissions, or sweep by permission name. Stated in the PR: cannot tell dead from 'not used by a built-in role', and names like database_access are reused across many view menus; explicit (view, permission) pairs make the object-level permissions unreachable.
    • Extend `migrate_roles()` to handle deletion with no successor. The author confirmed by reading it that an empty replacement tuple is silently never processed (gabotorresruiz reproduced this), so a small `delete_pvms()` helper was added and `migrate_roles`, `clean_perms` and FAB's `security_cleanup` were left untouched.
    • Rename all deprecated permissions to the successor named in `@deprecated(new_target=...)` (the author's own intermediate state: delete list reduced to 6, rename list up to 34). gabotorresruiz showed a role holding only the dead `can_testconn` would come out of the migration holding `Database.can_write` (create/edit/delete connections). After that thread `can_testconn`, `can_sqllab_viz`, `can_import_dashboards` and `can_add_slices` moved back to deletion.

    What happened

    Review changed the data materially. rebenitez1802 found that the delete list named `can_test_conn` where the real permission is `can_testconn`, and that the rename `can_get_or_create_table` targeted a name that never existed (the real one is `can_sqllab_table_viz`); both silently did nothing. The author then re-checked all 25 deletion candidates and reported that 19 had a live successor. gabotorresruiz showed a role holding only the dead `can_testconn` would gain `Database.can_write`, so four write-level renames went back to deletion. He also showed the policy was not pinned (re-adding that rename left 161 tests green); the author added `test_can_copy_dash_is_the_only_write_level_rename`. Three no-op Alembic merge migrations (2026-09-16, 09-17, 09-23) resolved head splits from master. UPDATING.md documents the split. The lists cover only individually verified permissions. +1152/-2 lines, 11 files. Merged 2026-09-23; not in a tagged release as of 2026-09-28 (latest tag 6.1.0, 2026-05-13).

    Evidence

    How this was checked: Read the PR body, all review, inline and issue comments, the diff (11 files), and the merged rename migration at 350dd0d, in which 29 keys are counted in PVM_MAP; PVM_LIST has 10 entries in the delete migration diff. The 25/19/6/34 figures are the author's statements in the 2026-09-15 comment; sqlite/Postgres upgrade runs are the reviewers' reported results and were not re-run.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · PR #44142 merged (merge commit 350dd0d)

  3. BHTTP-1FEATURED

    After sending ERROR, half-close and drain input for at most 1 s and 64 KiB

    ADOPTED

    Context

    If a program closes a TCP connection while unread data is pending, the operating system may send a reset, and the reset can reach the peer before the ERROR frame it just sent, destroying it. All 41 commits in the public repository are timestamped within 74 minutes (2026-09-22 00:10 to 01:24 +05:30) and the first ten share the times 00:10:25 to 00:10:26, so the history does not show how the design evolved over time; the reasoning below comes from a code comment, the specification and docs/ARCHITECTURE.md.

    Decision

    fail() writes the ERROR frame, then drain() calls CloseWrite (TCP FIN), sets a 1 s read deadline and copies at most 64 KiB (io.LimitReader) to io.Discard before the connection is closed. SPEC section 9 makes the FIN mandatory and the bounded drain a SHOULD with those reference values. c68262e added that the time limit runs from the moment the ERROR was sent, and that a peer that just sent an ERROR MUST treat a reset like EOF.

    Alternatives considered

    • Close immediately after writing ERROR. The comment above drainTime in server.go (362fc0a) and docs/ARCHITECTURE.md give the reason: closing with unread data pending can make the OS reset the connection, and the reset can destroy the ERROR before the peer reads it.
    • Drain without a bound. The same comment and docs/ARCHITECTURE.md: without both a time limit and a byte limit a hostile peer could keep the connection open forever by never stopping.

    What happened

    TestDrainIsBoundedInSizeAndTime (an endless sender and a silent peer, over net.Pipe, which has no CloseWrite, so the FIN step is not exercised there) and TestServerStopsTalkingToAPeerThatKeepsSending (real TCP) pin the bounds; go test -count=1 ./internal/server passes. The cost is that a connection that hit a fault stays open for up to 1 s.

    Evidence

    How this was checked: Read server.go, SPEC.md section 9 and the tests; go test -count=1 ./internal/server passes. Attribution: All 41 commits are authored by Ujjwaljain16 and none carries a Co-Authored-By trailer. Whether the specification and docs were AI-assisted cannot be determined from the repository; the owner should confirm before publishing.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit 362fc0a (server with drain); c68262e (spec wording)

  4. Apache SupersetFEATURED

    Treat a catalog qualifier as a self-reference only when the database is statically configured

    ADOPTED

    Context

    For engines with `supports_catalog = False` (such as MSSQL) Superset stores permissions and datasets without a catalog, but T-SQL and sqlglot still parse `abcm.dbo.temp` with `abcm` as a catalog. `raise_for_access` then looked up a permission like `[db].[abcm].[dbo]` that cannot exist and denied a query that only restated the connection's own database (issue #31406). Any change sits in an authorization path, so a wrong equality would grant access instead of denying it.

    Decision

    `raise_for_access` resolves the connection's statically configured database once per check and sets the parsed catalog to None only when it matches that name exactly on a non-catalog engine; the normalised catalog is used for both the catalog_perm/schema_perm check and the `SqlaTable.query_datasources_by_name` lookup. The database name comes from a new engine-spec hook, `get_catalog_from_engine_params` (default in db_engine_specs/base.py reads the URL database; the MSSQL override in mssql.py reads, in the order the real connection string is built: `Database=` inside `odbc_connect`, a `database` URL query parameter, the URL path database, then `connect_args["database"]`). Anything not statically determinable returns None and keeps the existing denial. The exact host/DSN-only URI from the original report stays denied and the PR was reworded to 'partially addresses #31406'.

    Alternatives considered

    • Compare the parsed catalog only to `database.url_object.database` (first revision of the PR). rebenitez1802's CHANGES_REQUESTED review showed that for `mssql+pyodbc://SuperSet:pw@abcm` SQLAlchemy yields host='abcm' and database=None, so the fix did not fire for the reporter's own connection, and the tests hid it by setting `url_object.database` on a mock.
    • Fall back to `database.database_name`. The author's comment says it is the Superset connection's display name and can be renamed independently of the real database, so using it as an authorization identity could turn a denied cross-database reference into an allowed one.
    • Resolve the login's default database by querying SQL Server. Adds live I/O to the authorization path; the author scoped it out and gabotorresruiz agreed live resolution has no place there.
    • Case-insensitive comparison of the qualifier and the database name (raised by a Bito bot review suggestion). Case sensitivity depends on the engine's configurable collation; a false negative falls back to today's denial, while a wrong case-fold would merge two distinct databases. The author kept the comparison exact and gabotorresruiz agreed; rebenitez1802 called the remaining false negative defensible if the PR was scoped honestly.

    What happened

    The path-form, `database` query-parameter, odbc_connect and connect_args forms now authorize self-referential qualifiers; a genuinely different database, or a wrong-case name, still fails. Review continued after approval: gabotorresruiz showed that a `?database=` URL query parameter overrides the path database in the connection string SQLAlchemy builds, so the merged hook checks `odbc_connect` first, then the query parameter, then the path, then `connect_args`. `Initial Catalog=` is deliberately not recognised: it is an OLEDB/ADO.NET keyword and the author reports that ODBC Driver 18 ignores it. Reviewers report that the new tests fail on the base commit while the denial guards pass (rebenitez1802: the two headline guards; gabotorresruiz: seven raise_for_access tests at the final head). Cost: the literal DSN-only URI from #31406 is still denied. sadpandajoe closed the issue on 2026-09-17, after the merge, inviting a new ticket if problems remain. +1317/-3 lines, 6 files. Not in a tagged release as of 2026-09-28.

    Evidence

    • pull requestapache/superset/pull/43974 ↗Description, author's design comments (2026-09-08, 09-09, 09-14, 09-16), a CHANGES_REQUESTED review by rebenitez1802 on 2026-09-08 (fix is a no-op for the reporter's URI; tests mask it), approvals by rebenitez1802 (09-09) and gabotorresruiz (09-11, 09-16). Merged by rusackas; merge commit 187e7d4d804555f2f7cd76ad49889a99493b91a7.
    • issueapache/superset/issues/31406 ↗Original report "Permission denied for sql access between databases" (opened 2024-12-11). The PR says it only partially addresses it; sadpandajoe closed it on 2026-09-17 citing the merge.
    • commit187e7d4 ↗Squash merge `fix(security): handle self-referential catalog qualifiers (#43974)`, 2026-09-16 (GitHub API).
    • filemssql.py@187e7d4 ↗`get_catalog_from_engine_params` and `_parse_odbc_connect_database` with the docstring stating the precedence rules and the fail-closed None result (read in the PR diff).

    How this was checked: Read the full PR body, both review threads, every human comment and inline comment, and the diff (gh pr diff 43974; files base.py, mssql.py, security/manager.py, three unit-test files). Test-failure claims are the reviewers' reported runs; the tests were not re-run locally.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · PR #43974 merged (merge commit 187e7d4)

  5. FuzeFEATURED

    Build the real image and run the real stack after mocks and SQLite missed critical defects

    ADOPTED

    Context

    By 15 September 2026 the backend had a 198-test suite (per the commit messages). The author's gaps.md records that all verification had run against in-memory SQLite (section 9) and a shared local virtualenv (section 11), and that Postgres-specific behaviour and the Docker build had only been reasoned about, not run. A production-readiness pass then built and ran the stack for real.

    Decision

    Three changes. (1) .github/workflows/docker-build.yml builds the production image on relevant pushes and pull requests, checks that /app/alembic.ini exists in the image, and imports the app inside it. (2) A disposable local rig: Postgres with pgvector and Redis in Docker, the real gunicorn/gevent command in a Linux container, seeded users with pre-minted JWTs, and scripts/locustfile.py wired to the heavy endpoints (the rig is not committed; the seed script and Locust file are). (3) scripts/e2e_smoke_test.py, a 23-check journey covering health, register, login, CORS, bookmark save through real RQ background processing, search, dashboard, recommendations, SSE ticket and auth edge cases. Per a74c00c it was run against hosted Postgres (Supabase) and Redis (Upstash).

    Alternatives considered

    • Keep relying on SQLite unit tests and a local virtualenv. gaps.md sections 9 and 11 say this was not a clean stand-in for the production image or for Postgres, and the defects listed below were not visible to it.

    What happened

    Defects the 198-test suite did not catch: (1) alembic.ini was never copied into the image, so the migration step in start.sh (added 2026-07-27) had probably failed on every boot behind its '|| echo Warning' fallback; this is the author's inference and the Space was not inspected. (2) requirements.txt did not resolve: a real build failed after 428 s with resolution-too-deep (unbounded mcp via scrapling[all]), and playwright==1.53.0 conflicted with scrapling's exact 1.56.0; fixed with mcp==1.24.0 and playwright==1.56.0. (3) alembic/env.py's statement timeout never applied, and CREATE INDEX CONCURRENTLY failed on a fresh Postgres. (4) The analyze_content() call after embedding raised TypeError on every worker job while RQ reported success. The workflow passed on all three commits. Stated limits: two gunicorn workers not load-tested; start.sh still soft-fails migrations; the first production migration run waits for a deploy, which is blocked while the Space is offline.

    Evidence

    • commit764e2c6 ↗Message lists the deploy blockers, docker-build.yml, and the pin fixes; 48 files.
    • commita74c00c ↗Message: content analysis silently broken for every bookmark, found by an end-to-end run against real Postgres and Redis.
    • filedocker-build.yml@491a221 ↗Builds the image, asserts /app/alembic.ini exists, imports the app inside it. Header names two defects it would have caught: the missing alembic.ini and an h2/hpack resolution error.
    • docgaps.md@491a221 ↗Sections 9, 11, 13 and 14 give the before and after.
    • filerequirements.txt@491a221 ↗mcp==1.24.0, playwright==1.56.0 pins present.
    • filee2e_smoke_test.py@a74c00c ↗The 23-check journey (23 check() calls counted).

    How this was checked: Read both commit messages and the relevant diffs, gaps.md sections 9-14, docker-build.yml; ran 'gh run list --workflow docker-build.yml' (read-only): success for 764e2c6, a74c00c, 491a221; confirmed the two pins in requirements.txt. Could not reproduce the build (Docker daemon not running on this machine). Attribution: Commits 764e2c6, a74c00c and 491a221 carry a 'Co-Authored-By: Claude Sonnet 5' trailer, and gaps.md is an AI-written first-person document (it addresses the repository owner as 'you').

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit 764e2c6 (with a74c00c the same day)

  6. FlashFlowFEATURED

    Failure classifier thresholds were fitted to six known outcomes; the docs now say so

    ADOPTED

    Context

    Stage 17 added 'flashflow report' with a four-label decision tree (STABLE, ACUTE_COLLAPSE, CHRONIC_COLLAPSE, RECOVERY_LIMITED) meant to reproduce Stage 15/16's characterization of six routing policies. It needed two numeric constants.

    Decision

    Gate on the bottleneck's dispatch share within its own congestion episode (concentrated if at least 1.2x fair share) and on committed work (severe if at least 10x capacity). An undrained queue is classed as collapse first; for every drained case committed work is then checked, whether or not the policy concentrated. The decision tree and both constants were adjusted iteratively until the tool reproduced Stage 15/16's published characterization of the six policies: five exactly, and weighted-round-robin as a documented refinement.

    Alternatives considered

    • Gate chronic vs acute on fraction of time over capacity (first version). Misclassified EWMA as RECOVERY_LIMITED: its slow drain (to about 6.85 s of 8 s) gives a fraction comparable to round-robin's.
    • Measure concentration from whole-run completed share. Gave EWMA about 0.8x fair share; completions undercount a backlogged target and the worst episode is not the run-long favorite.
    • Check committed work only on the concentrated branch. Misclassified P2C-load (committed work 4) as RECOVERY_LIMITED instead of STABLE.

    What happened

    The original doc and code comment said the constants were chosen up front and not tuned. An audit (9e24add describes its audit as 12 parallel agents) noted that the same document's bug narrative contradicted this; 38eefbf rewrote both to say the constants were calibrated on a fixed six-point sample with no held-out policy or scenario. 9e24add also fixed 'explain' printing 'traffic concentrated' for round-robin; tests now assert on the rendered text. A re-run of all six policies reproduced the documented labels and committed-work values (4, 163, 8, 97, 4, 71) but showed the code comment's '1.4x to 3.1x fair share for every non-round-robin policy' does not hold: EWMA is 4.37x, and P2C-load is 1.05x on two of three seeds, equal to round-robin's 1.053. P2C-load's STABLE label comes from the drained and low-severity gates, not the 1.2 threshold. A seventh policy or another scenario is untested.

    P99 latency and failure label for six routing policies

    One scenario: 5 targets, capacity 1, a flash-crowd workload, seeds 17000-17002. Lower is better, and the label is what the classifier says.

    • Weighted round-robin
      1.21 s · ACUTE_COLLAPSE · committed work 163
    • Least connections
      2.49 s · STABLE · committed work 8
    • Power of two choices
      2.68 s · STABLE · committed work 4
    • Adaptive
      2.88 s · ACUTE_COLLAPSE · committed work 71
    • Round-robin
      3.76 s · CHRONIC_COLLAPSE · committed work 4
    • EWMA
      4.40 s · ACUTE_COLLAPSE · committed work 97
    Source: P99 read from the dashboard's Compare tab; labels and committed work re-run from the report command in a clone

    Evidence

    How this was checked: Read internal/report/report.go and the Stage 17 doc. Built cmd/flashflow in a clone and ran the report for all six policies, then read concentration, committed work and drain time per seed from the JSON it wrote. Five of the six labels match the documented outcomes exactly; weighted round-robin is a documented refinement. The decisive commits (ec006c0, 38eefbf, 9e24add) are authored by Ujjwaljain16. The audit that prompted the correction was an AI-agent review, not an independent human review; the author made the resulting changes.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · doc first committed in ec006c0; correction in commit 38eefbf

  7. RecoveryOSFEATURED

    Take async advisory locks on a separate sync connection after one leaked permanently

    ADOPTED

    Context

    advisory_lock_async was written for AsyncSession callers (reconciliation, anomaly windows) that deliberately commit the session while still holding the lock. Postgres advisory locks taken with pg_advisory_lock belong to a session (connection), not to a transaction.

    Decision

    advisory_lock_async no longer uses the caller's session connection. It checks out its own psycopg2 connection from get_sync_engine() via asyncio.to_thread, takes and releases the lock there, catches BaseException (so task cancellation still rolls back), and shields the unlock call from cancellation. An earlier attempt that wrapped the session's async engine in a second AsyncEngine was replaced because it broke the test suite's per-test engine isolation.

    Alternatives considered

    • Lock and unlock on the caller's own AsyncSession. AsyncSession.commit() returns its DBAPI connection to the pool and may check out a different one. Found live: pg_advisory_unlock ran on a different backend pid, returned false without error, and the real holder went back to the pool still holding the lock, deadlocking every later caller of that key with no exception logged.
    • Wrap the session's async engine in a second AsyncEngine for the lock. Individually correct, but two tests later it hit the documented asyncpg event-loop failure that tests/integration/conftest.py rebuilds engines to avoid.

    What happened

    Four new tests in test_advisory_lock_async.py cover release after CancelledError with an aborted transaction, release after the wrapped block commits the session, lock and unlock never going through the caller's session, and release after a plain exception. The test file notes that the commit-based test can pass by chance on a quiet connection pool, so the never-through-the-session test is the deterministic guard. The in-process suite did not expose the bug; the commit message says it appeared when the demo endpoints were tested against a concurrent, multi-container deployment. The helper had been added four days earlier (cb317e9, 2026-08-29). Cost: every locked section now holds one extra pooled sync connection for its duration.

    Evidence

    • commit5655649 ↗Message describes the live reproduction (unlock on a different backend pid), the discarded second-AsyncEngine attempt, and the BaseException/shield changes.
    • filedatabase.py ↗advisory_lock_async docstring and implementation use get_sync_engine() plus asyncio.to_thread.
    • testtest_advisory_lock_async.py::test_lock_and_unlock_never_go_through_the_callers_session ↗Pins the separate-connection property.
    • commit6be61a3 ↗Earlier fix to the sync helper: rollback ran unconditionally and could discard a caller's uncommitted write; scoped to the exception path with a test that fails against the old code.
    • commitcb317e9 ↗Adds advisory_lock_async (2026-08-29) and wraps reconcile_pending_recovery and persist_anomaly_window in it; the first version of the helper that 5655649 later replaced.

    How this was checked: Read commit 5655649 and 6be61a3 messages and diffs, the current advisory_lock_async source, and the test names in test_advisory_lock_async.py. Did not reproduce the deadlock (needs a live multi-container Postgres). Attribution: Decisive commit 5655649 carries a 'Co-Authored-By: Claude Sonnet 5' trailer. All cited commits are authored by Ujjwaljain16. The commit message and docstring read as AI-assisted.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit 5655649

  8. RecoveryOSFEATURED

    Let the LLM break economic near-ties and raise risk flags, but never grant permission

    ADOPTED

    Context

    Money-moving decisions in RecoveryOS come from a deterministic propensity/EVI/policy chain. On 2026-08-29 an audit found the docs overclaimed what the LLM influenced, and b8ccb2b corrected them and added a test proving build_decision() never read diagnosis output, which meant deleting the whole diagnosis service would change no outcome.

    Decision

    Four days later, 3bf04ea gave the investigator a second output, a RecoveryRecommendation (closed six-action enum, confidence, closed-set risk_flags, rationale). orchestrator._apply_ai_fusion can use it in two ways, both behind ai_recommendation_fusion_enabled (default false). (1) Tie-break: it can pick among candidates that already cleared the EVI floor, are individually policy-ALLOWed, and lie within ai_tie_break_tolerance_bps (default 100, i.e. 1%) of the winner; 8e24eb6 later added a confidence floor of 0.5. (2) Escalation: AIRiskSignalEscalationRule, an ordinary policy rule, turns a non-empty risk_flags into ESCALATE. The recommendation has no amount, provider, ID or idempotency-key fields.

    Alternatives considered

    • Zero AI authority: LLM output is explanation only (the b8ccb2b position). It was kept for the core argmax and the 11 AI-blind rules, but the LLM then had no causal effect on outcomes; 3bf04ea deliberately superseded it for near-ties and risk flags. The old test was rewritten, and its docstring says so.
    • Let the LLM propose candidates or actions and check them afterwards. README section 16 gives this reason: if AI could propose novel candidates, 'AI may never create permission' would be unenforceable. It also states the AI never sees candidate EVI scores.

    What happened

    Structural tests hold the boundary: an AST walk fails if enqueue_recovery_job or process_job reference recommendation identifiers; the pure argmax and the 11 AI-blind policy rules are scanned for diagnosis and confidence identifiers; exactly one rule may reference ai_risk. The rule count is asserted exactly, so adding a rule fails the test until reviewed (78de026 updated it for MoneyExposureLimitRule). The TRD was corrected in 3bf04ea (new section 3.5, threat table) and again in 0456cfb (a RE-CORRECTED note that 'zero causality' was no longer true). Real-model evidence is thin: 2 real recommendations, no tie-break or escalation observed. The 0.5 floor is described in the commit and docs as fixed before any measurement; git history cannot confirm that. The headline benchmark is labelled whole-system lift, not AI-attributed lift.

    Evidence

    How this was checked: Read commits b8ccb2b, 3bf04ea, 8e24eb6; TRD section 3.5; the two test files; ran 'pytest tests/unit' (323 passed) and the execution-boundary and fusion tests in test_ai_recommendation_adversarial.py (4 passed). The integration-level AST tests in test_diagnosis_has_no_decision_authority.py were read but not run (they import DB fixtures). Attribution: Decisive commit 3bf04ea carries a 'Co-Authored-By: Claude Sonnet 5' trailer (it is also a 31-file, about 4,900-line commit that bundles the mission state machine). All cited commits are authored by Ujjwaljain16. The TRD and test docstrings are long AI-assisted prose.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit 3bf04ea (zero-authority position first pinned in b8ccb2b, 2026-08-29)

  9. RecoveryOSFEATURED

    Benchmark against a baseline with the same attempt budget and the same compliance rules

    ADOPTED

    Context

    The incremental-revenue number compares RecoveryOS's real outcomes with a counterfactual naive strategy on a simulator. The first baseline modelled one retry attempt while RecoveryOS could make several, and neither baseline checked the regulatory rules that block RecoveryOS.

    Decision

    The comparator was rebuilt in steps. 688cd1b made baseline and execution call one shared resolve_simulated_outcome(). 98f9387 gave the naive strategy the same attempt budget (later min(max_retries, mission_max_attempts), 0456cfb). 7c65faa then runs each baseline attempt through the real services.policy_engine.evaluate() chain with a fixed RETRY_NOW candidate, so only the compliance blocking is borrowed from RecoveryOS: propensity, EVI and action selection are not used, and the naive hopeless-failure filter stays. d11e0ae made outcome draws deterministic per (payment_id, attempt), because re-running one seed could give a different headline number. 7b9000d fixed time-based rules reading the real clock for synthetic payments (93% of one seed's BLOCKs).

    Alternatives considered

    • Single-attempt naive baseline (original). 98f9387: RecoveryOS gets up to max_retries attempts, so the gap mixed 'more attempts' with 'better decisions'. The same commit records that an earlier fair-baseline query summed the whole dataset against one payment and produced a nonsensical negative number.
    • Same-budget baseline that ignores compliance rules. Kept only as compliance_blind_fair_baseline_DIAGNOSTIC_ONLY: it may retry in NPCI peak windows and above RBI limits. RecoveryOS loses to it in all 5 seeds (mean minus 142,189 rupees), and the docs state the compliance-aware baseline, not this one, is the headline comparison.

    What happened

    The documented headline changed from +42,491.88 rupees (seed 42, single-attempt baseline) to +73,181.78 rupees (5 seeds, compliance-aware baseline). Limits visible in the code: the compliance-aware baseline stops at its first non-ALLOW verdict and does not reschedule, while RecoveryOS schedules re-evaluations. Both arms use the same per-(payment, attempt) draw, which plausibly explains why RecoveryOS recovers a strict superset (baseline_only = 0 in all seeds); this is an inference, and the artifact does not break down which mechanism produced the 22 to 43 extra recovered payments per seed. The artifact was generated on 2026-09-02, before MoneyExposureLimitRule was added on 2026-09-06, and was not regenerated. The benchmark runs on a simulator, as README section 17 states.

    Evidence

    • commit98f9387 ↗Introduces the same-attempt-budget baseline and states why the single-attempt baseline was unfair.
    • commit7c65faa ↗Compliance-aware comparator that reuses the policy engine unmodified.
    • commit7b9000d ↗Clock bug: 93% of one seed's BLOCKs traced to reading real wall-clock time.
    • commitd11e0ae ↗Non-reproducible outcome draws found and fixed; adds baseline_runs unique constraint.
    • filebaseline.py ↗Docstring for compute_and_persist_compliance_aware_baseline_run and the break-on-first-block loop.

    How this was checked: Read the four commit messages and the baseline.py compliance-aware function and loop end to end; read README sections 9, 10, 17; compared with the per-seed JSON (see the campaign investigation). Attribution: None of the cited commits (688cd1b, 98f9387, 7c65faa, 7b9000d, d11e0ae, 0456cfb) carry a Claude co-author trailer; all are authored by Ujjwaljain16. The repo as a whole is AI-assisted (16 of 139 commits carry the trailer; .claude/ is gitignored as AI tooling), so a general disclosure is still appropriate.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit 7c65faa (same-budget step in 98f9387, 2026-08-29)

  10. FuzeFEATURED

    Rejected a modular ML rewrite, then rebuilt a similar pipeline behind flags three days later

    PARTIAL

    Context

    In July 2026 code from a side branch (architecture-shift, last commit 2026-05-13) was being folded into integration-review. ML-CONVERGENCE-AUDIT.md (34ebbe5) and ARCHITECTURE-CONVERGENCE.md (665c370), both 2026-07-22, kept the 3,091-line unified_recommendation_orchestrator.py as canonical and rejected backend/ml/engines/, backend/ml/recommendation/ (including a cross-encoder re_ranker: 'High memory/CPU overhead; causes latency spikes on single-worker deployments'), the Supabase match_user_content RPC ('unneeded RPC dependencies') and a data_layer that 'bypasses' the repository and Unit of Work. cf9172d had brought over the Alembic framework and only the HNSW (0003), user-URL unique (0004) and embedding_metadata (0005) migrations.

    Decision

    On 2026-07-25 the same shapes were committed as new code. 1f4deb2 added a Pipeline plus Strategy layer (ml/recommendation/{domain,pipeline,retrieval,scorer,data_layer,shadow_evaluator}.py, ml/engines/{base,smart}_engine.py) with a pure-function scorer and a two-stage retriever: pgvector cosine ANN top-100 over saved_content.embedding (HNSW index from 0003), then scoring, with NULL-embedding rows appended. 691294b added Postgres functions search_bookmarks_semantic_v1 and search_bookmarks_hybrid_v1 (migration 0006) and a new /api/search/rpc-semantic endpoint behind the search_rpc flag, falling back to the existing SearchService. The legacy orchestrator stayed the serving path. ADR-001 to ADR-005 (a9ec4d1, same day) describe the layers; ADR-001 calls the orchestrator a 3,092-line 'god object'. No document mentions the 2026-07-22 audit.

    Alternatives considered

    • Keep only the monolithic orchestrator (position of the 2026-07-22 audit). Not kept. ADR-001 (2026-07-25) describes the orchestrator as a 3,092-line 'god object' to decompose; the repository has no note reconciling this with the audit's description of it as stable and canonical.
    • Supabase match_user_content.sql RPC. Held back in cf9172d as an 'unneeded RPC'; replaced by the project's own versioned Postgres functions in migration 0006.
    • Cross-encoder re-ranker from the side branch. Rejected in the audit for memory and CPU cost. The new design has no cross-encoder; ranking uses the pure-function scorer (ADR-004).
    • Existing ORM query using the pgvector <=> operator (SearchService.semantic_search). Kept as the fallback when the flag is off. It already orders by cosine distance inside Postgres, so the RPC's advantage is unproven: migration 0006 says it removes 'Python/SQLAlchemy round-trip overhead' but no comparison was measured.

    What happened

    The new code is only partly live. The orchestrator builds RecommendationPipeline() without a Unit of Work and, since RECOMMENDATION_SHADOW_MODE defaults to true, runs it after each uncached legacy request; its data layer logs 'initialized without UnitOfWork' and returns no candidates (verified: results []), so the shadow comparison sees nothing. Serving through the pipeline (RECOMMENDATION_PIPELINE_ENABLED, default false) would return an empty list. No code sets RecommendationRequest.query_embedding, which the ANN branch needs, so CandidateRetriever is unreachable and untested. The RPC functions have no tests, the endpoint has no frontend caller, and the flag defaults to false in code (runtime flag values cannot be checked). HNSW speed and recall are asserted only in the 0003 docstring; no EXPLAIN output or recall measurement is in the repo.

    Evidence

    • commit34ebbe5 ↗ML-CONVERGENCE-AUDIT.md rejects engines/, recommendation/, re_ranker and RPC search.
    • commitcf9172d ↗Brings only the HNSW/unique/metadata migrations; holds the RPC SQL as 'unneeded'.
    • commit1f4deb2 ↗Adds the pipeline, retriever, scorer, SmartEngine, shadow evaluator and golden tests (21 files, 2,108 insertions).
    • commit691294b ↗Adds migration 0006, rpc_search_service.py and a new POST /api/search/rpc-semantic endpoint gated by the search_rpc flag.
    • fileretrieval.py@491a221 ↗Two-stage ANN query and NULL-embedding fallback; the ANN branch needs request.query_embedding.
    • file0003_hnsw_indexes.py@491a221 ↗HNSW parameters and CONCURRENTLY rationale.
    • commita9ec4d1 ↗Adds ADR-001 to ADR-006 (2026-07-25, two minutes after 1f4deb2).
    • fileunified_recommendation_orchestrator.py@491a221 ↗Builds RecommendationPipeline() with no UnitOfWork; shadow mode default true, cutover default false.

    How this was checked: Read ML-CONVERGENCE-AUDIT.md, ARCHITECTURE-CONVERGENCE.md, cf9172d/1f4deb2/691294b diffs, retrieval.py, data_layer.py, orchestrator wiring (lines 2366-2440), grep for query_embedding assignments and CandidateRetriever usage (none outside retrieval.py/data_layer.py); ran RecommendationPipeline().run() in a Python shell: 'uow: None results: []'; verified backend/ml at 665c370 has no engines/ or recommendation/ directories; wc -l of the orchestrator before 1f4deb2 was 3,091. Attribution: All cited commits are authored by Ujjwaljain16 and none carries a Claude co-author trailer. The audit and ADR documents are written in a structured, emoji-marked style, but the repository does not say whether they were AI-assisted.

    Recorded at the time (a document or commit message states it) · reasons are inferred · commits 691294b and 1f4deb2 (reversing the audit in 34ebbe5 and cf9172d, 2026-07-22)

  11. TypeAheadXFEATURED

    Virtual nodes per Redis node set to 500 after trying 150 and 1000

    ADOPTED

    Context

    With 3 physical nodes the ring arcs are uneven unless each node is placed many times. The first ring used 150 virtual nodes per node. This was a course project, and the 150, 1000 and 500 stages all fall within under two hours of one day (66845c0 at 03:12 and b116d06 at 04:59 on 2026-06-10, +05:30).

    Decision

    At 66845c0 the phase-4 distribution doc reports that 150 vnodes gave ownership 39.6% / 28.6% / 31.7% and concludes VIRTUAL_NODES=1000 (Settings default 1000). In b116d06 the docs, scripts and Settings default were changed to 500: 'providing 90% of the load balancing benefits of 1000 virtual nodes, but consuming only 50% of the memory footprint and CPU routing cost'. backend/.env.example (changed in 05af6b4) and the Settings default are 500; the ConsistentHashRing constructor default in consistent_hash_ring.py is still 150 and applies only when no setting is passed.

    Alternatives considered

    • 150 virtual nodes. Measured ownership 39.6% / 28.6% / 31.7% (ring_analysis.py); redis-a owned about 6 points more than its fair share.
    • 1000 virtual nodes. Better balance (33.3% / 32.5% / 34.1%) on a ring of 3000 positions, twice the 1500 at 500. The doc's own estimate says the extra lookup cost is negligible, so the '90% of the benefit at 50% of the cost' argument in b116d06 is asserted rather than measured. b116d06 also left the sentence 'reduced the standard deviation of arc sizes by nearly 7x' in place, which matches 1000 (7.62e35 to 1.11e35), not 500 (2.40e35, about 3.2x).

    What happened

    Final ownership with 500: 33.86% / 34.58% / 31.57% (ring_analysis.py re-run, matching README). Rebalance moved 24.64% of keys at 1000 vnodes (66845c0 doc) and 26.18% at 500 (b116d06 doc), so 500 is a little further from the ideal 25%. A sweep over the deterministic ring gave ownership standard deviation of 25.87 (1 vnode), 6.02 (10), 0.40 (50), 4.64 (150), 1.28 (500) and 0.66 (1000) percentage points: balance is not monotonic in vnode count and one fixed ring is a single sample, so 50 happens to beat 500. Later doc figures (10 vnodes about 30% std, 150 about 10%) are not produced by any script in the repo; the sweep gives 6.02 and 4.64.

    Evidence

    • commit66845c0 ↗docs/phase4-distribution-analysis.md with the 150/500/1000 table; config default 1000.
    • commitb116d06 ↗Changes 1000 to 500 in config.py, scripts and docs with the 'sweet spot' rationale.
    • filering_analysis.py@66845c0 ↗Computes arc sizes and per-node ownership for 150, 500, 1000 vnodes.
    • file.env.example ↗VIRTUAL_NODES=500.
    • commit05af6b4 ↗Sets VIRTUAL_NODES=500 in backend/.env.example (that file was not part of b116d06).

    How this was checked: Ran scripts/ring_analysis.py (output: 150 -> 39.65/28.64/31.71, 500 -> 33.86/34.58/31.57, 1000 -> 33.38/32.51/34.11) and a separate sweep script over ConsistentHashRing (not committed to the repo); diffed b116d06 for the 1000 to 500 change. Attribution: Course project: the README does not say so, but docs/phase4-completion.md@66845c0 contains 'Viva Talking Points' and the phase briefs 3.md, 4.md and 5.md (committed, then deleted in the next phase) are written as instructions to a student ('most students will completely mess up'), so the work followed a written brief. All 15 commits are by Ujjwaljain16, none has a Co-Authored-By trailer, and 14 of 15 fall on one day (2026-06-10, 00:23 to 21:22 +05:30; the last is 2026-06-22).

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit b116d06 (settles on 500); 66845c0 (150 measured, 1000 chosen)

  12. SSE-ObservatoryFEATURED

    Proxy tickets for auth tokens were added, then abandoned by the client about 50 minutes later

    REVERTED

    Context

    Authorization tokens for the upstream SSE endpoint were passed as an ?auth= query parameter in the proxy URL (already the case in the February vite.config.ts proxy). The first Mar 7 proxy commit (3b6deed) introduced 'tickets': a client POSTs {url, token} to /api/sse/ticket and then opens /api/sse?ticket=... . The repository does not say why; the likely motive is keeping tokens out of URLs and logs, which is an inference. The app is deployed on Vercel, where an in-memory Map is not shared between function instances.

    Decision

    3b6deed added two ticket implementations in the same commit. server.js (Express) keeps tickets in a Map with a 30 s expiry and one-time use. api/sse/_ticket.js (Vercel) encrypts {url, token, exp} with AES-256-GCM so any function instance can decrypt it. 522838d replaced the random fallback secret with VERCEL_GIT_COMMIT_SHA and then a fixed string, with the comment 'to ensure consistency across Lambdas'. 7578190 added an origin check to the ticket endpoint. At 19:13, c343a28 changed obtainSSEProxyTicket() to stop calling /api/sse/ticket and return `/api/sse?url=...&auth=...` again, with the message 'switch to direct origin-locked proxy for better production stability'. Access control moved to an Origin/Referer check, and ticket decoding was kept as 'Legacy support'.

    Alternatives considered

    • Keep tokens in the query string (the February design). Replaced by tickets in 3b6deed and restored in c343a28. The only reason given is 'better production stability'; the repository does not say what failed.
    • In-memory ticket Map on every deployment. Used only in server.js. The Vercel functions use encrypted stateless tickets instead; the 522838d comment cites consistency across Lambdas, from which the lack of shared memory between instances is inferred.

    What happened

    At HEAD no client code calls the ticket endpoint: obtainSSEProxyTicket() builds a direct proxy URL, so auth tokens are again sent in the URL query string. The ticket code is still in the repository but is unused, and it was not hardened after being retired; it should be treated as dead code. Proxy access control now rests on the Origin/Referer header, which limits browsers but not non-browser clients. This is a known limitation of the design and is not fixed.

    Evidence

    • commit3b6deed ↗Adds tickets: a Map in server.js and AES-256-GCM in api/sse/_ticket.js (secret from PROXY_SECRET, random per-process fallback).
    • commit522838d ↗Secret changed to PROXY_SECRET || VERCEL_GIT_COMMIT_SHA || fixed string, comment 'ensure consistency across Lambdas'.
    • commitc343a28 ↗Removes the POST /api/sse/ticket call from src/utils/sseProxyUrl.ts, adds an Origin/Referer check to api/sse/index.js and keeps ticket decoding as 'Legacy support'.
    • fileserver.js@0e8c14b ↗Express ticket Map and /api/sse/ticket handler still present at HEAD.

    How this was checked: Read the diffs of 3b6deed, 522838d, 7578190 and c343a28 and the files at HEAD; searched src/, README.md and docs/ for uses of the ticket endpoint. Reasons: the c343a28 reason is stated; the token-leak motive for introducing tickets is inferred.

    Recorded at the time (a document or commit message states it) · reasons are inferred · commit c343a28 (19:13), following 3b6deed (18:22), 7578190 (19:04), 522838d (19:09)

  13. VitestFEATURED

    mergeTests: from a short extend() wrapper to merging fixture registrations directly (PR still open)

    PARTIAL

    Context

    Issue #9483 asked for a Playwright-style mergeTests to combine fixtures from several extended tests. In Vitest 4.1.0 a test built with test.extend() holds a TestFixtures object: a Map of registrations (each with scope, auto, deps and a `parent` link to the base implementation of the same-named fixture), a WeakMap of per-suite overrides used by test.override, and WeakMaps of file and worker contexts. Lookup for a suite walks up the suite chain to the nearest override. STATUS: PR #9662 is OPEN and unmerged. Maintainer sheremet-va posted seven CHANGES_REQUESTED reviews from 2026-02-15 to 2026-02-22; the contributor's last push is 2026-03-10 and no maintainer response follows.

    Decision

    Stage 1 (d189936, 2026-02-14; also the design pitched in the issue and still in the PR body): mergeTests(a, b) = a.extend(b's resolved fixtures), about 4 statements. Stage 2 (fa60ae8, 02-16): a loop calling currentTest.extend(next.getFixtures().toUserFixtures()), with a comment that overrides on the current test are dropped. Stage 3 (ddf97bf, 02-18, after the maintainer asked for a merge on TestFixtures): the loop passes the TestFixtures itself, and TestFixtures.extend gains an `instanceof TestFixtures` branch that copies registrations. Stage 4 (60bb43a, 03-09; head 30ddf39): mergeTests builds one Map itself with last-writer-wins, throws FixtureDependencyError for a different scope or auto option, keeps built-in fixture names from the first test, runs validateFixtures on the merged Map and wraps it in new TestFixtures(map); it no longer calls extend(). Types: six fixed overloads (1 to 6 arguments) instead of a variadic signature.

    Alternatives considered

    • Serialise the already-parsed registrations back to user fixtures (toUserFixtures) and replay them through .extend(). Maintainer sheremet-va (2026-02-15): 'why do we need to convert already converted fixtures into user definitions and then convert them back again?' He proposed a merge function on TestFixtures that accepts another Fixtures and iterates the registrations, overriding them.
    • Variadic generic signature constrained to readonly unknown[] with a structural context-extraction type via beforeEach. The author adopted it because TestAPI<any> rejected valid inputs (TestAPI is invariant in its context parameter); the maintainer accepted internal casts if the public API is strict and suggested writing many overloads, which the head implements.
    • Contributor's stated worry that a low-level merge would 'bypass the .extend() validation pipeline'. The maintainer asked why, since the original extend() calls already validated each test. The head adds its own scope, auto and validateFixtures checks in mergeTests instead.

    What happened

    Not merged, so nothing shipped; the record is about how the design moved. At the head: (1) The PR description, the issue comment, the docs ("equivalent to calling .extend() repeatedly"), the doc comment ("No new validation logic is introduced") and the contributor's blog post still describe the extend() chain, but mergeTests no longer calls extend() and adds its own validation; the `instanceof TestFixtures` branch in TestFixtures.extend is unused by the rest of the diff. (2) The head copies each item with its own `parent`; extend() would link a same-named override to the earlier registration, so a fixture that calls its own base could resolve differently (read from the code, not executed). (3) The diff also adds validateFixtures calls to extend() and override(). (4) On 02-18 the maintainer called the earlier version hard to review; on 02-21 he asked why chain.ts was touched (no longer in the diff). (5) Circular dependencies are documented as not detected at merge time.

    Evidence

    • pull requestvitest-dev/vitest/pull/9662 ↗Open PR: +1869 / -15 across 10 files and 29 commits; 7 CHANGES_REQUESTED reviews by maintainer sheremet-va (2026-02-15 to 2026-02-22), none dismissed; last contributor commit 2026-03-10.
    • issuevitest-dev/vitest/issues/9483 ↗Feature request; still open; contains the contributor's 2026-02-14 description of the extend-based approach.
    • commitd189936 ↗Stage 1: mergeTests as test.extend(testB.getFixtures().resolveFixtures()), 2026-02-14.
    • commitfa60ae8 ↗Stage 2: 'simplify mergeTests to use linear extension chain', 2026-02-16.
    • commit30ddf39 ↗Head of the PR on 2026-03-10; mergeTests builds a merged Map and does not call extend().
    • pull requestvitest-dev/vitest/pull/9662#discussion_r2809025114 ↗Maintainer objection to the serialise-and-replay approach.
    • filefixture.ts@v4.1.0 ↗TestFixtures holds a flat registrations Map (copied on extend()), an _overrides WeakMap looked up along the suite chain, and per-item `parent` links to a fixture's base implementation.
    • commitddf97bf ↗Stage 3, 2026-02-18: extend loop over TestFixtures objects, `instanceof TestFixtures` branch added to TestFixtures.extend.
    • commit60bb43a ↗Stage 4, 2026-03-09: mergeTests builds the merged Map itself and adds scope/auto conflict errors.
    • pull requestvitest-dev/vitest/pull/9662#pullrequestreview-3821150260 ↗Maintainer review of 2026-02-18: the rewritten implementation is hard to review and adds parent/ancestor tracking that registrations already cover.

    How this was checked: Read PR 9662 metadata, all 29 commit headlines, the full current diff, issue 9483, the PR issue comments and all 23 inline review comments; fetched suite.ts at commits d189936 and fa60ae8 and read the mergeTests bodies; fetched fixture.ts at v4.1.0 and read TestFixtures, get(), override() and parseUserFixtures(); read the contributor's blog post in src/data/blogPosts.ts (slug integration-complexity). Confirmed the PR is still OPEN via gh api on 2026-09-28. Did not build the PR branch or run its tests. Attribution: On 2026-02-22 the maintainer wrote that the tests looked AI-generated ("this AI instead just documents the wrong behaviour in the test").

    Reconstructed from code and history · reasons are stated in the repository · Commit ddf97bf (2026-02-18) and the author's PR comment of 2026-02-18 announcing direct registration merging, after the maintainer's review of 2026-02-15; reworked in 60bb43a (2026-03-09); head 30ddf39 (2026-03-10). PR #9662 is still open.

  14. AgentBrakeFEATURED

    Zod-validated YAML policy file, then a same-day catch-all fallback that disables every policy

    PARTIAL

    Context

    SUBMISSION.md states the reason for a config file: 'No more hardcoded env vars; policies are versionable artifacts.' The YAML plus Zod design is in the second commit of the repository (f98ec86); no environment-variable configuration for policies exists anywhere in the history, so a 'migration from env vars' is not shown by git. Before c8a3ea8 the no-file path was AgentBrakeConfigSchema.parse({}), which the schema of that moment (agent and policies required) would reject.

    Decision

    ConfigLoader.load() reads AGENT_BRAKE_CONFIG, then agent-brake.yml/.yaml/.json in the working directory, and validates with AgentBrakeConfigSchema.parse(). c8a3ea8 wraps each attempt in try/catch that logs 'trying next...', and at the end returns a hand-written object literal (agent 'safe-fallback-agent', empty limits and security) that is not passed through the schema. The same commit makes agent, policies, limits and security optional with defaults.

    Alternatives considered

    • Environment variables (the approach SUBMISSION.md says it moved away from). Rejected in the pitch text: not versionable. No code for it exists in the history.
    • Before c8a3ea8: throw on invalid config and use AgentBrakeConfigSchema.parse({}) when no file exists. The commit's own comment says the hard-coded default is 'to avoid Zod initialization errors'.

    What happened

    Fail-fast became fail-open. If the configuration file fails validation for any reason, the loader logs 'Failed to load config ... using safe defaults' and the proxy starts with no configured policies (BrakeProxy then applies only a default MaxToolCallsPolicy(10)), so calls that a valid configuration would block are forwarded. This was fixed in 373adcd (Sep 2026): an invalid or missing config now makes the proxy exit with an error unless AGENT_BRAKE_ALLOW_INVALID_CONFIG=1 is set, unknown keys are rejected, and denied_tools is enforced. Once the schema defaults were added, the catch-all was no longer needed to avoid the initialization error. The schema also accepts keys that nothing reads: denied_tools, global.on_violation, global.max_retries, budget.warn_threshold (the README sample sets 0.8, BudgetPolicy hard-codes 80 %), budget.currency, and agent.trust_level (logged only).

    Evidence

    • commitf98ec86 ↗Zod schema with YAML loader announced in the second commit (loader.ts blob is empty in this commit; content arrives in 933ba0e).
    • commitc8a3ea8 ↗Adds try/catch, 'trying next...' and the hard-coded fallback returned without schema.parse.
    • commit6fee097 ↗SUBMISSION.md: 'No more hardcoded env vars'.
    • fileloader.ts@0fc99c8 ↗Fallback literal after the loop.

    How this was checked: Read schema.ts and loader.ts across f98ec86, 933ba0e, f8835db, c8a3ea8 and HEAD; ran the proxy with a config containing one wrong type and checked the startup log and behaviour; grep for uses of the unused keys.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit c8a3ea8 (20:58, the pivot); original design f98ec86 (17:32) and SUBMISSION.md in 6fee097

  15. AgentBrakeFEATURED

    Circuit breaker and human approval shipped as policies with no signal path to drive them

    PARTIAL

    Context

    ROADMAP_V3.md (removed from the repository later the same day in 03de458) describes 'Interactive Sudo Mode': the policy 'pauses the request', returns an approval signal, and 'Human approves -> Request resumes'. The README lists 'Circuit Breaker: auto-cut connection if tools fail repeatedly' and 'Human-in-the-Loop: pause execution for approval via Slack/Webhook'. ada809d added both policies in one commit, together with tests/policies.test.ts (17 tests; the circuit-breaker tests call recordFailure() directly).

    Decision

    CircuitBreakerPolicy exposes recordFailure/recordSuccess and blocks a tool while its circuit is open. ApprovalPolicy stores a pending key (tool name + JSON arguments), returns action request_approval the first time, block ('Awaiting approval') on repeats, and exposes approve()/deny(). The proxy answers request_approval with an immediate JSON-RPC error -32001 (status 'pending') instead of holding the request. f98705f adds a WebhookNotifier with Slack buttons linking to `${approvalUrl}?action=approve|deny`. Both policies are constructed in src/proxy/index.ts.

    Alternatives considered

    • Hold the request open until a human answers (the roadmap flow). Not implemented; the code returns an error at once. No document says why.

    What happened

    Running the built proxy: six consecutive calls to a tool that always fails all reached the server (error -32603) with a breaker threshold of 3; none was short-circuited, because recordFailure() is called only from tests/policies.test.ts. For approval, a require_approval tool got -32001 on the first call and -32000 'Awaiting approval' on every retry. Nothing calls approve() or deny(), WebhookNotifier is never imported, and no server handles the approvalUrl links, so such tools stay blocked (fail-closed) and the advertised approval workflow does not exist. The README still lists both features. Later change (373adcd, Sep 2026): the proxy now parses server responses and reports real tool errors to the breaker, so it can trip. Approvals were not built; the README and code now say they are not implemented.

    Evidence

    How this was checked: grep for recordFailure, recordSuccess, .approve(, .deny(, WebhookNotifier across src, examples and tests; ran harness4.mjs against the built proxy.

    Reconstructed from code and history · reasons are inferred · commit ada809d (18:30); notifier in f98705f; intended flow in 8136b6f

  16. CampusSyncFEATURED

    Certificate OCR moved from Tesseract plus regex to one Gemini Vision call, no fallback

    ADOPTED

    Context

    The first API route (fa6484b, 2025-09-11) ran Tesseract.js on the uploaded file and returned the raw text; a comment says 'Naive extraction heuristics' but the code applies none. On 2025-09-23 (3c573ae) the author added Gemini 2.0 Flash to structure the OCR text, with a rule-based extractor (ocrExtract.ts) as the fallback, and also ran Tesseract in the browser on the upload page. The same commit adds a workaround comment for Tesseract worker paths breaking under the Next.js bundler.

    Decision

    The upload page now posts the file to /api/certificates/ocr-gemini. That route stores the file in Supabase Storage, sends the raw bytes as base64 inline data to gemini-2.0-flash-exp with a JSON-only prompt, and throws if no JSON object can be parsed. The old ocr/route.ts (Tesseract, PDF-to-image conversion, regex merge) is deleted in 980ef09. The page change is in 9f5427a, whose message covers only a delete-confirmation modal; that commit points at a route that is not yet in its tree, and the route is first committed in 980ef09. No commit message gives a reason for the switch. The nearest statement is SIMPLE-CERTIFICATE-GUIDE.md (added in 980ef09): the system was simplified 'to avoid all the complex OCR dependencies that were causing errors' (Jimp, Tesseract.js). That guide describes a different, text-only design that the shipped route does not implement, so the reason is inferred.

    Alternatives considered

    • Tesseract text (browser and server) then Gemini text structuring with regex fallback (the 3c573ae design). Removed in 9f5427a and 980ef09. The only reason in the repo is the dependency-error sentence in SIMPLE-CERTIFICATE-GUIDE.md; no accuracy comparison between the two approaches was found in the history.
    • pdf-parse text extraction plus a text-only LLM (described in SIMPLE-CERTIFICATE-GUIDE.md). Documented but not implemented: the 980ef09 tree has no page for it, and the ocr-gemini route sends the image or PDF bytes straight to Gemini.

    What happened

    One external dependency is now on the critical path: a missing Gemini key or an unparseable reply gives a 500 with no fallback. The fallback code (src/lib/ocr/llmExtractor.ts with fallbackExtraction, and src/lib/ocrExtract.ts, 615 lines) is used by nothing else at HEAD: only llmExtractor imports ocrExtract, and no file imports llmExtractor. A missing GEMINI_API_KEY is a critical failure in runtimeEnvCheck.ts, so in production the middleware returns 503 for every non-API page. tesseract.js ^6.0.1 is still in package.json and the README still advertises a 'Dual OCR Pipeline: Tesseract.js (local) + Google Gemini'. The route hardcodes gemini-2.0-flash-exp and ignores GEMINI_MODEL; a comment in envValidator.ts says that variable defaults to gemini-2.5-flash, but its getter defaults to gemini-2.0-flash-exp.

    Evidence

    • commit9f5427a ↗removes 'import Tesseract' and the browser OCR block from student/upload/page.tsx; page now calls /api/certificates/ocr-gemini, which does not exist in this commit's tree (commit message is about a delete modal)
    • commit980ef09 ↗deletes my-app/src/app/api/certificates/ocr/route.ts (Tesseract) and adds ocr-gemini/route.ts and SIMPLE-CERTIFICATE-GUIDE.md
    • commit3c573ae ↗the earlier design: Gemini structuring with rule-based fallback plus client-side Tesseract
    • fileroute.ts ↗no fallback path; throws on unparseable response
    • fileREADME.md ↗still claims Tesseract.js in the pipeline
    • fileSIMPLE-CERTIFICATE-GUIDE.md@980ef09 ↗source of the 'complex OCR dependencies that were causing errors' sentence; describes a text-only design that was not built

    How this was checked: git log -S'tesseract' and git show per commit filtered for tesseract lines; read ocr-gemini/route.ts and llmExtractor.ts in full; git grep for importers of LLMExtractor and ocrExtract at HEAD (none outside each other); git show 980ef09:my-app/SIMPLE-CERTIFICATE-GUIDE.md. Attribution: All cited commits are authored by Ujjwaljain16 and carry no Co-Authored-By trailer. The 980ef09 message ('Quality: 5.75/10 -> 9.8/10') and the guide read as AI-assisted output; this cannot be proven from the repo, so disclose generally rather than per commit.

    Reconstructed from code and history · reasons are inferred · commits 9f5427a (2025-10-14, browser stops running Tesseract) and 980ef09 (2025-10-15, server Tesseract route deleted, ocr-gemini route first tracked)

  17. CampusSyncFEATURED

    user_roles row-level security: self-referencing policies replaced by a SECURITY DEFINER function

    PARTIAL

    Context

    The role table (user_roles) needed two rules: a user reads their own row, and admins read and write every row. The first committed version of 001_create_user_roles.sql (1591252, 2025-09-12 01:43) checked admin status with a subquery on user_roles inside each policy on user_roles. POLICY-RECURSION-FIX.md (added in 3c573ae) describes the resulting Postgres error, 'infinite recursion detected in policy', and its cause: the policy queries the table that the policy protects. The next commit, 19 minutes later (d15304d), is titled as a fix for RLS policy recursion.

    Decision

    Final approach in the committed SQL (002_fix_user_roles_policies_v2.sql and 003_fix_recursion_completely.sql, both added in 3c573ae): an is_admin() function declared LANGUAGE plpgsql SECURITY DEFINER, so it reads user_roles without re-entering the policies, used by four admin policies (select, insert, update, delete) plus a 'users read own role' policy. POLICY-RECURSION-FIX.md states the reason: the function bypasses the policies and so breaks the recursion. Before that the table went through four other states: subquery policies (001 as first committed), JWT-claim policies (001 rewritten in d15304d), a second subquery form (002 v1), and RLS switched off (003_disable_rls_temporarily and 007_disable_user_roles_rls). The 007 file states its reasons for switching RLS off: the admin client must read roles without RLS, the policies caused infinite recursion, and role management is handled at the application level.

    Alternatives considered

    • Subquery on user_roles inside each admin policy (EXISTS in 001 as first committed in 1591252; IN subquery in 002_fix_user_roles_policies.sql, d15304d). Both select from user_roles inside a user_roles policy, the pattern POLICY-RECURSION-FIX.md names as the cause of the recursion error. The IN form was replaced eleven days later by the is_admin() version; the repo does not record whether it was run or how it failed.
    • Admin check via the JWT role claim: auth.jwt() ->> 'role' = 'admin' (001 rewritten in place in d15304d). The comment says it was used 'to avoid recursion', but 002 v1 in the same commit went back to a subquery. The repo does not say why the claim approach was dropped; it is not stated that the claim never matched.
    • Disable RLS on user_roles and enforce roles in application code with the service-role client (003_disable_rls_temporarily, 007_disable_user_roles_rls). Labelled temporary in 003_disable ('In production, you should re-enable RLS'). 007 (d0175ff, 2025-09-12) turns it off again with an application-level rationale; 002 v2 and 003_fix (3c573ae, 2025-09-23) re-enable it. The d0175ff commit message says role reads and writes go through the admin client.

    What happened

    Five policy names recur in every version; only the admin check changed. The setup docs tell the reader to paste the migration files into the Supabase SQL editor and the repo has no migration runner, so it cannot say which state was live: 007 (disable) sorts after 002 and 003_fix (enable) by filename. None of this SQL remains at HEAD: 003 and 007 were deleted on 2025-10-24 (3e030b5) and 001, 002, 002 v2 and 008 on 2025-11-07 (4561505), so the recursion history exists only in git. 28 API route files at HEAD still create the service-role client, which bypasses RLS.

    Evidence

    • commitd0175ff ↗2025-09-12: adds 007_disable_user_roles_rls.sql with the three-reason comment, and 008_create_get_user_role_function.sql (SECURITY DEFINER)
    • commit3c573ae ↗2025-09-23: adds 002_fix_user_roles_policies_v2.sql and 003_fix_recursion_completely.sql which re-enable RLS behind is_admin() SECURITY DEFINER (commit message is about OCR; the SQL is bundled into it)
    • file001_create_user_roles.sql@1591252 ↗first version: one own-row policy and four admin policies that each run an EXISTS subquery on user_roles itself
    • file003_disable_rls_temporarily.sql@d15304d ↗'Temporarily disable RLS to fix the recursion issue'
    • commit3e030b5 ↗2025-10-24: removes 003, 007, 013, 020, 029 from the repo as 'unnecessary mitigation/fix scripts'
    • file001_create_user_roles.sql@d15304d ↗001 rewritten 19 minutes later to use auth.jwt() ->> 'role', with the comment 'to avoid recursion'
    • filePOLICY-RECURSION-FIX.md@3c573ae ↗author's own explanation of the recursion and of the SECURITY DEFINER fix
    • file002_fix_user_roles_policies_v2.sql@3c573ae ↗is_admin() as LANGUAGE plpgsql SECURITY DEFINER used by four admin policies
    • commit4561505 ↗2025-11-07: removes 001, 002 v1, 002 v2 and 008 ('old migrations'); no user_roles SQL remains at HEAD

    How this was checked: Ran git log --all --diff-filter=A per file to date each migration; git show <sha>:path to read 001, 002 v1, 002 v2, 003 (both), 007, 008; counted CREATE POLICY per file with grep -ic; git grep -l createSupabaseAdminClient HEAD -- src/app/api gives 28 files. Attribution: All cited commits are authored by Ujjwaljain16 with no Co-Authored-By trailer. POLICY-RECURSION-FIX.md and the d15304d message read as AI-assisted; this cannot be proven from the repo.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit 3c573ae (adds 002_fix_user_roles_policies_v2.sql and 003_fix_recursion_completely.sql); earlier steps in d15304d and d0175ff, both 2025-09-12

  18. BHTTP-1

    A file that shrinks mid-transfer ends in ERROR(3), never a clean END_STREAM

    ADOPTED

    Context

    The RESPONSE announces content-length before the body streams. If the file changes size afterwards, the client has already been told a length. All 41 commits in the public repository are timestamped within 74 minutes (2026-09-22 00:10 to 01:24 +05:30); the server is in 362fc0a (00:10) and the test in 6218eda (00:26), so the history does not show the design evolving over time.

    Decision

    The server opens the file, takes the size from the open handle, and announces that in content-length. If the read comes up short after the RESPONSE is out, the server sends ERROR(3) on stream 0 and closes instead of ending the stream. An empty body ends on the RESPONSE frame itself, and a body that is an exact multiple of 16384 ends on its last full DATA frame; a trailing empty DATA frame is never sent.

    Alternatives considered

    • Stat the file by name, then read it later. docs/ARCHITECTURE.md: the announced length and the bytes sent could then come from different files if the file is replaced in between.
    • End an empty body with an empty DATA frame carrying END_STREAM. docs/ARCHITECTURE.md: one rule with no exceptions is easier to implement correctly than 'send an empty frame unless...'. SPEC 5.7 therefore ends an empty body on the RESPONSE and forbids a trailing empty DATA frame. This concerns framing rather than truncation and could be dropped.

    What happened

    TestFileThatShrinksMidTransferEndsWithAnErrorNotASilentTruncation creates a 48 MiB file, holds the client back so the server blocks mid-file, truncates it to zero, and asserts an ERROR frame with code PROTOCOL_ERROR (3) on stream 0, no END_STREAM on any DATA frame, and a closed connection. TestFileSizesAreFramedCorrectly covers sizes including 0, 1, 16383, 16384, 16385, 32768 and 1 MiB. The conformance fixture includes edge-16383, edge-16384 and edge-16385 files.

    Evidence

    How this was checked: Read fault_test.go, SPEC.md 5.7 and 11.4, docs/ARCHITECTURE.md; go test -count=1 ./internal/server passes. Attribution: All 41 commits are authored by Ujjwaljain16 and none carries a Co-Authored-By trailer. Whether the specification and docs were AI-assisted cannot be determined from the repository; the owner should confirm before publishing.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit 6218eda (fault-injection test); 362fc0a (server)

  19. FlashFlow

    Correct overstated research claims in place, citing the result files that contradict them

    ADOPTED

    Context

    After Stage 16 the README, the Stage 16 write-up and the landing page summarized the research. An audit of the earlier stages found statements that the committed result files did not support. The commit messages call it an independent audit; a related commit (9e24add) describes it as 12 parallel agents, so it was an AI-agent review, and the author made the corrections.

    Decision

    Fix each claim where it appeared and narrow it to what the data supports. For the two ledger-tracked claims (C20, C24) also record the correction in Stage16-ClaimLedger.md, citing the result file, rather than deleting the row. Documents that already scoped the claim correctly were left untouched.

    Alternatives considered

    • Leave the summary wording as published. 016-flagship-results.json contradicts 'Adaptive worst P99 in every seed': in seed 16000 EWMA (4399.88 ms) is worse than Adaptive (4072.11 ms). This is the status quo rather than an option the repo discusses.
    • Rename 'predictor' throughout Stage 15 and 16 documents. b2d9d30 judged this disproportionate: there the word is used in a rank-agreement sense; only the summary claims implying live use were changed.

    What happened

    (1) Per-seed P99 in 016-flagship-results.json: seed 16000 EWMA 4399.88 ms > Adaptive 4072.11; 16001 Adaptive 4732.39 vs EWMA 4717.98; 16002 Adaptive 4853.07 vs EWMA 4725.03, so Adaptive was worst of six in 2 of 3 seeds (recomputed from the JSON). (2) FindPeakEpisodeCongestionOnset needs the run's whole future, so committed backlog is a post hoc statistic, not a live predictor; a code comment now says so. (3) Stage 11's 'Adaptive wins 0/27' counted only sole wins, and round-robin's credit in tied configs came from an unstable sort.Slice; 'EWMA 85-99% in every heterogeneous config' was 57-99% (4 of 18 configs at 57-58%). (4) A claim that Stage 8's sampling rarely produces the severe, no-failure corner was measured at about 6.9%; a test with a 2-15% band guards it. (5) Stage 14's 'FALSIFIER FOUND' pre-dated its dedicated test, since 014A already held the reversal. Ledger claim C24 ('Adaptive is safe') is marked RETIRED.

    Evidence

    How this was checked: Recomputed the per-seed P99 ranking from the flagship result JSON and compared it with the commit message. Read the README as it stood before bbaecb1 (line 199 holds the false sentence), the other four commit messages and the C20/C24 ledger rows, and confirmed the rarity test exists and the package tests pass. The 57-99% range and the 6.9% rate were not recomputed. The cited commits are authored by Ujjwaljain16 and respond to an AI-agent audit; the original overclaims were also the author's.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit bbaecb1 (flagship claim); related fixes b2d9d30, d31d7c9, f367ca8, b90c1d4

  20. FlashFlow

    Add a minimal FIFO capacity model to the simulator after its flat model rewarded overload

    ADOPTED

    Context

    Stage 11's policy map showed EWMA beating Adaptive on mean latency under heterogeneity (15.90 ms vs 27.38 ms). Stage 11's own analysis said this was an artifact: RunWorld gave each target a fixed service time with no queueing, so sending 97.3% of requests to one target (utilization 1.09) cost nothing. The real engine was used as a cross-check because it has genuine concurrency.

    Decision

    Add TargetProfile.Capacity, where 0 or less means infinite (the old behavior, byte-for-byte), and otherwise per-target busy/queue state with deterministic FIFO waiting so CompletionRecord.Latency includes wait time. Add time-varying service time in the same commit. Then re-run the flagship comparison across Capacity 0/1/2/3.

    Alternatives considered

    • Accept the flat-model ranking as evidence that EWMA routes better. Stage11.md says taking it at face value would be an overclaim the model cannot support.
    • Build a general stochastic queueing or network simulator. Stage12.md: the stage 'was never a license to build a general-purpose network simulator'; the commit says it is deterministic discrete-event queueing, not an M/M/c simulation.
    • Use the real HTTP engine to study contention instead. No reason is stated in the repo for not doing so. Related facts: Stage12.md section 12 notes RealEngine never reads Capacity, so real and modelled contention are not like-for-like, and cmd/experiment-011f describes a real run as about 4 s of wall-clock.

    What happened

    The ranking flips only at one capacity: Capacity 0 EWMA 15.90 vs Adaptive 27.38 ms (EWMA wins); Capacity 1 EWMA 131.06 vs 27.93 ms (Adaptive wins); Capacity 2 16.36 vs 27.38 and Capacity 3 16.04 vs 27.38 (EWMA wins). 012-E: Adaptive faster in 12 of 12 traffic seeds, Cliff's delta 1.000, bootstrap CI on the mean gap [88.67, 107.58] ms. A re-run of 012-A and 012-E reproduced all of these (012-E point estimate 98.13 ms). The repo therefore reports that Adaptive's advantage exists only near the stability boundary, not across realistic capacities; a later commit (97285cd) also narrowed a related rho=1 claim to one scenario. Seven hand-computed contention tests were added, and the existing tests passed unchanged.

    Evidence

    How this was checked: Read the queue logic in world.go and the Capacity documentation in scenario.go, and compared the 012A JSON with the Stage12.md table. Re-ran experiments 012a and 012e in a clone (identical means; CI [88.67, 107.58]) and ran the internal/replay tests, which pass. The cited commits (bfec932, 9175a3f, 81a2817) are authored by Ujjwaljain16 and were not the product of an audit agent.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit bfec932 (model); finding in commit 9175a3f; robustness in 81a2817

  21. FlashFlow

    Feed real-engine policies from the proxy's own load and latency trackers

    ADOPTED

    Context

    Stage 11 compared virtual and real engines on the same policies. In the real engine, EWMA and Adaptive sent 100% of 300 requests to one target, and the target changed between fresh processes. The repo's stated goal was that the same policy code runs correctly in both engines.

    Decision

    Construct the ReverseProxy first with a nil selector, build the selector using the proxy's own LoadTracker() and LatencyTracker() through a new Trackers parameter on PolicySpec.New (zero value means build fresh, so the virtual engine is untouched), then attach it with SetSelector. Remove the earlier post-hoc bridge.

    Alternatives considered

    • Read the X-Selected-Edge response header after each request and feed latency back to the selector's own tracker (the Stage 11 fix in 052b894). Fixed latency only; calling OnDispatch/OnComplete back-to-back after the response would net to zero and never show real in-flight load. It was removed because it would double-count.
    • Treat the disagreement as a modeling-fidelity gap and document it. Stage 11 shows it was a wiring bug: policy.New's Instrumentation return value was discarded, so the selector read trackers nothing updated.

    What happened

    The defect was diagnosed from run-to-run variation in which target was locked and from an ablation: giving every request a unique key (removing cache affinity) still produced max_share 1.000. After the fix, the recorded 011-F shows max_share EWMA 0.973 real vs 0.973 virtual and Adaptive 0.500 vs 0.503. A re-run reproduced EWMA (0.973 in both engines, lowest p50 in both) and Adaptive's max_share (0.503), but real Adaptive's p50 was 30.2 ms against 15 ms virtual (the recorded file has 16.5 ms), so p50 agreement for Adaptive depends on timing. Stage11.md states that every earlier RealEngine result for load- or latency-aware policies reflected cold-start tie-breaking, not the intended logic. Two regression tests were added (EWMA prefers the fast real target; least-connections avoids the busy target) and pass.

    Evidence

    How this was checked: Read both commit messages, internal/engine/real.go lines 121-154 and Stage11.md section 10. Re-ran experiment 011f: maximum share and policy ranking agree between the engines, and p50 differs (a re-run gave 30.2 ms for the real engine against 15 ms virtual). Confirmed both regression tests exist and pass. The commits (052b894, b51eac0) are authored by Ujjwaljain16, and the defect was found by the author's own Program F rather than by an audit agent.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit b51eac0 (final fix); defect found in 052b894

  22. FlashFlow

    Run policy experiments on a single-threaded virtual-time event loop, not wall-clock time

    ADOPTED

    Context

    Through Stage 4 every experiment ran on real time and real HTTP. Stage 3's EWMA lock-in experiment gave a different traffic split on each of three real runs of identical targets, and Stage 4 needed a mock clock, a pre-reserved port and an artificial delay just to make timing repeatable. The question for Stage 5 was whether the same configuration and seed could yield the same execution history, cheaply.

    Decision

    Add internal/vtime: a heap-ordered EventQueue keyed on (virtual timestamp, insertion sequence) and an Engine that pops the earliest event, advances a MockClock to exactly that time, runs the callback and repeats, all on one goroutine. Domain code (cache, health registry, selectors) reads time only through the injected clock.Clock, so it ran under the engine unchanged.

    Alternatives considered

    • Keep using real wall-clock time and real HTTP, adding more controls (MockClock, fixed ports, widened race windows) per experiment. Stage 5 notes call the Stage 4 concessions 'the concrete, measured cost' of staying on real time; Experiment 003-D showed goroutine scheduling alone changed the outcome between runs.
    • Migrate all domain logic to virtual time. An audit of every time.Now/Sleep/After/Ticker call in internal/ found the state machines were already clock-injected; only the I/O scheduling layer was wall-clock bound, so only a driving engine was needed.
    • Model simulated concurrency with real goroutines and channels. internal/vtime/queue.go's package comment says this would reintroduce Go scheduler nondeterminism; overlapping requests are instead overlapping start/complete event pairs.

    What happened

    Determinism was tested by repetition: Experiment 005-B ran an identical 9-event scenario 50 times with identical traces, and a re-run gave 50 of 50 identical. An engine test shows that a single event scheduled 10 virtual minutes out is processed in under 100 ms of real time. The cost was a deliberately flat service model: 005-H shows upstream request counts matching the real engine (10/30/100) while virtual p99 stays at 100.0 ms and real p99 rises from 102.9 to 115.4 ms; the real-side figures are read from stored 004-C results, not re-run by 005-H. The missing contention later made Stage 11's flat-model ranking of EWMA over Adaptive an artifact (see contention-model-reverses-ewma-win).

    Evidence

    How this was checked: Read internal/vtime/queue.go, the engine commit message and docs/learning/005-virtual-time.md, and opened the three 003-D result JSONs (edge-a shares 94, 68.33 and 18.17 percent). Ran experiment 005b in a clone (all 50 runs identical) and 005h (numbers matched the recorded file, but 005-H reads the real-engine figures from stored 004-C JSON, so only the virtual half was re-executed). Package tests pass. Authored under Ujjwaljain16; the project was built with an AI coding assistant, so this early code should be treated as possibly AI-assisted (see the project page).

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit 6b0fe78 (engine); rationale in docs/learning/005-virtual-time.md first committed in 504f271

  23. Apache Superset

    Give dynamic native-filter option queries their own opt-in cache TTL setting

    ADOPTED

    Context

    With 'dynamically search all filter values', a native filter's option queries go through /api/v1/chart/data and therefore use the data cache timeout (issue #38219). With row-level security the dropdown kept old values until DATA_CACHE_CONFIG expired; force-refresh showed the right ones. The reporter wanted the shorter FILTER_STATE_CACHE_CONFIG timeout.

    Decision

    The merged change adds `NATIVE_FILTER_OPTIONS_CACHE_TIMEOUT` (default None) to superset/config.py and rewrites `QueryContextProcessor.get_cache_timeout()` with an explicit priority: custom_cache_timeout, then the new setting when the request is a native filter option query, then slice/dataset/database timeout, then DATA_CACHE_CONFIG, then CACHE_DEFAULT_TIMEOUT. The request is recognised by `native_filter_id` being set and `viz_type` starting with `filter_`. Values are compared with `is not None` so a configured 0 is honoured, and the config comment says to use -1, not 0, to disable caching.

    Alternatives considered

    • Use FILTER_STATE_CACHE_CONFIG['CACHE_DEFAULT_TIMEOUT'] for these queries (the PR's first version, still reflected in the PR title). The author's 2026-06-14 comment: these are still chart-data queries on the existing data cache, so a separate cache backend was the wrong abstraction; what is needed is an independent freshness policy.
    • Make the new behaviour the default. sadpandajoe asked to keep the current behaviour as default behind a config option because other code might override timeouts; the author agreed. rusackas twice noted that with default None the stale dropdown is not fixed until an operator sets it, and the author kept None for backward compatibility.
    • Detect filter queries with `not form_data.get('metrics')`. rusackas pointed out that nativeFilters/utils.ts sets `metrics: ['count']` for every native filter request, so the branch could never fire in production; the tests had used a synthetic payload. Detection was changed to native_filter_id plus the `filter_` viz_type prefix, and tests now use realistic payloads.

    What happened

    Operators can set a shorter TTL for filter options than for chart data, and a dataset-level timeout of, say, 24 hours no longer masks it because the new setting is checked before dataset/database timeouts. Nothing changes for existing installs until the setting is configured. The change also fixes the falsy-zero handling in the timeout chain. The author reports an end-to-end check on a real dynamic filter with the setting at 180 seconds: first search a cache miss, an immediate repeat a cache hit, and after deleting the row and waiting out the TTL a miss and rowcount 0 (PR comment of 2026-07-13; not independently re-run). Size +246/-4 lines in 3 files; opened 2026-03-27, merged 2026-07-24.

    Evidence

    • pull requestapache/superset/pull/38910 ↗Thread with sadpandajoe's backward-compatibility request, rusackas's finding about `metrics: ['count']`, and the author's redesign comment; merged by rusackas; merge commit 3ff5dbfe81e68662974ac4d15f5eb2de60ad40e1.
    • issueapache/superset/issues/38219 ↗'Dynamic query filters use data cache instead of filter state cache'; the author's reproduction and root-cause comment of 2026-03-27.
    • commit3ff5dbf ↗Squash merge `fix(native-filters): use FILTER_STATE_CACHE_CONFIG timeout for dynamic filter option queries (#38910)`, 2026-07-24 (GitHub API); the subject still names the abandoned mechanism.
    • filequery_context_processor.py (PR #38910) ↗`get_cache_timeout` docstring with the five-step precedence chain and `_is_native_filter_options_query` docstring explaining why `metrics` is not used.

    How this was checked: Read the PR body, all human comments, reviews and the final diff (gh pr diff 38910), the linked issue, and the merge commit metadata. The 180-second observation is the author's own comment and was not reproduced.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · PR #38910 merged (merge commit 3ff5dbf)

  24. MiniDB

    ARIES-style recovery without CLRs, with LSN equal to the byte offset in the WAL

    PARTIAL

    Context

    The project (a course capstone) includes crash recovery for a from-scratch engine. Full ARIES writes compensation log records (CLRs) while undoing; this project uses a simpler variant.

    Decision

    Recovery runs analysis, redo, undo. Analysis reads the last checkpoint LSN from checkpoint.meta (a fuzzy checkpoint of the active-transaction and dirty page tables, written via temp file and rename) and rebuilds both tables. Redo starts at the smallest recLSN and applies a record only if page.pageLsn < record.lsn. Undo follows each loser's prevLsn chain backwards, deleting inserted rows and restoring the before-image of deleted or updated ones, then appends an ABORT record. No CLRs are written; a code comment says they are 'intentionally' skipped and the README says 'for educational simplicity'. The LSN is the byte offset of the record in wal.log (LogManager.append), so undo can read any record by seeking to its LSN.

    Alternatives considered

    • Full ARIES with compensation log records. README section 8: not implemented 'for educational simplicity'.
    • Scan the log from the start to collect each loser's records (mentioned in a comment in CrashRecovery.ts). The same comment drops the scan: because the LSN is the physical byte offset, records can be read directly by position.

    What happened

    benchmarks/crash_recovery.ts (10,000 committed inserts, 500 uncommitted deletes, crash, recover twice) ends with 10,000 rows, and three crash-matrix tests plus two CrashRecovery unit tests pass. Undo progress is not logged: pages touched by undo get pageLsn = the current log tail, and only an ABORT record marks a loser as finished. Two gaps found in the audit: (1) recovery ends by writing a checkpoint that lists no active transactions; when a copy of the benchmark crashed again right after the first recovery, before pages were flushed, the second recovery found no loser, redid the 500 deletes and ended with 9,500 rows. (2) TxnManager.abort() (commit 8622bae) has a TODO for undo and only writes ABORT and releases locks, so a runtime abort rolls nothing back: an aborted insert and delete both stayed applied, before and after restart.

    Evidence

    • commit7b7c4ce ↗Adds CrashRecovery.ts (analysis/redo/undo, 'We intentionally SKIP Compensation Log Records'), LogManager.ts (LSN = byte offset), CheckpointManager.ts, and the crash tests and benchmark.
    • commit8622bae ↗TxnManager.abort contains 'TODO (Phase 6): Undo all changes' and never got the implementation; the contract in 93f9d9b says abort should undo.
    • testcrash_matrix.test.ts@7b7c4ce ↗Three tests: insert after WAL flush before page flush, delete before commit, commit record flushed but no clean commit.
    • fileCrashRecovery.ts@7b7c4ce ↗Analysis, redo and undo passes; comment "We intentionally SKIP Compensation Log Records (CLRs)"; ABORT appended after undo.

    How this was checked: Read CrashRecovery.ts, LogManager.ts, CheckpointManager.ts, TxnManager.ts and the crash tests; ran benchmarks/crash_recovery.ts and the full suite; wrote a script (decisions/minidb_experiments/zz_exp_abort.ts) that inserts and deletes inside a transaction, calls abort(), and selects: the aborted insert stayed and the aborted delete stayed applied, before and after a restart. Attribution: All cited commits are authored by Ujjwaljain16 (git shortlog: 21 commits, one author). The README lists a two-person team, and git history cannot show which parts each person wrote, so the record should say "commits by Ujjwaljain16; two-person course project". No Co-Authored-By trailers exist in the history; whether AI tools were used cannot be determined from the repository. Comments in CrashRecovery.ts are written as running first-person reasoning, which is consistent with, but not proof of, AI-assisted coding.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit 7b7c4ce (first non-stub CrashRecovery.ts, LogManager.ts, CheckpointManager.ts)

  25. MiniDB

    BufferPool flushes the log up to a page's pageLSN before writing that dirty page (steal / no-force)

    ADOPTED

    Context

    The interface contracts written in the first commits (93f9d9b) specify a steal / no-force buffer policy: dirty pages may be written before their transaction commits, and data pages need not be written at commit. That is only safe if the log record for a change reaches disk before the changed page does.

    Decision

    In BufferPool, every path that writes a dirty frame (eviction in fetchPage and newPage, flushPage, flushAll) first calls logManager.flush(frame.pageLsn) and only then diskManager.writePage. The frame's pageLsn and recLsn are set through setPageLsn(), which also writes the LSN into the page header; recLsn feeds getDirtyPageTable() for checkpoints (added in 7b7c4ce). COMMIT flushes only the log, not data pages (TxnManager.commit).

    Alternatives considered

    • Force data pages at commit / no-steal (the opposite policy). Not discussed anywhere in the repository. The IBufferPool and ILogManager contracts written in commit 93f9d9b specify the opposite: 'Undo rule (steal): log record flushed BEFORE dirty page written' and 'Redo rule (no-force): COMMIT log record flushed; data pages need not be'. No reason for choosing that policy is written down.

    What happened

    A test ('enforces WAL rule on eviction') sets pageLsn 42 on a dirty page, evicts it, and asserts the log manager was flushed to 42. The guarantee is only as good as its callers: fetchPage resets frame.pageLsn to INVALID_LSN when it loads a page, and no log flush happens for a frame whose pageLsn was never set. HeapFile logs and calls setPageLsn only when given an ExecContext. By code inspection, the B+ tree code (src/index) has no log manager or setPageLsn calls and recovery replays only INSERT/UPDATE/DELETE records with a heap RID, so index pages are outside this protocol. The crash tests never write a data page before the crash, so they do not exercise this rule.

    Evidence

    • commit8601d59 ↗BufferPool eviction path contains the 'WAL Rule: flush log up to victim's page LSN before evicting' comment and code.
    • commit93f9d9b ↗src/common/interfaces.ts states the steal / no-force WAL rules and the IBufferPool contract.
    • testBufferPool.test.ts@d075f71 ↗'enforces WAL rule on eviction' asserts lm.flushedLsn === 42 after a dirty page with LSN 42 is evicted.

    How this was checked: Read BufferPool.ts, HeapFile.ts (insertTuple/deleteTuple with ctx) and TxnManager.ts at HEAD; grep of src/index for log/setPageLsn found nothing; ran the jest suite (BufferPool tests pass). Attribution: All cited commits are authored by Ujjwaljain16 (git shortlog: 21 commits, one author). The README lists a two-person team, and git history cannot show which parts each person wrote, so the record should say "commits by Ujjwaljain16; two-person course project". No Co-Authored-By trailers exist in the history; whether AI tools were used cannot be determined from the repository.

    Recorded at the time (a document or commit message states it) · reasons are inferred · commit 8601d59 (implementation); the contract was written in commit 93f9d9b; test in commit d075f71

  26. MiniDB

    Row-level strict 2PL with a wait-for-graph detector that aborts the youngest transaction

    ADOPTED

    Context

    The engine needed serializable isolation between concurrent transactions and had to cope with lock cycles. Locks are taken per RID.

    Decision

    LockManager keeps a queue per RID with S/X modes and FIFO waiters; all locks are released only in TxnManager.commit/abort (strict 2PL). An S-to-X upgrade is queued at the front, and a second concurrent upgrader on the same RID is rejected immediately with 'Deadlock avoided'. DeadlockDetector runs every 100 ms (DEADLOCK_CHECK_INTERVAL_MS), builds a wait-for graph from the lock table, finds a cycle with a colored DFS, and aborts the transaction with the highest TxnId in the cycle (the youngest).

    Alternatives considered

    • MVCC. README 'Known Limitations' says the system uses strict 2PL instead of MVCC, accepting that readers block writers; Architecture.md gives the benefit as 'simple, deterministic serializability'.
    • Let a second S-to-X upgrader wait like any other request. Two upgraders on the same RID each hold S and wait for the other's S, an immediate deadlock; the code and the test 'rejects concurrent upgrades to prevent deadlock' fail fast instead.

    What happened

    Unit tests cover 2- and 3-transaction cycles and assert the youngest is aborted; an integration test asserts t2 (the younger) is the victim. Costs found on inspection and re-run: (1) every scanned row takes a row lock in both Volcano and vectorized scans, which the ablation in the record 'Vectorized engine missed a 5-10x goal' shows to be a large share of scan time; (2) the abort path does not undo the victim's writes (TxnManager.abort is a TODO); (3) benchmarks/strict_2pl_concurrency.ts, run directly with npx tsx, exits without printing the deadlock result because the detector's interval is unref()'d and nothing else keeps the event loop alive; with a keep-alive timer added it prints 'Cycle detected: 4 -> 3. Aborting TxnId=4' and 'Deadlock resolved!'.

    Evidence

    • commit8622bae ↗Adds LockManager, TxnManager, DeadlockDetector and their unit tests (2-txn and 3-txn deadlock, FIFO fairness, upgrade rejection).
    • testdeadlocks.test.ts@7b7c4ce ↗Asserts exactly one of two deadlocked transactions is aborted and that it is the younger; second test asserts the 'Deadlock avoided' error for concurrent upgrades.

    How this was checked: Read LockManager.ts, DeadlockDetector.ts, TxnManager.ts and the tests; ran the jest suite (all concurrency tests pass); ran strict_2pl_concurrency.ts twice as documented (silent exit, exit code 0) and once under a keep-alive wrapper (deadlock resolved, victim TxnId 4). Attribution: All cited commits are authored by Ujjwaljain16 (git shortlog: 21 commits, one author). The README lists a two-person team, and git history cannot show which parts each person wrote, so the record should say "commits by Ujjwaljain16; two-person course project". No Co-Authored-By trailers exist in the history; whether AI tools were used cannot be determined from the repository.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit 8622bae

  27. Lexis AI

    Make book ingestion resumable with per-stage checkpoint tables and job retry

    ADOPTED

    Context

    Ingestion turns an uploaded PDF into a concept graph with one Gemini call per chunk and one per concept pair, on a free tier of about 15 requests per minute (PROJECT_STATUS.md). The first version ran everything in one transaction with commits only at the end (e6d4198^), so any crash or rate-limit failure discarded all work.

    Decision

    Rewrite the orchestrator as stages that persist progress: chunks in batches of 20 (idempotent on conflict), raw_concepts per source chunk (commit every 20 chunk extractions; already-extracted chunks are skipped on resume), canonical concepts written once, relationship_candidates generated once and processed PENDING-first (commit every 50) with an evaluated_pairs table, and graph_build_jobs columns (current_stage, current_offset, retry_count, last_error, next_retry_at). The worker claims jobs with FOR UPDATE SKIP LOCKED and retries a failed job with exponential backoff up to 3 failures.

    Alternatives considered

    • Single transaction over the whole pipeline (e6d4198^). PROJECT_STATUS.md: 'no recovery on crash'; one failure loses every LLM call already paid for.

    What happened

    A crash now loses at most one uncommitted batch (up to 19 chunk extractions or 49 pair evaluations, plus chunks that produced no concepts, which leave no checkpoint row), not the whole run; canonicalisation still commits once at the end, and PROJECT_STATUS.md's 'zero LLM re-calls' is slightly generous. Retry delays are 60 s and 120 s (the code computes 30*2^n with n starting at 1), not the '30s -> 60s -> 120s' in the docs, and the third failure is permanent. The expensive part was not reduced: generate_pairs still returns every ordered pair (n(n-1) LLM calls; its docstring mentions heuristic filtering, the body is a 'mock heuristic'), so at the documented ~15 requests/min 100 concepts would take about 11 hours of relationship calls (arithmetic, not measured). Resumability has no automated test: tests/test_golden_book.py is a skeleton with its assertions commented out.

    Evidence

    • commite6d4198 ↗Rewrites ingestion_orchestrator.py (770 lines changed) and ingestion_worker.py; adds Alembic migration de54f380a89e_ingestion_batching.py.
    • commit09244cb ↗Parent-side version: generate_pairs over all ordered pairs, commits only at the end of process_job.
    • filePROJECT_STATUS.md@bade772 ↗Describes the checkpoint tables, retry policy and the before/after.
    • filegraph_builder.py@bade772 ↗generate_pairs adds (a,b) and (b,a) for every pair of concepts.

    How this was checked: git show e6d4198^:backend/app/services/ingestion_orchestrator.py; read the HEAD orchestrator stages, ingestion_worker.py (_backoff_seconds, mark_for_retry, FOR UPDATE SKIP LOCKED), graph_builder.py and tests/test_golden_book.py; git blame line counts; ran backend tests with a dummy GEMINI_API_KEY: 39 passed.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit e6d4198

    team (4 contributors: Ujjwaljain16 100 of 133 commits incl. 3 as 'Ujjwal Jain', abdurrahmaan11265 27, veekshitha Nelluru 5, Lavya 1; git blame gives ingestion_orchestrator.py 725 lines to Ujjwaljain16 and 12 to abdurrahmaan11265)Link to this record
  28. GitIssue

    Redis Streams with XAUTOCLAIM redelivery, a dead-letter stream and a stale-update guard

    ADOPTED

    Context

    GitHub webhooks are retried, arrive out of order, and a worker can die mid-event. The planning documents (week1.md, v1.md) set the requirements: a secure, lossless, idempotent ingestion pipeline with at-least-once delivery, no duplicated or corrupted rows, and newer updates winning over older ones.

    Decision

    The webhook (HMAC SHA-256, constant-time compare, empty secret rejects everything) does XADD to a stream. The worker uses a consumer group, ACKs only after process_event succeeds (in 9b1a20b the upsert; by 8eda41b also graph mapping), and each loop iteration first calls XAUTOCLAIM for entries idle over 30 s (WORKER_RECLAIM_IDLE_MS). On failure it does not ACK; when XPENDING's times_delivered reaches 5 it XADDs the message to github_events_dlq with a reason and then ACKs. The issue upsert has WHERE issues.updated_at <= EXCLUDED.updated_at, so an older redelivery cannot overwrite newer state.

    Alternatives considered

    • 'Do not ACK -> retry' with a plain XREADGROUP '>' loop (the sketch in week1.md). Inferred, not stated in the repo: reading with '>' only returns entries never delivered to the group, so an un-ACKed entry is not redelivered to a live consumer. The implementation adds XAUTOCLAIM and a delivery-count cutoff; TESTING_WEEK1.md (same commit as week1.md) already expects reclaim and dead-lettering.

    What happened

    Poison messages stop after 5 deliveries and are kept for inspection. In the current code the embedding job and the comment bot run as asyncio.create_task inside process_event, so a failure there is logged but not retried. When reclaimed messages exist the loop skips read_group for that iteration. Every worker test stubs reclaim_stale_messages to return nothing, so XAUTOCLAIM redelivery is never exercised, and the dead-letter test stubs the delivery count; no test runs against a real stream. The updated_at guard has Postgres integration tests that skip when no database is reachable. The 50 worker, webhook-signature and scoring tests pass.

    Evidence

    • commit9b1a20b ↗Adds redis_stream.py (xautoclaim, pending_delivery_count, push_dead_letter) and worker.py loop; store.py has the updated_at guard.
    • commitf252e57 ↗week1.md / v1.md plan for the queue and TESTING_WEEK1.md expectations.
    • fileworker.py@8eda41b ↗Reclaim before read, ACK after success, DLQ at worker_retry_max_attempts.
    • testtest_worker_processing.py@8eda41b ↗test_run_worker_does_not_ack_on_failure and test_run_worker_dead_letters_after_retry_threshold (mocked Redis).

    How this was checked: Read redis_stream.py, worker.py, webhook.py, store.py, config.py; ran `python -m pytest tests/test_worker_processing.py tests/test_webhook_signature.py tests/test_scoring_hybrid.py tests/test_scoring_signal.py` -> 50 passed in 27.95 s. Not run against Redis. Attribution: The design documents (week1.md, v1.md, week2.md, week4.5md) are pasted AI-assistant replies, so the design was planned with an AI assistant. The code commits are by Ujjwaljain16 with no Co-Authored-By trailer.

    Reconstructed from code and history · reasons are inferred · commit 9b1a20b (queue, worker and db modules); tests in 751352a

  29. SSE-Observatory

    Multiplex one EventSource per (url, token) across tabs in a SharedWorker, with a per-tab fallback

    ADOPTED

    Context

    The first version (7ddaced, 2026-02-10) opened one EventSource per browser tab inside useEventStream, so watching the same endpoint in several tabs opened several upstream connections. README.md section 9 (added in 76046c7, a few hours after the worker) states the concern as 'establishing 5 separate SSE connections drains server resources'.

    Decision

    src/workers/sharedSSEWorker.ts owns the EventSource. Connections are keyed by `${url}::${token}`. Each tab is a MessagePort; the worker broadcasts events to all ports, keeps the last 50 events (MAX_BUFFER) and sends them as a 'catchup' message to late joiners, and closes a connection 60 s after its last port leaves (zombie timer). One lastPing time per connection is refreshed by pings from any tab, and a 30 s sweep closes connections whose tabs stopped pinging. useSharedStream.ts is documented as a 'drop-in replacement' for useEventStream and falls back to the per-tab hook if SharedWorker is unavailable or the worker fails.

    Alternatives considered

    • One EventSource per tab (useEventStream, the original design). Kept only as the fallback path. It multiplies upstream connections by the number of tabs; README section 9 gives this as the reason for the change.

    What happened

    Measured end to end (see investigation sse-sharedworker-one-upstream-connection): three tabs on the same url and token produce one upstream request; a different token produces a second. The catch-up buffer is only 50 events, and the connection key includes the token in clear text. Reconnect logic still lives in the React hook (3 attempts, delay 2000 ms x attempt number), not in the worker, contrary to docs/technical/architecture.md, which says the worker handles 'Exponential Backoff reconnects'. The existing e2e test e2e/multiStream.spec.ts checks filter isolation between tabs but does not count upstream connections.

    Evidence

    • commit7ddaced ↗Original per-tab useEventStream with its own EventSource (2026-02-10).
    • commitf899859 ↗Adds src/workers/sharedSSEWorker.ts (connections Map, MAX_BUFFER=50, ZOMBIE_TIMEOUT_MS=60000, catchup message).
    • fileuseSharedStream.ts@0e8c14b ↗Header comment: drop-in replacement, falls back to useEventStream when SharedWorker is unsupported.
    • docREADME.md@0e8c14b ↗States the motivation.
    • testsharedSSEWorker.test.ts@0e8c14b ↗Six tests including zombie cleanup and 'should not duplicate events during reconnect storm'; all pass when run.

    How this was checked: Read sharedSSEWorker.ts, useSharedStream.ts and README section 9; checked the git history of both files; ran vitest (22 files, 103 passed, 1 skipped); ran the multi-tab measurement described in sse-sharedworker-one-upstream-connection. Attribution: The README wording ('exceedingly rare SharedWorker singleton') reads as AI-assisted; no Co-Authored-By trailers exist in the repository, so this cannot be proven.

    Reconstructed from code and history · reasons are stated in the repository · commit f899859 (06:47, one of many per-file commits made within two minutes on 2026-03-07, so the true writing date is unknown but between 2026-02-10 and 2026-03-07); the earlier per-tab hook is 7ddaced (2026-02-10)

  30. SSE-Observatory

    Replace the Vite dev-server SSE middleware with an Express server and Vercel functions

    ADOPTED

    Context

    To connect to third-party SSE endpoints the browser needs a CORS-free relay. On 2026-02-10 (a0c9c33 / 0c9d824) the relay existed twice: as a plugin inside vite.config.ts and as a plain-http server.js on port 3001 (the latter removed in 4a47828 on 2026-03-07). b9abae6 (06:46, 2026-03-07) still grew the Vite middleware to also serve the mock-session endpoints. The project had a vercel.json from its first commit, and the same evening's commits (a43c2a7, 25d610d, a25fbda) were Vercel deployment work; a dev-server plugin does not exist on Vercel.

    Decision

    c503d1a deletes the plugin from vite.config.ts and configures a Vite proxy of '/api' to the backend. 3b6deed adds a new Express server.js (helmet, cors, compression, express-rate-limit 100 requests / 15 min on /api, DNS-based private-IP check, mock session endpoints, static serving of dist/) and api/sse/*.js Vercel functions with the same proxy logic. Vite now only forwards /api/* to the backend.

    Alternatives considered

    • Keep the Vite middleware plugin (the state up to b9abae6). Not stated. Inferred from the same-day commits a43c2a7 ('persistent mock server via vercel kv'), 25d610d ('vercel config') and a25fbda ('optimize for vercel streaming'): a dev-server plugin cannot serve a deployed site.
    • Standalone plain-http server.js from the February version (removed in 4a47828). Replaced by the Express version; no reason recorded.

    What happened

    The proxy now exists in two implementations, server.js and api/sse/index.js, which differ (ticket scheme; server.js applies its origin and private-address checks only in production mode). Local development broke first: c4810f4 and bb1b61d ('solve ECONNREFUSED') changed the Vite target to http://127.0.0.1:3000 and bound Express to 127.0.0.1. The Vite proxy target is hard-coded to port 3000, and every EventSource, including URLs on localhost:4000, is routed through /api/sse (obtainSSEProxyTicket has no local-port exemption, unlike buildSSEProxyUrl). In one run where an unrelated process held port 3000, all 16 runnable e2e tests failed at 'Connected'; with server.js on the target port they passed. src/tests/proxy.test.ts is titled 'Vite Proxy & Architecture Simulation' but exercises a mock http server, not the proxy, and comments in sseProxyUrl.ts still call it the Vite proxy.

    Evidence

    • commitb9abae6 ↗Vite plugin still implements /api/sse and /api/mock/* (06:46, 2026-03-07).
    • commitc503d1a ↗Removes sseProxyPlugin from vite.config.ts; adds proxy '/api' -> :3000.
    • commit3b6deed ↗Adds Express server.js and api/sse/{index,ticket,_ticket}.js.
    • commitc4810f4 ↗Vite target localhost:3000 -> 127.0.0.1:3000 (ECONNREFUSED fix); bb1b61d does the same for the Express bind.
    • commit0c9d824 ↗February version: server.js and Vite plugin both present (a0c9c33 has the plugin in vite.config.ts).

    How this was checked: Read vite.config.ts at 45f7a64, b9abae6, c503d1a and HEAD; read server.js and api/sse/index.js at HEAD; ran playwright with Vite + server.js on :3055 (16 passed) and with the proxy pointing at a foreign process on :3000 (16 failed).

    Reconstructed from code and history · reasons are inferred · commits c503d1a (removes the Vite plugin, adds proxy '/api' to :3000) and 3b6deed (adds Express server.js), 18:21-18:22

  31. Vitest

    mergeTests: a large test suite the maintainer judged did not test what differs from extend()

    PARTIAL

    Context

    After the first review (2026-02-15) the contributor pushed 25 more commits over about three weeks, growing the PR to +1869 lines: 1,260 in test/cli/test/merge-tests.test.ts, 318 in test/core/test/merge-tests.test.ts and 57 in a type test, against 166 added lines of runtime source. At the version the maintainer last reviewed (f3a083c) the two runtime test files held about 1,327 lines; both were rewritten on 2026-03-09 and 03-10 and have had no maintainer response. STATUS: OPEN and unmerged; seven CHANGES_REQUESTED reviews stand.

    Decision

    Per the review record, the contributor tested the feature mainly by volume and variety of scenarios (the 2026-03-10 PR comment lists merge semantics, nested merges, overrides, dependency graphs, diamond inheritance, type inference, validation errors, scaling, lifecycle ordering, worker/file/test scopes, circular dependency detection, self-merge and 50+ fixture stress tests) rather than starting from the cases where mergeTests could differ from repeated extend(). This is the maintainer's characterisation; the contributor did not state the choice.

    Alternatives considered

    • Targeted cases the maintainer listed on 2026-02-21. Partly adopted after the maintainer's last review: the head adds a scope and auto conflict check in mergeTests and tests for a conflicting scope, same-name fixtures with different types, and cross-merge dependencies. The maintainer has not reviewed them. He had asked what happens when both tests define the same fixture, when fixtures depend on merged ones, when the same fixture has different types, and when it has a different scope or other settings.
    • Assert with toMatchInlineSnapshot instead of toContain. Maintainer: AGENTS.md forbids toContain for these tests because it hides stack traces and full output; tests is an array and the whole array should be asserted. The head converted many assertions but still has 15 toContain lines in the CLI test file.

    What happened

    Findings in the review thread (2026-02-21 and 02-22): (a) tests used toContain against output, which AGENTS.md forbids (15 toContain lines remain in the CLI test file at the head, next to 71 inline-snapshot assertions in the PR); (b) three comments on suite.ts say "This is not tested" about implementation branches; (c) "most of the current tests are basically the same test with different values" and only one used dependencies; (d) a test that "will always pass"; (e) a scope-mismatch test failed inside it.extend, not in mergeTests, which had no validation of its own; (f) a suspected `never` type to be checked with expectTypeOf. The final review says the tests document wrong behaviour and calls them AI-generated. On 2026-03-10 the contributor apologised, added scope and auto conflict checks and rewrote both test files; no maintainer reply follows. The thread supports one takeaway: the tests mostly exercised paths extend() already covers, while same-name fixtures across two tests went untested until the maintainer named them.

    Evidence

    How this was checked: Read all 23 inline review comments (author, path, line, commit) and the review states through gh api; read the diff to compute per-file line counts (gh pr view 9662 --json files); counted the source additions in fixture.ts, index.ts, suite.ts and public/index.ts (68 + 1 + 96 + 1 = 166). Did not run the tests. Attribution: The maintainer attributes the tests to AI in the 2026-02-22 review; the contributor's replies do not respond and the commits carry no AI trailer.

    Recorded at the time (a document or commit message states it) · reasons are inferred · PR #9662 (open): maintainer review comments of 2026-02-21 and 2026-02-22 on commits 1ad69af, bd0261a and f3a083c; last maintainer review 2026-02-22T15:31Z.

  32. AgentBrake

    Add per-argument regex allow/deny rules because a tool-name allowlist is too coarse

    ADOPTED

    Context

    ROADMAP_V3.md (8136b6f, 17:39; removed later the same day in 03de458) states the problem: 'AllowedTools: ["read_file"] is too broad. It allows reading /etc/passwd just as easily as /tmp/log.txt.' and calls the fix 'the showstopper feature' for the demo. Implementation followed about 12 minutes later (ed26678, 17:52).

    Decision

    GranularRuleSchema (zod) and GranularAccessPolicy: per tool, a list of rules with deny_if.arguments and allow_if.arguments, each a map argument-name -> regular expression string. A deny_if match, or an allow_if miss, returns the rule's action (default block; the sample config uses kill for secrets). Matching is String(value) against `new RegExp(pattern)` with no flags, evaluated per call.

    Alternatives considered

    • Tool-name allowlist only (AllowedToolsPolicy). Too broad, per the roadmap; kept alongside the new policy.

    What happened

    The policy matches case-sensitive regular expressions against the raw argument text, so it blocks the strings used in the demo but not equivalent variations (different letter case, spacing, path form or recipient domain). The bundled example configuration is illustrative, not a complete filter; this limitation is known and not fixed. The roadmap's example pattern '(?i)DROP TABLE' is not valid JavaScript regex syntax: an uncompilable pattern makes the policy throw, and the proxy forwarded a call when a policy throws (373adcd, Sep 2026, changed this: patterns are compiled and length-checked at startup, and a policy that throws blocks the call). A 'kill' action ends the whole proxy process on the first match.

    Evidence

    • commit8136b6f ↗ROADMAP_V3.md (deleted later the same day in 03de458) with the problem statement and a '(?i)DROP TABLE' example.
    • commited26678 ↗GranularAccessPolicy.ts and GranularRuleSchema.
    • fileenterprise-config.yml@0fc99c8 ↗Case-sensitive patterns such as '.*(DROP|DELETE|TRUNCATE|ALTER|UPDATE).*'.

    How this was checked: Ran the built proxy (node dist/src/proxy/index.js) against a mock tool server for the demo cases and read GranularAccessPolicy.ts and interceptor.ts. Attribution: ROADMAP_V3.md wording ('Why Unique?', 'showstopper feature') reads as AI-assisted; no Co-Authored-By trailers exist in the repository, so it cannot be proven.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit ed26678 (GranularAccessPolicy); design text in commit 8136b6f (ROADMAP_V3.md)

  33. SpentSmart

    Replay a scanned merchant QR's original parameters instead of rebuilding the UPI URL

    ADOPTED

    Context

    The prototype parsed a scanned QR and rebuilt a fresh upi://pay URL with URLSearchParams (pa, pn, am, cu, tn). Merchant QRs carry more parameters (mc, mid, tid, tr, sign, orgid, mode, purpose) and some are signed, so rebuilding drops or re-encodes fields. The docs of the time say some PSPs may reject such URLs.

    Decision

    parseUPIQRCode stores the original query parameters, still percent-encoded, in rawParams. buildUPIUrl treats a QR as a true merchant if it has any of mid/tid/tr/orgid/sign and replays those parameters (joined key=value, with %20 replaced by +). A QR without those signals but with mc, mode=02 or purpose=00 is treated as a 'pseudo-merchant' and rebuilt as a plain P2P URL (pa, pn, am, cu, tn, fresh tr). A non-upi:// QR (EMV / Bharat) is not parsed; the parser comment says UPI apps should handle it raw, but the scanner screen currently shows 'Invalid QR Code' for it.

    Alternatives considered

    • Rebuild every URL with URLSearchParams (constants/upi-config.ts@4057dde). docs/PRODUCTION_UPI_GUIDE.md: signed merchant QRs need exact replay 'byte-for-byte' to keep the sign field; docs/UPI_FIXES_APPLIED.md: mixing encoded and plain values causes double encoding. Both docs were deleted in e3c438b.
    • Parse EMV/Bharat QR in the app. docs/PRODUCTION_UPI_GUIDE.md: UPI apps already validate EMV TLV and CRC, so the app should pass the raw QR through. The pass-through itself is not implemented (the scanner rejects such codes).

    What happened

    According to docs/UPI_FIXES_APPLIED.md, the first version never returned rawParams, so every scan fell back to a rebuilt URL. Limits in the code: the parser drops empty-valued parameters and truncates values containing '=', the exact original query string (rawQuery) is never populated, and the %20 to + replacement means the replay is not byte for byte. The Google Pay path builds its own URL without rawParams. docs/PHONEPE_SECURITY_ANALYSIS.md reports PhonePe still declining a payment started from this app with the original URL; its explanation (PhonePe blocks third-party callers) is the doc's conclusion from manual tests and is not tested in the repo.

    Evidence

    • commitced0a20 ↗services/upi-parser.ts at this commit returns rawParams ('CRITICAL: Actually return rawParams!') and detects non-UPI (Bharat) QRs.
    • commit4057dde ↗Original constants/upi-config.ts (Ayush) builds the URL with URLSearchParams from pa/pn/am/cu/tn only.
    • commitcfb3c73 ↗constants/upi-config.ts: adds the strongMerchantSignals, isTrueMerchant / isPseudoMerchant branches and the %20 to + replacement in buildUPIUrl (2026-01-03).
    • filescanner.tsx@357d2df ↗A null result from parseUPIQRCode (including EMV / Bharat QR) shows the 'Invalid QR Code' alert.

    How this was checked: Read constants/upi-config.ts at 4057dde and HEAD, services/upi-parser.ts at ced0a20 and HEAD, and the two docs at 319b189. Attribution: The rationale docs (UPI_FIXES_APPLIED.md, PRODUCTION_UPI_GUIDE.md, PHONEPE_SECURITY_ANALYSIS.md, added in 319b189 and deleted in e3c438b) read as AI-assistant output addressed to the developer; the code changes are by Ujjwaljain16 with no Co-Authored-By trailer.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · docs/UPI_FIXES_APPLIED.md first committed in 319b189; code in ced0a20 and cfb3c73 (2026-01-03)

  34. SpentSmart

    Custom Kotlin Expo module with <queries> to find and launch UPI apps on Android 11+

    ADOPTED

    Context

    The prototype opened UPI apps with expo-linking: Linking.canOpenURL on a gpay:// URL, then on upi://pay. On Android 11+ package-visibility rules make canOpenURL/app discovery unreliable unless the app declares <queries>, and Expo Go cannot load a custom native module or change the merged manifest.

    Decision

    A local Expo module (modules/upi-intent, Kotlin) declares <queries> for the upi scheme plus 14 explicit UPI package names (815e2a9). The first Kotlin version (b89e812) exposes getUPIApps (a PackageManager query), launchAppByPackage and launchUPI (system chooser); launchUpiDirect, shareTo and shareBase64 (FileProvider, added through a config plugin) came on 2026-02-08 (f2f4e94). services/upi-app-launcher.ts uses the module in two chains: app discovery tries native getUPIApps and falls back to Linking.canOpenURL scheme probing; the chooser fallback tries native launchUPI, then expo-intent-launcher, then Linking. So the app still runs, with reduced behaviour, in Expo Go.

    Alternatives considered

    • expo-linking canOpenURL + gpay:// then upi://pay (the original prototype, services/upi-launcher.ts@4057dde). The manifest comment says explicit package visibility is 'the most reliable way to fix "App not installed" errors on Android 11+', and a code comment in upi-app-launcher.ts says canOpenURL on Android 11+ 'often returns false OR true mistakenly' without <queries>. Scheme probing also cannot enumerate installed UPI apps or target a package. It is kept only as a fallback.
    • expo-intent-launcher only (kept as a fallback tier). Not stated in the repo; inferred: it can start an intent but has no PackageManager query to list installed UPI apps, and the QR image sharing needs a FileProvider and native code.

    What happened

    Needs a development build (QRPaymentGenerator.tsx logs 'UpiIntent native module not available (Expo Go)'). The package list is hard-coded in the manifest, so a new UPI app needs an app update. The iOS side of the module is the unmodified Expo template (a 'hello' function), so this is Android-only. docs/CODEBASE_ANALYSIS.md claims a '100% success rate' for the native path; nothing in the repo measures that, and docs/PHONEPE_SECURITY_ANALYSIS.md (removed later in e3c438b) recorded PhonePe declining a payment started from this app even though the original QR URL was passed unchanged.

    Evidence

    • commit815e2a9 ↗Adds the <queries> manifest with the upi scheme and 14 package names, with the Android 11+ comment.
    • commitc5a0bcd ↗Adds services/upi-app-launcher.ts: native getUPIApps discovery with a Linking.canOpenURL fallback, and a chooser fallback of native launchUPI, then expo-intent-launcher, then Linking.
    • commit4057dde ↗The Ayush-authored prototype (tag lineage) that used Linking.canOpenURL, i.e. the approach that was replaced.
    • fileUpiIntentModule.kt@357d2df ↗AsyncFunctions launchAppByPackage, shareTo, shareBase64, launchUpiDirect, getUPIApps.
    • fileQRPaymentGenerator.tsx@357d2df ↗Comments 'This will fail in Expo Go' / 'requires dev build' show why the fallbacks exist.
    • commitb89e812 ↗First version of UpiIntentModule.kt: launchAppByPackage, getUPIApps and launchUPI.
    • commitf2f4e94 ↗Adds launchUpiDirect, shareTo, shareBase64 and the FileProvider config plugin (app.plugin.js).

    How this was checked: git show 815e2a9, c5a0bcd, 4057dde; read UpiIntentModule.kt, app.plugin.js, ios/UpiIntentModule.swift, services/upi-app-launcher.ts and QRPaymentGenerator.tsx at HEAD; git shortlog on the v2.01 tag lineage shows 59 Ujjwaljain16 / 12 Ayush commits. Attribution: docs/CODEBASE_ANALYSIS.md and docs/PHONEPE_SECURITY_ANALYSIS.md appear to be AI-assistant-written; the manifest comment and the code are the primary evidence. Code commits are by Ujjwaljain16 with no Co-Authored-By trailer; the prototype is by Ayush (author of 12 commits on 2025-12-28/29).

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit 815e2a9 (main lineage; same change is 32b17d2 on the v1.0.0/v2.0.0/v2.01 tag lineage)

  35. migrateDB

    Migration order comes from '-- DEPENDS:' comments, resolved by DFS with cycle detection

    ADOPTED

    Context

    Files are normally applied in filename order. A migration that needs a table created by a later-named file needs an explicit way to say so.

    Decision

    MigrationEngine.loadMigrationFiles sorts files by name (localeCompare), then DependencyResolver.resolveDependencies walks them in that order. It first throws a ValidationError if a listed dependency has no matching file. For each file it then places the names listed in '-- DEPENDS:' or '-- DEPENDS ON:' first, by recursion (post-order depth-first search), tracking a 'resolving' set to detect cycles and a 'resolved' set to skip repeats. A code comment labels this 'Topological sort'. The resolver is loaded with a lazy require inside the engine, marked with an eslint-disable comment. No commit or doc gives the reason for the design; the reason in the context is inferred.

    Alternatives considered

    • Filename ordering only. Needs no metadata but cannot place a migration after a later-named file without renaming. The repo does not record any other approach being weighed.

    What happened

    Reproduced on 1.0.1 with SQLite: files 01_c (DEPENDS: 03_a), 02_b and 03_a were applied as 03_a, 01_c, 02_b. A two-file cycle throws 'Circular dependency detected involving migration: 01_x.sql', naming only the first file reached, not the cycle path. The applied-history check compares database rows (read ORDER BY name) with the name-sorted file list, so DEPENDS reordering does not break later runs. That check does reject any new file that sorts before an applied one, with or without DEPENDS: adding 00_z.sql after the runs above fails with 'Migrations mismatch: found 01_c.sql in the DB but the next script in the filesystem is 00_z.sql.'

    Evidence

    How this was checked: Read the compiled resolver; installed the v1.0.1 tarball with better-sqlite3 in a temporary directory and ran migrate() on temporary directories (order, cycle); output pasted in the numbers of the investigation records. Attribution: Whether it was AI-assisted cannot be determined from the tarball. The reproduction was run during a 2026-09-28 audit.

    Reconstructed from code and history · reasons are inferred · file in package v1.0.0 (npm publish 2025-11-22), byte-identical in v1.0.1; the decision itself predates the first release

  36. Fuze

    Load the embedding model on first use, not at import, after OOM on a 512 MB host

    ADOPTED

    Context

    The Flask backend was deployed on Render, whose free web service has a 512 MB memory limit (stated in 2fe5988). utils/embedding_utils.py built the SentenceTransformer at import time (embedding_model = get_embedding_model()), and UnifiedDataLayer and the recommendations blueprint loaded it again at start-up. The commit messages describe out-of-memory failures at start-up and, in the following commits, gunicorn not binding its port in time for Render's health check; no deploy logs are in the repository.

    Decision

    Move model loading from import time to first use. get_embedding() calls get_embedding_model() on the first request that needs a vector (2fe5988). UnifiedDataLayer.embedding_model became a lazy property and the recommendations blueprint stopped loading it (33419b4). UniversalSemanticMatcher was made lazy and a bare root endpoint was added so gunicorn binds immediately (7fb60ed). Gunicorn workers went from 2 to 1 in Procfile and render.yaml (33419b4). Only the model weights are deferred: embedding_utils.py still imports sentence_transformers at module level. At HEAD the loader is a double-checked singleton behind a threading.RLock, with a DISABLE_EMBEDDINGS switch and a hash-based FallbackEmbeddingModel (384 dimensions, SHA-256 bag of words) when no real model loads.

    Alternatives considered

    • Keep eager loading at import (the previous behaviour). 2fe5988 says it exceeds the 512 MB limit at start-up.
    • Put a smaller model first in the fallback list (paraphrase-MiniLM-L3-v2, about 60 MB, ahead of all-MiniLM-L6-v2). Tried in 2fe5988 and reverted within four minutes in 583c5c4, which restores all-MiniLM-L6-v2 first and says the embeddings are unchanged; lazy loading alone was kept.
    • Keep two gunicorn workers. 33419b4 halves workers to save an estimated 200-300 MB (an estimate in the commit message, not a measurement).

    What happened

    Start-up no longer loads the model weights, and the port can bind quickly. The first embedding request pays the load; a code comment in 2fe9f50 expects '~6-7 seconds the first time', an expectation rather than a measurement. HEAD's get_embedding_model() first tries utils.production_optimizations.get_cached_embedding_model, a module that exists in no commit of any branch, so that branch always falls through its try/except. One test (backend/tests/test_utils.py) checks that repeated calls return the same object; nothing checks that importing the module leaves the model unloaded. Hugging Face Spaces work with 16 GB RAM (docs/DEPLOYMENT.md) began about a week later (first commit 1befff3, 2025-11-23), which removed the memory limit behind this change; that it ended the OOM problem is an inference.

    Evidence

    • commit2fe5988 ↗Message states the 512 MB limit and removes the import-time embedding_model = get_embedding_model().
    • commit33419b4 ↗Lazy property on UnifiedDataLayer; Procfile and render.yaml workers 2 to 1.
    • commit583c5c4 ↗Smaller-model-first ordering reverted within minutes; lazy loading kept.
    • commit7fb60ed ↗UniversalSemanticMatcher made lazy so gunicorn can bind for Render health checks.
    • fileembedding_utils.py@491a221 ↗RLock singleton, DISABLE_EMBEDDINGS, FallbackEmbeddingModel; import of utils.production_optimizations that does not exist.
    • filetest_utils.py@491a221 ↗test_get_embedding_model_singleton: the only test of the singleton; does not check import-time behaviour.

    How this was checked: Read commit messages and diffs of 2fe5988, 583c5c4, 33419b4, 7fb60ed, f360bba; read embedding_utils.py at HEAD; ran git log --all --name-only and searched for production_optimizations (only two .md summary files match, no .py); ran tests/test_embedding_utils.py (passes, 3 tests, none about lazy loading). Attribution: Authored by Ujjwaljain16; none of the cited commits carries a Claude co-author trailer.

    Recorded at the time (a document or commit message states it) · reasons are stated in the repository · commit 2fe5988 (follow-ups 33419b4 and 7fb60ed the same morning)