RecoveryOS
PROTOTYPERecovers failed payments: an LLM may recommend, but only a deterministic policy engine may act.
Razorpay Buildathon, Track 03 · Aug–Sep 2026
THE PROBLEM
Failed payments can often be recovered by retrying at the right time and channel, but a wrong retry costs money. The system separates diagnosis from authority: an AI can suggest what to do, but it cannot move money.
HOW IT WORKS
- An LLM investigator diagnoses each failed payment and may recommend an action. A deterministic policy and expected-value engine decides whether that action is allowed.
- An idempotent executor performs the action and replans from the outcome. Work is coordinated with Redis streams and workers on PostgreSQL, with a scheduler lease and reclaim.
- A Next.js dashboard shows a control tower, per-payment replanning, an audit explorer and experiment results.
THE PIPELINE
- Failed paymentArrives as a Razorpay webhook, verified with HMAC-SHA256.
- LLM investigatorDiagnoses the failure and may recommend an action. It cannot move money.
- Policy and expected-value engineDeterministic. Decides whether the action is allowed.
- Idempotent executorIdempotency key plus a PostgreSQL advisory lock. Replans from the outcome.
- Audit trail and dashboardControl tower, per-payment replanning and an audit explorer.
SCREENSHOTS
Captured from the project's own repository. Select an image to open it full size.
ENGINEERING EVIDENCE
- Execution is idempotent: an idempotency key plus a PostgreSQL advisory lock, with a unique-constraint backstop. An integration test races two real threads against it. idempotency.py ↗
- A test walks the syntax tree to prove that execution code cannot reference the AI recommendation. test_diagnosis_has_no_decision_authority.py ↗
- Razorpay webhook signatures are verified with HMAC-SHA256 and a constant-time comparison. webhooks.py ↗
- 555 test functions across unit, integration, evaluation and performance suites, run against real PostgreSQL sessions. CI runs lint, security gates, unit and integration jobs.
- Multi-seed simulator evaluation: across 5 independent 10,000-payment runs, mean incremental recovered revenue of ₹73,182 per run (95% CI ₹52,919 to ₹93,445), with the payment-level superset property holding on every seed. A script regenerates the result file. multi_seed_runner.py ↗
DECISIONS & INVESTIGATIONS
Four new tests in test_advisory_lock_async.py cover release after CancelledError with an aborted transaction, release after the wrapped block commits the session, lock and unlock never going through the caller's...
Structural tests hold the boundary: an AST walk fails if enqueue_recovery_job or process_job reference recommendation identifiers; the pure argmax and the 11 AI-blind policy rules are scanned for diagnosis and...
The documented headline changed from +42,491.88 rupees (seed 42, single-attempt baseline) to +73,181.78 rupees (5 seeds, compliance-aware baseline).
Incremental recovery versus the compliance-aware baseline was positive in all five seeds, and RecoveryOS's recovered payments were a strict superset of the baseline's each time (baseline_only = 0).
8,820 of val_random's 15,000 rows (58.8%) were verbatim copies of train rows, and 8,739 of test_scenario's 15,000 rows (58.3%) duplicated test_random rows.
WHAT ISN'T DONE
- Benchmarks run against a simulator I wrote, not real payment traffic.
- Real-model evidence is small: 4 payments and 2 real Gemini recommendations.
- Most of the measured lift comes from the deterministic engine, not the LLM. AI fusion is off by default.
- A message that always fails is retried without a cap, opt-out is not re-checked when a message is executed, and the /metrics endpoint is unauthenticated.
- No live demo (it needs PostgreSQL and Redis). Screenshots are in the repository.
- Built with an AI coding assistant: 16 of the 139 commits carry a Claude co-author trailer.
NEXT STEPS
Each one comes from a gap listed above. It says what fixing the gap would take; it is not a promise.
- Cap retries and add a dead-letter queue, so a message that always fails stops being retried.
- Re-check the customer's opt-out at execution time, not only when the decision is made.
- Put authentication on the /metrics endpoint.
- Evaluate against real payment traffic, or a larger set of real model recommendations, instead of only the simulator.
STACK
- Python
- FastAPI
- PostgreSQL
- Alembic
- Redis Streams
- Next.js
- Gemini
- Razorpay
- Prometheus
- Docker




