CORTEX
All projects

RecoveryOS

PROTOTYPE

Recovers failed payments: an LLM may recommend, but only a deterministic policy engine may act.

Razorpay Buildathon, Track 03 · Aug–Sep 2026

THE PROBLEM

Failed payments can often be recovered by retrying at the right time and channel, but a wrong retry costs money. The system separates diagnosis from authority: an AI can suggest what to do, but it cannot move money.

HOW IT WORKS

  • An LLM investigator diagnoses each failed payment and may recommend an action. A deterministic policy and expected-value engine decides whether that action is allowed.
  • An idempotent executor performs the action and replans from the outcome. Work is coordinated with Redis streams and workers on PostgreSQL, with a scheduler lease and reclaim.
  • A Next.js dashboard shows a control tower, per-payment replanning, an audit explorer and experiment results.

THE PIPELINE

  1. Failed paymentArrives as a Razorpay webhook, verified with HMAC-SHA256.
  2. LLM investigatorDiagnoses the failure and may recommend an action. It cannot move money.
  3. Policy and expected-value engineDeterministic. Decides whether the action is allowed.
  4. Idempotent executorIdempotency key plus a PostgreSQL advisory lock. Replans from the outcome.
  5. Audit trail and dashboardControl tower, per-payment replanning and an audit explorer.
Authority flows one way. The model's output is an input to the policy engine and never reaches the executor; a test walks the syntax tree to check that.

SCREENSHOTS

Captured from the project's own repository. Select an image to open it full size.

ENGINEERING EVIDENCE

  • Execution is idempotent: an idempotency key plus a PostgreSQL advisory lock, with a unique-constraint backstop. An integration test races two real threads against it. idempotency.py ↗
  • A test walks the syntax tree to prove that execution code cannot reference the AI recommendation. test_diagnosis_has_no_decision_authority.py ↗
  • Razorpay webhook signatures are verified with HMAC-SHA256 and a constant-time comparison. webhooks.py ↗
  • 555 test functions across unit, integration, evaluation and performance suites, run against real PostgreSQL sessions. CI runs lint, security gates, unit and integration jobs.
  • Multi-seed simulator evaluation: across 5 independent 10,000-payment runs, mean incremental recovered revenue of ₹73,182 per run (95% CI ₹52,919 to ₹93,445), with the payment-level superset property holding on every seed. A script regenerates the result file. multi_seed_runner.py ↗

WHAT ISN'T DONE

  • Benchmarks run against a simulator I wrote, not real payment traffic.
  • Real-model evidence is small: 4 payments and 2 real Gemini recommendations.
  • Most of the measured lift comes from the deterministic engine, not the LLM. AI fusion is off by default.
  • A message that always fails is retried without a cap, opt-out is not re-checked when a message is executed, and the /metrics endpoint is unauthenticated.
  • No live demo (it needs PostgreSQL and Redis). Screenshots are in the repository.
  • Built with an AI coding assistant: 16 of the 139 commits carry a Claude co-author trailer.

NEXT STEPS

Each one comes from a gap listed above. It says what fixing the gap would take; it is not a promise.

  • Cap retries and add a dead-letter queue, so a message that always fails stops being retried.
  • Re-check the customer's opt-out at execution time, not only when the decision is made.
  • Put authentication on the /metrics endpoint.
  • Evaluate against real payment traffic, or a larger set of real model recommendations, instead of only the simulator.

STACK

  • Python
  • FastAPI
  • PostgreSQL
  • Alembic
  • Redis Streams
  • Next.js
  • Gemini
  • Razorpay
  • Prometheus
  • Docker