GitIssue
IN PROGRESSDuplicate GitHub issue detection over signed webhooks and pgvector search. In progress: two of four planned stages are built.
Personal project · Mar 2026 · two days so far
THE PROBLEM
Large repositories accumulate duplicate issues. Find likely duplicates when an issue arrives.
HOW IT WORKS
- Webhooks are HMAC-verified and pushed onto a Redis Stream. Workers process them at-least-once and reclaim stalled messages, with a dead-letter stream for failures.
- Issues are embedded (384 dimensions) into PostgreSQL with an HNSW index alongside a full-text index, and scored with semantic, keyword, structural and label signals.
ENGINEERING EVIDENCE
- Redis Streams consumer groups with XAUTOCLAIM reclaim and a dead-letter stream. redis_stream.py ↗
- 115 tests collected across 20 files, including hybrid scoring, worker processing and webhook signatures.
DECISIONS & INVESTIGATIONS
The report (generated 2026-03-17) has 20 labelled suggestions: 1 true positive, 0 false positives, 1 related-not-duplicate and 18 cant_tell, all with labeled_by 'bootstrap-auto', an automated step whose script is not...
Poison messages stop after 5 deliveries and are kept for inspection.
WHAT ISN'T DONE
- Not complete. Ingestion and duplicate detection (weeks 1-2) are built and are what the evidence below covers. The evaluation set is tiny (20 labels), so precision or recall are not quoted.
- The repo also has code for an issue graph and cross-system sync (app/graph, app/sync) from a later, unfinished stage. It has not been evaluated, so it is not described here.
- No CI and not deployed.
NEXT STEPS
Each one comes from a gap listed above. It says what fixing the gap would take; it is not a promise.
- Finish and evaluate the issue-graph and cross-system sync stages, or drop the unfinished code.
- Grow the evaluation set past 20 labels before quoting precision or recall.
- Add CI.
STACK
- Python
- FastAPI
- Redis Streams
- PostgreSQL
- pgvector
- sentence-transformers
- Docker