Fuze
ONGOINGSemantic bookmark manager: pgvector search, a worker pipeline, and recommendations gated by regression tests.
Personal project, ongoing · Jul 2025 – present · 455 commits
THE PROBLEM
Saved bookmarks and posts pile up across tabs and apps and can't be searched by meaning.
HOW IT WORKS
- RQ workers on Redis extract content and compute 384-dimension MiniLM embeddings, which are stored in PostgreSQL with pgvector.
- Search uses HNSW indexes built concurrently through Alembic migrations and exposed as Postgres search functions.
- A rebuilt two-stage recommendation pipeline (ANN candidates, then re-ranking) exists behind feature flags, with shadow evaluation and CI gates. It is not serving production traffic yet.
THE PIPELINE
- ContentBookmarks and web clips.
- RQ workers on RedisExtract content and compute 384-dimension MiniLM embeddings.
- PostgreSQL with pgvectorHNSW indexes built concurrently through Alembic migrations.
- Search functionsExposed as Postgres search functions.
- Recommendations (behind flags)ANN candidates, then re-ranking, with shadow evaluation and CI gates.
ENGINEERING EVIDENCE
- HNSW indexes (m=16, ef_construction=64) built with CREATE INDEX CONCURRENTLY inside an Alembic migration. 0003_hnsw_indexes.py ↗
- A golden-set regression test in CI requires NDCG@10 and MRR of at least 0.85 on four queries. It compares paths that share one engine, so it is a regression guard, not a quality measurement. test_golden_dataset_regression.py ↗
- Reliability primitives: a circuit breaker, a distributed lock, an event and unit-of-work layer, and account lockout. circuit_breaker.py ↗
- About 207 test functions in 65 files, 7 GitHub Actions workflows, 10 migrations and 6 ADRs.
- Local benchmark (400 seeded rows, local PostgreSQL, 30 iterations): ANN search p50 0.64 ms and p99 1.02 ms; embedding generation on a cache miss p50 99 ms.
DECISIONS & INVESTIGATIONS
Defects the 198-test suite did not catch: (1) alembic.ini was never copied into the image, so the migration step in start.sh (added 2026-07-27) had probably failed on every boot behind its '|| echo Warning' fallback;...
At 75 concurrent users about 38% of authenticated requests returned 401 'Token has been revoked' for tokens that were never revoked (approximate figure from gaps.md and the commit message; no raw Locust output is...
The new code is only partly live.
The gate passes with a delta of exactly 0.0000 because both sides run the same class. golden_baseline_v1.json was generated by generate_golden_baseline.py, which runs SmartEngine; ADR-006 calls it the baseline of the...
Start-up no longer loads the model weights, and the port can bind quickly.
WHAT ISN'T DONE
- The hosted backend on Hugging Face Spaces is paused, so the live frontend cannot complete requests. Run it locally with the Quick Start.
- Some optimisation figures in the repository docs have no benchmark behind them, so they are not repeated here.
- In shadow mode the new recommendation pipeline currently returns an empty list, because no unit of work is passed to it. Turning on the cutover flag would serve that empty list.
- Parts were built with an AI coding assistant: some commits carry a Claude co-author trailer, and gaps.md is AI-written.
NEXT STEPS
Each one comes from a gap listed above. It says what fixing the gap would take; it is not a promise.
- Pass a unit of work to the new recommendation pipeline so shadow mode returns real results, before turning on the cutover flag.
- Back the optimisation figures in the repository docs with benchmarks, or remove them.
- Bring the hosted backend back so the live frontend works.
STACK
- Python
- Flask
- PostgreSQL
- pgvector
- Redis
- RQ
- SentenceTransformers
- React
- Alembic
- GitHub Actions