RR
← All work

2025 · Service

Course RAG

Retrieval over 40k pages of course material with hybrid search and a reranker, served under 500 ms.

Role

Backend + eval

Timeline

Jan — Mar 2025

Stack

FastAPI, Postgres, pgvector, bge-reranker

Problem

Students asked the same questions across lecture notes, slides, and forum threads, and keyword search returned the wrong lecture half the time.

Approach

Chunk by heading, embed with a small open model, and combine BM25 with vector search using reciprocal rank fusion. A cross-encoder reranks the top 40 before the answer is generated with citations.

schema.sql
create table chunks (
  id        bigserial primary key,
  doc_id    text not null,
  heading   text,
  body      text not null,
  embedding vector(768),
  tsv       tsvector generated always as (to_tsvector('english', body)) stored
);
create index on chunks using hnsw (embedding vector_cosine_ops);
create index on chunks using gin (tsv);

Architecture

architecture
  query ──► bm25 ─┐
                  ├─► rrf ──► reranker (top 40) ──► generator ──► answer + cites
  query ──► hnsw ─┘

Results

  • Answer accuracy on a 200-question held-out set rose from 61% to 84%.
  • Switching from a hosted embedding API to a local model cut monthly cost 40%.
  • p95 end-to-end latency: 410 ms at 200 QPS on one 4-vCPU box.

What I'd do next

Query rewriting for follow-up questions and a feedback loop that promotes chunks users mark as helpful.