← all projects
In progress · Two-node home lab

Ask the Library

self-hosted AI reading platform over 50,000 books

  • RAG
  • Vector Search
  • Embeddings
  • Self-hosted LLMs
  • Agentic Pipeline
  • Computer Vision
  • E2E Automation
  • Smoke Testing

What it involved

RAG
Answers are grounded in retrieved passages and stream back with citations from self-hosted models.
Vector Search
Elasticsearch kNN over 7.5M passages, with cold-cache latency cut from 17 s to 0.2 s.
Embeddings
Passage embeddings are computed locally for 50,700 books, with no cloud API.
Self-hosted LLMs
Ollama serves 14B, 3B and 1B models across a two-node home lab.
Agentic Pipeline
An ingestion agent gates new books with OCR heuristics, scan-boilerplate detection and an LLM sniff test.
Computer Vision
Cover scans are scored by Laplacian variance and a vision model, then regenerated with Stable Diffusion.
E2E Automation
Eleven Playwright specs, made deterministic by serving fixtures at the network layer.
Smoke Testing
A live suite against real infrastructure caught GPU contention starving search before users did.

About

A self-hosted RAG platform over a 50,700-book Project Gutenberg mirror — 7.5M passages, FastAPI backend, Elasticsearch kNN vector search on local embeddings, and streaming citation-grounded answers from self-hosted LLMs across a two-node home lab. Zero cloud APIs. Two React front-ends ship against the one API, an agentic ingestion pipeline gates what gets in, and a Stable Diffusion pipeline restores unusable cover scans. The testing on it is layered the way production systems need: a deterministic offline suite for speed, and a live smoke suite against real infrastructure that has already caught degradation before users did.

  • Python
  • FastAPI
  • Elasticsearch
  • kNN Vector Search
  • RAG
  • Ollama
  • React
  • TypeScript
  • Playwright
  • Docker
  • Stable Diffusion

How it's tested 8 checks

  • deterministic Playwright suite runs offline in under five seconds

    Eleven specs covering browse → search → detail → reader, made deterministic by intercepting at the network layer and serving fixtures. No infrastructure required, so it never flakes on a cold GPU or a slow index.

  • live smoke suite catches production degradation before users do

    A second layer pointed at real infrastructure — vector search, LLM generation, the media pipeline — with environment-switchable targets for post-deploy verification. It caught GPU contention starving the search embeddings, which the offline suite by design could not.

  • 50,700 books and 7.5M passages indexed for vector search

    Elasticsearch kNN over locally-computed embeddings — the retrieval layer the answers are grounded in.

  • answers stream from self-hosted models with citations, and no cloud API

    Streaming, citation-grounded Q&A served by Ollama across 14B/3B/1B tiers on a two-node home lab. Nothing is sent to a third-party API.

  • kNN cold-cache latency cut from 17s to 0.2s

    One of several infrastructure faults diagnosed end to end, alongside a CUDA/cuDNN conflict silently breaking GPU inference and request starvation from single-GPU contention.

  • ingestion agent gates new books behind three quality checks

    Public-domain acquisition from the Internet Archive filtered by OCR word-ratio heuristics, perceptual-hash detection of scanner boilerplate, and an LLM sniff test — with on-demand AI cleanup of the OCR text that survives.

  • unusable cover scans are detected and regenerated

    Covers scored by Laplacian variance and a vision model, then regenerated through Stable Diffusion on a local GPU — with a before/after debug page that recomputes the scores client-side so every automated keep-or-replace decision can be audited. It exposed two classifier blind spots that became fixes.

  • two React front-ends ship against one API

    A study-focused reader with themes, a two-page spread, a chapter-grounded chat panel and a vocabulary builder; plus a discovery UI with semantic search, token auth and saved lists.