Cur8
playlists for a moment, grounded in hand-curated crates and measured by an eval harness

The beta sign-up: making playlists is invite-only for now, but anyone can open a shared one. Use the arrows to see one.
What it involved
- RAG
- Each request is grounded in hand-curated crates of real recordings: the model picks crates, and the songs come from them.
- LLM Evaluation
- A two-tier harness: free deterministic checks on every change, then real model runs scored prompt by prompt.
- Golden Datasets
- 24 golden prompts list the right crates, the crates that must never fire, and artists the rules must block or keep.
- Hallucination Metrics
- Every suggested song is looked up in the catalog; any that can't be found is counted as invented.
- Guardrails
- Crate rules are enforced in code, and the harness still counts how often the model tried to break them.
- Model Comparison
- The same golden set runs against models from more than one provider by changing one flag.
- Data Enrichment
- Songs carry MusicBrainz credits and Last.fm tags, and the harness checks the AI's tags against those credits.
- Mobile-first QA
- Designed for phones first, down to the sideways tape-deck player and the compact portrait player.
About
Describe a moment and who it's for, and Cur8 returns a playlist of real songs behind one link that opens in Spotify, Apple Music or YouTube. The AI is only half of it. A retrieval layer I call ragTime grounds each request in hand-curated crates of real recordings, enriched with MusicBrainz facts and Last.fm tags, and crate rules (whose songs may never appear with whose) are enforced in code, not left to the prompt. Retrieval is over curated lists, not embeddings. Whether the model picked the right crate, invented songs or leaked a banned artist is answered by a two-tier eval harness with a golden set, and every run is saved so runs can be compared.
- TypeScript
- Next.js
- LLMs
- RAG
- LLM Evaluation
- MusicBrainz
- Spotify API
- Vercel
How it's tested 7 checks
a two-tier eval harness separates free deterministic checks from paid model runs
The offline tier runs on every change in about five seconds with no AI calls: keyword routing plus rule invariants (a crate never bans its own songs; nothing reached through the taste graph breaks a rule). The model tier makes the real calls and measures crate choice, invented or missing songs, and rule leaks per prompt.
24 golden prompts across 14 crates define what right looks like
Each case lists the crates that would be right, the crates that must never fire, artists the rules must block or keep, and what the model should read from the request (vocals, fame). A no-crate case asserts that nothing fires at all.
invented songs are counted by looking every suggestion up in the catalog
Every song the model suggests is resolved against the streaming catalog. A suggestion that can't be found is counted as invented or missing, so hallucination is a number per run rather than an impression.
rule leaks are blocked in code and still measured
The model can propose a banned artist; the code removes it before the listener sees it. The harness reports how often that happened, so a prompt change that makes the model lean on the guardrail shows up even though users never see the leak.
the same golden set runs against models from more than one provider
The engine is a flag on the harness, so model choice is a comparison on the same cases, not a guess. 13 saved runs so far, timestamped for diffing.
the AI's song tags are checked against MusicBrainz credits
The harness reports how often the tagger's vocal/instrumental label agrees with what each crate implies, and whether MusicBrainz performer credits back it up. It also counts duplicate compositions per crate, so one piece doesn't fill a playlist under five recordings.
36 unit tests cover the taste graph, catalog matching and facets
The rule graph, Spotify matching, song facets, YouTube link refresh and Apple Music key handling each have their own tests, run with Node's test runner.