173 lines
9.6 KiB
Markdown
173 lines
9.6 KiB
Markdown
# discovery-by-llm
|
|
|
|
Music discovery in a chat form: describe the moment, get verified Spotify
|
|
tracks with a reason per track, save the result as a playlist.
|
|
|
|
Spotify removed its recommendation and audio-intelligence endpoints for new
|
|
apps, so in this PoC the LLM fills that role. It interprets the request and
|
|
proposes candidates; Spotify verifies them and supplies the taste profile,
|
|
album art and playback. The eval suite compares the output against bare
|
|
Spotify search on the same questions.
|
|
|
|
> Chronological build log, including what was dropped and why (Dutch):
|
|
> [docs/logboek.md](docs/logboek.md).
|
|
|
|
## Running it
|
|
|
|
**Hosted instance** (credentials in the submission mail): a live deployment
|
|
bound to my own Spotify account, with real personalisation and playlist
|
|
writes. Spotify dev mode caps an app at 5 allowlisted users, so a reviewer
|
|
cannot log in with their own account; I can allowlist a reviewer account on
|
|
request.
|
|
|
|
**Without any keys:**
|
|
|
|
```
|
|
docker compose up --build
|
|
```
|
|
|
|
Demo mode replays recorded API cassettes through the same pipeline and
|
|
streaming path as live mode. The suggestion chips map to the recorded
|
|
scenarios; any other input replays the nearest scenario, with a banner
|
|
naming it. Playlist writes are simulated and labeled.
|
|
|
|
**Live with your own keys:** copy `.env.example` to `.env`, set
|
|
`APP_MODE=live`, a Spotify client id (redirect URI
|
|
`http://127.0.0.1:8888/callback`) and an Anthropic API key, then
|
|
`docker compose up --build` and log in via Spotify.
|
|
|
|
## How it works
|
|
|
|
```
|
|
query ──► call 1: intent + candidates ──► grounding ──► call 2: rerank ──► cards
|
|
(LLM, structured output) (Spotify (LLM, streamed
|
|
▲ search per track)
|
|
taste profile fan-out) │
|
|
(cached, /me/top + saved) │ playlist write
|
|
verified pool
|
|
```
|
|
|
|
- **Call 1** interprets the request (mood, activity, era, language,
|
|
familiarity) and proposes 35 candidate tracks, informed by a compressed
|
|
taste profile fetched once per session.
|
|
- **Grounding** resolves every candidate against Spotify search in a
|
|
bounded concurrent fan-out with early stop, a request deadline and a
|
|
name-to-id cache. Unverified candidates are dropped. Misses and
|
|
mismatches are logged as separate rates.
|
|
- **Call 2** reranks the verified pool and streams one justification per
|
|
track as NDJSON events; cards render as they arrive.
|
|
- Only track ids from the verified pool can reach the user. In live
|
|
testing at familiarity `new`, the model fabricated 12 of 32 candidates;
|
|
the matcher dropped all 12 and the user still received 15 verified
|
|
tracks.
|
|
- An honest limit on discovery depth: "new to you" means "not in your top
|
|
or saved tracks", and the intent prompt favors well-known tracks from
|
|
adjacent scenes because those verify reliably (chasing obscure work is
|
|
where the model starts inventing titles). A listener deep in those
|
|
scenes will recognize a fair share of the results. Going deeper needs a
|
|
repair loop and a better novelty signal; see the last section.
|
|
- Refinement turns rerank the existing pool; turn 2 makes zero Spotify
|
|
calls.
|
|
|
|
## Choices
|
|
|
|
- **Two LLM calls around a deterministic grounding step**, no agent loop:
|
|
predictable latency and every stage testable on its own.
|
|
- **Hand-rolled Spotify client** (~200 lines over httpx). The endpoint
|
|
surface is small and the retry, rate-limit and caching behavior is where
|
|
the engineering lives. Retry policy is per endpoint: bounded Retry-After
|
|
honor on reads, one token refresh, and no retries on playlist writes
|
|
(side effects would be ambiguous).
|
|
- **Ports and adapters.** Demo mode is a second implementation of the same
|
|
ports. The domain model never sees Spotify JSON.
|
|
- **NDJSON over a streamed POST.** The request has a body and EventSource
|
|
cannot POST. Typed events (metadata / track / warning / error / done)
|
|
with a strict client-side phase machine.
|
|
- **Degradation is explicit.** Below the grounding floor the app returns
|
|
what resolved plus a warning; quota exhaustion (QUOTA_EXCEEDED) is
|
|
recognized and never retried; live mode never falls back to fixtures.
|
|
- **Hosted seed session.** The public instance installs a session from an
|
|
escrowed refresh token at startup, so every visitor shares my account,
|
|
including playlist writes. HTTP basic auth at the edge is what makes
|
|
that acceptable. Playlists carry a fixed name prefix for bulk cleanup.
|
|
- **No embeddings.** Search is the only entry point into the catalog; an
|
|
index does not pay for itself inside a PoC.
|
|
|
|
## Performance
|
|
|
|
Time to first card is 20 to 30 seconds, measured through the UI and in the
|
|
eval suite; the spread is LLM latency. Streaming keeps the wait honest:
|
|
metadata arrives early, cards render one by one, progress states name what
|
|
is happening.
|
|
|
|
Under the hood:
|
|
|
|
- Search `limit` is capped at 10 since Feb 2026, so resolving 35 candidates
|
|
requires a fan-out: bounded concurrency, early stop once the rerank pool
|
|
is full, request deadline, partial results on quota exhaustion.
|
|
- Name-to-id resolution cache and a session-scoped taste-profile cache.
|
|
- Anthropic prompt caching between calls.
|
|
- Per-request counters (Spotify calls, cache hits, LLM tokens) in
|
|
structured logs, surfaced in the dev panel.
|
|
|
|
Bare search matches words; it has no notion of mood, era, energy or who is
|
|
asking. The pipeline answers from the listener's actual scenes. Across all
|
|
8 eval scenarios the pipeline delivered 11 to 15 verified tracks from 8 to
|
|
20 distinct artists; on the discovery scenario, all 15 were outside the
|
|
listener's top and saved tracks on Spotify. A sample (full table via
|
|
`eval/run_eval.py --baseline`; the eval checks are plain code assertions,
|
|
no LLM judging):
|
|
|
|
| "Energy for the gym" | Pipeline | Bare search |
|
|
| -------------------- | ---------------------------------------- | -------------------------------------------------------------- |
|
|
| 1 | Baddadan by Chase & Status, Bou, Flowdan | Redbone (with GloRilla) by Lil Baby |
|
|
| 2 | Tough Talk by Chase & Status, Kwengface | Rock That Body by Black Eyed Peas |
|
|
| 3 | Solar System by Sub Focus | Instigator by Inpatient, Ren, Chris Webby |
|
|
| 8 | DJ Turn It Up by Dimension | Gym Power Beat by Ultimate Fitness Playlist Power Workout Trax |
|
|
|
|
## Where to look
|
|
|
|
| Criterion | Where |
|
|
| -------------------------- | ------------------------------------------------------------------------------------------------------ |
|
|
| Code quality and structure | `backend/app/` (domain / ports / adapters / pipeline / api), `AGENTS.md` |
|
|
| AI application | `backend/app/prompts.py`, `backend/app/adapters/anthropic/llm.py`, `backend/app/pipeline/grounding.py` |
|
|
| Spotify integration | `backend/app/adapters/spotify/` (PKCE, retry policy, mapping) |
|
|
| Performance | `backend/app/pipeline/grounding.py` (fan-out, cache, early stop), `backend/app/observability/` |
|
|
| Streaming contract | `backend/app/api/recommendations.py`, `frontend/src/lib/recommendationStream.ts` |
|
|
| Evaluation | `eval/` (runner, scenarios, baseline arm, recorded fixtures) |
|
|
| Pragmatic cuts over time | `docs/logboek.md` |
|
|
|
|
## Method
|
|
|
|
About one hour of preparation (API research, a design spike) and about
|
|
eight hours of hands-on build time across three days. The build was
|
|
agent-assisted: coding agents implemented scoped tracks from briefs I
|
|
wrote, while I owned the architecture, the API contracts, every code
|
|
review, every merge and all live verification against the real APIs. Work
|
|
landed in reviewed batches, which is why commits arrive in bursts. The
|
|
logboek documents the process as it happened; the setup behind it (server,
|
|
Forgejo/GitHub mirror, deployment, how the agents are used) is mapped in
|
|
[docs/workflow.md](docs/workflow.md).
|
|
|
|
## What I cut, and what would come next
|
|
|
|
- A faster model on the intent call, which would cut most of the time to
|
|
first card; the grounding stage already measures the trade-off to watch
|
|
(fabrication rate), so the experiment is cheap to run safely.
|
|
- A guaranteed minimum grounding budget. The request deadline is anchored
|
|
at request start, so an unusually slow intent call can leave grounding
|
|
with no time at all; raising the deadline works around it (env-tunable),
|
|
but grounding deserves a floor of its own regardless of what call 1
|
|
spent.
|
|
- Deeper discovery. A repair loop that re-prompts the model for candidates
|
|
that failed verification would let the prompt chase less obvious work
|
|
without losing the pool, and a novelty signal beyond exact-id exclusion
|
|
(artist-level familiarity, recently played) would push results past
|
|
"famous in an adjacent scene".
|
|
- Refinement turns that re-ground; today turn 2 reranks the existing pool.
|
|
- Chat history persistence. Conversations are client-side only and the
|
|
server keeps no conversation state; sessions and caches are in-process,
|
|
Redis is the named production step.
|
|
- Playback queue control, an i18n framework (copy is already centralized),
|
|
per-stage byte snapshots in the eval, and a semantic index over grounded
|
|
candidates.
|