Local models only — cloud baseline (DeepSeek) excluded from the speed axis. Bars flagged ⚠ have suspected GPU-offload / inference issues (not representative of the model).
+{findings_panel}
▮ LEADERBOARD
diff --git a/prompts/aygea_test_battery.md b/prompts/aygea_test_battery.md
new file mode 100644
index 0000000..5127db6
--- /dev/null
+++ b/prompts/aygea_test_battery.md
@@ -0,0 +1,113 @@
+# 🎯 The Aygea Test — a multi-prompt battery drawn from your real projects
+
+## Why this exists
+
+The LFU-cache exam (`prompts/lfu_cache_prompt.txt`) is an excellent probe for
+**systems + async correctness** — O(1) data structures, locks, ACID, TTL. But it's
+one narrow axis. It tells you nothing about whether a model can do the work you
+*actually* do every day.
+
+So I surveyed `~/dev` (mewtwo) + `jirachi` and found your real workload clusters into a
+handful of archetypes. This battery mirrors them. Run each model against all five and
+you get a profile — "great at MCP, weak at real-time" — instead of a single score.
+
+## What your projects actually are (the evidence)
+
+From scanning `~/dev` + `jirachi`:
+
+| Archetype | Examples you have | What the code does |
+|---|---|---|
+| **MCP servers** (9!) | joplin-mcp, obsidian-mcp, mySupabaseMCP, project-rag, yt-video-summarizer-mcp, vault-mcp | tool defs, Zod/Pydantic schema validation, stdio/SSE/StreamableHTTP transport, input parsing |
+| **TTS / audio pipelines** | Chatterbox-TTS-Server, aygea-tts-app, vr-to-tts, ffxiv-tts, echokraut-bridge | external HTTP APIs, streaming responses, queueing, device/audio edge cases |
+| **Streaming / chat bridges** | aygeas-multistream, twitch-vod-to-youtube, twitch-discord-notifier, aygeas-chat-overlay | webhooks, OAuth, rate limits, real-time event handling |
+| **Data / API services** | project-rag, mySupabaseMCP, PostgresHA, aygeas-dashboard | SQL, connection pooling, pagination, REST/JSON |
+| **Automation / glue** | fix-tokens, notesCleanup, twitch-discord-notifier | cron-style tasks, idempotency, retries, partial-failure recovery |
+
+**Stack signal:** TypeScript/Node is dominant, Python second, async/await is in ~half of
+all files, try/catch is everywhere, Zod (`z.string`/`z.object`) and Pydantic (`BaseModel`)
+are your validation layer, Docker/compose is standard.
+
+The LFU exam tests *none* of that. These five prompts do.
+
+---
+
+## The 5 prompts
+
+Each is scoped to ~1 file, runnable, and gradable on the same 5-pillar / 100-pt rubric.
+Save each as `prompts/.txt` and feed it to the model.
+
+### 1. `mcp-server.txt` — Build an MCP tool server
+**Probe:** tool/schema correctness, transport, error handling. Your most common project.
+> Write a single-file MCP server (TypeScript `@modelcontextprotocol/sdk` OR Python `mcp`) that
+> exposes 3 tools against a JSONPlaceholder REST API:
+> `get_user(id)`, `list_posts_by_user(user_id, limit)`, `search_posts(query)`.
+> Each tool must: validate inputs with a schema (Zod or Pydantic), return typed results,
+> handle HTTP errors + timeouts gracefully (no silent failures), and not crash on bad input.
+> Run over stdio transport. Include 3 runnable tests (happy path, bad-id 404, malformed input).
+> No external state — pure stdlib + fetch/httpx + the MCP SDK.
+
+### 2. `tts-pipeline.txt` — Audio job queue with backpressure
+**Probe:** async queues, streaming, external-API resilience. Your TTS/audio shape.
+> Single-file async service (Python asyncio or Node) that accepts TTS "jobs" via an async
+> `submit(text, voice)` function, queues them, and processes them through a mock synthesizer
+> (`await mock_synthesize(text) -> bytes`, variable 50-300ms latency). Requirements:
+> bounded concurrency (max 4 in-flight), backpressure (reject when queue > 100), per-job
+> retry-on-failure (max 3, exponential backoff), a `drain()` that awaits all queued jobs,
+> and clean cancellation. Emit job lifecycle events to a callback. Include a 50-job stress
+> test proving the concurrency cap holds and no jobs are dropped on cancel.
+
+### 3. `webhook-bridge.txt` — Twitch/Discord event bridge
+**Probe:** webhook signature verification, rate limiting, idempotency. Your bridge shape.
+> Single-file HTTP service that receives Twitch EventSub webhooks (POST /webhook) and
+> forwards chat events to Discord via a mock webhook. Requirements:
+> HMAC-SHA256 signature verification of every request (reject 401 on mismatch), an
+> in-memory idempotency store keyed by the event id (skip replays within 5 min), a
+> token-bucket rate limiter capping Discord forwards to 5/sec, and graceful handling of
+> Discord 429 (read Retry-After, back off). No framework deps beyond a stdlib http server.
+> Include tests for: valid vs tampered signature, replayed event skipped, rate-limit trigger.
+
+### 4. `data-service.txt` — Paginated query service with a connection pool
+**Probe:** SQL, pooling, pagination, resource cleanup. Your data-service shape.
+> Single-file service wrapping a (mock) Postgres pool exposing:
+> `get_users(page, page_size)`, `get_user_with_posts(user_id)`, and a bulk
+> `relabel_users(id_label_pairs)`. Requirements: a real pooled-connection pattern (checkout /
+> return / leak-proof), parameterized queries (no string-interpolated SQL), correct
+> offset/limit pagination with a total-count, a transactional bulk update that rolls back on
+> any failure, and connection-checkout timeouts. Mock the DB; include tests proving: no
+> connection leak across 100 calls, pagination math, rollback on partial failure.
+
+### 5. `automation-glue.txt` — Idempotent batch job with retries
+**Probe:** idempotency, partial-failure recovery, observability. Your automation shape.
+> Single-file async batch processor that reads a list of "items", calls a flaky external
+> `process(item)` (fails ~20% randomly), and must: be idempotent (re-running resumes from a
+> checkpoint file, never reprocessing done items), retry failures with backoff (max 3),
+> write a progress checkpoint after each item, log a structured JSON summary at the end
+> (succeeded/failed/skipped counts + durations), and exit cleanly on SIGINT (flushing
+> checkpoint). Include a test that kills mid-run and proves resume skips completed items.
+
+---
+
+## How to score (reuse the existing rubric)
+
+Each prompt grades on the same 5 pillars (0–20 each, 100 total):
+1. **Complexity / correctness** — does it actually work, edge cases handled?
+2. **Async / concurrency** — locks, backpressure, cancellation, no races
+3. **Error handling** — no silent failures, retries, timeouts, graceful degradation
+4. **Resource / state safety** — connection leaks, idempotency, checkpoint integrity
+5. **Test integrity** — real assertions vs always-pass; do the tests catch the bugs above?
+
+> Note: pillars 3–5 map cleanly to your repeated patterns (try/catch everywhere,
+> retries, validation, "no silent failures" — your own recurring concern).
+
+## How the dashboard should evolve for this
+
+The current JSON schema assumes one prompt (`exam_prompt`). To support a battery:
+- Add `prompt_id` to each model entry (e.g. `"lfu"`, `"mcp"`, `"tts"`).
+- The leaderboard gets a **prompt filter** (default: show a model's average across all
+ prompts it has run).
+- A new **per-model radar across prompts** shows the profile ("strong at MCP, weak at async
+ pipelines") — the real value of a battery over a single exam.
+
+When you're ready to run these, tell me which prompt + model and I'll wire up grading the
+same way as the LFU set. The generator will need the `prompt_id` field + the filter; I can
+do that in one pass once you have ≥1 result from a second prompt.