diff --git a/CLAUDE.md b/CLAUDE.md index 73a4178..0976eae 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -4,10 +4,30 @@ A personal test harness for evaluating **local LLM models** running in **LM Studio** on an **Apple M3 Max MacBook Pro (48 GB unified memory)**. A fixed coding prompt (LFU cache) is fed to each model, the model's Python output is graded against a strict rubric, and the results are accumulated into a persistent JSON store and rendered into a standalone HTML dashboard. -This is **not** a git repo and is **not** an application with a build/run cycle. It is a folder of prompt files, model outputs, and generated artifacts. +This is **a git repo** (remote: `ssh://git@git.itsaygea.com:2222/admin/modelTesting.git`, branch `main`) deployed to Coolify, but it has **no runtime build/run cycle locally** — it's a folder of prompt files, model outputs, a JSON results store, and a Python generator that emits static HTML. ## Folder layout +``` +prompts/ + lfu_cache_prompt.txt # The fixed prompt given to every model (the "exam question") + grading.txt # The grader's instructions / rubric (read at session start) +outputs/ # Raw model outputs (the .py files each model produced) + -.py +data/ + benchmark_history.json # Persistent results store (SOURCE OF TRUTH — committed) +generate_dashboard.py # Reads the JSON → builds dashboard.html + pages/*.html +dashboard.html # GENERATED — gitignored (built on deploy, not committed) +pages/ # GENERATED — per-model detail pages, gitignored +Dockerfile # Coolify: python build stage → nginx static serve +DEPLOY.md # Gitea→Coolify deploy instructions +README.md +``` + +**Git policy:** source only is committed (`prompts/`, `outputs/`, `data/benchmark_history.json`, `generate_dashboard.py`, `Dockerfile`, docs). The generated `dashboard.html` + `pages/` are gitignored — Coolify rebuilds them on each push via the Dockerfile's `python3 generate_dashboard.py` step. If you ever want to commit the built HTML instead, uncomment those lines in `.gitignore`. + +**Remote:** `ssh://git@git.itsaygea.com:2222/admin/modelTesting.git` (SSH key auth as `admin` verified). Push to `main` → Coolify rebuilds. + ``` prompts/ lfu_cache_prompt.txt # The fixed prompt given to every model (the "exam question") @@ -22,7 +42,7 @@ dashboard.html # Generated standalone dashboard (dark-mode, Chart.js v ## The workflow (what to do when the user submits a model's output) -Follow the 4 steps in `prompts/grading.txt`: +Follow the 4 steps in `prompts/grading.txt`. **Captured metrics per model:** `tok_sec`, `total_tokens`, `ttft_sec` (ask the user for these when grading — can't be inferred from files), plus optional `speed_caveat` text (e.g. the Gemma4 GPU-offload note) and `tests_pass` (bool — always run the model's `.py` and record whether it crashes). 1. **Audit** the `.py` file against the 5 pillars (each 0–20 → 0–100 total): - Complexity violations (true O(1) — no heaps, `sorted()`, linear scans) @@ -32,7 +52,7 @@ Follow the 4 steps in `prompts/grading.txt`: - Test coverage integrity (real edge cases vs trivial always-pass asserts) 2. **Extract metrics**: model name, quant, tok/sec, verdict, archetype/best-for, critical bugs, patch code. 3. **Update `data/benchmark_history.json`** — append/update the model's entry (schema in `grading.txt`). Create the file if missing. -4. **Regenerate `dashboard.html`** — full overwrite, standalone single file, Chart.js (bar + radar), leaderboard table, collapsible audit cards, color-coded verdict badges (green 90+, blue/yellow 75–89, red <75). +4. **Regenerate the site** by running `python3 generate_dashboard.py` (reads `data/benchmark_history.json`, overwrites `dashboard.html` + every `pages/.html`). Style is locked: **cyberpunk-terminal** (dark `#0a0a0f` bg, neon cyan/magenta/lime accents, Fira Code/Sans, scanline+grid texture, glow-on-hover score bars). Dashboard = summary stats + Chart.js bar (score vs tok/sec) + radar (top-3 pillars) + leaderboard table linking to per-model detail pages. Each detail page = verdict/score header, metric tiles, pillar bars, "went right/wrong" columns, critical bugs, recommended use, and a refactored patch code block. `tests_pass`, cloud models (no tok/sec), and speed-caveats are all handled in the UI. The exam prompt (`lfu_cache_prompt.txt`) requires a **pure-stdlib Python 3.11+ async LFU cache** with frequency-bucket O(1), dual-layer TTL eviction, ACID-like transactions, and an `async def main()` test suite. Use it as the spec when judging correctness.