Grading (12 new entries, 18→30 total in benchmark_history.json): - KAT-Coder v2.5 Dev XL: lfu 49 / tts 52 / webhook 72 / rust 36 / automation 60 - Qwen3 Coder 30B: lfu 45 / tts 44 / webhook 62 / rust 54 / automation 58 - Qwen 3.6 35B-A3B uncensored (automation): 46 - Gemma 4 26B-A4B (data): 86 [tests pass] All "coder" models scored Critical Bugs across prompts — plausible-looking async code with fatal bugs (broken LFU eviction, in-flight cancel no-op, submit() raising instead of backpressuring, un-awaited async read-through). grade_run.py: switch from OpenAI-compat /v1/chat/completions (empty stats) to native /api/v1/chat — returns full stats incl. time_to_first_token_seconds. Verified on Gemma-26B (53.5 t/s, ttft 1.13). Two native-API gotchas handled: input (string) not messages; max_output_tokens not max_tokens (that 400s). New capture: outputs/gemma-4-26b-a4b-data.py (native-API run). Co-Authored-By: Claude <noreply@anthropic.com>
🧪 Local LLM Benchmark Suite
Grades local LLM models (run via LM Studio on an Apple M3 Max / 48 GB) on a strict systems-coding prompt — a pure-stdlib Python concurrent async LFU cache with TTL eviction and ACID transactions — and renders the results into a cyberpunk-terminal dashboard.
What's in here
prompts/
lfu_cache_prompt.txt # the exam prompt every model gets
grading.txt # the grader's rubric (5 pillars × 20 pts = 100)
outputs/ # raw model .py outputs (named <model>-<quant>.py)
data/
benchmark_history.json # persistent results store (source of truth)
generate_dashboard.py # reads the JSON → builds dashboard.html + pages/*.html
dashboard.html # generated — main leaderboard + charts (gitignored)
pages/ # generated — per-model detail pages (gitignored)
Workflow
-
Feed
prompts/lfu_cache_prompt.txtto a model in LM Studio. -
Save its output to
outputs/<model>-<quant>.py. -
Grade it (audit the 5 pillars, capture tok/sec + tokens + TTFT), and append its entry to
data/benchmark_history.json. Seeprompts/grading.txtfor the rubric. -
Regenerate the site:
python3 generate_dashboard.pyThis (re)writes
dashboard.htmland everypages/<model>.html.
Viewing locally
python3 -m http.server 8000
# open http://localhost:8000/dashboard.html
Deploying (Gitea + Coolify)
The repo holds source only (outputs/, data/, generate_dashboard.py, prompts/).
The generated dashboard.html and pages/ are gitignored — Coolify runs
python3 generate_dashboard.py as a build step, then serves the static files.
See DEPLOY.md for the exact Coolify service config.
Hardware
Apple M3 Max, 48 GB unified memory. Local inference via LM Studio. A cloud model (DeepSeek V4 Flash) is included as a quality baseline.