Files
adminandClaude f0281f2878 Initial benchmark suite: 8 graded models + cyberpunk dashboard generator
- prompts/: LFU cache exam + 5-pillar grading rubric
- outputs/: 8 model .py outputs (local + cloud baseline)
- data/benchmark_history.json: graded results (scores, metrics, bugs, patches)
- generate_dashboard.py: builds dashboard.html + pages/*.html from JSON
- Dockerfile + DEPLOY.md: Gitea→Coolify deploy (build-step, nginx static)
- .gitignore: generated HTML excluded (built on deploy)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-28 14:00:40 -07:00

54 lines
1.9 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 🧪 Local LLM Benchmark Suite
Grades local LLM models (run via **LM Studio** on an Apple M3 Max / 48 GB) on a strict
systems-coding prompt — a pure-stdlib Python **concurrent async LFU cache with TTL
eviction and ACID transactions** — and renders the results into a cyberpunk-terminal
dashboard.
## What's in here
```
prompts/
lfu_cache_prompt.txt # the exam prompt every model gets
grading.txt # the grader's rubric (5 pillars × 20 pts = 100)
outputs/ # raw model .py outputs (named <model>-<quant>.py)
data/
benchmark_history.json # persistent results store (source of truth)
generate_dashboard.py # reads the JSON → builds dashboard.html + pages/*.html
dashboard.html # generated — main leaderboard + charts (gitignored)
pages/ # generated — per-model detail pages (gitignored)
```
## Workflow
1. Feed `prompts/lfu_cache_prompt.txt` to a model in LM Studio.
2. Save its output to `outputs/<model>-<quant>.py`.
3. Grade it (audit the 5 pillars, capture tok/sec + tokens + TTFT), and append its
entry to `data/benchmark_history.json`. See `prompts/grading.txt` for the rubric.
4. Regenerate the site:
```bash
python3 generate_dashboard.py
```
This (re)writes `dashboard.html` and every `pages/<model>.html`.
## Viewing locally
```bash
python3 -m http.server 8000
# open http://localhost:8000/dashboard.html
```
## Deploying (Gitea + Coolify)
The repo holds **source only** (`outputs/`, `data/`, `generate_dashboard.py`, `prompts/`).
The generated `dashboard.html` and `pages/` are gitignored — Coolify runs
`python3 generate_dashboard.py` as a build step, then serves the static files.
See `DEPLOY.md` for the exact Coolify service config.
## Hardware
Apple M3 Max, 48 GB unified memory. Local inference via LM Studio.
A cloud model (DeepSeek V4 Flash) is included as a quality baseline.