Initial benchmark suite: 8 graded models + cyberpunk dashboard generator
- prompts/: LFU cache exam + 5-pillar grading rubric - outputs/: 8 model .py outputs (local + cloud baseline) - data/benchmark_history.json: graded results (scores, metrics, bugs, patches) - generate_dashboard.py: builds dashboard.html + pages/*.html from JSON - Dockerfile + DEPLOY.md: Gitea→Coolify deploy (build-step, nginx static) - .gitignore: generated HTML excluded (built on deploy) Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,53 @@
|
||||
# 🧪 Local LLM Benchmark Suite
|
||||
|
||||
Grades local LLM models (run via **LM Studio** on an Apple M3 Max / 48 GB) on a strict
|
||||
systems-coding prompt — a pure-stdlib Python **concurrent async LFU cache with TTL
|
||||
eviction and ACID transactions** — and renders the results into a cyberpunk-terminal
|
||||
dashboard.
|
||||
|
||||
## What's in here
|
||||
|
||||
```
|
||||
prompts/
|
||||
lfu_cache_prompt.txt # the exam prompt every model gets
|
||||
grading.txt # the grader's rubric (5 pillars × 20 pts = 100)
|
||||
outputs/ # raw model .py outputs (named <model>-<quant>.py)
|
||||
data/
|
||||
benchmark_history.json # persistent results store (source of truth)
|
||||
generate_dashboard.py # reads the JSON → builds dashboard.html + pages/*.html
|
||||
dashboard.html # generated — main leaderboard + charts (gitignored)
|
||||
pages/ # generated — per-model detail pages (gitignored)
|
||||
```
|
||||
|
||||
## Workflow
|
||||
|
||||
1. Feed `prompts/lfu_cache_prompt.txt` to a model in LM Studio.
|
||||
2. Save its output to `outputs/<model>-<quant>.py`.
|
||||
3. Grade it (audit the 5 pillars, capture tok/sec + tokens + TTFT), and append its
|
||||
entry to `data/benchmark_history.json`. See `prompts/grading.txt` for the rubric.
|
||||
4. Regenerate the site:
|
||||
|
||||
```bash
|
||||
python3 generate_dashboard.py
|
||||
```
|
||||
|
||||
This (re)writes `dashboard.html` and every `pages/<model>.html`.
|
||||
|
||||
## Viewing locally
|
||||
|
||||
```bash
|
||||
python3 -m http.server 8000
|
||||
# open http://localhost:8000/dashboard.html
|
||||
```
|
||||
|
||||
## Deploying (Gitea + Coolify)
|
||||
|
||||
The repo holds **source only** (`outputs/`, `data/`, `generate_dashboard.py`, `prompts/`).
|
||||
The generated `dashboard.html` and `pages/` are gitignored — Coolify runs
|
||||
`python3 generate_dashboard.py` as a build step, then serves the static files.
|
||||
See `DEPLOY.md` for the exact Coolify service config.
|
||||
|
||||
## Hardware
|
||||
|
||||
Apple M3 Max, 48 GB unified memory. Local inference via LM Studio.
|
||||
A cloud model (DeepSeek V4 Flash) is included as a quality baseline.
|
||||
Reference in New Issue
Block a user