Files
AygeaandClaude 2f99dd1e35 Grade muse-glimmer-28b full 7-prompt battery — new benchmark leader
7 entries (30→37 total). Muse Glimmer 28B (GGUF) avg 80.7 — the strongest
model in the benchmark, 6/7 prompts Minor Logic Flaws:
  lfu 76 | webhook 81 | automation 89 | rust 85 | data 88 | tts 58 | mcp 88

Standout results:
- rust 85 (KAT 36, Qwen3-Coder 54) — real tokio channels (mpsc::channel, not
  hallucinated mpsc::bounded), two-tier CancellationToken, zero clippy lints;
  one-line E0507 compile fix.
- automation 89 — first model to print a correct summary (98/2/0/100);
  atomic temp+fsync+rename checkpointing.
- data 88 edges out Gemma-26B's 86; mcp 88 sets the bar on a new prompt.
Only weak spot: tts 58 (backpressure raises instead of awaits, like Qwen3-Coder).

Captured via the native /api/v1/chat fix (real tok/sec + TTFT). Slow
deep-thinker: ~17-19 t/s, 5-9 min/prompt, ~5-9k tokens incl. reasoning.

Also gitignore checkpoint.json (automation test runtime artifact).

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-10 13:13:54 -07:00

26 lines
674 B
Plaintext

# --- Generated / build artifacts ---
# Dashboard HTML is generated by generate_dashboard.py from data/benchmark_history.json.
# Keep them OUT of git so the repo only holds source — the site regenerates on deploy.
# (If you'd rather commit the built HTML instead, comment out these lines.)
/dashboard.html
/pages/
# /data/benchmark_history.json # <-- uncomment to keep history local-only too
# --- Scratch / previews / tooling ---
*.png
/playwright-mcp/
/.playwright-mcp/
__pycache__/
*.pyc
.DS_Store
# --- Local env ---
.env
.env.*
*.local
# --- Webhook secret (NEVER commit) ---
.webhook.secret
# automation prompt test artifact (runtime-generated)
checkpoint.json