Commit Graph
6 Commits
Author SHA1 Message Date
adminandClaude b9f45a7c46 Gemma 4 26B-A4B head-to-head vs Qwen (4 prompts, via grade_run.py API)
Ran via tools/grade_run.py against LM Studio (no clipboard). Results:
  TTS:        80 Minor Flaws  PASSES (real N-worker concurrency) <- Qwen 49, didn't parse
  Rust:       72 Minor Flaws  COMPILES CLEAN (0 errs w/ deps)     <- Qwen 50, 7 real errors
  Webhook:    55 Critical     uses forbidden aiohttp (won't run)  <- Qwen 75, passed
  Automation: 48 Critical     SyntaxError (global-after-assign)   <- first run for both

DECISIVE head-to-head: Gemma generalizes where Qwen fails (TTS, Rust),
but Qwen beats it on stdlib-discipline prompts (webhook). The two are
COMPLEMENTARY local offloads, not redundant.

Fixed: grade_run.py extractor (markdown/prose wrapping, multi-fence lang
selection), TTFT-null handling in generator. TTFT capture from LM Studio
API still needs the right stats key (left null + noted).

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-28 18:42:58 -07:00
adminandClaude 241298b410 Grade Qwen 6-bit on webhook prompt: 75/100 Minor Flaws (runs, all 4 tests pass)
Fourth data point on Qwen 3.6 35B-A3B 6-bit MLX:
  LFU cache    82  (runs)
  Webhook     75  (runs, all tests pass) <- best non-LFU result
  TTS pipeline 49  (doesn't parse)
  Rust service 50  (7 compile errors)

Profile sharpens: Qwen 6-bit handles SINGLE-HANDLER logic well
(webhook HMAC/idempotency/rate-limit/429-backoff all correct) but fails
on multi-task orchestration (TTS) and typed/compiled langs (Rust).
Safe offload for HTTP/bridge/verification work; not for pipelines or Rust.

Added webhook prompt_id + pillar set. (Initial paste was mangled -
stripped '=' signs - re-pasted clean.)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-28 18:03:50 -07:00
adminandClaude e3d73de36a Grade Qwen 6-bit on Rust prompt: 50/100 Critical (does not compile, 7 errors)
Third data point on the same model:
  Qwen 3.6 35B-A3B 6-bit MLX:
    LFU cache    82  (runs clean)
    TTS pipeline 49  (doesn't parse)
    Rust service 50  (7 compile errors)

Profile is now sharp: strong on single-file Python data-structure/ACID
work; repeatedly ships non-compiling/non-parsing code on multi-task async
and typed-language prompts. Keep it on Python ACID tasks; do NOT offload
Rust or async-pipeline work.

Verified with cargo 1.94 (mpsc::bounded hallucination, ownership moves,
dead test override, broken remove_watcher). Lang field added for non-Python.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-28 17:53:15 -07:00
adminandClaude dd5724ed77 First multi-prompt result: Qwen 6-bit TTS = 49 (vs 82 LFU) + per-prompt schema
TTS grade for qwen3.6-35b-a3b-6bit-mlx: 49/100 Critical (same model that
scored 82 on LFU). File doesn't parse + bounded-concurrency is fake
(1 worker + inner semaphore = real concurrency 1). Per-task signal:
strong on data-structures, weak on async-pipeline work.

Schema: prompt_id + PILLARS_BY_PROMPT so each entry uses its own 5 pillars.
TODO_submission_tool.md sketches the grade-as-a-tool idea for later.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-28 17:43:38 -07:00
adminandClaude 8d2d5cd562 Grade 3 more models + dashboard v2 layout (quant/format as first-class)
New graded (11 total now):
  gemma4-26b-a4b-8bit-mlx    82  Minor Flaws  (tied top local; delta-based tx freq)
  qwen3.6-27b-8bit-mlx       78  Minor Flaws  (clean; anom. slow generation flagged)
  qwen3-coder-30b-6bit-mlx   50  Critical     (asyncio.Lock used with sync with -> crash)

Dashboard redesign:
  - Bar chart is now the full-width hero row (was cramped half-width)
  - 4 stat tiles squished 2x2 beside the radar up top
  - Quant + Format are dedicated columns in the leaderboard (MLX/GGUF/CLOUD chips)
  - New 'Format & Quant Showdown' panel: groups same-family variants so
    GGUF-vs-MLX and quant-depth comparisons are side by side
  - Bar-chart axis labels now include the quant so duplicate model names
    are distinguishable, with rotation for readability

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-28 16:51:50 -07:00
adminandClaude f0281f2878 Initial benchmark suite: 8 graded models + cyberpunk dashboard generator
- prompts/: LFU cache exam + 5-pillar grading rubric
- outputs/: 8 model .py outputs (local + cloud baseline)
- data/benchmark_history.json: graded results (scores, metrics, bugs, patches)
- generate_dashboard.py: builds dashboard.html + pages/*.html from JSON
- Dockerfile + DEPLOY.md: Gitea→Coolify deploy (build-step, nginx static)
- .gitignore: generated HTML excluded (built on deploy)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-28 14:00:40 -07:00