
AygeaandClaude
3507d33006
Grade 12-model backlog + fix native-API TTFT capture
Grading (12 new entries, 18→30 total in benchmark_history.json):
- KAT-Coder v2.5 Dev XL: lfu 49 / tts 52 / webhook 72 / rust 36 / automation 60
- Qwen3 Coder 30B: lfu 45 / tts 44 / webhook 62 / rust 54 / automation 58
- Qwen 3.6 35B-A3B uncensored (automation): 46
- Gemma 4 26B-A4B (data): 86 [tests pass]
All "coder" models scored Critical Bugs across prompts — plausible-looking
async code with fatal bugs (broken LFU eviction, in-flight cancel no-op,
submit() raising instead of backpressuring, un-awaited async read-through).
grade_run.py: switch from OpenAI-compat /v1/chat/completions (empty stats)
to native /api/v1/chat — returns full stats incl. time_to_first_token_seconds.
Verified on Gemma-26B (53.5 t/s, ttft 1.13). Two native-API gotchas handled:
input (string) not messages; max_output_tokens not max_tokens (that 400s).
New capture: outputs/gemma-4-26b-a4b-data.py (native-API run).
Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-28 21:17:08 -07:00
..
2026-07-28 21:17:08 -07:00
2026-07-28 18:42:58 -07:00
2026-07-28 18:42:58 -07:00
2026-07-28 18:42:58 -07:00
2026-07-28 18:42:58 -07:00
2026-07-28 16:51:50 -07:00
2026-07-28 14:00:40 -07:00
2026-07-28 14:00:40 -07:00
2026-07-28 21:17:08 -07:00
2026-07-28 14:00:40 -07:00
2026-07-28 19:22:41 -07:00
2026-07-28 19:22:41 -07:00
2026-07-28 19:22:41 -07:00
2026-07-28 19:22:41 -07:00
2026-07-28 19:22:41 -07:00
2026-07-28 14:00:40 -07:00
2026-07-28 19:22:41 -07:00
2026-07-28 19:22:41 -07:00
2026-07-28 19:22:41 -07:00
2026-07-28 19:22:41 -07:00
2026-07-28 19:22:41 -07:00
2026-07-28 16:51:50 -07:00
2026-07-28 16:51:50 -07:00
2026-07-28 14:00:40 -07:00
2026-07-28 17:53:15 -07:00
2026-07-28 17:43:38 -07:00
2026-07-28 18:03:50 -07:00
2026-07-28 14:00:40 -07:00
2026-07-28 20:21:13 -07:00
2026-07-28 14:00:40 -07:00