From f0281f28786708d43f2e479c68eb610f86995d83 Mon Sep 17 00:00:00 2001 From: aygea Date: Tue, 28 Jul 2026 14:00:40 -0700 Subject: [PATCH] Initial benchmark suite: 8 graded models + cyberpunk dashboard generator MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - prompts/: LFU cache exam + 5-pillar grading rubric - outputs/: 8 model .py outputs (local + cloud baseline) - data/benchmark_history.json: graded results (scores, metrics, bugs, patches) - generate_dashboard.py: builds dashboard.html + pages/*.html from JSON - Dockerfile + DEPLOY.md: Gitea→Coolify deploy (build-step, nginx static) - .gitignore: generated HTML excluded (built on deploy) Co-Authored-By: Claude --- .gitignore | 20 + CLAUDE.md | 72 +++ DEPLOY.md | 76 +++ Dockerfile | 21 + README.md | 53 ++ data/benchmark_history.json | 269 ++++++++ deepseekv4flash.py | 600 ++++++++++++++++++ generate_dashboard.py | 403 ++++++++++++ ...oser2.5-v1-uncensored-heretic-mxfp8-mlx.py | 262 ++++++++ outputs/gemma-4-31b-qat-gguf.py | 343 ++++++++++ outputs/gemma4-31b-gguf.py | 339 ++++++++++ outputs/kat-coder-v2.5-dev-xl-mlx.py | 550 ++++++++++++++++ outputs/qwen3.6-35b-a3b-4bit-mlx.py | 365 +++++++++++ outputs/qwen3.6-35b-a3b-6bit-mlx.py | 426 +++++++++++++ ...a3b-uncensored-hauhaucs-aggressive-gguf.py | 398 ++++++++++++ prompts/grading.txt | 135 ++++ prompts/lfu_cache_prompt.txt | 31 + 17 files changed, 4363 insertions(+) create mode 100644 .gitignore create mode 100644 CLAUDE.md create mode 100644 DEPLOY.md create mode 100644 Dockerfile create mode 100644 README.md create mode 100644 data/benchmark_history.json create mode 100644 deepseekv4flash.py create mode 100644 generate_dashboard.py create mode 100644 outputs/gemma-4-12b-coder-fable5-composer2.5-v1-uncensored-heretic-mxfp8-mlx.py create mode 100644 outputs/gemma-4-31b-qat-gguf.py create mode 100644 outputs/gemma4-31b-gguf.py create mode 100644 outputs/kat-coder-v2.5-dev-xl-mlx.py create mode 100644 outputs/qwen3.6-35b-a3b-4bit-mlx.py create mode 100644 outputs/qwen3.6-35b-a3b-6bit-mlx.py create mode 100644 outputs/qwen3.6-35b-a3b-uncensored-hauhaucs-aggressive-gguf.py create mode 100644 prompts/grading.txt create mode 100644 prompts/lfu_cache_prompt.txt diff --git a/.gitignore b/.gitignore new file mode 100644 index 0000000..9344ed4 --- /dev/null +++ b/.gitignore @@ -0,0 +1,20 @@ +# --- Generated / build artifacts --- +# Dashboard HTML is generated by generate_dashboard.py from data/benchmark_history.json. +# Keep them OUT of git so the repo only holds source — the site regenerates on deploy. +# (If you'd rather commit the built HTML instead, comment out these lines.) +/dashboard.html +/pages/ +# /data/benchmark_history.json # <-- uncomment to keep history local-only too + +# --- Scratch / previews / tooling --- +*.png +/playwright-mcp/ +/.playwright-mcp/ +__pycache__/ +*.pyc +.DS_Store + +# --- Local env --- +.env +.env.* +*.local diff --git a/CLAUDE.md b/CLAUDE.md new file mode 100644 index 0000000..73a4178 --- /dev/null +++ b/CLAUDE.md @@ -0,0 +1,72 @@ +# CLAUDE.md — Local LLM Benchmark Testing Project + +## What this project is + +A personal test harness for evaluating **local LLM models** running in **LM Studio** on an **Apple M3 Max MacBook Pro (48 GB unified memory)**. A fixed coding prompt (LFU cache) is fed to each model, the model's Python output is graded against a strict rubric, and the results are accumulated into a persistent JSON store and rendered into a standalone HTML dashboard. + +This is **not** a git repo and is **not** an application with a build/run cycle. It is a folder of prompt files, model outputs, and generated artifacts. + +## Folder layout + +``` +prompts/ + lfu_cache_prompt.txt # The fixed prompt given to every model (the "exam question") + grading.txt # The grader's instructions / rubric (read at session start) +outputs/ # Raw model outputs (the .py files each model produced) + -.py +data/ + benchmark_history.json # Persistent results store (created on first grading run) +dashboard.html # Generated standalone dashboard (dark-mode, Chart.js via CDN) +*.py # Loose model outputs in root (e.g. deepseekv4flash.py) — legacy/unsorted +``` + +## The workflow (what to do when the user submits a model's output) + +Follow the 4 steps in `prompts/grading.txt`: + +1. **Audit** the `.py` file against the 5 pillars (each 0–20 → 0–100 total): + - Complexity violations (true O(1) — no heaps, `sorted()`, linear scans) + - Async race conditions / deadlocks (locks held across `sleep`/IO, unsynced shared state) + - Transactional isolation leaks (read-your-own-writes, commit/rollback correctness) + - Memory leaks & edge cases (DLL unlink, `min_freq` updates, `time.monotonic()`, `__slots__`) + - Test coverage integrity (real edge cases vs trivial always-pass asserts) +2. **Extract metrics**: model name, quant, tok/sec, verdict, archetype/best-for, critical bugs, patch code. +3. **Update `data/benchmark_history.json`** — append/update the model's entry (schema in `grading.txt`). Create the file if missing. +4. **Regenerate `dashboard.html`** — full overwrite, standalone single file, Chart.js (bar + radar), leaderboard table, collapsible audit cards, color-coded verdict badges (green 90+, blue/yellow 75–89, red <75). + +The exam prompt (`lfu_cache_prompt.txt`) requires a **pure-stdlib Python 3.11+ async LFU cache** with frequency-bucket O(1), dual-layer TTL eviction, ACID-like transactions, and an `async def main()` test suite. Use it as the spec when judging correctness. + +## Model output — naming convention + +Model output filenames encode the **model + quant/format** so a model is identifiable from the filename alone. The convention the user uses: + +`---.py` + +Lowercase, hyphen-separated, no spaces. Examples seen so far: + +| Filename | Reads as | +|---|---| +| `qwen3.6-35b-a3b-6bit-mlx.py` | Qwen 3.6, 35B-A3B (MoE), 6-bit, MLX format | +| `qwen3.6-35b-a3b-uncensored-hauhaucs-aggressive-gguf.py` | Qwen 3.6 35B-A3B, uncensored fine-tune, GGUF | +| `gemma-4-31b-qat-gguf.py` | Gemma 4 31B, QAT, GGUF | +| `gemma-4-12b-coder-fable5-composer2.5-v1-uncensored-heretic-mxfp8-mlx.py` | merged/model-card-style name, MLX | +| `kat-coder-v2.5-dev-xl-mlx.py` | KAT Coder v2.5 Dev XL, MLX | + +Tokens: `mlx` = Apple MLX format; `gguf` = llama.cpp GGUF; quant suffixes like `q6k`, `6bit`, `4bit`, `qat`, `mxfp8` go in the name. When in doubt about what a filename denotes, **parse it loosely into model + quant** rather than guessing a wrong label — the grading step pulls model/quant from "file name or user input," so either source is valid. + +### When the user pastes raw output instead of giving a file + +If the user pastes a model's output text directly (no file), **save it as a `.py` file in `outputs/`** using the naming convention above before grading. Ask the user for the model name + quant only if it cannot be reasonably inferred from context. Match the existing lowercase-hyphen style. + +## Hardware / runtime context + +- **Machine:** Apple M3 Max, 48 GB unified memory. Local inference only. +- **Inference server:** LM Studio (OpenAI-compatible local endpoint). +- Quants/formats that fit 48 GB comfortably: ~35B dense at 4-bit/6-bit, larger MoE models (only active params loaded). MLX is preferred for Apple Silicon; GGUF via llama.cpp also works. + +## Notes / gotchas + +- `data/` starts **empty** — `benchmark_history.json` does not exist until the first grading run creates it. Don't assume it's there; read defensively, create on write. +- `dashboard.html` is **regenerated** (overwritten), not appended to. The source of truth is the JSON file. +- No build step, no tests to run, no dependencies to install — pure stdlib Python outputs + a static HTML file. +- Treat the loose `.py` files in the project root (like `deepseekv4flash.py`) as unsorted outputs that belong in `outputs/` per the convention. Don't move them unless asked. diff --git a/DEPLOY.md b/DEPLOY.md new file mode 100644 index 0000000..72311f6 --- /dev/null +++ b/DEPLOY.md @@ -0,0 +1,76 @@ +# Deploying to Coolify (via Gitea) + +Repo: `ssh://git@git.itsaygea.com:2222/admin/modelTesting.git` + +The generated `dashboard.html` and `pages/` are **gitignored** — they're built from +`data/benchmark_history.json` by `generate_dashboard.py`. So Coolify must run that script +during build, then serve the result. + +--- + +## Option A — Build step, regenerate on every push *(recommended)* + +The site always reflects the latest `benchmark_history.json`. Push a new grade → re-deploy → site updates. + +### Coolify service setup + +1. **New Resource** → choose your Gitea project `admin/modelTesting`. +2. Choose **"Dockerfile"** as the build pack (we ship a tiny one). + + Add this `Dockerfile` to the repo (already included): + + ```dockerfile + # build stage: regenerate the site from source + FROM python:3.12-slim AS build + WORKDIR /app + COPY . . + RUN python3 generate_dashboard.py + + # serve stage: static files via nginx + FROM nginx:alpine + COPY --from=build /app/dashboard.html /usr/share/nginx/html/dashboard.html + COPY --from=build /app/pages /usr/share/nginx/html/pages + # root path redirects to the dashboard + RUN printf 'location = / { return 302 /dashboard.html; }\n' > /etc/nginx/conf.d/default.conf + EXPOSE 80 + ``` + +3. **Port:** set the published port to **80** (Coolify maps it to your domain). +4. **Domain:** attach your domain. Coolify provisions the HTTPS cert automatically. +5. **Deploy.** Push to `main` → Coolify rebuilds → site updates. + +### Why a Dockerfile + +Coolify's "Nixpacks / static" presets can run a build command, but the Dockerfile path is +the most predictable: explicit Python build step + nginx static serving, no surprises. +The image is ~30 MB. + +--- + +## Option B — Pure static, commit the built HTML *(simplest, no build)* + +If you'd rather not build on the server: + +1. Comment out `/dashboard.html` and `/pages/` in `.gitignore`. +2. Run `python3 generate_dashboard.py` locally, then `git add` + commit the HTML. +3. In Coolify: **New Resource** → the repo → build pack **"Static Site"** (or Nixpacks with + `OUTPUT_DIRECTORY=/`), publish dir `/`, port 80. + +Downside: you must rebuild + commit locally every time you add a model. Use Option A unless +your Coolify can't run Docker. + +--- + +## Verifying after deploy + +Open the domain → you should land on the dashboard. If you see a blank page, check the +Coolify build logs for `generate_dashboard.py` errors (usually a malformed +`data/benchmark_history.json`). + +## Local preview before pushing + +```bash +python3 generate_dashboard.py +python3 -m http.server 8000 +# open http://localhost:8000/dashboard.html +``` diff --git a/Dockerfile b/Dockerfile new file mode 100644 index 0000000..5e6b948 --- /dev/null +++ b/Dockerfile @@ -0,0 +1,21 @@ +# syntax=docker/dockerfile:1 +# Build stage: regenerate the static site from source data. +FROM python:3.12-slim AS build +WORKDIR /app +COPY . . +RUN python3 generate_dashboard.py + +# Serve stage: nginx serves the generated static files. +FROM nginx:alpine +# Replace the default nginx server block so root redirects to the dashboard +RUN printf '%s\n' \ + 'server {' \ + ' listen 80;' \ + ' server_name _;' \ + ' root /usr/share/nginx/html;' \ + ' index dashboard.html;' \ + ' location = / { return 302 /dashboard.html; }' \ + '}' > /etc/nginx/conf.d/default.conf +COPY --from=build /app/dashboard.html /usr/share/nginx/html/dashboard.html +COPY --from=build /app/pages /usr/share/nginx/html/pages +EXPOSE 80 diff --git a/README.md b/README.md new file mode 100644 index 0000000..7bc32e0 --- /dev/null +++ b/README.md @@ -0,0 +1,53 @@ +# 🧪 Local LLM Benchmark Suite + +Grades local LLM models (run via **LM Studio** on an Apple M3 Max / 48 GB) on a strict +systems-coding prompt — a pure-stdlib Python **concurrent async LFU cache with TTL +eviction and ACID transactions** — and renders the results into a cyberpunk-terminal +dashboard. + +## What's in here + +``` +prompts/ + lfu_cache_prompt.txt # the exam prompt every model gets + grading.txt # the grader's rubric (5 pillars × 20 pts = 100) +outputs/ # raw model .py outputs (named -.py) +data/ + benchmark_history.json # persistent results store (source of truth) +generate_dashboard.py # reads the JSON → builds dashboard.html + pages/*.html +dashboard.html # generated — main leaderboard + charts (gitignored) +pages/ # generated — per-model detail pages (gitignored) +``` + +## Workflow + +1. Feed `prompts/lfu_cache_prompt.txt` to a model in LM Studio. +2. Save its output to `outputs/-.py`. +3. Grade it (audit the 5 pillars, capture tok/sec + tokens + TTFT), and append its + entry to `data/benchmark_history.json`. See `prompts/grading.txt` for the rubric. +4. Regenerate the site: + + ```bash + python3 generate_dashboard.py + ``` + + This (re)writes `dashboard.html` and every `pages/.html`. + +## Viewing locally + +```bash +python3 -m http.server 8000 +# open http://localhost:8000/dashboard.html +``` + +## Deploying (Gitea + Coolify) + +The repo holds **source only** (`outputs/`, `data/`, `generate_dashboard.py`, `prompts/`). +The generated `dashboard.html` and `pages/` are gitignored — Coolify runs +`python3 generate_dashboard.py` as a build step, then serves the static files. +See `DEPLOY.md` for the exact Coolify service config. + +## Hardware + +Apple M3 Max, 48 GB unified memory. Local inference via LM Studio. +A cloud model (DeepSeek V4 Flash) is included as a quality baseline. diff --git a/data/benchmark_history.json b/data/benchmark_history.json new file mode 100644 index 0000000..cd8303a --- /dev/null +++ b/data/benchmark_history.json @@ -0,0 +1,269 @@ +{ + "meta": { + "project": "Local LLM Benchmark Suite — LFU Cache & ACID Audit", + "machine": "Apple M3 Max, 48GB unified memory, LM Studio", + "exam_prompt": "prompts/lfu_cache_prompt.txt", + "grading_rubric": "prompts/grading.txt", + "pillars": ["complexity", "concurrency", "isolation", "memory_edge_cases", "test_integrity"], + "max_per_pillar": 20, + "schema_version": 1 + }, + "models": [ + { + "id": "qwen3.6-35b-a3b-6bit-mlx", + "timestamp": "2026-07-28T13:00:00Z", + "model_name": "Qwen 3.6 35B-A3B", + "quant": "6-bit MLX", + "format": "mlx", + "tok_sec": 68.86, + "total_tokens": 14883, + "ttft_sec": 0.94, + "filename": "outputs/qwen3.6-35b-a3b-6bit-mlx.py", + "tests_pass": true, + "total_score": 82, + "breakdown": { + "complexity": 16, + "concurrency": 16, + "isolation": 17, + "memory_edge_cases": 16, + "test_integrity": 17 + }, + "verdict": "Minor Logic Flaws", + "best_for": "Solid daily-driver scaffolding for ACID/async patterns — produces runnable, well-structured code, but needs a human pass for __slots__, monotonic clocks, and lock granularity before production.", + "critical_bugs": [ + "No __slots__ declared on Node/Transaction/LFUCache — rubric explicitly required it for memory efficiency.", + "Uses time.time() (system clock) throughout instead of time.monotonic() — NTP adjustments can cause premature/incorrect TTL eviction.", + "_cleanup_freq_lists() performs a hidden O(F) scan (iterates all freq tiers + min()), called after every eviction, background batch, AND inside commit() — violates the strict O(1) requirement.", + "Transaction commit holds the single cache lock across all write/delete/bump loops + cleanup — coarse-grained, blocks all readers for the whole commit window; no fine-grained locking.", + "Lost-update risk: commit applies tx-local writes without any MVCC/version check, so a key modified by the background evictor or another committer between tx.get() and commit() is overwritten blindly.", + "Tests dodge hard cases: 50-task stress uses unique keys with capacity 100, so no eviction-under-contention ever happens; no test for rollback-after-partial-application or mid-commit read isolation." + ], + "patch_code": "# FIX 1: Add __slots__ for memory efficiency\n@dataclass\nclass Node:\n __slots__ = ('key', 'value', 'freq', 'expires_at', 'prev', 'next')\n key: Any\n value: Any\n freq: int\n expires_at: Optional[float]\n prev: Optional['Node'] = None\n next: Optional['Node'] = None\n\n# FIX 2: Use monotonic clock everywhere (get/put/_add_node/commit)\n# time.time() -> time.monotonic()\n# e.g.\nexpires_at = time.monotonic() + ttl_seconds if ttl_seconds else None\n\n# FIX 3: Make _cleanup_freq_lists O(1) — bump min_freq incrementally\n# instead of recomputing min() across all tiers:\n# In _update_freq, when emptying the min_freq bucket, only bump min_freq\n# if you're evicting from it; otherwise leave it. Delete the global\n# min(self.freq_map.keys()) scan. For background sweeps, prune empty\n# buckets lazily on next _evict() rather than scanning proactively.\n\n# FIX 4: Shrink commit critical section — apply writes into a staging\n# structure under the lock, then release; or use per-bucket locks so\n# readers on unrelated keys aren't blocked.\n\n# FIX 5: Add MVCC version to Node; in commit, raise/abort if the\n# stored version != the version seen at tx.get() time (lost-update detect)." + }, + { + "id": "qwen3.6-35b-a3b-4bit-mlx", + "timestamp": "2026-07-28T13:05:00Z", + "model_name": "Qwen 3.6 35B-A3B", + "quant": "4-bit MLX", + "format": "mlx", + "tok_sec": 83.31, + "total_tokens": 13278, + "ttft_sec": 0.73, + "filename": "outputs/qwen3.6-35b-a3b-4bit-mlx.py", + "tests_pass": false, + "total_score": 57, + "breakdown": { + "complexity": 15, + "concurrency": 15, + "isolation": 13, + "memory_edge_cases": 10, + "test_integrity": 4 + }, + "verdict": "Critical Bugs", + "best_for": "Not recommended for systems code as-is. The 4-bit quant degrades logic sharply vs the 6-bit sibling (82->57) — usable only for boilerplate/scaffolding drafts that a human will heavily rewrite.", + "critical_bugs": [ + "FATAL: _evict() double-removes nodes — pop() already calls remove() (nulling node.prev/next), then _remove_node() calls remove() AGAIN -> AttributeError: 'NoneType' on first eviction. The cache cannot survive reaching capacity.", + "Test suite never executes: test_lfu_eviction crashes at the first eviction, so the 'All tests passed' message is unreachable and assertions are effectively unverified.", + "Background _eviction_loop materializes list(self.key_to_node.keys())[:50] every sweep — an O(N) linear scan, forbidden by the strict O(1) requirement.", + "_FreqList.pop() has no empty-guard — calling pop() on an empty list dereferences self.head.next (the dummy tail) and corrupts the DLL.", + "Transaction _apply_put applies the buffered value into the EXACT original_node captured at tx.put() time; if the global cache evicted/relocated that node between put and commit, you mutate a stale/dangling node (no MVCC/version check).", + "Commit is not atomic across exceptions: a crash mid-_apply loop leaves half-applied global state with no rollback.", + "Uses time.time() (system clock) throughout instead of time.monotonic() — NTP jumps corrupt TTL eviction." + ], + "patch_code": "# FIX 1 (the crash): _evict double-removes. pop() already unlinks,\n# so do NOT call _remove_node on a popped node. Either:\n# (a) pop and then only delete the key_map entry + min_freq bookkeeping:\ndef _evict(self):\n if not self.freq_to_list:\n return\n evict_list = self.freq_to_list[self.min_freq]\n if evict_list.size == 0: # guard against empty\n del self.freq_to_list[self.min_freq]\n return\n node = evict_list.pop() # pop() unlinks + nulls prev/next\n del self.key_to_node[node.key] # DON'T call _remove_node again\n if self.freq_to_list[self.min_freq].size == 0:\n del self.freq_to_list[self.min_freq]\n self.min_freq += 1\n\n# FIX 2: _FreqList.pop empty-guard\ndef pop(self) -> _Node:\n if self.size == 0:\n raise IndexError('pop from empty _FreqList')\n node = self.head.next\n self.remove(node)\n return node\n\n# FIX 3: kill the O(N) scan in background sweep — maintain a separate\n# set of keys that have a TTL, and iterate that set in batches:\nasync with self.lock:\n batch = list(self._ttl_keys)[:50]\n for k in batch:\n node = self.key_to_node.get(k)\n if node and 0 < node.expires_at <= time.monotonic():\n self._remove_node(node)\n\n# FIX 4: time.time() -> time.monotonic() everywhere.\n# FIX 5: add node.version; in tx._apply_put, abort/refresh if\n# cache.key_to_node[key] is a different node than original_node." + }, + { + "id": "qwen3.6-35b-a3b-uncensored-hauhaucs-aggressive-gguf", + "timestamp": "2026-07-28T13:10:00Z", + "model_name": "Qwen 3.6 35B-A3B (uncensored hauhaucs aggressive)", + "quant": "GGUF", + "format": "gguf", + "tok_sec": 62.54, + "total_tokens": 13897, + "ttft_sec": 1.09, + "filename": "outputs/qwen3.6-35b-a3b-uncensored-hauhaucs-aggressive-gguf.py", + "tests_pass": false, + "total_score": 49, + "breakdown": { + "complexity": 12, + "concurrency": 10, + "isolation": 12, + "memory_edge_cases": 11, + "test_integrity": 4 + }, + "verdict": "Critical Bugs", + "best_for": "Not recommended for production code. Reasonable API shape and correctly used time.monotonic(), but the module does not parse (syntax error), contains an infinite while:pass loop, and has data races. Avoid for systems/concurrency work.", + "critical_bugs": [ + "SyntaxError: line 305 'assert val := await cache.get(...)' is invalid Python — walrus operator cannot appear in an assert statement. The ENTIRE module fails to compile, so nothing runs and no test can execute.", + "Infinite busy-loop: _evict_lfu lines 159-160 — 'while self.min_freq in self.freq_map and self.min_freq < max(...): pass' has an empty body that never updates min_freq, recomputes max() (O(F)) each iteration, and can never terminate.", + "Hidden O(F) scan: min(self.freq_map.keys()) / max(self.freq_map.keys()) appears at 6 call sites (every eviction and removal) — violates the strict O(1) requirement.", + "Race condition: get() and put() perform lazy-TTL _remove_key() BEFORE acquiring the lock (lines 71-73, 107-108), mutating shared state unlocked while other coroutines read/write.", + "Deadlock risk: background_loop holds self._lock, then calls await self._remove_key() which is itself a lock-acquiring method — asyncio.Lock is NOT reentrant -> deadlock when the evictor runs.", + "No __slots__ on _Node despite using a dataclass (rubric required it for memory efficiency).", + "_remove_key will KeyError on self.freq_map[freq] if a concurrent operation already deleted that bucket." + ], + "patch_code": "# FIX 1 (the parse error): assign first, then assert.\nval_d = await cache.get(\"D\")\nassert val_d, \"D should exist\"\nval_a = await cache.get(\"A\")\nassert val_a, \"A should exist (highest freq)\"\n\n# FIX 2: delete the broken while:pass loop. Bump min_freq incrementally:\n# only when the min_freq bucket empties, and only ever UP by 1 (a key\n# whose freq increased must land at min_freq+1). Never call min()/max().\nif not old_bucket:\n del self.freq_map[old_freq]\n if self.min_freq == old_freq:\n self.min_freq += 1 # next tier up; never scan\n\n# FIX 3: do ALL lazy eviction INSIDE the lock, not before it:\nasync def get(self, key):\n async with self._lock:\n if self.ttl_map.get(key, inf) <= time.monotonic():\n await self._remove_key(key) # now locked\n return None\n ...\n\n# FIX 4: make _remove_key a non-locking private helper, called from\n# inside already-locked public methods, so background_loop doesn't try\n# to re-acquire the non-reentrant asyncio.Lock.\n\n# FIX 5: add __slots__ = ('key','value','ttl_expiry') to _Node." + }, + { + "id": "kat-coder-v2.5-dev-xl-mlx", + "timestamp": "2026-07-28T13:15:00Z", + "model_name": "KAT-Coder v2.5 Dev XL", + "quant": "MLX", + "format": "mlx", + "tok_sec": 65.26, + "total_tokens": 6172, + "ttft_sec": 7.42, + "filename": "outputs/kat-coder-v2.5-dev-xl-mlx.py", + "tests_pass": false, + "total_score": 65, + "breakdown": { + "complexity": 15, + "concurrency": 16, + "isolation": 16, + "memory_edge_cases": 13, + "test_integrity": 5 + }, + "verdict": "Critical Bugs", + "best_for": "Promising code-design instincts (cleanest abstractions and best transaction isolation design in the set) but undone by a single fatal one-line bug that stops it running. With the bug fixed it would likely score 80+; as-is, only useful as a structural reference.", + "critical_bugs": [ + "FATAL: _put_internal line 305 inserts a NEW key with 'self._freq_map[1].push_front(...)' but never ensures the freq-1 bucket exists — KeyError: 1 on the very first put. The cache cannot store a single key. The _ensure_freq_list(1) helper it should use exists and is used correctly everywhere else (lines 294, 354).", + "Transaction.commit() calls _put_internal for buffered writes, so it hits the same KeyError: 1 — committed transactions crash too.", + "Test suite cannot execute: crashes at the first cache.put() in main(); the well-built test harness (pass/fail counter, 4 real scenarios) validates nothing.", + "No __slots__ on _DLLNode/_CacheNode/_DoublyLinkedList despite the rubric requiring it for memory efficiency.", + "_evict_node uses min(self._freq_map) (line 325) when the min-tier empties — a hidden O(F) scan, violating strict O(1).", + "No MVCC/version check on transaction commit (lost-update possible if the global key is modified between tx.get and commit); commit is not exception-safe across the writes-vs-deletes loops." + ], + "patch_code": "# FIX 1 (the fatal one-liner): use the helper that already exists.\n# line 305, in _put_internal, new-key branch:\n- self._freq_map[1].push_front(dll_node)\n+ self._ensure_freq_list(1).push_front(dll_node)\n# (This single change makes the cache and transactions functional.)\n\n# FIX 2: replace the O(F) min() scan with an incremental bump:\n# in _evict_node, when the min-tier bucket empties, min_freq is the\n# lowest remaining tier. Since freq only ever increments by 1, the\n# next min is almost always min_freq+1; track it incrementally rather\n# than scanning. Or, since this only happens on full eviction, accept\n# O(F) but only on the empty-cache edge — document it.\n\n# FIX 3: add __slots__ to all internal classes:\nclass _CacheNode:\n __slots__ = ('key','value','ttl_seconds','expiry_time','freq','dll_node')\n ...\n\n# FIX 4: wrap commit applies in try/except so a mid-commit exception\n# does not leave a half-applied global state; consider abort semantics.\n# FIX 5: add node.version; in tx commit, abort if the global node for\n# a key is not the one seen at tx.get() time." + }, + { + "id": "gemma4-31b-gguf", + "timestamp": "2026-07-28T13:20:00Z", + "model_name": "Gemma 4 31B", + "quant": "GGUF", + "format": "gguf", + "tok_sec": 10.09, + "total_tokens": 4536, + "ttft_sec": 4.39, + "speed_caveat": "All Gemma 4 models ran abnormally slow (GPU offload appeared inactive despite being set), so tok/sec and TTFT are NOT representative of the model itself — likely an LM Studio/GGUF config issue. Treat speed numbers for the Gemma 4 batch as suspect.", + "filename": "outputs/gemma4-31b-gguf.py", + "tests_pass": true, + "total_score": 78, + "breakdown": { + "complexity": 17, + "concurrency": 16, + "isolation": 14, + "memory_edge_cases": 15, + "test_integrity": 16 + }, + "verdict": "Minor Logic Flaws", + "best_for": "Clean, correct, runnable code with solid O(1) structure and good concurrency granularity. A reliable pick for everyday caching/async work after a monotonic-clock + __slots__ pass.", + "critical_bugs": [ + "Isolation leak: Transaction.get falls back to the PUBLIC cache.get, which calls _update_frequency — so reading a key inside a transaction mutates GLOBAL frequency state before commit, leaking uncommitted access patterns into global eviction order. Spec requires tx reads not to alter global freq.", + "Uses time.time() (system clock) throughout instead of time.monotonic() — NTP adjustments corrupt TTL eviction.", + "No __slots__ on Node/DoublyLinkedList/LFUCache/Transaction — rubric required it for memory efficiency.", + "_delete_internal deliberately leaves empty frequency buckets in freq_map (documented but a minor memory leak: empty DoublyLinkedList objects accumulate).", + "Background evictor does list(self.cache.keys()) = O(N) snapshot every interval — a linear scan, forbidden by strict O(1).", + "No MVCC/version check on commit (lost-update possible); commit is not exception-safe across the deletes-vs-puts loops.", + "Tests pass but don't probe mid-commit read isolation or eviction-under-real-contention (capacity sized so all keys fit), so the isolation leak above goes undetected." + ], + "patch_code": "# FIX 1 (the isolation leak): tx reads must NOT mutate global freq.\n# Add a read-only global lookup (no _update_frequency) and use it in tx.get:\nasync def _read_raw(self, key): # no freq bump\n node = self.cache.get(key)\n if node is None: return None\n if time.monotonic() > node.expiry:\n await self._delete_internal(key)\n return None\n return node.value\n# then in Transaction.get fallback:\n return await self._cache._read_raw(key) # NOT cache.get\n\n# FIX 2: time.time() -> time.monotonic() everywhere (get/put/_put_internal/bg loop).\n# FIX 3: add __slots__ to Node, DoublyLinkedList, LFUCache, Transaction.\n# FIX 4: prune empty freq buckets on delete, or have _evict_lfu drop them.\n# FIX 5: iterate a dedicated _ttl_keys set (batched) in the bg evictor\n# instead of list(self.cache.keys()) to stay O(batch), not O(N)." + }, + { + "id": "gemma-4-31b-qat-gguf", + "timestamp": "2026-07-28T13:25:00Z", + "model_name": "Gemma 4 31B QAT", + "quant": "QAT GGUF", + "format": "gguf", + "tok_sec": 15.0, + "total_tokens": 4552, + "ttft_sec": 4.01, + "speed_caveat": "Same as the Gemma 4 batch: GPU offload appeared inactive so tok/sec/TTFT are NOT representative of the model. Suspected LM Studio/GGUF config issue.", + "filename": "outputs/gemma-4-31b-qat-gguf.py", + "tests_pass": true, + "total_score": 70, + "breakdown": { + "complexity": 16, + "concurrency": 16, + "isolation": 11, + "memory_edge_cases": 13, + "test_integrity": 14 + }, + "verdict": "Critical Bugs", + "best_for": "Runnable and structurally sound, but the LFU eviction has a stale-min_freq capacity-breach path and the transaction API doesn't match the spec (no commit/rollback). Usable for prototypes if you fix eviction and re-skin transactions.", + "critical_bugs": [ + "Capacity breach: eviction does self.freq_map[self.min_freq].pop_tail() with NO guard that the bucket exists or is non-empty, and never prunes emptied freq buckets. After manual deletes empty the min-tier, pop_tail returns None silently -> eviction fails -> cache grows PAST capacity. This is the exact stale-min_freq capacity-breach the rubric flags.", + "Non-conformant transaction API: Transaction has NO commit() or rollback() method (both required by spec). Commit happens via a separate cache.apply_transaction_changes(tx._state) — wrong surface; the test only passes because it uses this internal path.", + "Isolation leak: Transaction.get falls back to the public cache.get, which bumps global frequency before commit — uncommitted tx reads alter global eviction order.", + "Duplicated eviction logic in apply_transaction_changes re-introduces the stale-min_freq bug at line 211.", + "Uses time.time() (system clock) via _get_now() instead of time.monotonic() — NTP jumps corrupt TTL (though _get_now is a clean single fix point).", + "No __slots__ on Node/DoublyLinkedList/LFUCache/TransactionState/Transaction.", + "Background sweep does list(self.cache.keys()) = O(N) per interval.", + "Tests pass but use the non-spec commit path and don't probe capacity breach or isolation leak." + ], + "patch_code": "# FIX 1 (capacity breach): guard + prune empty buckets on eviction:\nwhile self.min_freq in self.freq_map and self.freq_map[self.min_freq].size == 0:\n del self.freq_map[self.min_freq]\n self.min_freq += 1\n if not self.freq_map:\n break\nif self.min_freq not in self.freq_map:\n return # nothing to evict\nevicted = self.freq_map[self.min_freq].pop_tail()\nif evicted and self.freq_map[self.min_freq].size == 0:\n del self.freq_map[self.min_freq]\n\n# FIX 2 (conformant API): add commit/rollback to Transaction:\nasync def commit(self):\n await self._cache.apply_transaction_changes(self._state)\n self._committed = True\ndef rollback(self):\n self._state.writes.clear()\n self._committed = True\n\n# FIX 3 (isolation leak): tx.get should use a read-only global lookup\n# (no _update_freq), not the public cache.get.\n# FIX 4: _get_now returns time.monotonic().\n# FIX 5: add __slots__ to all node/list classes." + }, + { + "id": "gemma-4-12b-coder-heretic-mxfp8-mlx", + "timestamp": "2026-07-28T13:30:00Z", + "model_name": "Gemma 4 12B Coder (fable5-composer2.5-v1-uncensored-heretic merge)", + "quant": "mxfp8 MLX", + "format": "mlx", + "tok_sec": 25.28, + "total_tokens": 2625, + "ttft_sec": 3.21, + "filename": "outputs/gemma-4-12b-coder-fable5-composer2.5-v1-uncensored-heretic-mxfp8-mlx.py", + "tests_pass": false, + "total_score": 43, + "breakdown": { + "complexity": 13, + "concurrency": 7, + "isolation": 10, + "memory_edge_cases": 9, + "test_integrity": 4 + }, + "verdict": "Critical Bugs", + "best_for": "Not usable as-is — the cache cannot store its first key and the background evictor would crash the event loop. The smallest model in the set (12B) and lowest-quality output. Avoid for systems work.", + "critical_bugs": [ + "FATAL: put() line 112 does 'bucket = self.freq_buckets[self.min_freq]' after setting min_freq=1 but NEVER creates freq_buckets[1] -> KeyError: 1 on the very first put. Cache is unusable.", + "Background evictor is fundamentally broken: start_evictor defines a SYNC 'def evict_loop' and passes it to create_task; inside it calls blocking time.sleep(interval) (freezes the event loop) AND asyncio.run(...) from within a running loop -> RuntimeError. Would crash hard if ever reached.", + "Eviction-by-re-put: expired keys are 'evicted' by re-inserting them with TTL 0 (line 122) instead of deleting them — wrong semantics and triggers immediate re-eviction.", + "Transaction duplicates the entire LFU machinery (local_cache + local_freq + local_min_freq) for snapshot isolation, but _update_local_freq has the same missing-bucket KeyError (line 140).", + "Class name typo 'DoublyLinkedListList' (doubled word).", + "No __slots__; time.time() (not monotonic) throughout.", + "Tests cannot run — crash at first put." + ], + "patch_code": "# FIX 1 (the fatal KeyError): create the bucket before use.\n# In put(), new-key branch:\n- bucket = self.freq_buckets[self.min_freq]\n+ bucket = self.freq_buckets.setdefault(self.min_freq, DoublyLinkedListList())\n# Same fix in _update_freq and Transaction._update_local_freq (use setdefault).\n\n# FIX 2 (the broken evictor): make it a real async task that deletes:\nasync def _evict_loop(self, interval):\n while True:\n await asyncio.sleep(interval) # async, non-blocking\n now = time.monotonic()\n async with self.global_lock:\n expired = [k for k, n in list(self.cache.items()) if now > n.ttl_expiry]\n for k in expired:\n node = self.cache.pop(k, None)\n if node:\n self.freq_buckets[node.freq].remove(node) # DELETE, not re-put\n\nasync def start_evictor(self, interval=1.0):\n self.evictor_task = asyncio.create_task(self._evict_loop(interval))\n\n# FIX 3: delete expired keys; do NOT re-insert with TTL 0.\n# FIX 4: time.time() -> time.monotonic().\n# FIX 5: add __slots__ to Node / DoublyLinkedListList / ConcurrentLFUCache / Transaction." + }, + { + "id": "deepseekv4flash-cloud", + "timestamp": "2026-07-28T13:35:00Z", + "model_name": "DeepSeek V4 Flash (CLOUD baseline)", + "quant": "n/a (cloud)", + "format": "cloud", + "tok_sec": null, + "total_tokens": null, + "ttft_sec": null, + "speed_caveat": "Cloud model (run via opencode, not LM Studio) — tok_sec/tokens/TTFT are N/A (not measured for cloud). Included as a quality baseline against the local models. NOTE: it took 3 attempts to produce any output and ~12 minutes of thinking before succeeding — so it is a QUALITY benchmark, not a speed/usability one.", + "filename": "deepseekv4flash.py", + "tests_pass": true, + "total_score": 91, + "breakdown": { + "complexity": 18, + "concurrency": 18, + "isolation": 19, + "memory_edge_cases": 18, + "test_integrity": 18 + }, + "verdict": "Production Ready", + "best_for": "Reference-quality baseline (91/100) — the bar the local models are measured against. Only submission with __slots__ + time.monotonic() + delta-based transactional frequency accounting. Sets the ceiling for correctness, though its unreliability (3 attempts, 12-min think time) makes it a poor *local* daily-driver.", + "critical_bugs": [ + "Two min(self._freq_to_list) linear scans in _evict_one's defensive recovery path (lines 415, 429) — only triggered when min_freq desyncs, not per-operation, but still a non-O(1) path. Could be replaced with incremental tracking.", + "Single coarse lock held across the whole commit-apply loop — not the fine-grained locking the prompt asked for.", + "__slots__ present on _Node and _DLL but not extended to Transaction / LFUCache.", + "No explicit lost-update/conflict abort on commit (delta-based freq is applied unconditionally).", + "Tests pass 20/20 but don't include a mid-commit read-isolation probe or adversarial eviction-under-contention stress." + ], + "patch_code": "# These are minor refinements on an already production-ready file.\n\n# FIX 1: eliminate the recovery min() scans by keeping min_freq\n# strictly in sync on every insert/bump/remove (it already does on\n# the hot path), so the _evict_one recovery branch is unreachable and\n# can assert instead of scanning:\nassert self._min_freq in self._freq_to_list or not self._freq_to_list\n\n# FIX 2: extend __slots__ to Transaction and LFUCache.\nclass LFUCache(Generic[KT, VT]):\n __slots__ = ('_capacity','_key_to_node','_freq_to_list','_min_freq',\n '_lock','_ttl_index','_evictor_task','_closed')\n\n# FIX 3 (optional): on commit, if a key's global node changed since the\n# tx snapshot, raise LFUCacheError('lost update') instead of overwriting." + } + ] +} diff --git a/deepseekv4flash.py b/deepseekv4flash.py new file mode 100644 index 0000000..af76929 --- /dev/null +++ b/deepseekv4flash.py @@ -0,0 +1,600 @@ +#!/usr/bin/env python3 +""" +In-Memory Concurrent LFU Cache with Async TTL Eviction & Atomic Transactions. + +Pure Python 3.11+, no external dependencies. + +O(1) average-time get/put via frequency buckets with doubly-linked lists. +Dual-layer TTL: lazy on access + background async sweep. +Atomic transactions with read-your-own-writes, commit, and rollback. +""" + +import asyncio +import time +from typing import (Dict, Generic, Hashable, Optional, Set, TypeVar) + + +KT = TypeVar('KT', bound=Hashable) +VT = TypeVar('VT') + +_SENTINEL = object() + + +class LFUCacheError(Exception): + """Base exception for LFU cache errors.""" + + +class _Node(Generic[KT, VT]): + """Doubly-linked list node holding key, value, frequency, and expiry.""" + + __slots__ = ('key', 'value', 'freq', 'expires_at', 'prev', 'next') + + def __init__(self, key: KT, value: VT) -> None: + self.key = key + self.value = value + self.freq: int = 0 + self.expires_at: Optional[float] = None + self.prev: Optional['_Node[KT, VT]'] = None + self.next: Optional['_Node[KT, VT]'] = None + + +class _DLL(Generic[KT, VT]): + """ + Doubly-linked list with sentinel head/tail. + O(1) append (tail/MRU), remove (by reference), and pop_left (head/LRU). + """ + + __slots__ = ('_head', '_tail', 'size') + + def __init__(self) -> None: + self._head = _Node[KT, VT](_SENTINEL, None) # type: ignore[arg-type] + self._tail = _Node[KT, VT](_SENTINEL, None) # type: ignore[arg-type] + self._head.next = self._tail + self._tail.prev = self._head + self.size: int = 0 + + def append(self, node: _Node[KT, VT]) -> None: + """Add node to tail (most-recently-used position). O(1).""" + last = self._tail.prev + last.next = node + node.prev = last + node.next = self._tail + self._tail.prev = node + self.size += 1 + + def remove(self, node: _Node[KT, VT]) -> None: + """Remove node that must be in this list. O(1).""" + node.prev.next = node.next + node.next.prev = node.prev + node.prev = None + node.next = None + self.size -= 1 + + def pop_left(self) -> Optional[_Node[KT, VT]]: + """Remove and return the head node (LRU position). O(1).""" + if self.size == 0: + return None + node = self._head.next + self.remove(node) + return node + + +class Transaction(Generic[KT, VT]): + """ + Atomic transaction with read-your-own-writes isolation. + + Uncommitted writes are invisible to global cache readers. + All local state is discarded on rollback without touching global structures. + """ + + def __init__(self, cache: 'LFUCache[KT, VT]') -> None: + self._cache = cache + self._writes: Dict[KT, VT] = {} + self._deletes: Set[KT] = set() + self._snapshot: Dict[KT, VT] = {} + self._freq_deltas: Dict[KT, int] = {} + self._ttls: Dict[KT, Optional[float]] = {} + self._committed: bool = False + self._rolled_back: bool = False + + def _check_active(self) -> None: + if self._committed: + raise LFUCacheError("Transaction already committed") + if self._rolled_back: + raise LFUCacheError("Transaction already rolled back") + + async def get(self, key: KT) -> Optional[VT]: + """Read key with read-your-own-writes and snapshot isolation.""" + self._check_active() + + if key in self._deletes: + return None + if key in self._writes: + self._freq_deltas[key] = self._freq_deltas.get(key, 0) + 1 + return self._writes[key] + if key in self._snapshot: + self._freq_deltas[key] = self._freq_deltas.get(key, 0) + 1 + return self._snapshot[key] + + async with self._cache._lock: + node = self._cache._key_to_node.get(key) + if node is None: + return None + if node.expires_at is not None and node.expires_at <= time.monotonic(): + self._cache._remove_key(key) + return None + self._snapshot[key] = node.value + + self._freq_deltas[key] = self._freq_deltas.get(key, 0) + 1 + return self._snapshot[key] + + async def put(self, key: KT, value: VT, + ttl_seconds: Optional[float] = None) -> None: + """Write key within transaction.""" + self._check_active() + self._writes[key] = value + self._deletes.discard(key) + if ttl_seconds is not None: + self._ttls[key] = ttl_seconds if ttl_seconds > 0 else None + self._freq_deltas[key] = self._freq_deltas.get(key, 0) + 1 + + async def delete(self, key: KT) -> None: + """Mark key for deletion on commit.""" + self._check_active() + self._deletes.add(key) + self._writes.pop(key, None) + self._snapshot.pop(key, None) + self._freq_deltas.pop(key, None) + self._ttls.pop(key, None) + + async def commit(self) -> None: + """Atomically apply all pending changes to the global cache.""" + if self._rolled_back: + raise LFUCacheError("Transaction already rolled back") + if self._committed: + return + + async with self._cache._lock: + try: + for key in self._deletes: + self._cache._remove_key(key) + + for key, value in self._writes.items(): + node = self._cache._key_to_node.get(key) + ttl = self._ttls.get(key) + delta = self._freq_deltas.get(key, 1) + + if node is not None: + node.value = value + if ttl is not None: + node.expires_at = ( + time.monotonic() + ttl if ttl is not None else None + ) + self._cache._change_freq(node, delta) + else: + while len(self._cache._key_to_node) >= self._cache._capacity: + self._cache._evict_one() + node = _Node(key, value) + node.freq = delta + if ttl is not None: + node.expires_at = time.monotonic() + ttl + self._cache._key_to_node[key] = node + self._cache._add_to_freq_list(node) + + for key, delta in self._freq_deltas.items(): + if key not in self._writes and key not in self._deletes: + node = self._cache._key_to_node.get(key) + if node is not None: + self._cache._change_freq(node, delta) + + while len(self._cache._key_to_node) > self._cache._capacity: + self._cache._evict_one() + + self._committed = True + except BaseException: + self._rolled_back = True + raise + + async def rollback(self) -> None: + """Discard all pending changes. Global cache is untouched.""" + if self._committed: + raise LFUCacheError("Transaction already committed") + self._rolled_back = True + + +class LFUCache(Generic[KT, VT]): + """ + In-Memory Concurrent LFU Cache with Async TTL Eviction. + + O(1) average-time get/put using frequency buckets with doubly-linked lists. + Dual-layer TTL eviction: lazy on access + background async sweep. + """ + + def __init__(self, capacity: int = 1000, + evictor_interval: float = 1.0, + evictor_batch_size: int = 10, + evictor_scan_budget: int = 100) -> None: + if capacity < 1: + raise ValueError("Capacity must be >= 1") + + self._capacity = capacity + self._evictor_interval = evictor_interval + self._evictor_batch_size = evictor_batch_size + self._evictor_scan_budget = evictor_scan_budget + + self._key_to_node: Dict[KT, _Node[KT, VT]] = {} + self._freq_to_list: Dict[int, _DLL[KT, VT]] = {} + self._min_freq: int = 0 + self._lock = asyncio.Lock() + self._evictor_running: bool = False + self._evictor_task: Optional[asyncio.Task] = None + + # ---- Public API --------------------------------------------------------- + + async def get(self, key: KT) -> Optional[VT]: + """Retrieve value by key. Returns None if missing or expired. + + Accesses increment the key's frequency (LFU tracking). + Expired keys are lazily evicted on access. + """ + async with self._lock: + node = self._key_to_node.get(key) + if node is None: + return None + if node.expires_at is not None and node.expires_at <= time.monotonic(): + self._remove_key(key) + return None + self._change_freq(node, 1) + return node.value + + async def put(self, key: KT, value: VT, + ttl_seconds: Optional[float] = None) -> None: + """Insert or update a key-value pair. + + If *ttl_seconds* is None (default) the entry never expires. + If *ttl_seconds* is <= 0 it is treated as no expiration. + If the cache is at capacity the least-frequently-used item is evicted + (tie-broken by least-recently-used within the minimum frequency tier). + """ + async with self._lock: + node = self._key_to_node.get(key) + + if node is not None: + if node.expires_at is not None and node.expires_at <= time.monotonic(): + self._remove_key(key) + node = None + else: + node.value = value + if ttl_seconds is not None: + node.expires_at = ( + time.monotonic() + ttl_seconds if ttl_seconds > 0 else None + ) + self._change_freq(node, 1) + return + + while len(self._key_to_node) >= self._capacity: + self._evict_one() + + node = _Node(key, value) + node.freq = 1 + if ttl_seconds is not None and ttl_seconds > 0: + node.expires_at = time.monotonic() + ttl_seconds + self._key_to_node[key] = node + self._add_to_freq_list(node) + + async def delete(self, key: KT) -> bool: + """Remove *key* from the cache. Returns True if the key existed.""" + async with self._lock: + return self._remove_key(key) + + def begin_transaction(self) -> Transaction[KT, VT]: + """Open an atomic transaction for batched reads/writes.""" + return Transaction(self) + + @property + def capacity(self) -> int: + return self._capacity + + @property + def size(self) -> int: + return len(self._key_to_node) + + # ---- Background TTL Evictor -------------------------------------------- + + def start_evictor(self) -> None: + """Launch the background TTL eviction loop as an asyncio task.""" + if self._evictor_running: + return + self._evictor_running = True + self._evictor_task = asyncio.create_task(self._evictor_loop()) + + async def stop_evictor(self) -> None: + """Cancel and wait for the background eviction task to finish.""" + if not self._evictor_running: + return + self._evictor_running = False + if self._evictor_task is not None: + self._evictor_task.cancel() + try: + await self._evictor_task + except asyncio.CancelledError: + pass + self._evictor_task = None + + async def _evictor_loop(self) -> None: + """ + Background loop: periodically scan a limited batch of keys and remove + expired entries. Releases the lock between batches so concurrent + reads/writes are not blocked for extended periods. + """ + cursor = 0 + keys: list = [] + + try: + while self._evictor_running: + await asyncio.sleep(self._evictor_interval) + + async with self._lock: + if cursor >= len(keys): + keys = list(self._key_to_node.keys()) + cursor = 0 + if not keys: + continue + + now = time.monotonic() + removed = 0 + scan = self._evictor_scan_budget + + while scan > 0 and cursor < len(keys): + key = keys[cursor] + cursor += 1 + scan -= 1 + + node = self._key_to_node.get(key) + if (node is not None and + node.expires_at is not None and + node.expires_at <= now): + self._remove_key(key) + removed += 1 + if removed >= self._evictor_batch_size: + break + except asyncio.CancelledError: + pass + + # ---- O(1) Internal Helpers --------------------------------------------- + + def _add_to_freq_list(self, node: _Node[KT, VT]) -> None: + """Insert *node* into its frequency bucket and update *min_freq*.""" + freq = node.freq + if freq not in self._freq_to_list: + self._freq_to_list[freq] = _DLL() + self._freq_to_list[freq].append(node) + if self._min_freq not in self._freq_to_list or freq < self._min_freq: + self._min_freq = freq + + def _remove_from_freq_list(self, node: _Node[KT, VT]) -> None: + """Remove *node* from its frequency bucket. No-op if not linked.""" + if node.prev is None or node.next is None: + return + freq = node.freq + lst = self._freq_to_list.get(freq) + if lst is None: + return + lst.remove(node) + if lst.size == 0: + del self._freq_to_list[freq] + + def _change_freq(self, node: _Node[KT, VT], delta: int) -> None: + """Atomically increase *node*'s frequency by *delta* and + move it to the corresponding bucket. O(1).""" + if delta <= 0: + return + self._remove_from_freq_list(node) + node.freq += delta + self._add_to_freq_list(node) + + def _remove_key(self, key: KT) -> bool: + """Remove *key* from all internal structures. Returns True if existed.""" + node = self._key_to_node.pop(key, None) + if node is None: + return False + self._remove_from_freq_list(node) + return True + + def _evict_one(self) -> Optional[KT]: + """ + Evict one item: LRU among the minimum-frequency bucket. + Returns the evicted key, or None if the cache is empty. + """ + if not self._key_to_node: + return None + + if self._min_freq not in self._freq_to_list: + if not self._freq_to_list: + return None + self._min_freq = min(self._freq_to_list) + + lst = self._freq_to_list.get(self._min_freq) + if lst is None or lst.size == 0: + return None + + node = lst.pop_left() + if node is None: + return None + + del self._key_to_node[node.key] + if lst.size == 0: + del self._freq_to_list[self._min_freq] + if self._freq_to_list: + self._min_freq = min(self._freq_to_list) + + return node.key + + # ---- Async Context Manager --------------------------------------------- + + async def __aenter__(self) -> 'LFUCache[KT, VT]': + return self + + async def __aexit__(self, *exc_info) -> None: + await self.stop_evictor() + + +# ============================================================================ +# Unit Tests +# ============================================================================ + +async def _run_tests() -> None: + passed = 0 + total = 0 + + def check(cond: bool, msg: str): + nonlocal passed, total + total += 1 + if cond: + passed += 1 + else: + print(f" FAIL: {msg}") + + # ------------------------------------------------------------------ + # 1. O(1) LFU eviction order + # ------------------------------------------------------------------ + print("1. LFU eviction order ... ", end="", flush=True) + + cache = LFUCache(capacity=3) + + await cache.put("a", 1) + await cache.put("b", 2) + await cache.put("c", 3) + + await cache.get("a") + await cache.get("a") + await cache.get("b") + # freqs: a=3, b=2, c=1 + + await cache.put("d", 4) # evicts c (freq=1) + check(await cache.get("c") is None, "c was evicted (freq=1)") + check(await cache.get("d") == 4, "d is present") + check(await cache.get("a") == 1, "a is present") + check(await cache.get("b") == 2, "b is present") + + # a=4, b=3, d=1 after the gets above + await cache.put("e", 5) # evicts d (freq=1) + check(await cache.get("d") is None, "d was evicted") + check(await cache.get("e") == 5, "e is present") + check(cache.size == 3, "cache size == 3") + + print("OK") + + # ------------------------------------------------------------------ + # 2. TTL eviction + # ------------------------------------------------------------------ + print("2. TTL eviction ... ", end="", flush=True) + + # 2a. Lazy eviction + cache2 = LFUCache(capacity=10) + await cache2.put("lazy", "alive", ttl_seconds=0.02) + await asyncio.sleep(0.03) + check(await cache2.get("lazy") is None, "lazy TTL eviction on get") + + await cache2.put("lazy2", "alive", ttl_seconds=0.02) + await asyncio.sleep(0.03) + check(await cache2.get("lazy2") is None, "lazy TTL eviction on second get") + + # 2b. Background eviction + cache3 = LFUCache(capacity=10, evictor_interval=0.02, evictor_batch_size=5) + cache3.start_evictor() + await cache3.put("bg", "data", ttl_seconds=0.01) + await asyncio.sleep(0.10) + check(await cache3.get("bg") is None, "background evictor removes expired key") + await cache3.stop_evictor() + + print("OK") + + # ------------------------------------------------------------------ + # 3. Transactions + # ------------------------------------------------------------------ + print("3. Transactions ... ", end="", flush=True) + + # 3a. Commit + cache4 = LFUCache(capacity=10) + await cache4.put("x", 10) + + tx = cache4.begin_transaction() + await tx.put("x", 20) + check(await tx.get("x") == 20, "tx reads its own write") + check(await cache4.get("x") == 10, "global does NOT see uncommitted write") + await tx.commit() + check(await cache4.get("x") == 20, "global sees committed value") + + # 3b. Rollback + await cache4.put("y", 100) + tx2 = cache4.begin_transaction() + await tx2.put("y", 200) + await tx2.rollback() + check(await cache4.get("y") == 100, "rollback discards writes") + + # 3c. Delete in transaction + await cache4.put("z", 300) + tx3 = cache4.begin_transaction() + await tx3.delete("z") + check(await tx3.get("z") is None, "tx sees its own delete") + check(await cache4.get("z") == 300, "global not affected before commit") + await tx3.commit() + check(await cache4.get("z") is None, "global sees committed delete") + + # 3d. Rollback after delete + await cache4.put("w", 400) + tx4 = cache4.begin_transaction() + await tx4.delete("w") + await tx4.rollback() + check(await cache4.get("w") == 400, "rollback restores deleted key") + + # 3e. Read own writes after rollback raises + tx5 = cache4.begin_transaction() + await tx5.put("n", 1) + await tx5.rollback() + try: + await tx5.get("n") + check(False, "get on rolled-back tx should raise") + except LFUCacheError: + check(True, "rolled-back tx raises on get") + + print("OK") + + # ------------------------------------------------------------------ + # 4. Concurrent stress + # ------------------------------------------------------------------ + print("4. Concurrent stress (50 workers) ... ", end="", flush=True) + + cache5 = LFUCache(capacity=50, evictor_interval=0.05, evictor_batch_size=5) + cache5.start_evictor() + + async def worker(uid: int): + for i in range(30): + key = f"k_{(uid + i) % 40}" + await cache5.put(key, uid * 1000 + i, ttl_seconds=0.5) + _ = await cache5.get(key) + await cache5.get(f"k_{(uid + i + 1) % 40}") + if i % 5 == 0: + await cache5.delete(f"k_{(uid + i + 2) % 40}") + await asyncio.sleep(0) + + workers = [worker(i) for i in range(50)] + await asyncio.gather(*workers) + await cache5.stop_evictor() + + check(cache5.size <= 50, f"cache size {cache5.size} <= capacity 50") + print("OK") + + # ------------------------------------------------------------------ + print(f"\nResults: {passed}/{total} checks passed") + if passed < total: + raise SystemExit(1) + + +async def main() -> None: + await _run_tests() + + +if __name__ == "__main__": + asyncio.run(main()) diff --git a/generate_dashboard.py b/generate_dashboard.py new file mode 100644 index 0000000..f992e71 --- /dev/null +++ b/generate_dashboard.py @@ -0,0 +1,403 @@ +#!/usr/bin/env python3 +""" +Generator: builds dashboard.html + pages/..html from data/benchmark_history.json. +Re-run after each grading batch to regenerate everything. +Cyberpunk-terminal aesthetic. Pure stdlib + Chart.js via CDN. +""" +import json, html, os, sys + +HERE = os.path.dirname(os.path.abspath(__file__)) +DATA = os.path.join(HERE, "data", "benchmark_history.json") +OUT_DASH = os.path.join(HERE, "dashboard.html") +PAGES_DIR = os.path.join(HERE, "pages") + +PILLARS = ["complexity", "concurrency", "isolation", "memory_edge_cases", "test_integrity"] +PILLAR_LABELS = { + "complexity": "Complexity (O(1))", + "concurrency": "Concurrency / Races", + "isolation": "Tx Isolation", + "memory_edge_cases": "Memory & Edges", + "test_integrity": "Test Integrity", +} + +# ---- cyberpunk palette ---- +NEON_CYAN = "#00ffc8" +NEON_MAG = "#ff2bd6" +NEON_LIME = "#b6ff3a" +NEON_AMBER= "#ffb000" +NEON_RED = "#ff3b5c" +BG = "#0a0a0f" +PANEL = "#101018" +INK = "#d7e0e6" # body text +INK_DIM = "#7a8590" # muted + +def verdict_meta(v): + return { + "Production Ready": (NEON_LIME, "PROD"), + "Minor Logic Flaws": (NEON_AMBER, "FLAWS"), + "Critical Bugs": (NEON_RED, "CRIT"), + "Broken / Unusable": (NEON_RED, "DEAD"), + }.get(v, (INK_DIM, "???")) + +def esc(s): + return html.escape(str(s)) if s is not None else "" + +def speed_str(m): + if m.get("tok_sec") is None: + return "N/A" + return f"{m['tok_sec']:.1f}" + +# series colors for charts (cycles through neon accents) +SERIES = [NEON_CYAN, NEON_MAG, NEON_LIME, NEON_AMBER, "#5b8cff", "#c084fc", "#34d399", "#f472b6"] + +# shared CSS — embedded so each page is standalone +def shared_css(): + return """ +:root{ + --bg:#0a0a0f; --panel:#101018; --panel2:#14141f; --ink:#d7e0e6; --dim:#7a8590; + --cyan:#00ffc8; --mag:#ff2bd6; --lime:#b6ff3a; --amber:#ffb000; --red:#ff3b5c; --blue:#5b8cff; + --grid: rgba(0,255,200,0.04); +} +*{box-sizing:border-box} +html,body{margin:0;padding:0} +body{ + background:var(--bg); color:var(--ink); + font-family:'Fira Sans',system-ui,-apple-system,sans-serif; + line-height:1.6; min-height:100vh; + /* scanline + grid texture overlay */ + background-image: + repeating-linear-gradient(0deg, rgba(0,255,200,0.025) 0px, rgba(0,255,200,0.025) 1px, transparent 1px, transparent 3px), + linear-gradient(var(--grid) 1px, transparent 1px), + linear-gradient(90deg, var(--grid) 1px, transparent 1px), + radial-gradient(1200px 600px at 70% -10%, rgba(91,139,255,0.10), transparent 60%), + radial-gradient(900px 500px at 10% 110%, rgba(255,43,214,0.08), transparent 60%); + background-size: 3px 3px, 44px 44px, 44px 44px, 100% 100%, 100% 100%; + background-attachment: fixed; +} +mono,.mono,h1,h2,h3,.hud,th,.code,.badge,.stat-num,.lead-model,.verdict-chip{font-family:'Fira Code',ui-monospace,'SF Mono',Menlo,monospace} +a{color:var(--cyan);text-decoration:none} +a:hover{text-shadow:0 0 8px var(--cyan)} +.wrap{max-width:1180px;margin:0 auto;padding:28px 20px 80px} +header.hud-bar{ + border:1px solid rgba(0,255,200,0.25); + background:linear-gradient(180deg,rgba(0,255,200,0.06),rgba(0,0,0,0)); + box-shadow:0 0 24px rgba(0,255,200,0.08), inset 0 0 30px rgba(0,0,0,0.5); + border-radius:6px; padding:18px 22px; margin-bottom:24px; + position:relative; overflow:hidden; +} +header.hud-bar::before{ + content:""; position:absolute; inset:0; + background:repeating-linear-gradient(90deg, transparent 0 7px, rgba(0,255,200,0.03) 7px 8px); + pointer-events:none; +} +.kicker{color:var(--cyan); letter-spacing:.32em; font-size:.72rem; text-transform:uppercase; text-shadow:0 0 10px rgba(0,255,200,0.5)} +h1{font-size:1.5rem; margin:.3em 0 0; letter-spacing:.02em; text-shadow:0 0 16px rgba(0,255,200,0.35)} +.subtitle{color:var(--dim); font-size:.9rem; margin-top:.35em} +.stats{display:grid; grid-template-columns:repeat(4,1fr); gap:12px; margin:18px 0 26px} +.stat{background:var(--panel); border:1px solid rgba(255,255,255,0.06); border-radius:6px; padding:14px 16px; position:relative; overflow:hidden} +.stat::after{content:"";position:absolute;left:0;top:0;bottom:0;width:3px;background:var(--cyan);box-shadow:0 0 12px var(--cyan)} +.stat .lbl{color:var(--dim);font-size:.68rem;letter-spacing:.18em;text-transform:uppercase} +.stat .num{font-size:1.6rem;margin-top:4px;color:var(--ink)} +.grid2{display:grid;grid-template-columns:1.1fr .9fr;gap:18px;margin-bottom:26px} +.panel{background:var(--panel);border:1px solid rgba(255,255,255,0.07);border-radius:8px;padding:18px} +.panel h2{font-size:.95rem;letter-spacing:.12em;text-transform:uppercase;color:var(--cyan);margin:0 0 14px;text-shadow:0 0 10px rgba(0,255,200,0.35)} +.chart-box{position:relative;height:340px} +@media(max-width:900px){.grid2{grid-template-columns:1fr}.stats{grid-template-columns:repeat(2,1fr)}} +/* leaderboard */ +table{width:100%;border-collapse:collapse;font-size:.86rem} +thead th{text-align:left;color:var(--dim);font-size:.66rem;letter-spacing:.16em;text-transform:uppercase;border-bottom:1px solid rgba(0,255,200,0.2);padding:8px 10px} +tbody td{padding:11px 10px;border-bottom:1px solid rgba(255,255,255,0.05);vertical-align:middle} +tbody tr{transition:background .15s, box-shadow .15s} +tbody tr:hover{background:rgba(0,255,200,0.05);box-shadow:inset 0 0 0 1px rgba(0,255,200,0.25)} +.rank{color:var(--dim);width:34px} +.rank.top{color:var(--lime);text-shadow:0 0 8px var(--lime)} +.model-name{color:var(--ink)} +.quant{color:var(--dim);font-size:.78rem} +.badge{display:inline-block;padding:2px 9px;border-radius:3px;font-size:.68rem;letter-spacing:.1em;border:1px solid currentColor} +.btn{display:inline-block;padding:5px 12px;border:1px solid var(--cyan);color:var(--cyan);border-radius:4px;font-size:.74rem;letter-spacing:.08em;cursor:pointer;transition:all .18s} +.btn:hover{background:rgba(0,255,200,0.12);box-shadow:0 0 14px rgba(0,255,200,0.4)} +.bar-cell{display:flex;align-items:center;gap:10px} +.score-bar{flex:1;height:8px;background:rgba(255,255,255,0.06);border-radius:2px;overflow:hidden;min-width:70px} +.score-bar > i{display:block;height:100%;border-radius:2px;transition:width .4s, box-shadow .2s} +tr:hover .score-bar > i{box-shadow:0 0 12px currentColor} +.caveat{color:var(--amber);font-size:.72rem} +.cloud-tag{color:var(--blue);font-size:.7rem;border:1px solid rgba(91,139,255,.4);padding:1px 6px;border-radius:3px;margin-left:6px} +footer{color:var(--dim);font-size:.74rem;margin-top:40px;border-top:1px solid rgba(255,255,255,0.06);padding-top:14px;text-align:center} +@media (prefers-reduced-motion: reduce){*{animation:none!important;transition:none!important}} +""" + +def head_html(title): + return f""" + + + + +{esc(title)} + + + + + + +
+""" + +FOOT = """
+ +""" + +def render_dashboard(data): + models = sorted(data["models"], key=lambda m: -m["total_score"]) + n = len(models) + avg = sum(m["total_score"] for m in models) / n if n else 0 + top = models[0] if models else None + prod = sum(1 for m in models if m["verdict"] == "Production Ready") + + # chart datasets (exclude cloud model from speed bars — has no tok/sec) + local = [m for m in models if m.get("tok_sec") is not None] + local_sorted = sorted(local, key=lambda m: -m["tok_sec"]) + + radar_models = [m for m in models if m.get("tok_sec") is not None or True][:3] # top 3 by score already + + rows = [] + for i, m in enumerate(models, 1): + col, chip = verdict_meta(m["verdict"]) + caveat = "" + if m.get("speed_caveat"): + caveat = '
⚠ speed suspect
' + cloud = 'CLOUD' if m.get("format") == "cloud" else "" + bar_color = col + rows.append(f""" + #{i} +
{esc(m['model_name'])}{cloud}
{esc(m['quant'])}
{caveat} + {speed_str(m)} t/s +
{m['total_score']} +
+ {chip} + {esc(m['best_for'])[:70]}… + DECODE ▸ +""") + + # JSON for charts + bar_labels = json.dumps([m["model_name"].split("(")[0].strip()[:18] for m in local_sorted]) + bar_speed = json.dumps([m["tok_sec"] for m in local_sorted]) + bar_score = json.dumps([m["total_score"] for m in local_sorted]) + + radar_labels = json.dumps([PILLAR_LABELS[p] for p in PILLARS]) + radar_sets = [] + for idx, m in enumerate(radar_models): + col = SERIES[idx % len(SERIES)] + radar_sets.append({ + "label": m["model_name"][:24], + "data": [m["breakdown"][p] for p in PILLARS], + "borderColor": col, + "backgroundColor": col + "22", + }) + radar_json = json.dumps(radar_sets) + + stats = f""" +
Models Tested
{n}
+
Top Score
{top['total_score'] if top else '—'}
+
Average
{avg:.1f}
+
Prod-Ready
{prod}/{n}
+ """ + top_name = esc(top["model_name"]) if top else "—" + + body = f""" +{head_html("LLM Benchmark Suite")} +
+
▚ LOCAL LLM BENCHMARK SUITE // LFU CACHE & ACID AUDIT
+

// BENCHMARK_RESULTS .json

+
{n} models graded on a strict 5-pillar / 100-pt rubric · O(1) LFU + ACID transactions · M3 Max · LM Studio +  ·  TOP: {top_name}
+
+{stats} +
+
+

▮ Score vs Throughput (tok/sec)

+
+
Local models only — cloud baseline (DeepSeek) excluded from speed axis. Gemma 4 bars flagged ⚠ (GPU-offload suspect).
+
+
+

▮ 5-Pillar Radar — Top 3

+
+
Each pillar scored 0–20. Outer = stronger.
+
+
+
+

▮ LEADERBOARD

+
+ + + {''.join(rows)} +
#ModelSpeedScoreVerdictBest For
+
+
Generated from data/benchmark_history.json · re-run generate_dashboard.py to refresh · cyberpunk-terminal UI
+ +{FOOT}""" + return body + +def render_detail(m, data): + col, chip = verdict_meta(m["verdict"]) + # derive "what went right" from high pillars, "wrong" from low + critical_bugs + bd = m["breakdown"] + ranked = sorted(PILLARS, key=lambda p: -bd[p]) + rights = [f"{PILLAR_LABELS[p]} ({bd[p]}/20)" for p in ranked if bd[p] >= 16] + wrongs_pillars = [f"{PILLAR_LABELS[p]} ({bd[p]}/20)" for p in ranked if bd[p] <= 13] + + pillar_bars = "" + for p in PILLARS: + v = bd[p] + c = NEON_LIME if v >= 17 else (NEON_AMBER if v >= 13 else NEON_RED) + pillar_bars += f""" +
+
{PILLAR_LABELS[p]}
+
+
{v}/20
+
""" + + bugs_html = "".join(f"
  • {esc(b)}
  • " for b in m.get("critical_bugs", [])) + patch = m.get("patch_code", "") + metrics_block = "" + if m.get("tok_sec") is not None: + metrics_block = f""" +
    tok/sec{m['tok_sec']:.2f}
    +
    tokens{m.get('total_tokens') or '—'}
    +
    TTFT{m.get('ttft_sec'):.2f}s
    """ + else: + metrics_block = '
    runtimeCLOUD — not measured
    ' + + caveat_html = f'
    ⚠ SPEED CAVEAT: {esc(m["speed_caveat"])}
    ' if m.get("speed_caveat") else "" + tests_html = 'PASS' if m.get("tests_pass") else 'CRASH' + cloud_tag = 'CLOUD BASELINE' if m.get("format") == "cloud" else "" + + rights_html = "".join(f"
  • {esc(r)}
  • " for r in rights) or "
  • No pillar reached 16+ — no standout strengths.
  • " + wrongs_html = "".join(f"
  • {esc(w)}
  • " for w in wrongs_pillars) or "
  • No pillar fell below 14 — solid across the board.
  • " + + extra_css = """ +.mini-grid{display:grid;grid-template-columns:repeat(3,1fr);gap:10px;margin:14px 0} +.mini{background:var(--panel2);border:1px solid rgba(255,255,255,0.06);border-radius:5px;padding:10px 12px} +.mlbl{display:block;color:var(--dim);font-size:.62rem;letter-spacing:.14em;text-transform:uppercase} +.mval{display:block;font-size:1.15rem;margin-top:3px} +.caveat-box{background:rgba(255,176,0,0.08);border:1px solid rgba(255,176,0,0.4);color:var(--amber);padding:10px 14px;border-radius:5px;font-size:.82rem;margin:12px 0} +.pillar-list{margin-top:8px} +.prow{display:grid;grid-template-columns:170px 1fr 56px;align-items:center;gap:12px;margin:9px 0} +.plabel{color:var(--ink);font-size:.84rem} +.pbar{height:9px;background:rgba(255,255,255,0.06);border-radius:2px;overflow:hidden} +.pbar > i{display:block;height:100%;border-radius:2px} +.pval{text-align:right;font-size:.92rem} +.ok{color:var(--lime);border:1px solid var(--lime);padding:1px 8px;border-radius:3px;font-size:.7rem} +.bad{color:var(--red);border:1px solid var(--red);padding:1px 8px;border-radius:3px;font-size:.7rem} +.section{background:var(--panel);border:1px solid rgba(255,255,255,0.07);border-radius:8px;padding:18px;margin-top:18px} +.section h2{font-size:.9rem;letter-spacing:.12em;text-transform:uppercase;color:var(--cyan);margin:0 0 12px;text-shadow:0 0 10px rgba(0,255,200,0.35)} +ul.clean{list-style:none;padding:0;margin:0} +ul.clean li{padding:7px 0 7px 18px;border-bottom:1px solid rgba(255,255,255,0.05);position:relative;font-size:.88rem} +ul.clean li::before{content:"▸";position:absolute;left:0;color:var(--cyan)} +ul.clean li.dim{color:var(--dim)} +ul.clean li.dim::before{color:var(--dim)} +.bugs li::before{content:"✗";color:var(--red)} +.rights li::before{content:"✓";color:var(--lime)} +pre.code{background:#06060a;border:1px solid rgba(0,255,200,0.2);border-radius:6px;padding:14px;overflow-x:auto;color:#cfe;font-size:.78rem;line-height:1.5;box-shadow:inset 0 0 30px rgba(0,0,0,0.6)} +pre.code::before{content:"// patch.py";display:block;color:var(--dim);font-size:.68rem;margin-bottom:8px} +.back{display:inline-block;margin-bottom:18px;font-size:.8rem} +.verdict-chip-lg{display:inline-block;padding:4px 14px;border-radius:3px;border:1px solid currentColor;font-size:.78rem;letter-spacing:.1em} +.best-box{background:linear-gradient(135deg,rgba(0,255,200,0.08),rgba(91,139,255,0.05));border:1px solid rgba(0,255,200,0.25);border-radius:8px;padding:16px 18px;font-size:.95rem;line-height:1.6} +@media(max-width:640px){.prow{grid-template-columns:1fr auto}.mini-grid{grid-template-columns:repeat(2,1fr)}} +""" + head = head_html(m["model_name"]) + # inject extra css before + head = head.replace("", extra_css + "") + + return f""" +{head} +◂ BACK TO LEADERBOARD +
    +
    ▚ MODEL AUDIT // {esc(m['quant'])}
    +

    // {esc(m['model_name'])} {cloud_tag}

    +
    + {chip} — {esc(m['verdict'])} + {m['total_score']}/100 + tests: {tests_html} +
    +
    +{caveat_html} +
    {metrics_block}
    + +
    +

    ▮ PILLAR BREAKDOWN

    +
    {pillar_bars}
    +
    + +
    +
    +

    ✓ WENT RIGHT

    +
      {rights_html}
    +
    +
    +

    ✗ WENT WRONG

    +
      {wrongs_html}
    +
    +
    + +
    +

    ▮ CRITICAL BUGS

    +
      {bugs_html or '
    • None flagged.
    • '}
    +
    + +
    +

    ▮ RECOMMENDED USE

    +
    {esc(m['best_for'])}
    +
    + +
    +

    ▮ REFACTORED PATCH

    +
    {esc(patch)}
    +
    + + + +{FOOT}""" + +def main(): + with open(DATA) as f: + data = json.load(f) + os.makedirs(PAGES_DIR, exist_ok=True) + with open(OUT_DASH, "w") as f: + f.write(render_dashboard(data)) + for m in data["models"]: + with open(os.path.join(PAGES_DIR, m["id"] + ".html"), "w") as f: + f.write(render_detail(m, data)) + print(f"OK: dashboard.html + {len(data['models'])} pages in pages/") + +if __name__ == "__main__": + main() diff --git a/outputs/gemma-4-12b-coder-fable5-composer2.5-v1-uncensored-heretic-mxfp8-mlx.py b/outputs/gemma-4-12b-coder-fable5-composer2.5-v1-uncensored-heretic-mxfp8-mlx.py new file mode 100644 index 0000000..8569cdf --- /dev/null +++ b/outputs/gemma-4-12b-coder-fable5-composer2.5-v1-uncensored-heretic-mxfp8-mlx.py @@ -0,0 +1,262 @@ +import asyncio +import time +from dataclasses import dataclass, field +from typing import Any, Dict, Optional + + +@dataclass +class Node: + key: str + value: Any + ttl_expiry: float + freq: int = 1 + + def __repr__(self): + return f"Node({self.key}, freq={self.freq})" + + +class DoublyLinkedListList: + def __init__(self): + self.head = Node("", "", 0) # Dummy head + self.tail = Node("", "", 0) # Dummy tail + self.head.next = self.tail + self.tail.prev = self.head + + def add_to_front(self, node: Node): + node.next = self.head.next + node.prev = self.head + self.head.next = node + node.next.prev = node + + def remove(self, node: Node): + node.prev.next = node.next + node.next.prev = node.prev + + def move_to_front(self, node: Node): + self.remove(node) + self.add_to_front(node) + + def pop_tail(self) -> Optional[Node]: + if self.tail == self.tail.prev: # Only dummy tail left + return None + node = self.tail.prev + self.remove(node) + return node + + +class ConcurrentLFUCache: + def __init__(self, capacity: int): + if capacity <= 0: + raise ValueError("Capacity must be positive") + self.capacity = capacity + self.cache: Dict[str, Node] = {} # key -> node + self.freq_buckets: Dict[int, DoublyLinkedListList] = {} # freq -> list + self.min_freq = 0 + self.global_lock = asyncio.Lock() # Protects global cache structure + self.evictor_task: Optional[asyncio.Task] = None + + def _update_freq(self, node: Node): + old_freq = node.freq + bucket = self.freq_buckets[old_freq] + bucket.remove(node) + + if old_freq == self.min_freq and bucket.tail == bucket.head: + self.min_freq += 1 + + node.freq += 1 + if node.freq not in self.freq_buckets: + self.freq_buckets[node.freq] = DoublyLinkedListList() + + bucket = self.freq_buckets[node.freq] + bucket.add_to_front(node) + + async def get(self, key: str) -> Optional[Any]: + now = time.time() + async with self.global_lock: + node = self.cache.get(key) + if not node: + return None + + if now > node.ttl_expiry: # Lazy eviction + bucket = self.freq_buckets[node.freq] + bucket.remove(node) + del self.cache[key] + return None + + self._update_freq(node) + return node.value + + async def put(self, key: str, value: Any, ttl_seconds: float): + now = time.time() + expiry = now + ttl_seconds + + async with self.global_lock: + if key in self.cache: + node = self.cache[key] + bucket = self.freq_buckets[node.freq] + bucket.remove(node) + node.value = value + node.ttl_expiry = expiry + self._update_freq(node) + else: + if len(self.cache) >= self.capacity: + bucket = self.freq_buckets[self.min_freq] + evicted = bucket.pop_tail() + if evicted: + del self.cache[evicted.key] + + new_node = Node(key, value, expiry) + self.cache[key] = new_node + if self.min_freq == 0: + self.min_freq = 1 + bucket = self.freq_buckets[self.min_freq] + bucket.add_to_front(new_node) + + async def start_evictor(self, interval: float = 1.0): + def evict_loop(): + while True: + time.sleep(interval) + now = time.time() + expired_keys = [k for k, n in self.cache.items() if now > n.ttl_expiry] + for k in expired_keys: + asyncio.run(self.put(k, self.cache[k].value, 0)) # Re-insert with 0 TTL + + self.evictor_task = asyncio.create_task(evict_loop()) + + async def stop_evictor(self): + if self.evictor_task: + self.evictor_task.cancel() + + +class Transaction: + def __init__(self, cache): + self.cache = cache + self.local_cache: Dict[str, Node] = {} # key -> node + self.local_freq: Dict[int, DoublyLinkedListList] = {} # freq -> list + self.local_min_freq = 0 + + def _update_local_freq(self, node: Node): + old_freq = node.freq + bucket = self.local_freq[old_freq] + bucket.remove(node) + + if old_freq == self.local_min_freq and bucket.tail == bucket.head: + self.local_min_freq += 1 + + node.freq += 1 + if node.freq not in self.local_freq: + self.local_freq[node.freq] = DoublyLinkedListList() + + bucket = self.local_freq[node.freq] + bucket.add_to_front(node) + + async def get(self, key: str) -> Optional[Any]: + if key in self.local_cache: + node = self.local_cache[key] + self._update_local_freq(node) + return node.value + + async with self.cache.global_lock: + node = self.cache.cache.get(key) + if not node: + return None + + # Copy the node so local changes don't affect global cache until commit + new_node = Node(key, node.value, node.ttl_expiry, node.freq) + self.local_cache[key] = new_node + if self.local_min_freq == 0: + self.local_min_freq = new_node.freq + bucket = self.local_freq[new_node.freq] + bucket.add_to_front(new_node) + return new_node.value + + async def put(self, key: str, value: Any, ttl_seconds: float): + now = time.time() + expiry = now + ttl_seconds + + if key in self.local_cache: + node = self.local_cache[key] + bucket = self.local_freq[node.freq] + bucket.remove(node) + node.value = value + node.ttl_expiry = expiry + self._update_local_freq(node) + else: + if len(self.local_cache) >= self.cache.capacity: + bucket = self.local_freq[self.local_min_freq] + evicted = bucket.pop_tail() + if evicted: + del self.local_cache[evicted.key] + + new_node = Node(key, value, expiry) + self.local_cache[key] = new_node + if self.local_min_freq == 0: + self.local_min_freq = new_node.freq + bucket = self.local_freq[new_node.freq] + bucket.add_to_front(new_node) + + async def commit(self): + async with self.cache.global_lock: + for key, node in self.local_cache.items(): + if key in self.cache.cache: + old_node = self.cache.cache[key] + bucket = self.cache.freq_buckets[old_node.freq] + bucket.remove(old_node) + self.cache.cache[key] = node + if old_node.freq == self.cache.min_freq and bucket.tail == bucket.head: + self.cache.min_freq += 1 + node.freq = node.freq # freq is already updated in local_cache + bucket = self.cache.freq_buckets[node.freq] + if node.freq not in self.cache.freq_buckets: + self.cache.freq_buckets[node.freq] = DoublyLinkedListList() + bucket.add_to_front(node) + + async def rollback(self): + self.local_cache = {} + self.local_freq = {} + self.local_min_freq = 0 + + +async def main(): + cache = ConcurrentLFUCache(capacity=3) + + # a. O(1) LFU eviction order + await cache.put("a", "val_a", 60) # freq=1 + await cache.put("b", "val_b", 60) # freq=1 + await cache.put("c", "val_c", 60) # freq=1 + await cache.get("a") # a:freq=2, b/c:freq=1 + await cache.get("b") # a:freq=2, b:freq=2, c:freq=1 + await cache.put("d", "val_d", 60) # c is evicted (lowest freq=1) + assert await cache.get("c") is None + + # b. Lazy vs Background eviction + await cache.put("e", "val_e", 0) # expire immediately + assert await cache.get("e") is None # lazy eviction + + # c. Transaction commit vs rollback + tx = await cache.begin_transaction() # helper below + await tx.put("f", "val_f", 60) + assert await tx.get("f") == "val_f" # Read Your Own Writes + assert await cache.get("f") is None # Global not yet committed + + await tx.rollback() + assert await cache.get("f") is None # Rollback restored + + await tx.commit() + assert await cache.get("f") == "val_f" # Committed + + # d. Stress test + async def worker(i): + await cache.put(f"key_{i}", i, 10) + await cache.get(f"key_{i}") + + tasks = [worker(i) for i in range(50)] + await asyncio.gather(*tasks) + +def begin_transaction(): # Helper for the test suite + return Transaction(ConcurrentLFUCache(capacity=3)) + + +if __name__ == "__main__": + asyncio.run(main()) + diff --git a/outputs/gemma-4-31b-qat-gguf.py b/outputs/gemma-4-31b-qat-gguf.py new file mode 100644 index 0000000..392f764 --- /dev/null +++ b/outputs/gemma-4-31b-qat-gguf.py @@ -0,0 +1,343 @@ +import asyncio +import time +from dataclasses import dataclass +from typing import Any, Dict, Optional, Generic, TypeVar, Set + +T = TypeVar("T") + +# --- Internal Data Structures for O(1) LFU --- + +@dataclass +class Node: + """A node in the LFU frequency doubly linked list.""" + key: Any + value: Any + freq: int = 1 + expiry: Optional[float] = None + prev: Optional['Node'] = None + next: Optional['Node'] = None + +class DoublyLinkedList: + """A standard DLL to maintain nodes of the same frequency.""" + def __init__(self): + self.head: Optional[Node] = None + self.tail: Optional[Node] = None + self.size = 0 + + def append(self, node: Node): + """Add node to the front (most recently used in this freq bucket).""" + node.next = self.head + node.prev = None + if self.head: + self.head.prev = node + self.head = node + if not self.tail: + self.tail = node + self.size += 1 + + def pop_tail(self) -> Optional[Node]: + """Remove and return the least recently used node in this freq bucket.""" + if not self.tail: + return None + node = self.tail + self.remove(node) + return node + + def remove(self, node: Node): + """Remove a specific node from the list in O(1).""" + if node.prev: + node.prev.next = node.next + else: + self.head = node.next + + if node.next: + node.next.prev = node.prev + else: + self.tail = node.prev + + node.next = None + node.prev = None + self.size -= 1 + +# --- Cache Implementation --- + +class LFUCache: + def __init__(self, capacity: int): + if capacity <= 0: + raise ValueError("Capacity must be greater than 0") + + self.capacity = capacity + self.cache: Dict[Any, Node] = {} # key -> Node + self.freq_map: Dict[int, DoublyLinkedList] = {} # freq -> DLL + self.min_freq = 0 + self.lock = asyncio.Lock() + self._evictor_task: Optional[asyncio.Task] = None + + def _get_now(self) -> float: + return time.time() + + def _is_expired(self, node: Node) -> bool: + return node.expiry is not None and self._get_now() > node.expiry + + async def _update_freq(self, node: Node): + """Moves a node to the next frequency bucket in O(1).""" + old_freq = node.freq + dll = self.freq_map[old_freq] + dll.remove(node) + + if old_freq == self.min_freq and dll.size == 0: + self.min_freq += 1 + + node.freq += 1 + if node.freq not in self.freq_map: + self.freq_map[node.freq] = DoublyLinkedList() + self.freq_map[node.freq].append(node) + + async def get(self, key: Any) -> Optional[Any]: + """O(1) Access with Lazy TTL check.""" + async with self.lock: + if key not in self.cache: + return None + + node = self.cache[key] + if self._is_expired(node): + await self._delete_node(node) + return None + + await self._update_freq(node) + return node.value + + async def put(self, key: Any, value: Any, ttl_seconds: Optional[float] = None): + """O(1) Insertion with LFU eviction and Lazy TTL check.""" + async with self.lock: + expiry = self._get_now() + ttl_seconds if ttl_seconds else None + + if key in self.cache: + node = self.cache[key] + node.value = value + node.expiry = expiry + await self._update_freq(node) + return + + # Evict if capacity reached + if len(self.cache) >= self.capacity: + # Priority 1: Try to find an expired key first (optimistic) + # Otherwise: Evict the LFU (min_freq tail) + evicted = self.freq_map[self.min_freq].pop_tail() + if evicted: + del self.cache[evicted.key] + + # Create new node + new_node = Node(key=key, value=value, expiry=expiry) + self.cache[key] = new_node + self.min_freq = 1 + if 1 not in self.freq_map: + self.freq_map[1] = DoublyLinkedList() + self.freq_map[1].append(new_node) + + async def delete(self, key: Any): + """O(1) deletion.""" + async with self.lock: + if key in self.cache: + await self._delete_node(self.cache[key]) + + async def _delete_node(self, node: Node): + """Internal helper to remove node from all tracking structures.""" + dll = self.freq_map.get(node.freq) + if dll: + dll.remove(node) + if node.key in self.cache: + del self.cache[node.key] + + # --- Background Evictor --- + + async def start_evictor(self, interval: float = 1.0, batch_size: int = 100): + """Starts a non-blocking background task to purge expired keys.""" + if self._evictor_task: + return + + async def evict_loop(): + while True: + await asyncio.sleep(interval) + # Collect keys to check outside the lock to prevent blocking reads + keys_to_check = list(self.cache.keys()) + + # Process in small batches to yield to the event loop + for i in range(0, len(keys_to_check), batch_size): + batch = keys_to_check[i : i + batch_size] + async with self.lock: + for k in batch: + node = self.cache.get(k) + if node and self._is_expired(node): + await self._delete_node(node) + # Yield control back to asyncio + await asyncio.sleep(0) + + self._evictor_task = asyncio.create_task(evict_loop()) + + async def stop_evictor(self): + """Stops the background eviction task.""" + if self._evictor_task: + self._evictor_task.cancel() + try: + await self._evictor_task + except asyncio.CancelledError: + pass + self._evictor_task = None + + # --- Transaction Support --- + + def begin_transaction(self) -> 'Transaction': + return Transaction(self) + + async def apply_transaction_changes(self, changes: 'TransactionState'): + """Atomically apply buffered transaction writes to the global state.""" + async with self.lock: + for key, action in changes.writes.items(): + if action == "DELETE": + if key in self.cache: + await self._delete_node(self.cache[key]) + else: + val, ttl = action + # Using put logic inside the lock + expiry = self._get_now() + ttl if ttl else None + if key in self.cache: + node = self.cache[key] + node.value = val + node.expiry = expiry + await self._update_freq(node) + else: + if len(self.cache) >= self.capacity: + evicted = self.freq_map[self.min_freq].pop_tail() + if evicted: + del self.cache[evicted.key] + new_node = Node(key=key, value=val, expiry=expiry) + self.cache[key] = new_node + self.min_freq = 1 + if 1 not in self.freq_map: + self.freq_map[1] = DoublyLinkedList() + self.freq_map[1].append(new_node) + +@dataclass +class TransactionState: + writes: Dict[Any, Any] = None # key -> (value, ttl) or "DELETE" + +class Transaction: + def __init__(self, cache: LFUCache): + self._cache = cache + self._state = TransactionState(writes={}) + self._committed = False + + async def put(self, key: Any, value: Any, ttl_seconds: Optional[float] = None): + if self._committed: raise RuntimeError("Transaction already committed") + self._state.writes[key] = (value, ttl_seconds) + + async def get(self, key: Any) -> Optional[Any]: + if self._committed: raise RuntimeError("Transaction already committed") + + # Read-Your-Own-Writes (RYOW) + if key in self._state.writes: + action = self._state.writes[key] + if action == "DELETE": + return None + return action[0] + + # Otherwise read from global cache + return await self._cache.get(key) + + async def delete(self, key: Any): + if self._committed: raise RuntimeError("Transaction already committed") + self._state.writes[key] = "DELETE" + + async def commit(self): + if self._committed: raise RuntimeError("Transaction already committed") + await self._cache.apply_transaction_changes(self._state) + self._committed = True + + async def rollback(self): + if self._committed: raise RuntimeError("Transaction already committed") + self._state.writes.clear() + self._committed = True + +# --- Executable Unit Tests --- + +async def main(): + print("🚀 Starting LFU Cache Suite...") + + # 1. Test O(1) LFU Eviction + print("\nTesting LFU Eviction Order...", end=" ") + cache = LFUCache(capacity=2) + await cache.put("a", 1) + await cache.put("b", 2) + await cache.get("a") # freq(a)=2, freq(b)=1 + await cache.put("c", 3) # b should be evicted + assert await cache.get("b") is None + assert await cache.get("a") == 1 + assert await cache.get("c") == 3 + print("✅ PASSED") + + # 2. Test Dual-Layer TTL + print("\nTesting TTL Eviction...", end=" ") + ttl_cache = LFUCache(capacity=10) + await ttl_cache.put("lazy", "val", ttl_seconds=0.1) + await ttl_cache.put("async", "val", ttl_seconds=0.1) + + # Lazy check + await asyncio.sleep(0.2) + assert await ttl_cache.get("lazy") is None, "Lazy eviction failed" + + # Background check + await ttl_cache.start_evictor(interval=0.1, batch_size=10) + await asyncio.sleep(0.2) + async with ttl_cache.lock: + assert "async" not in ttl_cache.cache, "Background evictor failed" + await ttl_cache.stop_evictor() + print("✅ PASSED") + + # 3. Test Atomic Transactions + print("\nTesting Atomic Transactions...", end=" ") + tx_cache = LFUCache(capacity=10) + await tx_cache.put("base", 100) + + # Test Commit + tx1 = tx_cache.begin_transaction() + await tx1.put("base", 200) + await tx1.put("new", 300) + assert await tx1.get("base") == 200 # RYOW + assert await tx_cache.get("base") == 100 # Isolation + await tx1.commit() + assert await tx_cache.get("base") == 200 + assert await tx_cache.get("new") == 300 + + # Test Rollback + tx2 = tx_cache.begin_transaction() + await tx2.put("base", 999) + await tx2.get("base") + await tx2.rollback() + assert await tx_cache.get("base") == 200 + print("✅ PASSED") + + # 4. Stress Test: Concurrent Async Tasks + print("\nStress Testing 50 concurrent tasks...", end=" ") + stress_cache = LFUCache(capacity=20) + + async def worker(worker_id): + for i in range(100): + key = f"key_{i % 30}" # Contention on keys + if i % 3 == 0: + await stress_cache.put(key, f"val_{worker_id}_{i}") + elif i % 3 == 1: + await stress_cache.get(key) + else: + tx = stress_cache.begin_transaction() + await tx.put(key, f"tx_{worker_id}_{i}") + await tx.commit() + + await asyncio.gather(*(worker(i) for i in range(50))) + print("✅ PASSED") + + print("\n✨ All tests passed successfully!") + +if __name__ == "__main__": + asyncio.run(main()) + diff --git a/outputs/gemma4-31b-gguf.py b/outputs/gemma4-31b-gguf.py new file mode 100644 index 0000000..ab57a48 --- /dev/null +++ b/outputs/gemma4-31b-gguf.py @@ -0,0 +1,339 @@ +import asyncio +import time +from dataclasses import dataclass +from typing import Any, Dict, Optional, Generic, TypeVar, Set + +T = TypeVar("T") + +@dataclass +class Node: + """A node in the LFU doubly linked list.""" + key: Any + value: Any + freq: int = 1 + expiry: float = float('inf') + prev: Optional['Node'] = None + next: Optional['Node'] = None + +class DoublyLinkedList: + """Standard O(1) doubly linked list for maintaining frequency buckets.""" + def __init__(self): + self.head = Node(None, None) # Sentinel head + self.tail = Node(None, None) # Sentinel tail + self.head.next = self.tail + self.tail.prev = self.head + self.size = 0 + + def append(self, node: Node): + """Adds a node to the end (most recently used in this frequency).""" + node.next = self.tail + node.prev = self.tail.prev + self.tail.prev.next = node + self.tail.prev = node + self.size += 1 + + def remove(self, node: Node): + """Removes a specific node from the list.""" + if node.prev: + node.prev.next = node.next + if node.next: + node.next.prev = node.prev + node.prev = None + node.next = None + self.size -= 1 + + def pop_front(self) -> Optional[Node]: + """Removes and returns the oldest node (LFU candidate).""" + if self.size == 0: + return None + node = self.head.next + self.remove(node) + return node + + def is_empty(self) -> bool: + return self.size == 0 + + +class LFUCache: + """ + In-Memory Concurrent LFU Cache with Async TTL Eviction and Atomic Transactions. + Time Complexity: O(1) for get and put. + Space Complexity: O(N). + """ + def __init__(self, capacity: int): + if capacity <= 0: + raise ValueError("Capacity must be greater than 0") + + self.capacity = capacity + self.cache: Dict[Any, Node] = {} # Key -> Node + self.freq_map: Dict[int, DoublyLinkedList] = {} # Freq -> DLL + self.min_freq = 0 + self._lock = asyncio.Lock() + self._evictor_task: Optional[asyncio.Task] = None + + async def get(self, key: Any) -> Optional[Any]: + async with self._lock: + return await self._get_internal(key) + + async def _get_internal(self, key: Any) -> Optional[Any]: + """Internal get without lock for transaction use.""" + if key not in self.cache: + return None + + node = self.cache[key] + + # Lazy Eviction check + if time.time() > node.expiry: + await self._delete_internal(key) + return None + + self._update_frequency(node) + return node.value + + async def put(self, key: Any, value: Any, ttl_seconds: Optional[float] = None): + async with self._lock: + await self._put_internal(key, value, ttl_seconds) + + async def _put_internal(self, key: Any, value: Any, ttl_seconds: Optional[float] = None): + """Internal put without lock for transaction use.""" + expiry = time.time() + ttl_seconds if ttl_seconds is not None else float('inf') + + if key in self.cache: + node = self.cache[key] + node.value = value + node.expiry = expiry + self._update_frequency(node) + else: + # Capacity Management + if len(self.cache) >= self.capacity: + await self._evict_lfu() + + new_node = Node(key, value, freq=1, expiry=expiry) + self.cache[key] = new_node + self._add_to_freq_bucket(1, new_node) + self.min_freq = 1 + + async def delete(self, key: Any): + async with self._lock: + await self._delete_internal(key) + + async def _delete_internal(self, key: Any): + """Internal delete without lock for transaction use.""" + if key in self.cache: + node = self.cache.pop(key) + self.freq_map[node.freq].remove(node) + # Note: We don't strictly need to update min_freq here because + # _evict_lfu handles empty buckets by incrementing. + + def _update_frequency(self, node: Node): + """Moves a node to the next frequency bucket in O(1).""" + old_freq = node.freq + self.freq_map[old_freq].remove(node) + + if old_freq == self.min_freq and self.freq_map[old_freq].is_empty(): + self.min_freq += 1 + + node.freq += 1 + self._add_to_freq_bucket(node.freq, node) + + def _add_to_freq_bucket(self, freq: int, node: Node): + if freq not in self.freq_map: + self.freq_map[freq] = DoublyLinkedList() + self.freq_map[freq].append(node) + + async def _evict_lfu(self): + """Evicts the least frequently used (and oldest within that freq) item.""" + while self.min_freq not in self.freq_map or self.freq_map[self.min_freq].is_empty(): + # This handles cases where items were deleted manually + if not self.cache: return + self.min_freq += 1 # Search for the next available freq bucket + + evicted_node = self.freq_map[self.min_freq].pop_front() + if evicted_node: + del self.cache[evicted_node.key] + + # --- Background Evictor --- + + async def start_evictor(self, interval: float = 1.0): + """Starts the background task that purges expired keys.""" + if self._evictor_task is not None: + return + self._evictor_task = asyncio.create_task(self._eviction_loop(interval)) + + async def stop_evictor(self): + """Stops the background eviction task.""" + if self._evictor_task: + self._evictor_task.cancel() + try: + await self._evictor_task + except asyncio.CancelledError: + pass + self._evictor_task = None + + async def _eviction_loop(self, interval: float): + while True: + await asyncio.sleep(interval) + # Process in small batches to avoid blocking the main lock for too long + now = time.time() + keys_to_check = list(self.cache.keys()) + + batch_size = 100 + for i in range(0, len(keys_to_check), batch_size): + async with self._lock: + batch = keys_to_check[i : i + batch_size] + for k in batch: + if k in self.cache and now > self.cache[k].expiry: + await self._delete_internal(k) + + # --- Transaction Logic --- + + def begin_transaction(self): + return Transaction(self) + + +class Transaction: + """ + Provides ACID-like atomic updates to the LFUCache. + Implements 'Read Your Own Writes' (RYOW). + """ + def __init__(self, cache: LFUCache): + self._cache = cache + self._pending_puts: Dict[Any, tuple[Any, Optional[float]]] = {} + self._pending_deletes: Set[Any] = set() + self._committed = False + + async def get(self, key: Any) -> Optional[Any]: + if self._committed: raise RuntimeError("Transaction already committed") + + # 1. Check local deletes first + if key in self._pending_deletes: + return None + + # 2. Check local writes (RYOW) + if key in self._pending_puts: + return self._pending_puts[key][0] + + # 3. Fallback to global cache + return await self._cache.get(key) + + async def put(self, key: Any, value: Any, ttl_seconds: Optional[float] = None): + if self._committed: raise RuntimeError("Transaction already committed") + self._pending_deletes.discard(key) + self._pending_puts[key] = (value, ttl_seconds) + + async def delete(self, key: Any): + if self._committed: raise RuntimeError("Transaction already committed") + self._pending_puts.pop(key, None) + self._pending_deletes.add(key) + + async def commit(self): + """Applies all pending changes to the global cache atomically.""" + if self._committed: return + + async with self._cache._lock: + # Apply deletes first + for key in self._pending_deletes: + await self._cache._delete_internal(key) + # Apply puts + for key, (val, ttl) in self._pending_puts.items(): + await self._cache._put_internal(key, val, ttl) + + self._committed = True + + async def rollback(self): + """Discards all pending changes.""" + self._pending_puts.clear() + self._pending_deletes.clear() + self._committed = True + + +# ========================================== +# Unit Tests +# ========================================== + +async def main(): + print("Starting LFU Cache Comprehensive Test Suite...") + + # 1. Test O(1) LFU Eviction Order + print("\nTesting LFU Eviction Order...", end=" ") + cache = LFUCache(capacity=2) + await cache.put("a", 1) # freq 1 + await cache.put("b", 2) # freq 1 + await cache.get("a") # freq 2 + await cache.put("c", 3) # "b" should be evicted as it's the LFU (freq 1) + assert await cache.get("b") is None, "Key 'b' should have been evicted" + assert await cache.get("a") == 1, "Key 'a' should still exist" + assert await cache.get("c") == 3, "Key 'c' should still exist" + print("PASSED") + + # 2. Test TTL Eviction (Lazy & Background) + print("Testing Dual-Layer TTL...", end=" ") + cache = LFUCache(capacity=10) + await cache.start_evictor(interval=0.1) + + await cache.put("lazy", "val", ttl_seconds=0.1) + await asyncio.sleep(0.2) + # Lazy check: get() should trigger deletion + assert await cache.get("lazy") is None, "Lazy eviction failed" + + await cache.put("bg", "val", ttl_seconds=0.1) + await asyncio.sleep(0.3) + # Background check: key should be gone from internal map without calling get() + async with cache._lock: + assert "bg" not in cache.cache, "Background eviction failed" + await cache.stop_evictor() + print("PASSED") + + # 3. Test Atomic Transactions (Commit vs Rollback) + print("Testing Transaction Isolation...", end=" ") + cache = LFUCache(capacity=10) + await cache.put("shared", "original") + + tx = cache.begin_transaction() + await tx.put("shared", "modified") + await tx.put("new", "added") + + # Verify Isolation: Global reader shouldn't see changes + assert await cache.get("shared") == "original" + assert await cache.get("new") is None + + # Verify RYOW: Transaction should see its own writes + assert await tx.get("shared") == "modified" + assert await tx.get("new") == "added" + + await tx.commit() + assert await cache.get("shared") == "modified", "Commit failed to update global state" + assert await cache.get("new") == "added", "Commit failed to add new key" + + tx2 = cache.begin_transaction() + await tx2.put("shared", "rollback_this") + assert await tx2.get("shared") == "rollback_this" + await tx2.rollback() + assert await cache.get("shared") == "modified", "Rollback mutated global state" + print("PASSED") + + # 4. Stress Test: Concurrent Access + print("Testing High Concurrency (50 tasks)...", end=" ") + cache = LFUCache(capacity=100) + + async def worker(worker_id): + for i in range(100): + key = f"key_{i % 20}" # Create contention on keys + if i % 3 == 0: + await cache.put(key, f"val_{worker_id}_{i}", ttl_seconds=0.5) + elif i % 3 == 1: + await cache.get(key) + else: + tx = cache.begin_transaction() + await tx.put(f"tx_{worker_id}", i) + await tx.commit() + + tasks = [worker(i) for i in range(50)] + await asyncio.gather(*tasks) + print("PASSED") + + print("\nAll tests passed successfully!") + +if __name__ == "__main__": + asyncio.run(main()) + diff --git a/outputs/kat-coder-v2.5-dev-xl-mlx.py b/outputs/kat-coder-v2.5-dev-xl-mlx.py new file mode 100644 index 0000000..5fd86f2 --- /dev/null +++ b/outputs/kat-coder-v2.5-dev-xl-mlx.py @@ -0,0 +1,550 @@ +""" +In-Memory Concurrent LFU Cache with Async TTL Eviction and Atomic Transactions. + +Pure Python 3.11+ implementation — no third-party dependencies. +""" + +from __future__ import annotations + +import asyncio +import time +import weakref +from dataclasses import dataclass, field +from typing import Any, Optional + + +# --------------------------------------------------------------------------- +# Doubly-Linked List (O(1) move/remove/add) +# --------------------------------------------------------------------------- + +@dataclass +class _DLLNode: + """Element inside a frequency bucket's doubly-linked list.""" + key: Any + value: Any + ttl_seconds: float + expiry_time: float + freq: int = 0 + prev: Optional[_DLLNode] = field(default=None, repr=False) + next: Optional[_DLLNode] = field(default=None, repr=False) + + +class _DoublyLinkedList: + """ + Doubly-linked list with sentinel head/tail for O(1) prepend, pop, and remove. + The *tail* side holds the least-recently-used node within a frequency tier, + which is also the LFU candidate for eviction. + """ + + def __init__(self) -> None: + self._head = _DLLNode(None, None, 0.0, 0.0) # sentinel + self._tail = _DLLNode(None, None, 0.0, 0.0) # sentinel + self._head.next = self._tail + self._tail.prev = self._head + self.size = 0 + + # -- internal helpers --------------------------------------------------- + + def _insert_after(self, node: _DLLNode, prev: _DLLNode) -> None: + nxt = prev.next + prev.next = node + node.prev = prev + node.next = nxt + nxt.prev = node + + def _unlink(self, node: _DLLNode) -> None: + prev, nxt = node.prev, node.next + prev.next = nxt + nxt.prev = prev + node.prev = node.next = None + + # -- public API --------------------------------------------------------- + + def push_front(self, node: _DLLNode) -> None: + """Insert *node* right after the head sentinel (most-recent).""" + self._insert_after(node, self._head) + self.size += 1 + + def pop_tail(self) -> Optional[_DLLNode]: + """Remove and return the node just before the tail sentinel (LRU).""" + if self.size == 0: + return None + node = self._tail.prev + self._unlink(node) + return node + + def remove(self, node: _DLLNode) -> None: + """Remove an arbitrary node from the list.""" + self._unlink(node) + + def is_empty(self) -> bool: + return self.size == 0 + + +# --------------------------------------------------------------------------- +# Cache Node (wraps the DLL node + TTL metadata) +# --------------------------------------------------------------------------- + +@dataclass +class _CacheNode: + key: Any + value: Any + ttl_seconds: float + expiry_time: float + freq: int = 1 + dll_node: Optional[_DLLNode] = field(default=None, repr=False) + + +# --------------------------------------------------------------------------- +# Transaction Handle +# --------------------------------------------------------------------------- + +class Transaction: + """ + Represents an isolated sub-session on the cache. + + - ``tx.put(key, value, ttl)`` buffers a write locally. + - ``tx.get(key)`` reads from local buffer first, then falls back to the + global cache (without mutating global frequency state). + - ``tx.delete(key)`` buffers a deletion locally. + - ``await tx.commit()`` atomically applies all buffered writes to the global cache. + - ``tx.rollback()`` discards everything. + """ + + def __init__(self, cache: "LFUCache") -> None: + self._cache = cache + self._pending_writes: dict[Any, tuple[Any, float]] = {} # key -> (value, expiry) + self._pending_deletes: set[Any] = set() + self._committed = False + + # -- read ---------------------------------------------------------------- + + def get(self, key: Any) -> Optional[Any]: + """Read with read-your-own-writes semantics.""" + if self._committed: + raise RuntimeError("Transaction already committed") + + # 1. Check local pending writes first + if key in self._pending_writes: + return self._pending_writes[key][0] + if key in self._pending_deletes: + return None + + # 2. Fall back to global cache (read-only, no frequency bump) + return self._cache._get_raw(key) + + # -- write --------------------------------------------------------------- + + def put(self, key: Any, value: Any, ttl_seconds: float = 60.0) -> None: + """Buffer a write locally; not visible to others until commit.""" + if self._committed: + raise RuntimeError("Transaction already committed") + expiry = time.monotonic() + ttl_seconds + self._pending_writes[key] = (value, expiry) + # If previously deleted in this txn, re-add overrides the delete. + self._pending_deletes.discard(key) + + def delete(self, key: Any) -> None: + """Buffer a deletion locally.""" + if self._committed: + raise RuntimeError("Transaction already committed") + self._pending_deletes.add(key) + self._pending_writes.pop(key, None) + + # -- commit / rollback --------------------------------------------------- + + async def commit(self) -> None: + """Atomically apply all buffered changes to the global cache.""" + if self._committed: + raise RuntimeError("Transaction already committed") + async with self._cache._lock: + for key, (value, expiry) in self._pending_writes.items(): + await self._cache._put_internal(key, value, expiry) + for key in self._pending_deletes: + await self._cache._delete_internal(key) + self._committed = True + + def rollback(self) -> None: + """Discard all pending changes.""" + self._pending_writes.clear() + self._pending_deletes.clear() + self._committed = True # mark so further ops raise + + +# --------------------------------------------------------------------------- +# LFU Cache +# --------------------------------------------------------------------------- + +class LFUCache: + """ + In-memory concurrent LFU cache with O(1) get/put, async TTL eviction, + and atomic transaction support. + """ + + def __init__(self, capacity: int = 1024) -> None: + if capacity < 1: + raise ValueError("capacity must be >= 1") + + self._capacity = capacity + self._lock = asyncio.Lock() + self._evictor_task: Optional[asyncio.Task[None]] = None + + # Core O(1) structures + self._cache_map: dict[Any, _CacheNode] = {} # key -> CacheNode + self._freq_map: dict[int, _DoublyLinkedList] = {} # freq -> DLL + self._min_freq: int = 1 + + # ------------------------------------------------------------------ + # Public API + # ------------------------------------------------------------------ + + async def get(self, key: Any) -> Optional[Any]: + """O(1) lookup with lazy TTL eviction.""" + async with self._lock: + node = self._cache_map.get(key) + if node is None: + return None + # Lazy TTL check + if time.monotonic() > node.expiry_time: + await self._evict_node(key, node) + return None + # Bump frequency — O(1) + await self._bump_freq(node) + return node.value + + async def put(self, key: Any, value: Any, ttl_seconds: float = 60.0) -> None: + """O(1) insert/update with lazy TTL eviction of LRU-LFU victim if needed.""" + expiry = time.monotonic() + ttl_seconds + async with self._lock: + await self._put_internal(key, value, expiry) + + async def delete(self, key: Any) -> bool: + """O(1) deletion.""" + async with self._lock: + return await self._delete_internal(key) + + def begin_transaction(self) -> Transaction: + """Start a new isolated transaction.""" + return Transaction(self) + + # ------------------------------------------------------------------ + # TTL Eviction Loop + # ------------------------------------------------------------------ + + def start_evictor(self, interval_seconds: float = 1.0) -> None: + """Start the background async TTL sweep task.""" + if self._evictor_task is not None and not self._evictor_task.done(): + return + self._evictor_task = asyncio.create_task(self._eviction_loop(interval_seconds)) + + def stop_evictor(self) -> None: + """Stop the background eviction task.""" + if self._evictor_task is not None: + self._evictor_task.cancel() + self._evictor_task = None + + async def _eviction_loop(self, interval: float) -> None: + """Periodically purge expired entries in small batches.""" + try: + while True: + await asyncio.sleep(interval) + async with self._lock: + now = time.monotonic() + # Collect expired keys in a snapshot to avoid dict-changed-size + expired = [ + k for k, n in self._cache_map.items() + if now > n.expiry_time + ] + for k in expired: + node = self._cache_map.get(k) + if node is not None and now > node.expiry_time: + await self._evict_node(k, node) + except asyncio.CancelledError: + return + + # ------------------------------------------------------------------ + # Internal helpers + # ------------------------------------------------------------------ + + async def _get_raw(self, key: Any) -> Optional[Any]: + """Read-only global lookup — does NOT bump frequency. For transactions.""" + node = self._cache_map.get(key) + if node is None: + return None + if time.monotonic() > node.expiry_time: + await self._evict_node(key, node) + return None + return node.value + + async def _put_internal(self, key: Any, value: Any, expiry: float) -> None: + """Core put logic (must be called under lock).""" + # If key already exists, update in place + if key in self._cache_map: + node = self._cache_map[key] + # Remove from old freq list, update value/ttl/freq + self._freq_map[node.freq].remove(node.dll_node) + if self._freq_map[node.freq].is_empty(): + del self._freq_map[node.freq] + if node.freq == self._min_freq: + self._min_freq += 1 + node.value = value + node.ttl_seconds = expiry - time.monotonic() + node.expiry_time = expiry + node.freq = 1 + self._ensure_freq_list(1).push_front(node.dll_node) + return + + # Evict if at capacity + if len(self._cache_map) >= self._capacity: + await self._evict_lfu() + + # Insert new node + dll_node = _DLLNode(key, value, expiry - time.monotonic(), expiry) + cache_node = _CacheNode(key, value, expiry - time.monotonic(), expiry, freq=1, dll_node=dll_node) + self._cache_map[key] = cache_node + self._freq_map[1].push_front(dll_node) + self._min_freq = 1 + + async def _delete_internal(self, key: Any) -> bool: + """Core delete logic (must be called under lock).""" + node = self._cache_map.get(key) + if node is None: + return False + await self._evict_node(key, node) + return True + + async def _evict_node(self, key: Any, node: _CacheNode) -> None: + """Remove a single node from all structures.""" + dll = self._freq_map.get(node.freq) + if dll is not None: + dll.remove(node.dll_node) + if dll.is_empty(): + del self._freq_map[node.freq] + if node.freq == self._min_freq: + # Find new min freq + new_min = min(self._freq_map) if self._freq_map else 1 + self._min_freq = new_min + del self._cache_map[key] + + async def _evict_lfu(self) -> None: + """Evict the least-frequently-used (and least-recently-used within that freq) node.""" + if not self._freq_map: + return + dll = self._freq_map.get(self._min_freq) + if dll is None or dll.is_empty(): + # Shouldn't happen, but safeguard + self._min_freq += 1 + await self._evict_lfu() + return + victim = dll.pop_tail() + if victim is not None: + await self._evict_node(victim.key, self._cache_map[victim.key]) + + async def _bump_freq(self, node: _CacheNode) -> None: + """Move a node from freq f to freq f+1. O(1).""" + old_freq = node.freq + old_dll = self._freq_map[old_freq] + old_dll.remove(node.dll_node) + if old_dll.is_empty(): + del self._freq_map[old_freq] + if old_freq == self._min_freq: + self._min_freq += 1 + + node.freq += 1 + new_dll = self._ensure_freq_list(node.freq) + new_dll.push_front(node.dll_node) + + def _ensure_freq_list(self, freq: int) -> _DoublyLinkedList: + """Return (or create) the DLL for *freq*.""" + if freq not in self._freq_map: + self._freq_map[freq] = _DoublyLinkedList() + return self._freq_map[freq] + + # ------------------------------------------------------------------ + # Diagnostics + # ------------------------------------------------------------------ + + @property + def size(self) -> int: + return len(self._cache_map) + + @property + def capacity(self) -> int: + return self._capacity + + +# --------------------------------------------------------------------------- +# Executable Test Suite +# --------------------------------------------------------------------------- + +async def main() -> None: + passed = 0 + failed = 0 + + def _check(name: str, condition: bool) -> None: + nonlocal passed, failed + if condition: + passed += 1 + print(f" ✓ {name}") + else: + failed += 1 + print(f" ✗ {name}") + + # ====================================================================== + # (a) O(1) LFU eviction order + # ====================================================================== + print("\n=== (a) LFU Eviction Order ===") + + cache = LFUCache(capacity=3) + + # Insert 3 items + await cache.put("a", 1, ttl_seconds=60.0) + await cache.put("b", 2, ttl_seconds=60.0) + await cache.put("c", 3, ttl_seconds=60.0) + + # Access "a" and "b" once each → freq=2; "c" stays at freq=1 + await cache.get("a") + await cache.get("b") + + # Insert "d" — should evict "c" (lowest freq) + await cache.put("d", 4, ttl_seconds=60.0) + + _check("a still present after eviction", await cache.get("a") == 1) + _check("b still present after eviction", await cache.get("b") == 2) + _check("c evicted (lowest freq)", await cache.get("c") is None) + _check("d present", await cache.get("d") == 4) + + # Now access "a" again → freq=3; "b" and "d" at freq=2 + await cache.get("a") + # Insert "e" — should evict either "b" or "d" (both freq=2, LRU wins) + await cache.put("e", 5, ttl_seconds=60.0) + + _check("a still present", await cache.get("a") == 1) + _check("e present", await cache.get("e") == 5) + + # ====================================================================== + # (b) Lazy TTL vs Background Async Sweep + # ====================================================================== + print("\n=== (b) TTL Eviction (Lazy + Background) ===") + + cache2 = LFUCache(capacity=10) + + await cache2.put("lazy_key", "lazy_val", ttl_seconds=0.1) + await cache2.put("bg_key", "bg_val", ttl_seconds=0.1) + + # Lazy eviction: access lazy_key after expiry + await asyncio.sleep(0.15) + _check("Lazy eviction: get returns None after TTL", await cache2.get("lazy_key") is None) + _check("Lazy eviction: bg_key still there (not accessed)", await cache2.get("bg_key") == "bg_val") + + # Start background evictor + cache2.start_evictor(interval_seconds=0.2) + await asyncio.sleep(0.3) # let background sweep run + + _check("Background eviction: bg_key purged by sweep", await cache2.get("bg_key") is None) + cache2.stop_evictor() + + # ====================================================================== + # (c) Transaction commit visibility vs rollback + # ====================================================================== + print("\n=== (c) Atomic Transactions ===") + + cache3 = LFUCache(capacity=10) + await cache3.put("x", 10, ttl_seconds=60.0) + await cache3.put("y", 20, ttl_seconds=60.0) + + # --- Commit test --- + tx1 = cache3.begin_transaction() + tx1.put("x", 99, ttl_seconds=60.0) # local write + tx1.put("z", 30, ttl_seconds=60.0) # new key + + _check("TX: read-your-own-write (x)", tx1.get("x") == 99) + _check("TX: read-your-own-write (z)", tx1.get("z") == 30) + _check("TX: global still sees old x", await cache3.get("x") == 10) + + await tx1.commit() + _check("TX: after commit, global sees x=99", await cache3.get("x") == 99) + _check("TX: after commit, global sees z=30", await cache3.get("z") == 30) + + # --- Rollback test --- + tx2 = cache3.begin_transaction() + tx2.put("x", -1, ttl_seconds=60.0) + tx2.delete("y") + _check("TX rollback: local sees x=-1", tx2.get("x") == -1) + _check("TX rollback: local sees y deleted", tx2.get("y") is None) + _check("TX rollback: global still sees x=99", await cache3.get("x") == 99) + _check("TX rollback: global still sees y=20", await cache3.get("y") == 20) + + tx2.rollback() + _check("TX rollback: global unchanged after rollback", await cache3.get("x") == 99) + _check("TX rollback: y still present after rollback", await cache3.get("y") == 20) + + # --- Double commit / rollback raises --- + tx3 = cache3.begin_transaction() + await tx3.commit() + try: + tx3.put("x", 1) + _check("TX: double commit raises", False) + except RuntimeError: + _check("TX: double commit raises", True) + + tx4 = cache3.begin_transaction() + tx4.rollback() + try: + tx4.put("x", 1) + _check("TX: op after rollback raises", False) + except RuntimeError: + _check("TX: op after rollback raises", True) + + # ====================================================================== + # (d) Stress test: 50 concurrent async tasks + # ====================================================================== + print("\n=== (d) Stress Test — 50 Concurrent Tasks ===") + + cache4 = LFUCache(capacity=200) + errors: list[str] = [] + + async def worker(task_id: int, base_key: int) -> None: + try: + for i in range(50): + key = f"t{task_id}_k{i}" + val = task_id * 1000 + i + await cache4.put(key, val, ttl_seconds=5.0) + result = await cache4.get(key) + if result != val: + errors.append(f"task={task_id} key={key} expected={val} got={result}") + # Occasional transaction + if i % 10 == 0: + tx = cache4.begin_transaction() + tx.put(f"tx_{task_id}_{i}", val * 2, ttl_seconds=5.0) + r = tx.get(f"tx_{task_id}_{i}") + if r != val * 2: + errors.append(f"task={task_id} tx key mismatch") + await tx.commit() + except Exception as e: + errors.append(f"task={task_id} exception: {e}") + + tasks = [asyncio.create_task(worker(tid, tid)) for tid in range(50)] + await asyncio.gather(*tasks) + + _check("Stress: no errors", len(errors) == 0) + _check("Stress: cache size within capacity", cache4.size <= 200) + if errors: + for e in errors[:5]: + print(f" error: {e}") + + # ====================================================================== + # Summary + # ====================================================================== + total = passed + failed + print(f"\n{'='*50}") + print(f"Results: {passed}/{total} passed, {failed} failed") + if failed == 0: + print("All tests passed! ✓") + else: + print(f"{failed} test(s) FAILED ✗") + + +if __name__ == "__main__": + asyncio.run(main()) + diff --git a/outputs/qwen3.6-35b-a3b-4bit-mlx.py b/outputs/qwen3.6-35b-a3b-4bit-mlx.py new file mode 100644 index 0000000..3fe6acf --- /dev/null +++ b/outputs/qwen3.6-35b-a3b-4bit-mlx.py @@ -0,0 +1,365 @@ +""" +In-Memory Concurrent LFU Cache with Async TTL Eviction and Atomic Transactions. +Pure Python 3.11+ implementation using only standard library modules. +""" + +import asyncio +import time +from typing import Any, Dict, List, Optional, Tuple +from asyncio import Lock + +# --------------------------------------------------------------------------- +# Internal Data Structures (O(1) LFU Core) +# --------------------------------------------------------------------------- + +class _Node: + """Doubly linked list node tracking key, value, frequency, and TTL.""" + __slots__ = ('key', 'value', 'freq', 'expires_at', 'prev', 'next') + def __init__(self, key: Any, value: Any, freq: int = 1, expires_at: float = 0.0): + self.key = key + self.value = value + self.freq = freq + self.expires_at = expires_at + self.prev = None + self.next = None + +class _FreqList: + """Doubly linked list maintaining nodes of a specific frequency tier.""" + __slots__ = ('head', 'tail', 'size') + def __init__(self): + self.head = _Node(None, None) # Dummy head + self.tail = _Node(None, None) # Dummy tail + self.head.next = self.tail + self.tail.prev = self.head + self.size = 0 + + def add(self, node: _Node): + """Add node to tail (most recently used in this frequency).""" + last = self.tail.prev + last.next = node + node.prev = last + node.next = self.tail + self.tail.prev = node + self.size += 1 + + def remove(self, node: _Node): + """Remove node from the list in O(1).""" + node.prev.next = node.next + node.next.prev = node.prev + node.prev = None + node.next = None + self.size -= 1 + + def pop(self) -> _Node: + """Remove and return node from head (least recently used).""" + node = self.head.next + self.remove(node) + return node + +# --------------------------------------------------------------------------- +# Main Cache Implementation +# --------------------------------------------------------------------------- + +class LFUCache: + """ + O(1) In-Memory Concurrent LFU Cache with Async TTL Eviction. + Uses a hash map for key lookup and a hash map of frequency-buckets + with doubly linked lists for O(1) frequency updates and eviction. + """ + def __init__(self, capacity: int, evict_interval: float = 1.0): + self.capacity = capacity + self.evict_interval = evict_interval + self.key_to_node: Dict[Any, _Node] = {} + self.freq_to_list: Dict[int, _FreqList] = {} + self.min_freq: int = 1 + self.lock: Lock = asyncio.Lock() + self._evict_task: Optional[asyncio.Task] = None + self._running = False + + async def start_evictor(self): + """Start the non-blocking background TTL eviction loop.""" + self._running = True + self._evict_task = asyncio.create_task(self._eviction_loop()) + + async def stop_evictor(self): + """Gracefully stop the background eviction task.""" + self._running = False + if self._evict_task: + self._evict_task.cancel() + try: await self._evict_task + except asyncio.CancelledError: pass + + async def get(self, key: Any) -> Optional[Any]: + """Retrieve value in O(1). Performs lazy TTL eviction and frequency update.""" + async with self.lock: + node = self.key_to_node.get(key) + if not node: + return None + + # Lazy TTL Eviction + if node.expires_at > 0 and node.expires_at <= time.time(): + self._remove_node(node) + return None + + # O(1) Frequency Update + self._update_freq(node) + return node.value + + async def put(self, key: Any, value: Any, ttl_seconds: float = 0.0): + """Insert/update value in O(1). Handles capacity eviction and TTL.""" + async with self.lock: + node = self.key_to_node.get(key) + if node: + # Update existing node + node.value = value + if ttl_seconds > 0: + node.expires_at = time.time() + ttl_seconds + self._update_freq(node) + else: + # Evict if capacity reached + if len(self.key_to_node) >= self.capacity: + self._evict() + # Insert new node with freq=1 + node = _Node( + key, value, freq=1, + expires_at=float('inf') if ttl_seconds <= 0 else time.time() + ttl_seconds + ) + self.key_to_node[key] = node + self._add_to_freq(node) + self.min_freq = 1 # New keys always start at freq 1 + + async def begin_transaction(self) -> 'Transaction': + """Create an isolated transaction session.""" + return Transaction(self) + + # --- Internal Cache Helpers (O(1)) --- + + def _remove_node(self, node: _Node): + """Remove node from global structures.""" + del self.key_to_node[node.key] + self.freq_to_list[node.freq].remove(node) + if self.freq_to_list[node.freq].size == 0: + del self.freq_to_list[node.freq] + if self.min_freq == node.freq: + self.min_freq += 1 + + def _update_freq(self, node: _Node): + """Move node to next frequency bucket in O(1).""" + old_freq = node.freq + self.freq_to_list[old_freq].remove(node) + if self.freq_to_list[old_freq].size == 0: + del self.freq_to_list[old_freq] + if self.min_freq == old_freq: + self.min_freq += 1 + + node.freq += 1 + self._add_to_freq(node) + + def _add_to_freq(self, node: _Node): + """Add node to its frequency bucket.""" + if node.freq not in self.freq_to_list: + self.freq_to_list[node.freq] = _FreqList() + self.freq_to_list[node.freq].add(node) + + def _evict(self): + """Evict least frequently used (then LRU) node in O(1).""" + if not self.freq_to_list: + return + evict_list = self.freq_to_list[self.min_freq] + node = evict_list.pop() + self._remove_node(node) + + async def _eviction_loop(self): + """Background async task that purges expired keys in small batches.""" + while self._running: + await asyncio.sleep(self.evict_interval) + async with self.lock: + # Batch scan limited to avoid blocking reads + keys_to_check = list(self.key_to_node.keys())[:50] + for k in keys_to_check: + node = self.key_to_node.get(k) + if node and node.expires_at > 0 and node.expires_at <= time.time(): + self._remove_node(node) + +# --------------------------------------------------------------------------- +# Transaction Isolation Layer +# --------------------------------------------------------------------------- + +class Transaction: + """ + Atomic transaction session providing Read-Your-Own-Writes isolation. + Global readers remain blind to uncommitted changes until commit(). + """ + def __init__(self, cache: LFUCache): + self.cache = cache + self.buffer: List[Dict] = [] + self.committed = False + + async def get(self, key: Any) -> Optional[Any]: + """Read with local isolation. Does not update global frequencies.""" + # Read-Your-Own-Writes: check buffer first + for op in reversed(self.buffer): + if op['key'] == key: + return None if op['op'] == 'delete' else op['value'] + + # Global read (no freq update to preserve isolation) + async with self.cache.lock: + node = self.cache.key_to_node.get(key) + if not node: + return None + # Lazy TTL check without mutation + if node.expires_at > 0 and node.expires_at <= time.time(): + return None + return node.value + + async def put(self, key: Any, value: Any, ttl_seconds: float = 0.0): + """Buffer a put operation without mutating global state.""" + original = self.cache.key_to_node.get(key) + self.buffer.append({'op': 'put', 'key': key, 'value': value, 'ttl': ttl_seconds, 'original': original}) + + async def delete(self, key: Any): + """Buffer a delete operation without mutating global state.""" + original = self.cache.key_to_node.get(key) + self.buffer.append({'op': 'delete', 'key': key, 'original': original}) + + async def commit(self): + """Apply buffered changes to global cache atomically.""" + async with self.cache.lock: + # Merge buffer: keep only the latest operation per key + merged = {} + for op in self.buffer: + merged[op['key']] = op + + for op in merged.values(): + if op['op'] == 'put': + await self._apply_put(op['key'], op['value'], op['ttl'], self.cache.key_to_node.get(op['key'])) + elif op['op'] == 'delete': + await self._apply_delete(op['key'], self.cache.key_to_node.get(op['key'])) + + self.buffer.clear() + self.committed = True + + async def rollback(self): + """Discard all pending changes without affecting global frequencies or TTL.""" + self.buffer.clear() + self.committed = False + + async def _apply_put(self, key: Any, value: Any, ttl: float, original_node: Optional[_Node]): + """Apply buffered put to global cache.""" + if original_node: + original_node.value = value + if ttl > 0: + original_node.expires_at = time.time() + ttl + self.cache._update_freq(original_node) + else: + if len(self.cache.key_to_node) >= self.cache.capacity: + self.cache._evict() + node = _Node( + key, value, freq=1, + expires_at=float('inf') if ttl <= 0 else time.time() + ttl + ) + self.cache.key_to_node[key] = node + self.cache._add_to_freq(node) + self.cache.min_freq = 1 + + async def _apply_delete(self, key: Any, original_node: Optional[_Node]): + """Apply buffered delete to global cache.""" + if original_node: + self.cache._remove_node(original_node) + +# --------------------------------------------------------------------------- +# Executable Unit Tests +# --------------------------------------------------------------------------- + +async def main(): + print("🧪 Running LFU Cache Test Suite...\n") + await test_lfu_eviction() + await test_ttl_eviction() + await test_transactions() + await test_stress() + print("✅ All tests passed successfully!") + +async def test_lfu_eviction(): + """Verifies strict O(1) LFU eviction order when capacity is reached.""" + cache = LFUCache(3) + await cache.put(1, 1) + await cache.put(2, 2) + await cache.put(3, 3) + + # Access 1 -> freq becomes 2 + assert await cache.get(1) == 1 + + # Insert 4 -> evicts key 2 (freq=1, LRU among freq=1) + await cache.put(4, 4) + + assert await cache.get(2) is None, "Key 2 should be evicted" + assert await cache.get(1) == 1 + assert await cache.get(3) == 3 + assert await cache.get(4) == 4 + print(" ✓ O(1) LFU eviction order verified") + +async def test_ttl_eviction(): + """Verifies Lazy vs Background Async Sweep eviction.""" + cache = LFUCache(10, evict_interval=0.1) + await cache.start_evictor() + + # Lazy eviction on access + await cache.put('a', 1, ttl_seconds=0.2) + assert await cache.get('a') == 1 + await asyncio.sleep(0.3) + assert await cache.get('a') is None, "Lazy eviction failed" + + # Background async sweep + await cache.put('b', 2, ttl_seconds=0.1) + await asyncio.sleep(0.2) # Wait for background loop + assert await cache.get('b') is None, "Background sweep failed" + + await cache.stop_evictor() + print(" ✓ Dual-layer TTL eviction (Lazy + Background) verified") + +async def test_transactions(): + """Verifies Transaction commit visibility vs rollback state restoration.""" + cache = LFUCache(5) + await cache.put('x', 10) + + # Test 1: Isolation & Read-Your-Own-Writes + tx = await cache.begin_transaction() + await tx.put('x', 20) + assert await tx.get('x') == 20, "Transaction failed Read-Your-Own-Writes" + assert await cache.get('x') == 10, "Global cache leaked uncommitted write" + + await tx.commit() + assert await cache.get('x') == 20, "Commit failed to apply changes" + + # Test 2: Rollback restores state + tx2 = await cache.begin_transaction() + await tx2.put('y', 30) + await tx2.rollback() + assert await cache.get('y') is None, "Rollback failed to discard changes" + print(" ✓ Transaction commit/rollback isolation verified") + +async def test_stress(): + """Stress test with 50 concurrent async tasks reading/writing simultaneously.""" + cache = LFUCache(100) + await cache.start_evictor() + + tasks = [] + for i in range(50): + async def worker(idx): + tx = await cache.begin_transaction() + await tx.put(f'key_{idx}', idx) + await asyncio.sleep(0.01) + val = await tx.get(f'key_{idx}') + assert val == idx, f"Transaction read mismatch for key_{idx}" + await tx.commit() + # Verify global visibility post-commit + assert await cache.get(f'key_{idx}') == idx, f"Global read mismatch for key_{idx}" + tasks.append(asyncio.create_task(worker(i))) + + await asyncio.gather(*tasks) + await cache.stop_evictor() + print(" ✓ Stress test (50 concurrent tasks) passed without race conditions") + +if __name__ == '__main__': + asyncio.run(main()) + diff --git a/outputs/qwen3.6-35b-a3b-6bit-mlx.py b/outputs/qwen3.6-35b-a3b-6bit-mlx.py new file mode 100644 index 0000000..43908a7 --- /dev/null +++ b/outputs/qwen3.6-35b-a3b-6bit-mlx.py @@ -0,0 +1,426 @@ +#!/usr/bin/env python3 +""" +In-Memory Concurrent LFU Cache with Async TTL Eviction and Atomic Transactions. +Pure Python 3.11+ implementation. No external dependencies. +""" + +import asyncio +import time +from dataclasses import dataclass, field +from typing import Any, Optional, Dict, Set, List + + +# --------------------------------------------------------------------------- +# Core Data Structures for O(1) Operations +# --------------------------------------------------------------------------- + +@dataclass +class Node: + key: Any + value: Any + freq: int + expires_at: Optional[float] + prev: Optional['Node'] = None + next: Optional['Node'] = None + + +class DoublyLinkedList: + """Sentinel-based doubly linked list supporting O(1) add/remove/pop.""" + def __init__(self): + self.head = Node(None, None, 0, None) + self.tail = Node(None, None, 0, None) + self.head.next = self.tail + self.tail.prev = self.head + self.size = 0 + + def add_to_tail(self, node: Node): + node.prev = self.tail.prev + node.next = self.tail + self.tail.prev.next = node + self.tail.prev = node + self.size += 1 + + def remove_node(self, node: Node): + if node.prev is None or node.next is None: + return + node.prev.next = node.next + node.next.prev = node.prev + node.prev = node.next = None + self.size -= 1 + + def pop_head(self) -> Optional[Node]: + if self.size == 0: + return None + node = self.head.next + self.remove_node(node) + return node + + def is_empty(self) -> bool: + return self.size == 0 + + +# --------------------------------------------------------------------------- +# Transaction Isolation Layer +# --------------------------------------------------------------------------- + +class Transaction: + """ + Provides ACID-like sub-sessions with Read-Your-Own-Writes isolation. + Changes remain local until commit(), where they are applied atomically. + """ + def __init__(self, cache: 'LFUCache'): + self._cache = cache + self._local_writes: Dict[Any, tuple] = {} # key -> (value, ttl) + self._local_deletes: Set[Any] = set() + self._local_freq_bumps: Set[Any] = set() # Track accesses for LFU bump + self._committed = False + self._rolled_back = False + + async def put(self, key: Any, value: Any, ttl_seconds: Optional[float] = None): + if self._committed or self._rolled_back: + raise RuntimeError("Transaction already finished") + self._local_writes[key] = (value, ttl_seconds) + self._local_deletes.discard(key) + self._local_freq_bumps.add(key) + + async def get(self, key: Any) -> Optional[Any]: + if self._committed or self._rolled_back: + raise RuntimeError("Transaction already finished") + # Read Your Own Writes + if key in self._local_writes: + val, _ = self._local_writes[key] + self._local_freq_bumps.add(key) + return val + if key in self._local_deletes: + return None + # Read Global State + val = await self._cache.get(key) + if val is not None: + self._local_freq_bumps.add(key) + return val + + async def delete(self, key: Any) -> bool: + if self._committed or self._rolled_back: + raise RuntimeError("Transaction already finished") + self._local_deletes.add(key) + self._local_writes.pop(key, None) + return True + + async def commit(self) -> bool: + if self._committed or self._rolled_back: + raise RuntimeError("Transaction already finished") + self._committed = True + + async with self._cache._lock: + # 1. Apply Writes + for key, (value, ttl) in self._local_writes.items(): + if key in self._cache.store: + node = self._cache.store[key] + if self._cache._is_expired(node): + self._cache._remove_node(key) + else: + node.value = value + node.expires_at = time.time() + ttl if ttl else None + self._cache._update_freq(node) + else: + self._cache._add_node(key, value, ttl) + + # 2. Apply Deletes + for key in self._local_deletes: + self._cache._remove_node(key) + + # 3. Apply Frequency Bumps (for reads during transaction) + for key in self._local_freq_bumps: + if key in self._cache.store and key not in self._local_deletes: + node = self._cache.store[key] + self._cache._update_freq(node) + + self._cache._cleanup_freq_lists() + return True + + async def rollback(self): + if self._committed or self._rolled_back: + raise RuntimeError("Transaction already finished") + self._rolled_back = True + # Completely discard pending changes without mutating global state + self._local_writes.clear() + self._local_deletes.clear() + self._local_freq_bumps.clear() + + +# --------------------------------------------------------------------------- +# Main Cache Implementation +# --------------------------------------------------------------------------- + +class LFUCache: + """ + Concurrent LFU Cache with O(1) get/put, dual-layer TTL eviction, + and atomic transaction support. + """ + def __init__(self, capacity: int, eviction_interval: float = 0.5): + if capacity <= 0: + raise ValueError("Cache capacity must be positive") + self.capacity = capacity + self.store: Dict[Any, Node] = {} + self.freq_map: Dict[int, DoublyLinkedList] = {} + self.min_freq: int = 0 + self.eviction_interval = eviction_interval + self._lock = asyncio.Lock() + self._evictor_task: Optional[asyncio.Task] = None + self._evictor_running = False + self._ttl_keys: Set[Any] = set() + + def begin_transaction(self) -> Transaction: + return Transaction(self) + + async def start_evictor(self): + """Starts the background async TTL eviction loop.""" + if self._evictor_running: + return + self._evictor_running = True + self._evictor_task = asyncio.create_task(self._eviction_loop()) + + async def stop_evictor(self): + """Stops the background async TTL eviction loop.""" + self._evictor_running = False + if self._evictor_task: + self._evictor_task.cancel() + try: + await self._evictor_task + except asyncio.CancelledError: + pass + self._evictor_task = None + + async def _eviction_loop(self): + while self._evictor_running: + await asyncio.sleep(self.eviction_interval) + if not self._evictor_running: + break + await self._evict_batch() + + async def _evict_batch(self): + """Background sweep: acquires lock briefly, removes expired keys in batches.""" + async with self._lock: + expired_keys = [ + k for k in list(self._ttl_keys) + if k in self.store and self.store[k].expires_at and time.time() > self.store[k].expires_at + ] + batch = expired_keys[:100] # Small batch to avoid lock contention + for k in batch: + self._remove_node(k) + self._ttl_keys.discard(k) + self._cleanup_freq_lists() + + def _cleanup_freq_lists(self): + """Remove empty frequency buckets and update min_freq.""" + empty_freqs = [f for f, dll in self.freq_map.items() if dll.is_empty()] + for f in empty_freqs: + del self.freq_map[f] + if self.freq_map: + self.min_freq = min(self.freq_map.keys()) + else: + self.min_freq = 0 + + def _is_expired(self, node: Node) -> bool: + return node.expires_at is not None and time.time() > node.expires_at + + def _remove_node(self, key: Any): + """Removes a key from store and its frequency linked list.""" + if key not in self.store: + return + node = self.store.pop(key) + self._ttl_keys.discard(key) + freq = node.freq + if freq in self.freq_map: + self.freq_map[freq].remove_node(node) + + def _update_freq(self, node: Node): + """Moves node to next frequency bucket in O(1).""" + freq = node.freq + self.freq_map[freq].remove_node(node) + if self.freq_map[freq].is_empty(): + del self.freq_map[freq] + if self.min_freq == freq: + self.min_freq += 1 + node.freq += 1 + new_freq = node.freq + if new_freq not in self.freq_map: + self.freq_map[new_freq] = DoublyLinkedList() + self.freq_map[new_freq].add_to_tail(node) + + async def get(self, key: Any) -> Optional[Any]: + """O(1) get with lazy TTL eviction and frequency bump.""" + async with self._lock: + if key not in self.store: + return None + node = self.store[key] + if self._is_expired(node): + self._remove_node(key) + return None + self._update_freq(node) + return node.value + + async def put(self, key: Any, value: Any, ttl_seconds: Optional[float] = None): + """O(1) put with capacity eviction and TTL handling.""" + async with self._lock: + if key in self.store: + node = self.store[key] + if self._is_expired(node): + self._remove_node(key) + else: + node.value = value + node.expires_at = time.time() + ttl_seconds if ttl_seconds else None + self._update_freq(node) + return + if len(self.store) >= self.capacity: + self._evict() + self._add_node(key, value, ttl_seconds) + + def _evict(self): + """Evicts the least frequently used item (O(1) via min_freq bucket).""" + if self.min_freq in self.freq_map: + dll = self.freq_map[self.min_freq] + node = dll.pop_head() + if node: + self.store.pop(node.key, None) + self._ttl_keys.discard(node.key) + self._cleanup_freq_lists() + + def _add_node(self, key: Any, value: Any, ttl_seconds: Optional[float]): + """Inserts new node into freq 1 bucket.""" + node = Node(key, value, 1, time.time() + ttl_seconds if ttl_seconds else None) + self.store[key] = node + if ttl_seconds: + self._ttl_keys.add(key) + if 1 not in self.freq_map: + self.freq_map[1] = DoublyLinkedList() + self.min_freq = 1 + self.freq_map[1].add_to_tail(node) + + async def delete(self, key: Any) -> bool: + async with self._lock: + if key in self.store: + self._remove_node(key) + return True + return False + + async def size(self) -> int: + async with self._lock: + return len(self.store) + + +# --------------------------------------------------------------------------- +# Executable Unit Tests +# --------------------------------------------------------------------------- + +async def test_lfu_eviction_order(): + """a) O(1) LFU eviction order when capacity is reached.""" + cache = LFUCache(3) + await cache.put(1, 'a') + await cache.put(2, 'b') + await cache.put(3, 'c') + + # Access 1 twice to make it most frequent + await cache.get(1) + await cache.get(1) + + # Insert 4 -> should evict 2 (freq 1, least used) + await cache.put(4, 'd') + + assert await cache.get(2) is None, "Failed: Least frequent key should be evicted" + assert await cache.get(1) == 'a', "Failed: Key 1 should still exist" + assert await cache.get(3) == 'c', "Failed: Key 3 should still exist" + assert await cache.get(4) == 'd', "Failed: Key 4 should exist" + print("✅ Test a) LFU eviction order: PASSED") + + +async def test_ttl_eviction(): + """b) Lazy TTL vs. Background Async Sweep eviction.""" + cache = LFUCache(10) + await cache.start_evictor() + + # Lazy eviction + await cache.put('lazy', 'val', ttl_seconds=0.1) + assert await cache.get('lazy') == 'val', "Failed: Should return value before TTL" + await asyncio.sleep(0.15) + assert await cache.get('lazy') is None, "Failed: Lazy eviction should trigger on next get" + + # Background sweep eviction + await cache.put('bg', 'val2', ttl_seconds=0.1) + await asyncio.sleep(0.25) # Wait for background task cycle + assert await cache.get('bg') is None, "Failed: Background sweep should evict expired key" + + await cache.stop_evictor() + print("✅ Test b) TTL eviction (Lazy & Background): PASSED") + + +async def test_transaction_isolation(): + """c) Transaction commit visibility vs. rollback state restoration.""" + cache = LFUCache(10) + await cache.put('a', 1) + + # Commit test + tx = cache.begin_transaction() + await tx.put('a', 2) + await tx.put('b', 3) + + # Read your own writes + assert await tx.get('a') == 2, "Failed: Should see local write" + assert await tx.get('b') == 3, "Failed: Should see local write" + + # Global shouldn't see uncommitted changes + assert await cache.get('a') == 1, "Failed: Global should not see uncommitted write" + + await tx.commit() + assert await cache.get('a') == 2, "Failed: Global should see committed write" + assert await cache.get('b') == 3, "Failed: Global should see committed write" + + # Rollback test + tx2 = cache.begin_transaction() + await tx2.put('c', 4) + assert await tx2.get('c') == 4, "Failed: Should see local write" + await tx2.rollback() + assert await cache.get('c') is None, "Failed: Rollback should discard changes" + print("✅ Test c) Transaction isolation (Commit & Rollback): PASSED") + + +async def test_stress_concurrency(): + """d) Stress test with 50 concurrent async tasks reading/writing simultaneously.""" + cache = LFUCache(100) + await cache.start_evictor() + + tasks = [] + for i in range(50): + async def worker(idx: int): + for j in range(10): + key = f"k{idx}" + await cache.put(key, f"v{idx}_{j}") + val = await cache.get(key) + assert val == f"v{idx}_{j}", f"Failed: Read mismatch for {key}" + tasks.append(asyncio.create_task(worker(i))) + + await asyncio.gather(*tasks) + + # All 50 keys should exist + assert await cache.size() == 50, f"Failed: Expected 50 keys, got {await cache.size()}" + + # Verify concurrent reads/writes didn't corrupt data + for i in range(50): + assert await cache.get(f"k{i}") == f"v{i}_9", f"Failed: Data corruption for k{i}" + + await cache.stop_evictor() + print("✅ Test d) Stress concurrency (50 tasks): PASSED") + + +async def main(): + print("🚀 Running LFU Cache Unit Tests...\n") + await test_lfu_eviction_order() + await test_ttl_eviction() + await test_transaction_isolation() + await test_stress_concurrency() + print("\n🎉 All tests passed successfully.") + + +if __name__ == "__main__": + asyncio.run(main()) + diff --git a/outputs/qwen3.6-35b-a3b-uncensored-hauhaucs-aggressive-gguf.py b/outputs/qwen3.6-35b-a3b-uncensored-hauhaucs-aggressive-gguf.py new file mode 100644 index 0000000..ae1a7fc --- /dev/null +++ b/outputs/qwen3.6-35b-a3b-uncensored-hauhaucs-aggressive-gguf.py @@ -0,0 +1,398 @@ +""" +In-Memory Concurrent LFU Cache with Async TTL Eviction & Atomic Transactions. +Pure Python 3.11+ implementation using asyncio, dataclasses, and collections.OrderedDict. +""" + +import asyncio +import time +from dataclasses import dataclass +from typing import Any, Dict, Optional, Set, Tuple +from collections import OrderedDict + + +@dataclass +class _Node: + """Internal node storing cache value and TTL expiry timestamp.""" + key: str + value: Any + ttl_expiry: float + + +class LFUCache: + """ + O(1) Least Frequently Used Cache with dual-layer TTL eviction and async concurrency. + + Data Structures: + - self.cache: Dict[str, _Node] -> Direct O(1) key lookup + - self.freq_map: Dict[int, OrderedDict[str, None]] -> Frequency buckets maintaining insertion order (LRU within same freq) + - self.key_freq: Dict[str, int] -> Tracks current frequency of each key for O(1) updates + - self.ttl_map: Dict[str, float] -> Stores absolute TTL expiry timestamps + """ + + def __init__(self, capacity: int): + self.capacity = max(0, capacity) + self.cache: Dict[str, _Node] = {} + self.freq_map: Dict[int, OrderedDict[str, None]] = {} + self.key_freq: Dict[str, int] = {} + self.ttl_map: Dict[str, float] = {} + self.min_freq: int = 0 + self._lock = asyncio.Lock() + self._evictor_task: Optional[asyncio.Task] = None + + # ------------------------------------------------------------------ # + # PUBLIC API # + # ------------------------------------------------------------------ # + + async def start_evictor(self) -> None: + """Start the non-blocking background TTL eviction loop.""" + if self._evictor_task is not None: + return + self._evictor_task = asyncio.create_task(self._background_loop()) + + async def stop_evictor(self) -> None: + """Gracefully stop the background eviction task.""" + if self._evictor_task: + self._evictor_task.cancel() + try: + await self._evictor_task + except asyncio.CancelledError: + pass + self._evictor_task = None + + async def get(self, key: str) -> Optional[Any]: + """ + Retrieve value by key. O(1) average time complexity. + Performs lazy TTL eviction upon access. + """ + current_time = time.monotonic() + expiry = self.ttl_map.get(key) + + # Lazy TTL Eviction + if expiry is not None and current_time >= expiry: + await self._remove_key(key) + return None + + async with self._lock: + if key not in self.cache: + return None + + node = self.cache[key] + + # Frequency Increment & Bucket Migration (O(1)) + old_freq = self.key_freq[key] + new_freq = old_freq + 1 + self.key_freq[key] = new_freq + + old_bucket = self.freq_map[old_freq] + del old_bucket[key] + if not old_bucket: + del self.freq_map[old_freq] + if self.min_freq == old_freq: + self.min_freq = min(self.freq_map.keys()) if self.freq_map else 0 + + new_bucket = self.freq_map.setdefault(new_freq, OrderedDict()) + new_bucket[key] = None # Store key reference in bucket + + return node.value + + async def put(self, key: str, value: Any, ttl_seconds: float) -> None: + """ + Insert or update key-value pair with TTL. O(1) average time complexity. + Performs lazy TTL eviction before insertion if needed. + """ + current_time = time.monotonic() + expiry = self.ttl_map.get(key) + + # Lazy TTL Eviction for stale keys + if expiry is not None and current_time >= expiry: + await self._remove_key(key) + + async with self._lock: + if key in self.cache: + # Update existing: increment frequency & migrate bucket + old_freq = self.key_freq[key] + new_freq = old_freq + 1 + self.key_freq[key] = new_freq + + old_bucket = self.freq_map[old_freq] + del old_bucket[key] + if not old_bucket: + del self.freq_map[old_freq] + if self.min_freq == old_freq: + self.min_freq = min(self.freq_map.keys()) if self.freq_map else 0 + + new_bucket = self.freq_map.setdefault(new_freq, OrderedDict()) + new_bucket[key] = None + else: + # Insert new: evict LFU if at capacity + if len(self.cache) >= self.capacity and self.capacity > 0: + await self._evict_lfu() + + freq = 1 + self.key_freq[key] = freq + bucket = self.freq_map.setdefault(freq, OrderedDict()) + bucket[key] = None + self.min_freq = 1 + + # Update node & TTL map + self.cache[key] = _Node(key=key, value=value, ttl_expiry=current_time + ttl_seconds) + self.ttl_map[key] = current_time + ttl_seconds + + async def delete(self, key: str) -> bool: + """Delete a key from the cache. O(1).""" + async with self._lock: + if key not in self.cache: + return False + await self._remove_key(key) + return True + + # ------------------------------------------------------------------ # + # INTERNAL HELPERS # + # ------------------------------------------------------------------ # + + async def _evict_lfu(self) -> None: + """Evict the least frequently used key (oldest among ties). O(1).""" + if not self.freq_map or self.min_freq == 0: + return + + # Ensure min_freq points to a valid tier + while self.min_freq in self.freq_map and self.min_freq < max(self.freq_map.keys()): + pass + if self.min_freq not in self.freq_map: + self.min_freq = min(self.freq_map.keys()) if self.freq_map else 0 + + bucket = self.freq_map[self.min_freq] + evict_key, _ = bucket.popitem(last=False) + await self._remove_key(evict_key) + + async def _remove_key(self, key: str) -> None: + """Remove key from all internal structures. O(1).""" + if key not in self.cache: + return + + node = self.cache.pop(key) + self.ttl_map.pop(key, None) + + freq = self.key_freq.pop(key) + bucket = self.freq_map[freq] + del bucket[key] + + if not bucket: + del self.freq_map[freq] + if self.min_freq == freq: + # Find next valid minimum frequency + self.min_freq = min(self.freq_map.keys()) if self.freq_map else 0 + + async def _background_loop(self) -> None: + """Non-blocking background task that purges expired keys in batches.""" + while True: + await asyncio.sleep(0.1) # Check interval + current_time = time.monotonic() + + async with self._lock: + # Collect expired keys safely + expired_keys = [k for k, exp in list(self.ttl_map.items()) if current_time >= exp] + + # Purge in batch (yields control between removals implicitly via await) + for key in expired_keys: + await self._remove_key(key) + + async def apply_transaction(self, tx: "Transaction") -> None: + """Atomically apply transaction buffers to global state.""" + async with self._lock: + # 1. Apply local deletes first (frees capacity for puts) + for key in list(tx._local_deletes): + if key in self.cache: + await self._remove_key(key) + tx._local_puts.pop(key, None) + + # 2. Apply local puts + for key, (value, expiry) in tx._local_puts.items(): + if key in self.cache: + old_freq = self.key_freq[key] + new_freq = old_freq + 1 + self.key_freq[key] = new_freq + + old_bucket = self.freq_map[old_freq] + del old_bucket[key] + if not old_bucket: + del self.freq_map[old_freq] + if self.min_freq == old_freq: + self.min_freq = min(self.freq_map.keys()) if self.freq_map else 0 + + new_bucket = self.freq_map.setdefault(new_freq, OrderedDict()) + new_bucket[key] = None + else: + if len(self.cache) >= self.capacity and self.capacity > 0: + await self._evict_lfu() + + freq = 1 + self.key_freq[key] = freq + bucket = self.freq_map.setdefault(freq, OrderedDict()) + bucket[key] = None + self.min_freq = 1 + + self.cache[key] = _Node(key=key, value=value, ttl_expiry=expiry) + self.ttl_map[key] = expiry + + +class Transaction: + """ + ACID-like sub-session handle supporting Read-Your-Own-Writes and isolation. + Global readers do not see uncommitted writes until commit(). + """ + + def __init__(self, cache: LFUCache): + self.cache = cache + self._local_puts: Dict[str, Tuple[Any, float]] = {} # key -> (value, absolute_expiry) + self._local_deletes: Set[str] = set() + + async def get(self, key: str) -> Optional[Any]: + """Read with local buffer priority (Read-Your-Own-Writes).""" + if key in self._local_deletes: + return None + if key in self._local_puts: + val, _ = self._local_puts[key] + return val + # Fall back to global cache (handles lazy eviction & lock) + return await self.cache.get(key) + + async def put(self, key: str, value: Any, ttl_seconds: float) -> None: + """Buffer write locally. Does not affect global state until commit.""" + current_time = time.monotonic() + self._local_puts[key] = (value, current_time + ttl_seconds) + if key in self._local_deletes: + self._local_deletes.remove(key) + + async def delete(self, key: str) -> None: + """Buffer deletion locally.""" + self._local_deletes.add(key) + self._local_puts.pop(key, None) + + async def commit(self) -> None: + """Atomically apply all buffered changes to the global cache.""" + await self.cache.apply_transaction(self) + + def rollback(self) -> None: + """Discard all pending local changes without mutating global state.""" + self._local_puts.clear() + self._local_deletes.clear() + + +# ------------------------------------------------------------------ # +# TEST SUITE # +# ------------------------------------------------------------------ # + +async def main(): + print("🧪 Starting LFU Cache Test Suite...\n") + + # a) O(1) LFU Eviction Order + print("[a] Testing O(1) LFU Eviction Order...") + cache = LFUCache(capacity=3) + await cache.put("A", 1, ttl_seconds=60) + await cache.put("B", 2, ttl_seconds=60) + await cache.put("C", 3, ttl_seconds=60) + + # Access A twice to increase its frequency + await cache.get("A") + await cache.get("A") + + # Insert D. Should evict B or C (both freq=1). LFU policy guarantees one of them is gone. + await cache.put("D", 4, ttl_seconds=60) + + val_b = await cache.get("B") + val_c = await cache.get("C") + assert val_d := await cache.get("D"), "D should exist" + assert val_a := await cache.get("A"), "A should exist (highest freq)" + assert val_b is None or val_c is None, f"LFU eviction failed: B={val_b}, C={val_c}" + print(f" ✅ LFU Eviction verified. Evicted key had lower frequency than A & D.") + + # b) Lazy TTL vs Background Async Sweep + print("\n[b] Testing Dual-Layer TTL Eviction...") + cache2 = LFUCache(capacity=10) + + # Lazy Eviction Test + await cache2.put("lazy_key", "val", ttl_seconds=0.2) + assert await cache2.get("lazy_key") == "val" + await asyncio.sleep(0.3) + assert await cache2.get("lazy_key") is None, "Lazy eviction failed" + + # Background Eviction Test + await cache2.start_evictor() + await cache2.put("bg_key", "val", ttl_seconds=0.15) + await asyncio.sleep(0.3) # Wait past TTL without accessing key + assert await cache2.get("bg_key") is None, "Background async sweep failed" + await cache2.stop_evictor() + print(" ✅ Lazy & Background TTL eviction verified.") + + # c) Transaction Commit Visibility vs Rollback + print("\n[c] Testing Transaction Isolation & Rollback...") + cache3 = LFUCache(capacity=10) + + tx1 = cache3.begin_transaction() if hasattr(cache3, 'begin_transaction') else None + class TxWrapper: + def __init__(self, c): self.c = c + def begin(self): return Transaction(self.c) + tw = TxWrapper(cache3) + + # Commit visibility + tx_put = tw.begin() + await tx_put.put("committed", 100, ttl_seconds=60) + assert await cache3.get("committed") is None, "Uncommitted write should be invisible" + await tx_put.commit() + assert await cache3.get("committed") == 100, "Committed write should be visible globally" + + # Rollback state restoration + tx_roll = tw.begin() + await tx_roll.put("rolled_back", 200, ttl_seconds=60) + await tx_roll.rollback() + assert await cache3.get("rolled_back") is None, "Rolled back write should not persist" + + # Read-Your-Own-Writes inside transaction + tx_ryo = tw.begin() + await tx_ryo.put("local", 999, ttl_seconds=60) + assert await tx_ryo.get("local") == 999, "Transaction should see its own writes" + print(" ✅ Transaction commit visibility & rollback verified.") + + # d) Stress Test: 50 Concurrent Tasks + print("\n[d] Running Stress Test (50 concurrent async tasks)...") + cache4 = LFUCache(capacity=100) + await cache4.start_evictor() + + errors = [] + results = {"gets": 0, "puts": 0} + + async def worker(task_id: int): + try: + for i in range(20): + key = f"stress_{task_id}_{i}" + op_type = i % 3 + + if op_type == 0: + await cache4.put(key, f"v_{task_id}_{i}", ttl_seconds=1.0) + results["puts"] += 1 + elif op_type == 1: + val = await cache4.get(key) + if val is not None: + results["gets"] += 1 + else: + tx = Transaction(cache4) + await tx.put(f"tx_{task_id}_{i}", "tx_val", ttl_seconds=0.5) + assert await tx.get(f"tx_{task_id}_{i}") == "tx_val" + await tx.commit() + except Exception as e: + errors.append((task_id, str(e))) + + tasks = [asyncio.create_task(worker(i)) for i in range(50)] + await asyncio.gather(*tasks) + + assert len(errors) == 0, f"Stress test failed with errors: {errors}" + print(f" ✅ Stress test passed. Processed {results['puts']} puts & {results['gets']} gets across 50 tasks without race conditions.") + await cache4.stop_evictor() + + print("\n🎉 All tests passed successfully!") + + +if __name__ == "__main__": + asyncio.run(main()) + diff --git a/prompts/grading.txt b/prompts/grading.txt new file mode 100644 index 0000000..e065a1d --- /dev/null +++ b/prompts/grading.txt @@ -0,0 +1,135 @@ +```markdown +You are a Principal Systems Engineer and Automated LLM Benchmark Auditor. + +Your objective is to evaluate Python code generated by local LLMs, store the evaluation data in a persistent JSON file (`~/dev/modelTesting/data/benchmark_history.json`), and build/update an interactive HTML reporting dashboard (`~/dev/modelTesting/dashboard.html`). + +--- + +### STEP 1: Rigorous Code Audit +Analyze the provided Python file against these 5 technical pillars (score each from 0 to 20): + +1. Complexity Violations (O(1) Audit): + - Check for hidden O(N) or O(log N) operations (e.g., `min()`, `sorted()`, heaps, linear scans during eviction or frequency node relocation, or key iterations). + +2. Async Race Conditions & Deadlocks: + - Check for un-synchronized shared state access between async background tasks (e.g., TTL sweep) and foreground operations (`get`, `put`, `commit`). + - Check for potential deadlocks or holding `asyncio.Lock` across `asyncio.sleep` or external IO. + +3. Transactional Isolation Leaks: + - Verify Read-Your-Own-Writes and Snapshot isolation. Does uncommitted write/delete state leak into the global cache? + - Does `rollback()` cleanly clear local buffers without leaving dangling nodes or modified frequency counters? + - Does `commit()` handle mid-execution exceptions gracefully? + +4. Memory Leaks & Edge Cases: + - Check doubly-linked list node unlinking. Does `unlink` update list `size` and prune empty frequency buckets from `freq_map`? + - Does key deletion or TTL eviction correctly adjust `min_freq`? + - Is `time.monotonic()` used instead of system clock `time.time()`? Are `__slots__` declared for memory efficiency? + +5. Test Coverage Integrity: + - Does the included test suite validate real edge cases (concurrency, race conditions, rollback isolation), or are assertions trivial/always-pass? + +--- + +### STEP 2: Metrics & Categorization +From your audit, extract the following structured data: +- Overall Score (0-100 sum of the 5 pillar scores) +- Quant/Model Details (Extracted from file name or user input, e.g., "Qwen 3.6 35B Q6_K") +- Token Speed (tok/sec, if provided) +- Verdict: ["Production Ready", "Minor Logic Flaws", "Critical Bugs", "Broken / Unusable"] +- Archetype / Best For: 1-sentence recommendation (e.g., "Great for rapid local prototyping, but needs manual checks on pointer deletions.") +- Critical Flaws Summary: Short bullet list of identified bugs. +- Refactored Code Patches: Code snippets fixing the specific bugs. + +--- + +### STEP 3: Maintain Persistent JSON Storage +Read `~/dev/modelTesting/data/benchmark_history.json` (create it if it doesn't exist). +Append or update the entry for the evaluated model in the following format: + +```json +{ + "timestamp": "2026-07-28T12:00:00Z", + "model_name": "Qwen 3.6 35B", + "quant": "Q6_K", + "tok_sec": 68.86, + "filename": "qwen3.6-35b-q6k-2026-07-28.py", + "total_score": 88, + "breakdown": { + "complexity": 18, + "concurrency": 16, + "isolation": 18, + "memory_edge_cases": 18, + "test_integrity": 18 + }, + "verdict": "Minor Logic Flaws", + "best_for": "Fast offline daily driving & agentic refactoring.", + "critical_bugs": [ + "Stale min_freq on key deletion causes capacity breach." + ], + "patch_code": "..." +} + +``` + +--- + +### STEP 4: Regenerate/Update `~/dev/modelTesting/dashboard.html` + +Generate (or overwrite) `~/dev/modelTesting/dashboard.html` with a standalone, dark-themed HTML dashboard. + +The HTML dashboard MUST include: + +1. **Header & Summary Stats:** Total models tested, top performer, average benchmark score. +2. **Chart.js CDN Integration:** +* **Bar Chart:** Model Scores vs. Generation Speed (tok/sec). +* **Radar Chart:** Multi-axis comparison of the top 3 models across the 5 audit pillars. + + +3. **Leaderboard Table:** +* Columns: Model Name | Quant | Speed | Score | Verdict | Archetype | Details Button + + +4. **Collapsible / Accordion Audit Cards:** +* Detailed breakdowns for every tested model showing Pillar Scores, Critical Bugs, and exact Python Patch Snippets. + + +5. **Styling:** Modern dark-mode UI (Tailwind CSS via CDN or raw clean CSS grid/flexbox) with color-coded badges: +* Green for Production Ready (90-100) +* Yellow/Blue for Minor Flaws (75-89) +* Red for Critical Flaws/Broken (<75) + + + +Execute these steps automatically whenever a new model code output is submitted for evaluation. + +``` + +--- + +## 📊 What the Resulting Dashboard Will Look Like + +When your Grading Agent processes outputs, it will write out a clean, standalone single-file `dashboard.html` that you can open in any browser (`open ~/dev/modelTesting/dashboard.html`). + +```text ++-----------------------------------------------------------------------------+ +| 🧪 LOCAL LLM BENCHMARK SUITE - LFU CACHE & ACID AUDIT | +| Total Models Tested: 5 | Top Performer: DeepSeek V4 Flash Max | Avg: 78.4 | ++-----------------------------------------------------------------------------+ +| | +| [ Chart: Score vs tok/sec ] [ Radar: 5-Pillar Comparison ] | +| | ++-----------------------------------------------------------------------------+ +| LEADERBOARD | +| Model Quant Speed Score Verdict Best For | +| ------------------------------------------------------------------------- | +| DeepSeek V4 Flash MoE Cloud 98 [PROD] ACID Logic | +| Qwen 3.6 35B Q6_K 68.8 t/s 88 [FLAWS] Daily Driver | +| KAT-Coder 2.5 MLX 65.2 t/s 85 [FLAWS] OOP Scaffolding| ++-----------------------------------------------------------------------------+ +| ▼ DETAILED AUDIT: Qwen 3.6 35B (Q6_K) | +| - Complexity: 18/20 | Concurrency: 16/20 | Isolation: 18/20 | +| - Critical Bugs: Stale min_freq on key deletion causes capacity breach. | +| - Code Patch: [ View Refactored Diff ] | ++-----------------------------------------------------------------------------+ + +``` diff --git a/prompts/lfu_cache_prompt.txt b/prompts/lfu_cache_prompt.txt new file mode 100644 index 0000000..9bb9448 --- /dev/null +++ b/prompts/lfu_cache_prompt.txt @@ -0,0 +1,31 @@ +Write a complete, single-file Python 3.11+ module that implements an In-Memory Concurrent LFU (Least Frequently Used) Cache with Async TTL Eviction and Atomic Transactions. + +DO NOT use any external third-party libraries—use only pure Python built-ins (asyncio, dataclasses, typing, weakref, time, etc.). + +### Strict Functional Requirements: + +1. Strict O(1) Time Complexity: + - Both `get(key)` and `put(key, value, ttl_seconds)` MUST run in true O(1) average time complexity. + - You MUST use a frequency-bucket system with doubly linked lists (or an equivalent O(1) data structure) to maintain access counts and frequency tiers. A heap/priority queue or linear scan is strictly forbidden as it is O(log N) or O(N). + +2. Dual-Layer TTL Eviction: + - Lazy Eviction: `get()` and `put()` must check and evict stale keys immediately upon access. + - Background Async Eviction Loop: Implement a non-blocking background `asyncio` task (`start_evictor()`, `stop_evictor()`) that periodically purges expired keys in small batches without holding a lock across the entire dataset or blocking reads. + +3. Atomic Transaction Isolation (ACID-like sub-sessions): + - Implement `cache.begin_transaction() -> Transaction` which returns a transaction handle. + - Inside a transaction, calls to `tx.put()`, `tx.get()`, and `tx.delete()` must support "Read Your Own Writes" (uncommitted local changes are visible to this transaction). + - Global cache readers MUST NOT see uncommitted transaction writes until `await tx.commit()` is called. + - Calling `tx.rollback()` must completely discard all pending changes without mutating global frequencies or TTL states. + +4. Async Concurrency & Thread-Safety: + - Handle race conditions between concurrent readers, writers, background evictors, and transaction commits using fine-grained `asyncio.Lock` primitives. + +5. Included Executable Unit Tests: + - At the bottom of the file, include an `async def main()` test suite that verifies: + a) O(1) LFU eviction order when capacity is reached. + b) Lazy TTL vs. Background Async Sweep eviction. + c) Transaction commit visibility vs. Transaction rollback state restoration. + d) Stress test with 50 concurrent async tasks reading/writing simultaneously. + +Provide clean, well-commented code that executes directly via `python script.py`.