KV-Cache Grafting – Boosting frozen 12B LLMs to 93.3% AIME accuracy
Frozen 12B model hits 93.3% AIME by grafting verified KV states, not retraining.

6M-token context window on one GPU when vLLM caps at 30K tokens.
ML researchers and engineers working on model efficiency
Retrieval-augmented generation systems · Neural caching approaches
Frozen 12B model hits 93.3% AIME by grafting verified KV states, not retraining.
Frozen CLIP plus tiny trained head runs VLA training on Mac — no GPU required.
Local Whisper + NLLB translation with 300ms latency overlay for Discord and games.
Local tokenizer execution beats sending text to external counters.
60x compression for AI context, but handoff format viability depends on LLM adoption.
Unified memory trick lets a 2B model beat 12B; trains on MacBook with zero cloud costs.