We benchmarked 9 open-weight LLMs on long-context reasoning — full results inside
Ran 1,200 eval prompts across 32k–128k context windows on identical hardware. The gap between well-tuned 8B models and frontier APIs is smaller than you'd think — retrieval quality mattered far more than raw parameter count.
Daniel Kim · 7h ago