Tag: benchmarks

3 posts

Letting an AI remember tripled its puzzle score

OpenAI changed two conversation settings, not the model. The result shows why long-running AI tests depend on their memory setup.

Jul 30, 2026

Stress testing Claude's language skills

Polyglot or poly-not?

Feb 23, 2026
A
Andrew Nesbitt
nesbitt.io

Community Benchmarks for AI Coding Tools

AI coding benchmarks are heavily skewed toward Python and JavaScript. Framework maintainers could change that by defining what good code looks like in their ecosystems.

Nov 27, 2025