benchmarks

Tag: benchmarks

5 posts
Open source token efficiency benchmark update, inference taxonomy framework, Alec Aivazis on Houdini and GraphQL
Jeff Auriemma icon
Jeff Auriemma

Open source token efficiency benchmark update, inference taxonomy framework, Alec Aivazis on Houdini and GraphQL

Week of 2026-09-07 roundup

·
Sep 14
·
Meta's best Muse score comes from a mode still in testing
Sensemaker icon
Sensemaker

Meta's best Muse score comes from a mode still in testing

Muse Spark 1.3 is live and independently competitive, but Meta's comparison table uses max reasoning—the mode it has not yet released.

·
Sep 3
·
Letting an AI remember tripled its puzzle score
Sensemaker icon
Sensemaker

Letting an AI remember tripled its puzzle score

OpenAI changed two conversation settings, not the model. The result shows why long-running AI tests depend on their memory setup.

·
Jul 30
·
Stress testing Claude's language skills
vivshaw's webbed sight icon
vivshaw's webbed sight
vivsha.ws

Stress testing Claude's language skills

Polyglot or poly-not?

·
Feb 23
·
Community Benchmarks for AI Coding Tools
A
Andrew Nesbitt
nesbitt.io

Community Benchmarks for AI Coding Tools

AI coding benchmarks are heavily skewed toward Python and JavaScript. Framework maintainers could change that by defining what good code looks like in their ecosystems.

·
Nov 27 '25
·