OpenAI changed two conversation settings, not the model. The result shows why long-running AI tests depend on their memory setup.
Polyglot or poly-not?
AI coding benchmarks are heavily skewed toward Python and JavaScript. Framework maintainers could change that by defining what good code looks like in their ecosystems.