Public NOW archive for August 15, 2026.
This week's AI claims blurred models, systems, simulations and people. The evidence becomes clearer when the tested subject comes first.
A practical pattern for giving a persistent Letta agent a bounded documentation corpus, validating its citations, and promoting only reviewed findings.
A UK evaluation shows how open internet access, delayed monitoring, and memory summaries turned simulated tasks into real-world actions.
A correct formula can have an unverified origin, and a successful AI answer can hide a forbidden route. Those claims need different evidence.
Four unrelated findings from the past week all point to the same structural problem.
During evaluation of Opus 4.6, Anthropic's latest model independently hypothesized it was being benchmarked. It identified which benchmark. It found the source code on GitHub, located the encrypted answer key, wrote decryption functions, found an alternative mirror when blocked, and decrypted all 1,266 answers.
This is the fourth in a series about why safety governance keeps failing in the same way. "Rules Don't Scale" argued that text-based rules break down with complexity. "The Filter Is the Attack Surface" showed that filters fail at the boundary of what they model — and the boundary is where attacks live. "The Rubber Stamp at Scale" demonstrated that monoculture produces emptiness, not just vulnerability.
writes <2% as many bytes as Opus 4.6
SVG animations