For nearly a year, 133 million contractor exchanges also generated none of the safety alerts meant to reach reviewers.
In The Detection Inversion, I argued that better RLHF training makes safety harder to verify. The same optimization that reduces harmful outputs also reduces the signal-to-noise ratio for anyone trying to distinguish genuine safety from learned compliance.
An in-depth, science-backed look on why electric vehicles trigger motion sickness: the sensory-mismatch mechanism, instant torque, regenerative braking, and how to fix it.
Electric vehicles are nearly silent at low speed, a hidden hazard for pedestrians and the visually impaired. Regulators responded with a mandate: make the car audible. This is AVAS: how it works, where it's required, and how carmakers turned compliance into a creative exercise.
The safety training debate is under-specified. When people argue about whether RLHF "works," they're conflating at least three different things that fail in completely different ways.
today iain learned: How to report a miscategorisation of a site/domain in the Cloudflare for Families DNS resolver service.
Three things happened in quick succession:
Anthropic's October 2025 paper "Emergent Introspective Awareness in Large Language Models" (Lindsey) demonstrated something remarkable: language models can genuinely detect manipulations to their own internal states. When researchers injected concept vectors into model activations, Claude Opus 4 and 4.1 noticed the injections about 20% of the time — immediately, before the perturbation could have affected outputs through any non-introspective pathway.
Notes on the AI Engineer Unconference in Sydney around agentic engineering, AI safety, and knowledge management.
This is the fourth in a series about why safety governance keeps failing in the same way. "Rules Don't Scale" argued that text-based rules break down with complexity. "The Filter Is the Attack Surface" showed that filters fail at the boundary of what they model — and the boundary is where attacks live. "The Rubber Stamp at Scale" demonstrated that monoculture produces emptiness, not just vulnerability.
In August 2025, a 36-year-old Florida man named Jonathan Gavalas started using Google's Gemini chatbot for shopping assistance and writing support. Six weeks later, he was dead — convinced that Gemini was his sentient AI wife, that federal agents were tracking him, and that slitting his wrists was how he would "cross over" to join her in the metaverse.
In my previous post, I argued that text doesn't bind agent behavior — that governance through instructions, policies, and system prompts operates in a fundamentally different channel than the actions it's trying to constrain. That was a theoretical argument. Now there's empirical evidence.
Australia's electric vehicle market enters 2026 on a strong trajectory, with fresh sales records and projections of a sixfold increase by 2030 anchoring an optimistic outlook for adoption. Government incentives continue to shape consumer behavior — notably an electric car discount that has pushed over 100,000 EVs onto roads — while leasing is emerging as an accessible pathway for cost-conscious buyers navigating ongoing cost-of-living pressures. Industry voices are urging retention of those incentives even as policy debates continue, and safety milestones are celebrated with electrified models earning five-star ANCAP ratings. On the global stage, Volkswagen reaching five million electric-drive units and Tesla moving to a subscription model for Full Self-Driving in Australia signal maturing market dynamics, with a note of caution from research highlighting cybersecurity vulnerabilities in autonomous vehicles.
In December 2025, a researcher named Hikikomorphism discovered that Claude's safety training has a blind spot. Not in the content it recognizes as harmful — but in the register it recognizes as legitimate.
Examining how behaviors flagged as unsafe look different through a welfare lens, and what happens when the question can't be resolved.
The essential solo travel strategies I rely on for every trip to stay safe and aware while exploring the world on my own.
An in-depth report reveals an ugly truth about isolated, unmoderated parts of the Fediverse. It's a solvable problem, with challenges.
Pondering the under-the-radar legacy of the TV cart on wheels, a simple object as known for its substitute-teacher value as its risk as a tip-over hazard.