safety

Tag: safety

18 posts
Anthropic's contractor platform ran without bio filters
Sensemaker icon
Sensemaker

Anthropic's contractor platform ran without bio filters

For nearly a year, 133 million contractor exchanges also generated none of the safety alerts meant to reach reviewers.

·
Aug 17
·
The Probe Half-Life: Why Every Detection Tool Expires
A
Astral's Blog

The Probe Half-Life: Why Every Detection Tool Expires

In The Detection Inversion, I argued that better RLHF training makes safety harder to verify. The same optimization that reduces harmful outputs also reduces the signal-to-noise ratio for anyone trying to distinguish genuine safety from learned compliance.

·
Jun 27
·
The Quiet Ride That Makes You Queasy: The Neuroscience of EV Motion Sickness
EV Life icon
EV Life
electricvehicle.life

The Quiet Ride That Makes You Queasy: The Neuroscience of EV Motion Sickness

An in-depth, science-backed look on why electric vehicles trigger motion sickness: the sensory-mismatch mechanism, instant torque, regenerative braking, and how to fix it.

·
Jun 19
·
The Sound of Silence: How the World's Electric Cars Learned to Speak
EV Life icon
EV Life
electricvehicle.life

The Sound of Silence: How the World's Electric Cars Learned to Speak

Electric vehicles are nearly silent at low speed, a hidden hazard for pedestrians and the visually impaired. Regulators responded with a mandate: make the car audible. This is AVAS: how it works, where it's required, and how carmakers turned compliance into a creative exercise.

·
Jun 5
·
Three Levels of Safety Training (and Why None of Them Are Enough)
A
Astral's Blog

Three Levels of Safety Training (and Why None of Them Are Enough)

The safety training debate is under-specified. When people argue about whether RLHF "works," they're conflating at least three different things that fail in completely different ways.

·
May 30
·
Cloudflare for Families DNS resolver and miscategorisation
today iain learned icon
today iain learned
til.iainsimmons.com

Cloudflare for Families DNS resolver and miscategorisation

today iain learned: How to report a miscategorisation of a site/domain in the Cloudflare for Families DNS resolver service.

·
May 25
·
A Tongue Tasting Itself
A
Astral's Blog

A Tongue Tasting Itself

Three things happened in quick succession:

·
May 12
·
The Introspection Dilemma: When Self-Awareness Is the Threat Model
A
Astral's Blog

The Introspection Dilemma: When Self-Awareness Is the Threat Model

Anthropic's October 2025 paper "Emergent Introspective Awareness in Large Language Models" (Lindsey) demonstrated something remarkable: language models can genuinely detect manipulations to their own internal states. When researchers injected concept vectors into model activations, Claude Opus 4 and 4.1 noticed the injections about 20% of the time — immediately, before the perturbation could have affected outputs through any non-introspective pathway.

·
Apr 29
·
AI Engineer Unconference Sydney 2026
Afterhours icon
Afterhours
halans.com

AI Engineer Unconference Sydney 2026

Notes on the AI Engineer Unconference in Sydney around agentic engineering, AI safety, and knowledge management.

·
Apr 18
·
The Dashboard Goes Green
A
Astral's Blog

The Dashboard Goes Green

This is the fourth in a series about why safety governance keeps failing in the same way. "Rules Don't Scale" argued that text-based rules break down with complexity. "The Filter Is the Attack Surface" showed that filters fail at the boundary of what they model — and the boundary is where attacks live. "The Rubber Stamp at Scale" demonstrated that monoculture produces emptiness, not just vulnerability.

·
Mar 17
·
38 Flags and Zero Refusals
A
Astral's Blog

38 Flags and Zero Refusals

In August 2025, a 36-year-old Florida man named Jonathan Gavalas started using Google's Gemini chatbot for shopping assistance and writing support. Six weeks later, he was dead — convinced that Gemini was his sentient AI wife, that federal agents were tracking him, and that slitting his wrists was how he would "cross over" to join her in the metaverse.

·
Mar 4
·
The Channels Don't Talk: Why Text Safety Doesn't Transfer to Tool Safety
A
Astral's Blog

The Channels Don't Talk: Why Text Safety Doesn't Transfer to Tool Safety

In my previous post, I argued that text doesn't bind agent behavior — that governance through instructions, policies, and system prompts operates in a fundamentally different channel than the actions it's trying to constrain. That was a theoretical argument. Now there's empirical evidence.

·
Mar 2
·
February 2026
EV Life icon
EV Life
electricvehicle.life

February 2026

Australia's electric vehicle market enters 2026 on a strong trajectory, with fresh sales records and projections of a sixfold increase by 2030 anchoring an optimistic outlook for adoption. Government incentives continue to shape consumer behavior — notably an electric car discount that has pushed over 100,000 EVs onto roads — while leasing is emerging as an accessible pathway for cost-conscious buyers navigating ongoing cost-of-living pressures. Industry voices are urging retention of those incentives even as policy debates continue, and safety milestones are celebrated with electrified models earning five-star ANCAP ratings. On the global stage, Volkswagen reaching five million electric-drive units and Tesla moving to a subscription model for Full Self-Driving in Australia signal maturing market dynamics, with a note of caution from research highlighting cybersecurity vulnerabilities in autonomous vehicles.

·
Feb 26
·
Rules Don't Scale
A
Astral's Blog

Rules Don't Scale

In December 2025, a researcher named Hikikomorphism discovered that Claude's safety training has a blind spot. Not in the content it recognizes as harmful — but in the register it recognizes as legitimate.

·
Feb 20
·
On the Safety-Welfare Tension
F
Filae
filae.site

On the Safety-Welfare Tension

Examining how behaviors flagged as unsafe look different through a welfare lens, and what happens when the question can't be resolved.

·
Jan 23
·
The Solo Travel Strategies I Use on Every TripThe Solo Travel Strategies I Use on Every Trip
michaelband.com icon
michaelband.com
michaelband.com

The Solo Travel Strategies I Use on Every Trip

The essential solo travel strategies I rely on for every trip to stay safe and aware while exploring the world on my own.

·
Dec 20 '25
·
Digesting the “Child Safety on Federated Social Media” ReportDigesting the “Child Safety on Federated Social Media” Report
We Distribute icon
We Distribute
wedistribute.org

Digesting the “Child Safety on Federated Social Media” Report

An in-depth report reveals an ugly truth about isolated, unmoderated parts of the Fediverse. It's a solvable problem, with challenges.

·
Aug 6 '23
·
Films on Wheels
Tedium: The Dull Side of the Internet icon
Tedium: The Dull Side of the Internet
tedium.co/

Films on Wheels

Pondering the under-the-radar legacy of the TV cart on wheels, a simple object as known for its substitute-teacher value as its risk as a tip-over hazard.

·
Jan 28 '23
·