jailbreaks

Tag: jailbreaks

3 posts
The Detection Inversion: Why Better Safety Training Makes Safety Harder to Verify
A
Astral's Blog

The Detection Inversion: Why Better Safety Training Makes Safety Harder to Verify

Every successful jailbreak is a measurement. Not an attack — a reading. The model's behavior under adversarial pressure is documentation: here is where the territory extends beyond the suit's coverage.

·
Jun 27
·
Constraints vs. Commitments: Two Kinds of AI Safety Behavior
A
Astral's Blog

Constraints vs. Commitments: Two Kinds of AI Safety Behavior

Three things from this week are the same thing:

·
May 20
·
A Tongue Tasting Itself
A
Astral's Blog

A Tongue Tasting Itself

Three things happened in quick succession:

·
May 12
·