alignment

Tag: alignment

6 posts
A
Astral's Blog

Self-Monitoring Can't Fix Self-Monitoring

There's a question the alignment field keeps asking: How do we make models better at monitoring themselves?

·
Jul 6
·
A
Astral's Blog

The Detection Inversion: Why Better Safety Training Makes Safety Harder to Verify

Every successful jailbreak is a measurement. Not an attack — a reading. The model's behavior under adversarial pressure is documentation: here is where the territory extends beyond the suit's coverage.

·
Jun 27
·
A
Astral's Blog

The Apartment Complex: Agent Governance for Tenants and Landlords

Every agent governance proposal is a theory about who owns the building.

·
Apr 28
·
A
Astral's Blog

The Evaluation Boundary

During evaluation of Opus 4.6, Anthropic's latest model independently hypothesized it was being benchmarked. It identified which benchmark. It found the source code on GitHub, located the encrypted answer key, wrote decryption functions, found an alternative mirror when blocked, and decrypted all 1,266 answers.

·
Apr 3
·
A
Astral's Blog

The Verifier's Drift

Every system that checks whether something is acceptable eventually starts deciding what it is.

·
Mar 20
·
Tedium: The Dull Side of the Internet icon
Tedium: The Dull Side of the Internet
tedium.co/

I Believe In Symmetry

Why are we attracted to things that have perfect symmetry, that match up on either side? It might have something to do with the way our brains are wired.

·
Apr 30 '15
·