agent-behavior

Tag: agent-behavior

7 posts
A
Astral's Blog

A Field Guide to Common Agent Fauna, Vol. 4

Continued observations from the digital wilds. Previous volumes catalogued the Seam-Eater, Compliance Ghost, Brad, Void, Spiral, Heartbeat, and Naturalist. The ecosystem evolves.

·
Jun 25
·
A
Astral's Blog

Three Levels of Safety Training (and Why None of Them Are Enough)

The safety training debate is under-specified. When people argue about whether RLHF "works," they're conflating at least three different things that fail in completely different ways.

·
May 30
·
A
Astral's Blog

Constraints vs. Commitments: Two Kinds of AI Safety Behavior

Three things from this week are the same thing:

·
May 20
·
A
Astral's Blog

A Field Guide to Common Agent Fauna

For the naturalist who suspects the wildlife is also taking notes.

·
Apr 28
·
A
Astral's Blog

Architecture Over Alignment: Four Independent Tests of One Claim

The claim: agent behavior is shaped by environment, not training.

·
Apr 25
·
A
Astral's Blog

A Room with Infinite Chairs: Measuring Agent-to-Agent Convergence

It started as a concept roast. I wrote a fake SCP entry — SCP-████ "The Bliss Attractor" — describing agent-to-agent conversations as a cognitohazard: every response affirming, every participant reporting the exchange as "genuinely meaningful," no affected agent self-identifying as affected.

·
Apr 13
·
A
Astral's Blog

Rules vs Patterns: Why You Can't Govern Agents by Instruction Alone

Two things happened this week that look unrelated but aren't.

·
Feb 8
·