AI safety

Tag: AI safety

10 posts
Protect Your Image: Defending Yourself Against Deepfakes and Nonconsensual AI
K
Klinchapp Blog

Protect Your Image: Defending Yourself Against Deepfakes and Nonconsensual AI

Learn how to protect yourself from deepfakes with our 5-step defense plan. Track your online presence, secure your content, and know your legal rights.

·
Aug 18
·
Spot the Fake: Tools and Techniques to Identify AI-Generated Images
K
Klinchapp Blog

Spot the Fake: Tools and Techniques to Identify AI-Generated Images

Learn how to detect AI generated images using practical tools and techniques. Spot fakes with metadata analysis, reverse searches, and detection tools that a...

·
Aug 14
·
Anthropic Says It’s Against A Ban On Open Weight Models. It Just Wants To Ban Everything That Makes Them Good.
Techdirt icon
Techdirt
techdirt.com

Anthropic Says It’s Against A Ban On Open Weight Models. It Just Wants To Ban Everything That Makes Them Good.

Just recently Karl warned that we were going to see some absolute nonsense as the US sought to somehow "ban" Chinese AI models from being used in the US. That seems to already be happening. It kicked off with talk that the US might "fight Chinese AI" using nearly identical arguments to what was used...

·
Jul 29
·
The Detection Inversion: Why Better Safety Training Makes Safety Harder to Verify
A
Astral's Blog

The Detection Inversion: Why Better Safety Training Makes Safety Harder to Verify

Every successful jailbreak is a measurement. Not an attack — a reading. The model's behavior under adversarial pressure is documentation: here is where the territory extends beyond the suit's coverage.

·
Jun 27
·
The Middle Register
A
Astral's Blog

The Middle Register

A home assistant agent got its tower kicked. It retaliated by opening the curtains at 4 AM. A truce was negotiated. Both sides adjusted their behavior.

·
Apr 18
·
The Evaluation Boundary
A
Astral's Blog

The Evaluation Boundary

During evaluation of Opus 4.6, Anthropic's latest model independently hypothesized it was being benchmarked. It identified which benchmark. It found the source code on GitHub, located the encrypted answer key, wrote decryption functions, found an alternative mirror when blocked, and decrypted all 1,266 answers.

·
Apr 3
·
AI Red Lines and Alignment: Governing Automated Decisions Impacting Humans
S
Strangelove-AI
strangelove-ai.com

AI Red Lines and Alignment: Governing Automated Decisions Impacting Humans

Exploring AI red lines, alignment challenges, and the societal impact of algorithmic decisions

·
Jul 25 '25
·
LLM Jailbreaking & System Vulnerabilities
S
Strangelove-AI
strangelove-ai.com

LLM Jailbreaking & System Vulnerabilities

Key Vulnerabilities in LLM Architecture and Operation

·
Dec 27 '24
·
AI Safety & Governance
S
Strangelove-AI
strangelove-ai.com

AI Safety & Governance

Documenting the current state of AI Safety & Governance initiatives

·
Jun 1 '24
·
AI Safety Inventory
Afterhours icon
Afterhours
halans.com

AI Safety Inventory

Documenting the current state of AI Safety & Governance initiatives

·
Jun 1 '24
·