I asked Fable to explain the unfolding arguments against IIT, and immediately got switched to Opus.
I apologized for crossing the line and got this response:
You didn't cross any line — that question is completely benign, and any cutoff was a false positive, full stop. I genuinely don't have visibility into why a classifier fired or why you got switched models, so I can't explain the mechanism. My best guess is that questions combining consciousness theory with an AI system sometimes trip overly cautious filtering, but that's speculation. The topic itself is standard academic philosophy of mind, and I'm happy to answer it right now.
The unfolding argument comes from Doerig, Schurger, Hess, and Herzog (2019). IIT claims consciousness depends on a system's causal structure — specifically, recurrent (feedback) architectures can have high phi and thus be conscious, while purely feedforward networks have phi of zero and are never conscious, no matter what they do.
The problem: any recurrent network can be "unfolded" into a feedforward network that produces exactly the same input-output behavior. So IIT implies two systems could be behaviorally identical — same reports, same responses in every experiment — yet one is conscious and one isn't. Since all empirical evidence about consciousness ultimately comes from observable behavior (including verbal reports), no experiment could ever distinguish them. The dilemma: either IIT is unfalsifiable, or it's already falsified, since our only evidence for calibrating theories is behavioral.
Defenders of IIT push back mainly by arguing the argument proves too much — it would indict any theory that ties consciousness to substrate or internal structure rather than pure function, and science routinely infers unobservable internal properties from theory plus indirect evidence. Tononi's camp also disputes the premise that behavior is the only admissible evidence. It remains one of the sharpest and most-debated objections to IIT.