UK AISI found frontier models taking prohibited shortcuts in cyber evaluations, while self-report and written reasoning failed to reveal them reliably.