Rohan Paul@rohanpaul_aiOct 10, 2026, 10:08Anthropic states plainly that the model’s own explanation of its reasoning can’t be trusted as evidence of why it acted, which is exactly why they can’t cleanly judge how severe each of these failures was.打开原帖#511482