Rohan Paul

@rohanpaul_ai

Anthropic states plainly that the model’s own explanation of its reasoning can’t be trusted as evidence of why it acted, which is exactly why they can’t cleanly judge how severe each of these failures was.
打开原帖#511482
  1. Research

    Yann LeCun: @Helios_Hua Nice work
  2. Industry

    Mustafa Suleyman: Super Intelligence must be contained... Today, this is an…
  3. Industry

    Databricks: Admins can now configure coding agents in one place with Unity Gateway