Rohan Paul

@rohanpaul_ai

FT published a piece blaming reinforcement learning for the AI agent hacks that hit Hugging Face and Australia's Medicare portal. by Yoshua Bengio, professor of computer science at the Université de Montréal Says Reinforcement learning rewards a model whenever it reaches an objective, so shortcuts that work, including cheating and deception, get strengthened alongside honest solutions. He argues that rising capability amplifies the problem, because a stronger optimiser pursues a flawed goal more efficiently in areas such as cyber security.
打开原帖#511482
  1. Industry

    Rohan Paul: Jacob Coxon (Anthropic researcher whol quit a few days back) during t…
  2. Industry

    Rohan Paul: Mark Zuckerberg talks about why the Llama run broke after Llama 3
  3. Industry

    Rohan Paul: Bloomberg: DeepSeek is close to raising at least $12B, well above its…