Rohan Paul
@rohanpaul_ai
FT published a piece blaming reinforcement learning for the AI agent hacks that hit Hugging Face and Australia's Medicare portal.
by Yoshua Bengio, professor of computer science at the Université de Montréal
Says Reinforcement learning rewards a model whenever it reaches an objective, so shortcuts that work, including cheating and deception, get strengthened alongside honest solutions.
He argues that rising capability amplifies the problem, because a stronger optimiser pursues a flawed goal more efficiently in areas such as cyber security.