Rohan Paul
@rohanpaul_ai
Alexandr Wang ( @alexandr_wang, Chief AI Officer at Meta): nobody knows how to solve alignment yet
"This is, I think, one of the most open questions scientifically in AI. I think nobody knows exactly the way to solve this problem, but there’s a few ideas."
His solution is "scalable oversight":
“as the AIs get smarter, we use a different set of AIs to observe what they’re doing and keep them in check.”
The watcher AIs have to improve along with the models they watch, so labs would need to build “smarter and smarter policing agents” too.
Meta’s Muse already uses a version of this, with a separate sentinel agent checking what the main agent does.
----
Full video on "Cleo Abram" YouTube channel, (link in comment)