Ethan Mollick

@emollick

A couple years later, with GPT-6. https://underdark-duel.netlify.app Apparently the mindflayer has a "101-bone skeletal rig" (This is 5E 2014 rules, and I didn't see major errors, though the mind flayer might have had more creative tactics. Also sorry for the repeated errors in posting) > I propose the Encounter Test as a nerdy benchmark standard for AI. Ask an AI to simulate an encounter between two D&D creatures & see how long it takes to mess up. > > Drow vs. mind flayer: GPT-4o does best, Gemini is cute. Outcomes similar (I am sure better prompting would help)
打开原帖#511482
  1. Research

    will depue: no world models, only autoregressive video + mistakes
  2. Products

    Robert Scoble: agent demos vs multi-step reliability
  3. Research

    François Chollet: AGI must produce more than its inputs