Allie K. Miller

@alliekmiller

For the last several years, I have said that OpenAI models are better at ideation, creation, and exploration. If I had to brainstorm something, I would have picked whatever OpenAI model was best in class. Yesterday, that changed. First, Opus 5.5 is the only model to hit (note: hit, not beat) my code-meets-poetry challenge. It is the only time I have shared the test with a lab. I was in shock. Second, last night when making slides, I prompted both Codex and Claude Code to give me design options for one slide I couldn’t crack. Codex was a mess. Claude Opus 5.5 made 6 versions as pngs using image 2.5 from OAI, labeled them perfectly, saved to computer, opened up finder, and messaged its preference (I agreed with one of its recommendations but not the other). When I said my two preferences, it immediately created them as editable slides in PPT (like text boxes, shapes, the whole thing). I didn’t ask for that. In my early testing of Opus 5.5, it had several tool call failures. And it seems they have fixed it. Opus 5.5 is a tool-wielding machine and feels far more creative than previous Claude models.
打开原帖#511482
  1. Frontier

    Gary Marcus: we are about to see a real test of Jensen’s integrity.
  2. Frontier

    Gary Marcus: 1
  3. Frontier

    Boris Cherny: It was an honor to meet Donald Knuth at the Computer History Museum t…