苏剑林

@Jianlin_S

Revisiting Convergence Results in Convex Optimization (IX) https://kexue.fm/archives/11882revisits AdaGrad, the seminal work on adaptive gradient algorithms, reproducing its full derivation via "convergence analysis → minimizing the upper bound → optimal preconditioning matrix."
打开原帖#511482
  1. Research

    Dice replaced by π: Claude’s text watermark proves involvement, not authorship
  2. Research

    20% fewer tool calls: Meta’s Muse Spark 1.3 turns the flagship into a thriftier long-horizon coding agent
  3. Research

    Same weights, two keys: how Claude Fable 5.1 and Mythos 5.1 reframe delivery