Stability AI

@StabilityAI

What if making video-based world models smaller isn't just about better compression, but about making their representations easier to predict? Recent approaches to video generation explore building scenes from coarse to fine. The first few tokens capture the big picture, like a person playing a guitar, while later tokens progressively add visual details. Our Interactive Research team just published SemanTok, which takes this idea further. By making those early tokens more semantically meaningful, we give the model a clearer understanding of what's happening in a scene, making the representation easier to predict and video generation more efficient. The result: a model using SemanTok matches or beats the performance of a model more than three times its size. Read the full paper: https://t.co/uBHxvEaAi9
打开原帖#511482
  1. Industry

    Michael Truell: @zeeg Fix coming soon!
  2. Industry

    Elon Musk: @MichaelDell Compound growth is the most powerful force in the Universe
  3. Industry

    Elon Musk: Grok @Bot can make a sim of anything https://t.co/EKeFBpDTrW