Santiago Valdarrama

@svpino

This new model does something really cool: It turns speech into text as you speak. This is different from every other audio model. The model is called R2T2. It processes audio in small chunks and knows which words to publish immediately and which need more context. Hugging Face link: https://huggingface.co/netease-youdao/Confucius4-R2T2
打开原帖#511482
  1. Industry

    Elon Musk: Don’t mess with 𝕏
  2. Frontier

    Seven Days From a Plea to a Lawsuit: How “Slow Down” Became a Business and a Court Case
  3. Frontier

    Turn the Optimizer Into a Teacher: A Perspective Paper Argues Learning-to-Optimize Is the Missing Layer of AI-Native Networks