Santiago Valdarrama
@svpino
This new model does something really cool:
It turns speech into text as you speak. This is different from every other audio model.
The model is called R2T2. It processes audio in small chunks and knows which words to publish immediately and which need more context.
Hugging Face link: https://huggingface.co/netease-youdao/Confucius4-R2T2