Rohan Paul

@rohanpaul_ai

One of the stranger assumptions in AI is that the same model architecture should work equally well for thinking and interacting. A voice model has a weird job, it cannot just produce the right answer. It has to keep up with a person while the sequence keeps getting longer. In voice AI "speech" may not the right abstraction for the hard part, rather continuous state could be it. Here, Cartesia's founder talking how they came into the space from sequence modeling rather than speech research, which is probably why they focused so heavily on state space models. --- (Full video on “The Neon Show” YT channel, link in comment)
打开原帖#511482
  1. Industry

    Elon Musk: Important to use Grok 4.7 with our Build harness for the best results…
  2. Industry

    Alexandr Wang: we are excited for muse to be partnering deeply with @Shopify to enab…
  3. Industry

    Clément Delangue: lots of them! Reality is that LLM APIs are far from being the best wa…