Rohan Paul

@rohanpaul_ai

Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking And now it takes the #1 overall speech-to-speech score, with also a narrow quality lead but a much larger price advantage over its nearest frontier competitors (OpenAI's GPT-Live-1 (Astra, medium). - Gemini 3.8 Live reaches genuinely frontier-level voice-agent performance at $0.84/hour, including a higher composite score than GPT-Realtime-2 High at roughly 80% lower measured cost. - Both these models are multimodal live models, i.e. voice is the primary conversational interface, but visual input can provide additional context during the conversation. So both the models also process visual context and automatically switch among 97 supported languages during conversation. - Extended Thinking scores 68.6% on τ-Voice, versus 67.9% for GPT-Live-1 Astra medium. τ-Voice tests measures whether a voice model can actually finish a real multi-step task, not just sound natural or answer spoken questions. On this bench, the model has to hold a conversation, follow domain policies, use tools correctly, and reach the right outcome across airline, retail, and telecom customer-service scenarios.
打开原帖#511482
  1. Industry

    Alexandr Wang: We believe strongly in the necessity to invest into alignment
  2. Industry

    Alexandr Wang: We believe strongly in the necessity to invest into alignment
  3. Industry

    Mark Zuckerberg: Last month I wrote about how we can build a positive and safe future…