Rohan Paul

@rohanpaul_ai

Google dropped EmbeddingGemma 2 for on-device multimodal AI, under an Apache 2.0 license > puts text, code, images, audio and video into 1 searchable space on phones. gives phones a missing piece: a way to understand and search your own stuff without sending it to a server. > 740M parameters, uses the Gemma 4 architecture. > Its parts are modular, so a text-only app needs just 270M parameters, while a 170M vision encoder and a 300M audio encoder load only when needed. > On a Pixel 11 Pro, quantized text weights take about 191MB of active RAM, and the full multimodal model takes about 567MB. > The context window grows 4x to 8K tokens, enough for roughly 5.5 minutes of audio, 29 images or 58 video frames in 1 input. > Code search gained most, with the MTEB Code score rising from 68.76 to 78.68 while multilingual text scores held level. > Google also claims top sub-1B results on audio and vision benchmarks and wins over some specialist models twice its size,
打开原帖#511482
  1. Industry

    Greg Brockman: towards acceleration of scientific discovery and improving quality of…
  2. Research

    François Chollet: Does G "emerge" from math + code RLVR?
  3. Research

    François Chollet: What if the jagged frontier is mainly math + code?