Rohan Paul
@rohanpaul_ai
Google dropped EmbeddingGemma 2 for on-device multimodal AI, under an Apache 2.0 license
> puts text, code, images, audio and video into 1 searchable space on phones.
gives phones a missing piece: a way to understand and search your own stuff without sending it to a server.
> 740M parameters, uses the Gemma 4 architecture.
> Its parts are modular, so a text-only app needs just 270M parameters, while a 170M vision encoder and a 300M audio encoder load only when needed.
> On a Pixel 11 Pro, quantized text weights take about 191MB of active RAM, and the full multimodal model takes about 567MB.
> The context window grows 4x to 8K tokens, enough for roughly 5.5 minutes of audio, 29 images or 58 video frames in 1 input.
> Code search gained most, with the MTEB Code score rising from 68.76 to 78.68 while multilingual text scores held level.
> Google also claims top sub-1B results on audio and vision benchmarks and wins over some specialist models twice its size,