Loading...
At its annual I/O conference, Google shocked attendees with a live demonstration of Gemini Ultra 2.0. The new model is truly multimodal from the ground up, capable of processing live video feeds, spatial audio, and continuous text input simultaneously without any perceivable lag.
The demonstration showcased a user wearing smart glasses, with Gemini acting as a real-time interpreter, spatial guide, and context-aware assistant. For example, the AI successfully identified a broken mechanical part via the video feed and verbally guided the user through the repair process step-by-step.
"We are moving from conversational AI to ambient AI," declared Sundar Pichai. "Gemini Ultra 2.0 understands the world exactly as you do—through sight, sound, and context."
The model will be available to enterprise developers next month, with consumer rollout slated for late 2026.