We may earn compensation from reviewed products. Learn about our editorial policies.
Google’s annual I/O developer conference was yesterday and as expected, artificial intelligence took centre stage.
Google unveiled Gemini 1.5 Flash, the Veo text-to-video generator, and launched updates and more that will enable further integration for its AI capabilities across its entire ecosystem, from Android to Google Workspace.
Here’s what we learned.
1. Gemini 1.5 Flash and Pro
Google kicked off the event by introducing Gemini 1.5 Flash, a new addition to its large language model (LLM) family. Gemini AI can understand images and text and offer creative and problem solving solutions.
Set to be the fastest Gemini model served through the API, this iteration offers a more cost-effective alternative to the Gemini 1.5 Pro. Gemini 1.5 Flash has allegedly been optimized for speed and efficiency, making it an attractive choice for developers and businesses.
The existing Gemini 1.5 Pro has also undergone significant enhancements. The Pro was already capable of scanning an image and understanding it. For example, you could take a picture of a maths problem and Gemini Pro would give you a step by step solution.
Now it’s been upgraded, across a range of benchmarks, including translation, reasoning, coding, and more. Notably, Gemini 1.5 Pro now offers a one-million-token context window, allowing users to leverage the model’s prowess on larger bodies of work, such as lengthy documents and extensive email threads.
2. Gemini Nano
Google also announced at I/O 2024 that they’ve updated their Gemini Nano model, which has now been expanded to include multimodal capabilities. Starting with Pixel devices, Gemini Nano will be able to understand text, visual and audio inputs.
Basically, you’ll be able to point your camera at an object and have Gemini instantly identify and provide information about it, or have the AI assistant transcribe and comprehend your spoken queries in real-time. But in a more powerful way than the updated Google search engine (see below).
This integration of Gemini Nano’s multimodal prowess with Android’s accessibility features, such as TalkBack, will benefit users with visual impairments or low vision. This is because they will now benefit from AI-powered descriptions of unlabeled images.
3. Gemini Advanced
Google has also upgraded its Gemini Advanced subscription tier. In addition to the previously announced perks, such as access to the latest Gemini models, Gemini Advanced users can get new features.
One of the additions is Gemini Live, a mobile-centric experience that allows users to engage in conversations with the AI assistant, complete with a selection of natural-sounding voices and the ability to interrupt the dialogue at any point. Gemini Live’s integration with Project Astra, Google’s vision for the future of AI assistants, further enhances its contextual awareness, enabling it to understand and respond to visual cues in real time.
Gemini Advanced subscribers will also be able to create their own personalized “Gems” – custom versions of the Gemini assistant tailored to specific tasks or personas, akin to OpenAI’s GPT Store.
The idea is it can be more than one thing: a coding partner, a meal planner, or a creative writing muse.
4. Google Search
The company’s new Search Generative Experience (SGE) is set to roll out to all users in the United States, providing them with AI-powered overviews that offer concise, conversational answers to their queries.
Google is also introducing AI-organized search results, which use natural language processing to generate unique headlines and summaries that better suit the user’s needs.
So you can search for dining options, vacation inspiration, or the latest entertainment releases, and the search results will be tailored to your specific interests and preferences.
Google is also expanding the capabilities of its video search feature, allowing users to capture footage of an object or scene and then query the AI to provide relevant information. Basically you can point your smartphone at a vintage record player and instantly receive step-by-step instructions on how to use it.
5. Veo Text-to-Video Generator
In the realm of generative AI, the company unveiled Veo, a text-to-video model, capable of generating high-quality 1080p videos that exceed a minute in length. Veo’s natural language understanding allows it to better interpret user prompts, resulting in video outputs that more closely align with the intended vision.
Veo is set to incorporate cinematic techniques, such as time-lapse and aerial shots, further enhancing the visual appeal of the generated content. Users can also leverage Veo’s editing capabilities to refine and customize the output, making it a powerful tool for content creators and visual storytellers.
While Veo is currently available only to select creators through Google’s VideoFX experimental platform, the company has promised to bring its capabilities to YouTube Shorts and other platforms in the future.
6. Imagen 3 Text-to-Image Generation
On the heels of Veo’s debut, the I/o developer conference launched the latest iteration of its Imagen text-to-image generator, Imagen 3.
This model boasts significant improvements in image quality, with greater detail and fewer artifacts, resulting in more realistic outputs.
Imagen 3’s enhanced natural language understanding allows it to better interpret user prompts, ensuring that the generated images more closely align with the intended subject matter and artistic vision. This will help those looking for a logo design, infographic creation, and visual branding.
Similarly to the above, Imagen 3 is currently available only to select creators through Google’s Image FX platform, the company has promised to make it more widely accessible through its Vertex AI service in the near future.
7. SynthID: Watermarking the Future of Generative AI
Google has expanded its SynthID technology to cover not just images, but also text and video.
With SynthID, any content generated by Google’s AI models, including Veo and Imagen 3, will bear a watermark that identifies it as AI-generated.
This feature aims to provide transparency and help combat the potential misuse of these powerful technologies, ensuring that creators and consumers alike can make informed decisions about the content they encounter online.
8. Ask Photos
There’s a new “Ask Photos” feature, powered by Gemini, which seeks to make navigating your photo gallery smoother.
Users can now pose natural language queries to their Google Photos library, such as “Show me photos of my daughter learning to swim,” and the AI will automatically surface the relevant images, saving time and frustration.
As the “Ask Photos” feature evolves, users can expect to see even more advanced capabilities, such as the ability to identify specific objects, locations, or events within their photo collections.
9. Android & Gemini
Google Assistant is being replaced by Gemini.
Gemini, the company’s flagship AI assistant, will soon take over the familiar Google Assistant, becoming the default AI agent across Android devices.
But the integration of Gemini into Android goes far beyond a simple name change. The AI assistant will be deeply embedded throughout the operating system, providing multimodal support and context-aware capabilities.
One particularly noteworthy feature is the enhancement of Android’s “Circle to Search” functionality. Users will now be able to circle text, images, or even equations on their device’s screen, and Gemini will provide step-by-step guidance on how to solve the problem or find the relevant information.
10. Gemini for Google Workspace
Google Workspace is also incorporating more Gemini.
The Gemini side panel in key Workspace apps, including Gmail, Docs, Sheets, and Slides, will be upgraded to leverage the advanced capabilities of Gemini 1.5 Pro.
This integration means that users can now tap into Gemini’s one-million-token context window to summarize lengthy email threads, generate targeted responses, and even pull relevant data from spreadsheets – without leaving the Workspace environment.
But the Gemini-powered enhancements don’t stop there. Gmail users can now leverage features like “Summarize,” “Gmail Q&A,” and “Contextual Smart Reply,” which leverage Gemini’s natural language processing to provide intelligent, personalized assistance with email management and communication.
11. Project Astra
Underlying many of Google’s AI announcements at I/O 2024 is Project Astra.
This project aims to move beyond the traditional voice-based interactions to a more multimodal experience.
Gemini Live, the immersive conversational feature introduced for Gemini Advanced, is a prime example of Project Astra’s influence. By enabling users to engage with Gemini through a variety of input and output modalities, including video and natural language, the assistant should become a more intuitive and powerful tool for tackling everyday tasks and queries.
Moreover, the integration of Gemini Nano’s multimodal capabilities into Android’s accessibility features, such as TalkBack, demonstrates Google’s commitment to making AI-powered assistance accessible to users with diverse needs and abilities.
What do you think?
And thats what’s new at Google!
What do you think about all this AI? Let us know in the comments down below!