We may earn compensation from reviewed products. Learn about our editorial policies.
There have been whispers. But now it seems to be true. Google Gemini is coming soon.
In fact, if reports of a December 2023 release are to be believed, very soon indeed. And that means we’ve reached the precipice of yet another huge leap in generative AI technology; the dawn of a multimodal era.
What does that very boring and robotic-sounding word mean you ask? Well, it means that several companies have developed AI tools that can communicate in more than one format. So far, if you’ve used ChatGPT or Google Bard, you’ll have only been able to type commands. And they will only have replied to you via text.
Now, that’s about to change.
But, I hear you cry in desperation, Siri has been able to talk and write to us for years, what’s all the fuss about?!? Well, let’s have a look at what these latest rumblings from Google and OpenAI mean.
What is Google Gemini?
The short answer? A multimodal generative AI tool.
That means Google Gemini will be able to receive and distribute information beyond standard textual formats. The rumours are that Gemini will be the first AI tool to be able to communicate using images.
We’re not exactly sure about the details, but we can assume that means Gemini will be able to comprehend things like graphs and explain them or describe a photo back to a user.
There is no suggestion that Google Gemini will be able to create original images from user commands, but this probably isn’t far away (once the wildly dangerous legal minefields have been thoroughly defused).
Even so, this is a big development. Getting an AI model to understand text is relatively simple. You can, in theory, teach an AI every word in existence, and then the most likely combinations of words to form certain sentences. Images are a different beast with an infinite amount of different variations.
And that means you need a huge database of content to train AI models on. Luckily for Google, as the foremost data harvester in the world, they have all the image data they need.
Gemini is also rumoured to be coming for GitHub’s Copilot with their own version of a coding assistant. And in general, they look to have vastly improved their Large Language Model used for Bard, so much so that it could surpass GPT-4. Now that would be a turn-up for the books.
But something tells me a certain company won’t take this lying down…
What about ChatGPT?
Like a toxic video game studio, OpenAI is reportedly crunching hard to try and beat the release of Google Gemini. Although, Google has already released it to a select number of businesses for early access.
OpenAI had already previewed multimodal capabilities when it released GPT-4, but stopped short of integrating them into their publicly available toolkit. Those features were going to be released soon anyway, probably early next year, but Google’s imminent release has put the rocket boosters on GPT-4.
For the past year, OpenAI has dominated the generative AI space. Their LLM has been the best since it was released and no one has come close. Google Bard, for example, is noticeably worse at performing simple tasks like suggesting blog post titles.
And, although backed by Microsoft, OpenAI was still battling a much stronger foe in Google. So, to give up their seat on the AI throne would throw away years of hard work. Sure, people may still use their service, but Google is far more ubiquitous as a company. Businesses and people know Google’s brand more and are therefore more likely to use their products.
For now, no one knows who will be the first to release an AI with multimodal capabilities. My money is on Google, who, back at Google I/O 2023, made it clear that they saw AI as the future of tech, and therefore their company.
But is it more important to be the first or the best?
Where’s Apple AI…and goodbye Bard?
When it comes to tech, we have historically thought of Apple as the producer of the best products.
Want the best phone? Get an iPhone. Want the best laptop? Get a Macbook Pro.
Although this precedent has changed recently, especially in terms of phones where the Samsung Z Fold 5 and Google Pixel Fold have taken the biscuit, Apple is still expected to be the market leader in whatever it puts its mind to.
But Apple hasn’t so much as lifted a finger over generative AI, the most defining bit of tech of the 21st Century. I have just one question; why?!
I’m not saying they could do better or need to do better. But for a company as savvy as Apple, surely they want a slice of the ever-growing AI pie? It seems an obvious area to invest in, even just to future-proof your devices for when generative AI becomes more pervasive.
Of course, they have Siri, technically a multimodal AI. But Siri is neither generative nor very good anymore. It’s not even the best AI assistant (Alexa is much better). Siri can tell you a recipe, send a text or whatever simple command you wish. But it is infantile compared to GPT-4.
And what of Google Bard? Google only released their first generative AI chat model a few months ago. But it seems like they are going to consign it to irrelevance already which, honestly, is no great loss. Bard will probably just become Gemini or the other way around?
But that’s a minor detail. The point is, Google has demonstrated levels of ambition and innovation in the last five years that Apple hasn’t seen since, well, the first iPhone. And that makes no sense!
Perhaps Apple has an ace up its sleeve. But I think they’ve missed the boat, train, plane and now rocket ship.
Is multimodal AI a good thing?
In a technical sense, the answer is obviously yes. Looking beyond civilian use, image-related AI capabilities will benefit us all, just look at the world of medicine for example, where AI is already helping doctors to detect cancerous tumours.
But I asked the above question in relation to something that is already plaguing generative AI; the law.
Nobody has comprehensively tried to regulate AI yet. And yet, there are a myriad of legal and ethical issues springing up all over the place.
Authors suing for plagiarised work, writers striking to prevent the use of AI, unauthorised use of personal data, the list goes on and on. Adding multimodal capabilities only furthers the chances of legal challenges, especially when generative AI image creation is inevitably added to Google Gemini and ChatGPT.
Anybody who has a job involved with producing visual media will not like what is going on here. And who can blame them?
There is only one response to such concerns as always. AI’s encroachment on life is inevitable, it’s too beneficial to stop it, but can we strike a balance to regulate it?
One thing’s for sure, Google Gemini is about to make that debate a whole lot more complicated.
Google Gemini Final Thoughts
On the whole, I’m looking forward to the release of Google Gemini.
I don’t really care who out of Google and OpenAI gets there first, I care about who does it best. And if Google releases something nicely polished and actually useful, which Bard really wasn’t, then it will be a seminal moment for tech development and society in general.
So I wait with genuine excitement to play with what should be a free service from Google (if they are serious about their whole democratising AI initiative).
But I can never get truly giddy about AI like I want to, not until someone steps up to put some much-needed regulation in place.