- AI hallucination is all part of the package, says Snowflake CEO.
- Certain situations require 100% accuracy so AI inherently unsuitable.
- Models lose their ‘magic’ if right all the time, apparently.
Reducing hallucinations from AI models is the goal of many big providers of the technology, yet if we insist on 100% accuracy, an AI will never produce any answers.
So says Snowflake CEO, Sridhar Ramaswamy, speaking on The Logan Bartlett Show (see video below).
AIs not given leeway to make mistakes, AKA hallucinate, would “become very, very worried about making mistakes and they will say, ‘I don’t know the context’ to everything,” according to Anthropic founder Jared Kaplan. And the opinion of Sam Altman, the mind behind OpenAI, is that a fully accurate AI “won’t have the magic that people like so much.”
This acceptance of inaccurate answers to queries is something that the vast majority of users of models simply don’t expect. And in most cases, people aren’t aware that responses occasionally riddled with mistakes are all part of the deal.
Yet the snake-oil of AI is constantly added to the marketing material of any technology solution aimed at businesses or consumers. Whether or not the claims that there really is an AI element are true, in every case the advertisers would at least like the ‘powered by AI’ to be true. In 2024, it’s getting hard to find a piece of software that’s not been touched by AI.
AI hallucination in your business finances
Ramaswamy said that for applications like analysing financial data, AI tools simply “can’t make mistakes,” (as in, cannot be permitted to). The models that are often bolted-on to existing solutions to help businesses ‘make better decisions’ are inaccurate by nature. The double kicker in the tail, the CEO said, is that there’s no way of knowing which results produced by an AI contain errors, and how many errors the AI has been engineered to be allowed to make.
“No one publishes hallucination rates on these on their models or on their solution,” Ramaswamy said, which suggests that those marketing their models are keeping quiet about the inherent necessity for models to make things up. Without the ability to fabricate to produce answers, there would be no answers at all. Yet users are not made aware of this before signing-up for the latest technology.
“I don’t think the […] AI industry […] does itself any favours by simply not talking about things like hallucination rates,” he said.
Backwards engineering solutions to AI hallucination
The tech industry’s solution to unreliable results from AI models is to engineer guard rails which either correct suspect output, or add censorship layers; either at the point of inference (when queries are posed and answers formulated) or adding code to tweak results and/or warn users about the veracity of responses they are about to receive.
Google Gemini Pro, for example, hallucinates 7.7% of the time, yet its “relentless innovation” is apparently capable of, “unlocking the ability to accurately process large-scale documents, [and] thousands of lines of code.” Processing material is one thing; answering questions accurately about what’s in the material is another.
Anthropic’s Claude-2 has an error rate of 17.4%. “Think of Claude as a friendly, enthusiastic colleague or personal assistant who can be instructed in natural language to help you with many tasks,” the company’s website announcement read, when the model was released in 2023. An “enthusiastic colleague” may be great to have around the office, but you wouldn’t want to trust their judgement in anything important given the knowledge they’re wrong around one time in five. And good advice might be not to let them write code for, say, an autonomous vehicle.
Deflating the bubble
Over the course of the next few years, AI models will gain in power as the technology and methods powering them are refined. The same processes of refinement will likely be applied to the number of practical use-cases for AI models, falling from its current envisaged infinite number to just a few, mostly as the limitations of the technology become clearer (or vendors of AIs stop making exaggerated claims about their dubious wares).
In the same way that blockchain was touted as the technology everything could leverage around 2018-2019, AI’s current wave-cresting will last only as long as no-one actually relies on it to do anything meaningful. The ensuing realisation will hopefully show the contexts in which an inherently and necessarily-flawed set of algorithms is appropriate.
“The insidious thing about hallucinations is not that the model is getting 5% of the answers wrong, it’s that you don’t know which 5% is wrong, and that’s […] a trust issue,” Ramaswamy said.
Author
- View all posts
Joe Green is a writer based in Bristol, UK. He acquired his first computer with dial-up modem in 1992 and has worked in the tech industry since 2000. He writes and podcasts, specialising in open-source, networking, cybersecurity, software development and online privacy.