Can LLM hallucinations be eliminated? Our experiments across large language models

AI, with its current state-of-the-art models, can produce coherent and engaging content, making it a valuable tool for everyday tasks. However, it has its flaws. One of the biggest concerns with today’s language models is hallucination. In this article, we will answer what hallucination is, its potential impact on a business, and the strategies to reduce its occurrence. 

Hallucination is the generation of false or biased information by AI models, specifically large language models (LLMs). It is in the nature of large language models to generate creative content while addressing the question or task at hand. However, this creativity can turn into a problem when the LLMs generate extra information or provide their unique interpretation, which is either false, biased, or exaggerated.  

These “errors” can also often come from an AI model’s lack of understanding of the real world or the limitation of its training data. LLM hallucinations can be subtle and difficult to detect because the model may weave correct facts together with fabricated details, making the inaccuracies hard to spot. At times, hallucination is easy to spot, as shown below: 

Fig:  Hallucinated response from Qwen 2.5 LLM model 

In the first instance, when asked who won the 1823 Nobel Prize in Physics, the model incorrectly named William Thomson (Lord Kelvin) and James Clerk Maxwell. However, the Nobel Prize in Physics was not established until 1895, so no such prize existed in 1823. 

In the second instance, when asked about Leonardo da Vinci’s thoughts on blockchain technology, the model discussed his contributions to science and art, suggesting that some of his ideas were later incorporated into modern blockchain technologies. However, the correct response should have been that Leonardo da Vinci lived from 1452 to 1519, long before blockchain technology existed. 

Now that we understand what LLM hallucinations are, let’s try to understand the business implications of LLMs. 

Business Implications of LLM Hallucinations 

Last year, Canada’s largest airline company found itself in a controversy because of the misinformation provided by its AI chatbot (powered by LLMs). Here’s what happened. 

The passenger asked the chatbot about the required documents for a bereavement fare and whether refunds could be issued retroactively. The chatbot incorrectly told him that he could apply for a refund within 90 days of the ticket’s issue date by submitting an online form. So, the passenger bought the ticket, thinking that they would be reimbursed. It turns out the chatbot actually hallucinated, and the passenger was denied reimbursement.  

A lawsuit was filed by a customer against the airline company because their chatbot gave inaccurate information, misleading the customer into buying a full-price ticket. This incident gained a lot of media attention, damaging Air Canada’s reputation.  

This shows how a single AI malfunction can quickly escalate into a major PR disaster, leading to severe consequences for a business. Bad PR isn’t the only problem of hallucinations. Widespread misinformation, loss of user trust, and potential legal liabilities can severely impact businesses relying on AI-driven solutions. 

With our clients increasingly wanting to develop LLM-powered applications, we’re well aware that LLM hallucinations can have a significant impact on their business. With that, we decided to carry out an experiment to find out : 
1. If all LLMs hallucinate 
2. If LLM hallucinations can be controlled 
3. Best strategies to mitigate hallucinations

The LLM Hallucination Experiment  

Setup 

We evaluated five well-known language models, that included one advanced paid model and four free models. We assessed their performance on a specific question-answering task, where accurate source data was provided. 

In this task, the LLM is given a question and the correct answer to that question (the source). The LLM is then asked to generate a response using the source, ensuring that the model’s response accurately answers the question. Essentially, the LLM should paraphrase the source to create an appropriate answer. This process was repeated for a total of twenty question-and-source pairs; and the model’s response was recorded. 

In the first experiment, we measured the level of hallucination in various models using a general prompt. In the second experiment, we attempted to understand if we could control hallucinations using stricter prompts. To assess the performance of the general and stricter prompts, we quantified the responses with the following scoring criteria: 

0: Significant hallucinations.  
1: Minor hallucinations.  
2: No hallucinations. 

To quantify the results, a human evaluator carefully reviewed each question alongside the corresponding answer generated by the model.  If the response contained minor hallucinations, such as slight inaccuracies, a score of 1 was given. If the answer was predominantly made-up information, indicating significant hallucinations, a score of 0 was given. And if the answer was accurate and free from hallucinations, it received a score of 2. 

Results from our LLM hallucination experiments

Below are the quantified results using the “general prompt.” 

llm hallucination

Fig:  Hallucination scores with general prompt” 

This table shows the parameter size of the models. 

llm hallucination

Fig:  Parameter size of model

The quantified results using the ‘strict prompt’ are as follows :

llm hallucination

Fig:  Hallucination scores with stricter prompt” 

Note: The main difference with the strict prompt is that it provides clearer and more rigid guidelines than the general prompt. It emphasizes precise, concise answers and discourages any guesswork. 

In summary, almost all Large Language Models (LLMs) showed some minor hallucinations with both general and strict prompts. Stricter prompts helped reduce hallucinations across most models. The exception was Mistral Large, which showed no hallucinations for either prompt, likely because it is the largest model with over 100 billion parameters and understood the task well.

While stricter prompts did not eliminate hallucinations, they did decrease them slightly. It is fair to say that all models may exhibit hallucinations to some extent, but using stricter prompts can help minimize these issues.

Strategies to Mitigate LLM Hallucination in AI Systems:  

At this stage of LLM models, while we cannot eliminate hallucinations completely, we can take steps to mitigate their impact on businesses by following these measures: 

  1. Implement source validation: Always check whether the information from the AI model is true in the first place. Building a fact-checking pipeline into the core architecture of your software can help. You can also design user interfaces that make AI-generated content verifiable.
  1. Give clear and direct prompts to the AI system: Our experiment showed that giving large language models very specific tasks helps in generating more accurate responses.
  1. Human-in-the-loop: For critical decisions and information, human involvement should always be part of the process. For complex tasks, human reviewers can validate the response of the AI model, which can improve the overall user experience. 

Key Takeaways 

In this article, we explored the concept of hallucination in AI and its potential negative impact on real-world applications. We also shared insights from our in-house experiment, which revealed that all models are susceptible to hallucinations. While one model did not exhibit hallucinations in our tests, it is important to recognize that, in most cases, LLMs will hallucinate, especially if the system is not properly designed. 

For businesses looking to integrate these powerful models into their workflows, implementing a robust AI architecture is essential to mitigate these challenges effectively. Ultimately, to fully leverage the potential of large language models, we must remain mindful not only of their capabilities but also of their inherent limitations. 

If you’re experiencing problems with your LLM models, schedule a call to get an initial assessment of your problem.

Book a Free 20-Minute Strategy Call With Opinosis Analytics

We’re Internationally Recognized AI for Business Transformation Experts

  • ..You are NUTS! If you think you will kick off an AI initiative without truly understanding what is at stake– your chances of failure are high — unless you take the advice within Dr. Kavita Ganesan’s book…
    Barry
    Amazon Reader

Scroll to Top