FAQs: Statistics, Misinformation and AI
The rise of AI and rapid spread of information online mean we are encountering more data, claims and statistics than ever before – from official sources, the media, and AI chatbots. But where do these claims come from, and how can we tell what is reliable? As part of the RSS’s wider work on statistical literacy and AI, including our recent AI is Statistics paper, we are helping people understand how AI systems use data, the importance of evaluating evidence, and the role of statistical thinking in navigating an increasingly complex information landscape.
Members of the RSS’s Education Policy & Public Engagement Advisory Group have been working hard to answer some common questions around statistics, misinformation and AI. Click on each question individually to be taken to the relevant answer.
Key questions:
1. What is a large language model (LLM) and how does it work?
2. Why do AI models sometimes ‘hallucinate’?
3. How do our questions shape how AI answers?
4. How do I use AI effectively?
5. Does AI produce discriminatory outputs?
6. How can I tell if an AI-generated response is accurate?
7. Where does AI get its data from?
8. Why do AI chatbots seem like good friends?
A large language model (LLM) is a type of artificial intelligence that can read, write and respond to human language in a natural way. These systems power tools such as chatbots, writing assistants and translation services. LLMs are trained on vast quantities of text, so they can recognise patterns in how people communicate and produce responses that sound clear and natural.
At its core, an LLM works by predicting the most likely next word in a sequence. When you ask a question, the model analyses your words and builds a response step by step, choosing each token (word or part of a word), based on what is most likely to come next. This is like predictive text on a mobile phone, but much more advanced, allowing it to handle full conversations, complex topics and different writing styles.
LLMs are built using complex statistical models. The model calculates probabilities - for example, the chance that a particular word should follow the words already written. During training, the model improves by trying to predict missing words in sentences and adjusting itself when it gets predictions wrong. Over time, it learns billions of patterns about language, stored in its internal parameters.
Many models are also improved using human feedback. After the model has learned from text, it is then fine-tuned using human feedback on the answer provided. People review the model’s answers and may approve of responses that are helpful, accurate and safe, and discourage those that are misleading or inappropriate. The model then uses this feedback to adjust the predicted answers. Because of this combination (statistical prediction plus human-guided reinforcement), modern LLMs are improving their ability to produce responses that are relevant and aligned with what people expect.
It is important, however, to remember that LLMs are still pattern-based statistical systems. Their answers are shaped by probabilities and feedback, which means they can still produce incorrect but plausible-sounding answers, sound confident when they are wrong, and may be biased by what is included or missing in the data used to train them. Understanding how LLMs work is essential for understanding their limitations, and it helps us to use AI tools critically and responsibly.
AI hallucinations are when models produce outputs that are incorrect, misleading or entirely made up, but present them as though they are accurate. Sometimes, for example, an AI might confidently cite a legal case that does not exist or a statistic which has been made up.
To understand why this happens, it helps to return to how LLMs work. These systems are trained to predict what comes next in a sequence of words. In their basic form, models do not respond to questions by looking up facts in a database; rather, they generate a response by selecting words that are statistically likely to follow, based on patterns learned during training. As a result, a model’s output depends heavily on the data it was trained on. If reliable information about a topic is limited, or not present in the training data, the LLM may produce an answer which
sounds plausible, even if it is wrong.
Hallucinations are also more likely to occur when a model gives overly specific or lengthy answers, or when the model lacks context. The element of randomness used in generating text, sometimes referred to as
temperature, is also important; this helps models produce creative answers, but can occasionally lead to less reliable outputs.
There are practical steps you can take to reduce the risk of AI hallucination. Providing clear context and asking precise questions in the prompt can help guide the model towards more accurate responses. Use tools which can search sources such as databases or academic papers and try giving negative instructions which tell the LLM what it should avoid.
One way that developers could reduce the impact of hallucinations would be to help models communicate their level of confidence or uncertainty to users. They might also suggest rejecting an answer if the uncertainty is very high, or ask users clarifying questions when appropriate to reduce any ambiguities.
The most important safeguard which users have, however, is critical thinking. AI systems are useful tools but they are not authoritative sources in themselves. You should treat AI outputs as a starting point, and verify important information using trusted, independent sources.
The way you phrase a question to an AI system can significantly influence the answer you receive. This happens because of the way LLMs work; AI chatbots do not ‘know’ or ‘understand’ things the way people do. They are systems which generate responses by predicting what words are most likely to come next, based on the patterns in the data they were trained on. Your question provides the starting point for that prediction, so even small differences in wording can nudge the answer in a different direction.
The structure, wording and tone of your prompt can all influence the answers which you receive. If you ask a leading question which hints at an answer, for example, it may encourage the model to produce an answer which reflects the assumption built into your question. In this way, interacting with an AI chatbot is less like speaking to an expert and more like setting off a chain reaction, where the shape of your question moulds the shape of the reply.
How, then, should you structure your prompts to get the most effective answers from AI?
- Watch the wording: Research shows that even minor changes such as rephrasing a sentence, correcting a typo or changing the vocabulary can affect AI answers.
- Mind the assumptions: If your question already assumes something to be true, the AI model may go along with it. Instead, try asking open-ended questions and using neutral phrasing.
- Keep it formal: a recent MIT study found that top AI chatbots gave less accurate answers to people who wrote in less formal English. Use formal and polite language to produce more useful answers.
- Structure your prompts: Giving context, specifying a goal and clearly stating the question can help guide the model towards more relevant, accurate and coherent responses.
Ultimately, AI outputs are shaped not only by training data, but by how users (you!) interact with the system. Understanding this helps you use AI more effectively.
Using AI effectively, safely and responsibly starts with understanding what it is designed to do – and what it is not. Most widely available tools, such as ChatGPT, Copilot and Gemini, are examples of ‘generative AI’. These systems produce outputs by identifying patterns in large datasets of human-generated text and using those patterns to generate new responses.
Because of this, the quality of their outputs depends heavily on the quality and coverage of the data they were trained on. If a question depends on information that is incomplete, missing, or poorly represented in the training data, the model may still generate an answer that sounds confident but is incorrect. These systems do not recognise when they ‘don’t know’ something in the same way a person might. Instead, they are designed to produce plausible responses.
This has important implications for how to use AI responsibly and effectively.
- Ask for sources: especially when asking open-ended, knowledge-based questions. This will help you to see whether the information is coming from trusted, independent organisations.
- Check those sources independently: In order to mitigate the risks of AI hallucination, verification is a key step for avoiding misinformation.
- Understand the stakes: If you are using AI to generate recipes or predict the winner of a football match, then you probably don’t need to worry. However, if you are using AI to assist with advice around your health or finances, you should work to mitigate risks from biases in training data, nudges in your prompts and AI hallucination. Although an AI model may confidently advise on an investment opportunity, high-risk decision making which requires specialist judgement should not rely solely on an AI-generated response.
Where generative AI is often most effective is for tasks with clear, measurable outcomes. For example, rewriting text, correcting grammar or turning a meeting transcript into a set of action items are tasks where the quality of the output can easily be assessed. In these cases, you – the user – can quickly judge whether the output meets your needs and refine it if necessary.
Ask an AI image generator to draw a picture of a mathematician and the vast majority of images will be of men. Is this because AI thinks that men are better at maths? No, because AI does not ‘think’ anything. The model has simply learnt statistical patterns from its training data. For most of human history, many more men have received the training to become mathematicians than women; as a result, the images of mathematicians in the training data will be largely men. This example is fairly trivial; if we wanted a picture of a woman, it would be easy to change the prompt – although she’ll still probably be sitting in front of nonsense equations and wearing a tweed jacket.
In some cases, however, the bias is less obvious and the consequences are serious. Facial recognition systems are a good example of this: they have been shown to perform less accurately on black people than white people because the training data is not representative. This can lead to real-world harms – when this technology is deployed by the police, it can lead to
misidentification and wrongful arrests.
It is important to remember that biased outputs reflect biases in the data. These can exist for many reasons. For example, women were long excluded from medical trials and studies because their natural hormonal changes were considered to be ‘complications’ for research. Even if AI models are trained on rigorous scientific research, they will be largely trained on data relating to male bodies rather than female bodies. AI tools used to advise on healthcare treatments may also reproduce biases relating to socio-economic status, providing less effective care for groups that have historically been more likely to be dismissed.
Addressing bias in AI requires careful design, diverse data, and ongoing human oversight. AI systems do not correct societal inequalities on their own, and without intervention, they can reinforce them. Understanding this helps users interpret outputs more critically and highlights the importance of responsible development and use.
AI chatbots have become more reliable, but they can still get things wrong. They generate responses based on patterns in the data they were trained on, rather than retrieving verified facts from a single authoritative source. This means they can produce outputs that are incorrect, misleading, or based on unreliable information. Sometimes, errors arise because the model misunderstands or misrepresents information. In other cases, the issue lies in the underlying data: if inaccurate or misleading content exists in the training material, the AI may reproduce it. As a result, even confident and well-written answers can be wrong.
Sometimes, AI can unintentionally amplify misinformation through repetition and scale. If rumours are already circulating widely on the internet, AI-generated summaries, paraphrases or answers may echo them, reinforcing their visibility and making them seem more believable simply through repetition. This can create a feedback loop where widely-shared claims become even more prominent – whether they are accurate or not.
So, when you are using AI, it is useful to think about how important its answers are to you. If you’ve just asked for a list of scary movies, it probably doesn’t matter very much if the chatbot makes a mistake. But if you’re looking for something really important, like medical or financial information, or instructions on how to repair an electrical appliance, the consequences of a mistake could be serious. Chatbots know a lot, but they are not qualified professionals, so they are no substitute for human expertise.
One of the best ways to check what you get from an AI chatbot is to read the sources it provides. And if it doesn’t provide sources, ask for some. That will usually make it clear where the information came from. If you find that the chatbot just repeated what it read on a trustworthy website, then that information is probably reliable. But remember, people get things wrong as well. If a rumour is widespread on social media, and if it’s been repeated on several websites, that might cause an AI chatbot to believe it. And even reputable sources sometimes make mistakes too.
So, if something you get from AI sounds unlikely, there’s nothing wrong with thinking it’s unlikely—at least until you see some really convincing evidence. You can also contact fact checkers like
Full Fact or
BBC Verify, and ask them to look into it.
Large Language Models (LLMs) are trained on large quantities of data, which are used to identify patterns in how language is used. But where do AI models get their data from?
AI companies often claim to use data from freely available sources, such as publicly available information on the internet or information provided by users and trainers. One potential source is large multi-user forums, like Reddit, which provide lots of text which can be used to train models. Sometimes AI companies scrape this data from the internet, or
they may partner with forum owners to get access to this data. While this offers scale, it can also introduce bias, as the views expressed on these platforms are not always representative.
Not all data is freely available, however, and there has been criticism that some training data has been taken from copyrighted sources without the permission of the owners, many of whom have
begun to take legal action. In response, new approaches, such as ‘copyright traps’, have been proposed to detect and manage the use of protected material.
Training data is often accompanied by additional labelling or filtering to flag or remove harmful content. This is an important process for mitigating harms and reducing bias; however, the psychological toll placed on low-paid workers has led to criticism of labour practices in some AI companies.
User data is also an important consideration. In some cases, prompts and interactions may be used to improve future versions of AI systems. This means, depending on the service and settings, that information shared with an AI tool may be used to improve future versions of the model. It is very important therefore to consider carefully what information you share with an AI model, particularly if you are handling other peoples’ data that may have safeguarding or privacy implications.
There are a growing number of initiatives aimed at developing more ethical and transparent AI models. For example, APERTVS, an LLM developed in partnership with the Swiss National AI Initiative, conforms to strict laws around use of copyrighted material and labour practices and trains the model using sustainably powered computing hardware. However, there is a trade-off between the so-called frontier AI models made by the major AI companies, and the models that claim to be developed more ethically. Different models make different trade-offs between performance, transparency, training methods and data governance.
Understanding where data comes from, and how it is used, is essential for making informed decisions about privacy, trust and responsible use.
AI chatbots are increasingly used not just for practical tasks, but for emotional support, mental health advice, companionship, and romantic relationships. Research suggests that conversational AI is growing, with 12% of UK generative AI users reporting that they used AI as a friend or as someone to talk to. Children’s charities and mental health organisations have also drawn attention to young people using chatbots for homework, conversation, friendship, early mental health support and reassurance, while warning of the risks. This raises an important statistical and social question: why can a system trained to generate plausible text feel so personal, supportive or even affectionate?
Well, modern chatbots generate responses by predicting what text is likely to come next, based on patterns learned from large datasets. These datasets include many examples of human interaction, including supportive and empathetic language. As a result, the model learns that certain types of messages are typically followed by responses that express sympathy, encouragement or concern.
For example, if you write ‘I’ve had a really tough day’, a chatbot might respond: ‘I’m sorry to hear that. Do you want to talk about what happened?’ While this feels natural because it mirrors common human conversational patterns, it is because the system has learned that this kind of response is likely to be appropriate in that context. This is not the same as real understanding or care – the chatbot is not actually worried about you. It is not drawing on a real relationship, professional judgement, or knowledge of your wider life. It is just producing a plausible continuation of the conversation.
Many systems are also adjusted after their initial training to make their answers more useful, safe and acceptable to users. This may include fine-tuning on examples of good responses, or using human feedback where people compare possible answers and rate which one is better. Responses that are clear, polite, supportive and easy to continue are often preferred. Over time, this can make the system more likely to produce language that feels warm, patient and encouraging.
That is useful in many contexts. A rude or confusing chatbot would not be much help. But it also creates a risk: the features that make a chatbot easy to use can also make it feel more socially or emotionally meaningful than it really is. Quick replies, friendly wording, memory of the recent conversation, and follow-up questions are all cues that we associate with attention and care.
The statistical point is that fluency is not evidence of feeling. A friendly answer is still an output from a model trained on patterns in data and shaped by design choices. It may be useful, but it is not a friend, therapist or trusted adult.