AI is increasingly becoming part of how people live and work. We spoke to RSS Head of Policy Jonathan Everett about the ways AI is disrupting traditional research, posing challenges for regulators, and even rewriting the philosophy of science. From our recently published papers to our engagement with decision-makers, he reports on the Society’s work to help the world keep pace.
AI has become a major area of work for the RSS. What's driving that focus?
Our major focus has been making the case that AI is a fundamentally statistical technology, and that this has implications for how we should effectively and ethically engage with it. That was the purpose of our AI Task Force's
AI is Statistics position paper, which we released in March.
One of the themes in that paper is evaluation. Why is that so important?
In
AI is Statistics, we set out some of the ways AI evaluation is currently falling short. Models are often evaluated on datasets that don't reflect real-world conditions, assessed only before deployment, or benchmarked in ways that fail to capture how they will actually be used.
Put simply, we need a better approach to evaluation.
Before joining the RSS, you worked as a philosopher of science. How does that background influence the way you think about AI?
Philosophy of science is all about trying to understand how science works. Traditionally, a big part of that has been around understanding reproducibility. One of the most interesting things (to me, anyway) about AI is that our understanding of reproducibility may have to change. And that could impact how we think about science itself.
Traditionally, scientific progress has relied on empirical predictions being tested through reproducible experiments. In data-led research, that has generally meant that, if researchers use the same data and the same methods, they should get the same results.
This rests on three assumptions: that methods are transparent and inspectable; that other researchers can apply those methods and obtain the same results; and that the underlying process itself is repeatable.
When AI is introduced into research, those assumptions can no longer be guaranteed.
What changes when AI becomes part of the research process?
If research relies on an AI model, particularly an advanced foundation model, it may not be possible to provide a fully inspectable account of how that model arrived at a particular output.
Even if another researcher uses precisely the same model, data and prompts, reproducing the results may depend on having access to the same level of computing power as the original researcher. If the model is proprietary, they may not even be able to use the same model in the first place.
And even when researchers have access to the same model and equivalent computing resources, AI outputs are probabilistic and may vary between users.
So we need to rethink what we mean by reproducibility.
I think so.
As AI becomes embedded in research, reproducibility can no longer be understood simply as obtaining exactly the same result from the same data and methods. The more important question becomes whether findings can be shown to be reliable, stable and trustworthy despite the probabilistic nature of the tools used to produce them.
That, in turn, means developing measures of reproducibility that can be reported as part of the research process. Developing those measures is fundamentally statistical.
What does the RSS think should happen next?
In
AI Regulation Needs Statistics, we identify a regulatory gap.
Existing research governance places a strong emphasis on transparency and good practice, but much of it is still rooted in the traditional understanding of reproducibility. We argue that regulators and research funders need to be more proactive in encouraging good practice when AI is used in research.
What's particularly interesting from the RSS perspective is that statistics sits at the centre of that change. Statistics is both powering the technology that's driving these challenges and providing the tools needed to demonstrate that AI-powered research is trustworthy.
And the regulators aren’t currently equipped to deal with this gap…
No, they’re dealing with a very complex landscape where AI models may be developed by one company, deployed by another and then used by individuals who have no oversight of how outputs are produced.
Responsibility is dispersed, which means we need a system-wide approach to evaluation.
How has the RSS explored those challenges in practice?
We've looked at case studies in particular sectors, including healthcare, education and financial services, to understand the challenges that are arising and identify broader lessons for government and regulators.
Those lessons are set out in our report
AI Regulation Needs Statistics.
Can you give an example?
Health research is particularly useful to think about.
AI is increasingly used as a research tool to analyse datasets, and that makes it harder to guarantee that health data is truly anonymised. Traditional anonymisation assumes that removing direct identifiers is enough to prevent individuals being recognised. But AI can detect complex patterns across clinical, demographic and genomic data that may enable re-identification, even when obvious identifiers have been removed.
As a result, whether someone is identifiable can no longer be treated as a fixed property of a dataset. It depends on the data itself, the AI models applied to it and the wider data environment. Evaluating whether an AI system is safe to use therefore becomes deeply context-sensitive and highly statistical in nature.
And how are you hoping to tackle these issues?
Our work is now focused on engaging with the new government and reaching out to regulators to identify the most effective ways of making the case for the changes we'd like to see, including ensuring teams have the statistical skills they need.
We're also keen to help connect these organisations with the expertise of RSS members.
Regulation is one side of the equation. What about public understanding?
That's the other major strand of our work on AI.
We think it's important that people understand the statistical nature of AI. Even when using a chatbot, it's important to recognise that it isn't responding in the same way a human would. It's predicting the most likely or optimal response.
To support this, we've published
a series of FAQs explaining how AI works, why it can produce misinformation and how people can be alert to it.
If you have questions or comments about our work, get in touch at policy@rss.org.uk.