Call for Evidence: AI testing, evaluation and assurance in the UK

| Resume a previously saved form
Resume Later

In order to be able to resume this form later, please enter your email and choose a password.

Password must contain the following:
  • 12 Characters
  • 1 Uppercase letter
  • 1 Lowercase letter
  • 1 Number
  • 1 Special character

The NPL Centre for AI Measurement is seeking evidence to inform how it supports the UK's emerging AI assurance market and the development of rigorous, comparable methods for testing and evaluating AI systems. We are interested both in evidence about wider challenges across the UK's AI assurance ecosystem and, where relevant, your own organisation's needs and potential interest in future support or collaboration.


For the purpose of this call, AI assurance refers to the process of measuring, evaluating and communicating the trustworthiness of AI systems and their components.


This questionnaire focuses primarily on the technical foundations of assurance: the measurement science, methods, metrics, tools and infrastructure used to test and evaluate AI. Broader governance, legal, organisational and regulatory aspects are important, but are not the primary focus of this call. For more information, please visit the Call for Evidence page.

Part 1: About you






You answer will be used to tailor questions you see in Part 2.

Part 2: Challenges, unmet needs and opportunities

This section gathers evidence about the underlying challenges and unmet needs in AI testing and evaluation. Part 2A is for all respondents and asks about challenges across the wider UK ecosystem. Part 2B is optional and asks about needs, opportunities or interests specific to you or your organisation.

Part 2A: Ecosystem-wide challenges and capability gaps


Please identify, where you can, the types of AI, sectors, use cases, risk domains or evaluation capabilities affected.


Part 2B: Organisation-specific needs and opportunities

This section is optional and asks about your own organisation's needs or interests. Questions are tailored to the primary perspective you selected in Part 1.

Supply side

For organisations developing or providing AI testing, evaluation or assurance products and services.


Please describe the capability, the AI systems or use cases it would apply to, and the unmet market or user need it would address.

Please identify the most significant barriers, such as access to data, models, expertise, real-world testing environments, standards, validation or funding.

Demand side

For organisations developing, procuring or deploying AI


Please describe the use case, the type of AI system involved and its intended context of use, and how the testing or evaluation gap is affecting your ability to proceed. 


Intermediary

For research, standards, regulatory or other ecosystem organisations


For example, this could include research expertise, datasets, standards expertise or programme delivery. 

Part 3: Intervention design

We are exploring a range of ways a publicly funded programme could support the development and piloting of scientifically rigorous, repeatable and comparable AI testing and evaluation techniques underpinning AI assurance. The types of interventions below are illustrative examples to prompt your thinking, not a fixed or final set. We would like to understand which you think would be most useful, which you might take part in, and what conditions would need to be in place. We would also like to hear about approaches we haven't listed.


a) Collaborative research sprints: a series of short, intensive events bringing assurance providers, deployers and researchers together to tackle a defined assurance challenge.


b) Matched testing pilot: pairing an assurance provider with a deploying organisation to conduct testing and evaluation of a deployment-ready AI use case, contributing towards codification of best practice.


c) Innovation challenge fund: a competitive grant where teams apply to solve a defined problem developed by a sponsor (e.g. governmental body or business) in testing and evaluation, with winning proposals funded to develop and trial their solution, contributing towards the codification of best practice


d) AI assurance methods accelerator: giving third-party AI assurance providers dedicated access to a network of technical experts from NPL and strategic partner organisations to workshop methodological challenges, refine and validate emerging assurance approaches and build evidence on their effectiveness, repeatability and limitations.





Please describe it, including what it would produce.

For example, as an assurance provider, as a deploying organisation hosting a test, as a research or delivery partner, or as an advisor.


For example, types of AI, AI testing and evaluation techniques, or risk domains.

Part 4: Closing questions

NPL does not intend to publish individual responses to this Call for Evidence. We expect to publish an aggregated summary of the insights gathered, which may take the form of a report or a blog post. We will not attribute findings to individual respondents or organisations.

This would acknowledge that you contributed to the evidence base only. We will not attribute any particular views or findings to you as a result of your inclusion in this list.





Your contact details will be used only for the purposes you have selected above and will be handled in accordance with the NPL Privacy Notice.


To find out how we use and manage your data please read our privacy notice and terms and conditions.


© NPL Management Limited 2024 | Hampton Road, Teddington, Middlesex, TW11 0LW | Tel: 020 8977 3222