European flag
>
The Potential of Ai in Policymaking, Scientific Advice and Evidence Review: Literature Review

The Potential of Ai in Policymaking, Scientific Advice and Evidence Review: Literature Review

7 August 2026
SAPEA
doi:10.5281/zenodo.21414001
The Potential of Ai in Policymaking, Scientific Advice and Evidence Review: Literature Review

Summary

Scope and structure of this report

  • This introductory literature review is designed to provide input to SAPEA’s Work Package 7, The use of AI in the Scientific Advice Mechanism. The work has been conducted by Cardiff University and was undertaken iteratively over the period April 2025 to February 2026.
  • The scope is on how AI might be used in policymaking, scientific advice and in scientific processes related to evidence review. Section 1 covers a timeline and definitions of AI, and the potential of AI to support policymaking and science advice. It also looks at considerations for the practical implementation of AI in a government/policy context. This section also covers AI governance, including international legislation, ethical guidelines and frameworks. Section 2 examines several key tasks that are undertaken as part of evidence review in the SAM; these are peer review, literature searching, literature screening, data extraction and synthesis. The section also sets out published guidelines and frameworks on the use of AI in systematic evidence review.

Main points and findings

Timeline and definitions

  • The timeline for AI goes back to the 1950s, when Alan Turing proposed the famous Turing Test and the term “artificial intelligence” was used for the first time.
  • There is no single universally accepted definition of AI. A commonly cited definition is from the OECD, which defines AI as a “machine-based system that can infer from the input it receives, how to generate outputs that can influence physical or virtual environments. AI systems vary in their level of autonomy and adaptiveness”.

Potential benefits of AI in policymaking

  • The literature suggests that AI has the potential to increase the productivity and efficiency of government and policymaking throughout the policy cycle. AI could help with agenda setting (for example, analysing large datasets to identify critical challenges, priorities or trends); policy development (by providing evidence-based insights, predicting the impact of policies, simulating policy options or modelling scenarios); policy implementation and continuous improvement (by enabling quicker evaluation and feedback loops).

Potential challenges of AI in policymaking

  • Risks include possible bias and discrimination against certain viewpoints or groups; provision of information that is incorrect or flawed; lack of transparency and accountability; negative impacts of poor system design; breaches of intellectual property rights; concerns over data security; issues around current geopolitics, the dominance of big tech and role of certain state actors.

Risk assessment

  • Risk assessments are crucial; policymakers should be transparent about their use of AI, particularly when decisions carry significant potential impact or high risk.

Implementation of AI in policymaking

  • The evidence suggests that most use of AI in government focuses on internal operations and public service delivery, with limited use in policy development and evaluation. Many AI projects are at the planning or pilot stages, rather than ready for implementation and scaling.
  • When considering the implementation of AI, it is important to assess the problem being addressed and why AI is being considered; the resources and competencies that would be required; what the risks are and how these are mitigated; and how performance and usage of AI will be measured. A cost-benefit analysis should be conducted.

AI governance

  • AI should be ‘human-centric’, prioritising human needs and societal good; building and maintaining public trust are vital.
  • In ensuring that the design and deployment of AI are responsible, ethical and trustworthy, both ‘soft’ (standards and principles) and ‘hard’ laws (statutes and legislation, regulations and directives) are either already being implemented or under development/revision. A number of these are described in the document, covering areas like responsible use, transparency, accountability, privacy and data governance, intellectual property rights, diversity and inclusion, and environmental sustainability.

Examples of European initiatives

  • We describe European-level initiatives, including the published case study of the JRC’s AI tool, GPT@JRC, which is being rolled out at the European Commission, where use cases are classified according to their technical complexity, the AI expertise of the requesting team and the resource investment required.

AI in science advice

  • We found limited evidence on the use of AI in science advice. The literature indicates that AI may have potential uses, for example, supporting evidence aggregation and drafting text. However, there are challenges. For example, the ‘black box’ nature of AI can make it unclear how science advice has been produced, and therefore opacity is an issue. The limited evidence from actual use cases showed gaps in performance between AI and human experts, as well as the potential for bias towards certain types of knowledge and societal groups.
  • Dealing with the challenges requires both advisers and policymakers to apply a high degree of AI literacy and critical thinking skills, whilst also adhering to ethically based frameworks and guidelines.

Potential of AI to support evidence review

  • In terms of the potential opportunities and challenges of using AI in a range of tasks related to systematic review, published overviews of the literature indicate that results are mixed and inconsistent.
  • We conducted literature searches to assess the merit of using AI tools across a range of specific tasks associated with evidence review in the SAM. These were:
Peer review
  • The literature suggests that AI could be suitable for certain, limited tasks (e.g. screening manuscripts for relevance to a journal’s scope or checking language) but is less suitable for tasks such as providing detailed feedback to authors. Current AI tools appear to lack the necessary in-depth expertise and fail to understand the subject area sufficiently. There are also issues around a lack of transparency, accountability and potential bias.
Literature searching
  • Evidence from the literature suggests that AI should not be relied on exclusively for searching, given that the search function in AI tools (such as Elicit) is found to be sub-optimal in terms of performance. By finding additional material, AI tools may be a useful check and/or can complement a traditional literature search performed by a human although there is a trade-off between this and the investment of time and other resources.
Literature screening
  • The literature indicates that AI tools do show potential for timesaving, especially when the screening criteria are simple and the review is non-exhaustive. However, performance declines as the complexity of screening increases. The investment time in training the AI, and the need for effective prompt engineering (instructions given to the AI) are issues to be considered.
Literature search, data extraction and synthesis
  • Although the published evidence appeared limited, the available evidence suggests that AI can be prone to hallucinations (information that is untrue or inaccurate). It can often struggle to access full text content, due to copyright issues. AI may be a useful adjunct in finding additional relevant literature but cannot replace human systematic reviewers at the present time.

Guidelines and frameworks

  • The number of AI-generated systematic evidence reviews that appear convincing but are flawed is likely to grow, potentially undermining the use of evidence review in areas like policymaking. Important guidelines and frameworks such as RAISE are being developed for the use of AI in systematic evidence review, which we summarise in detail in the report.

Overall conclusions

  • Although AI has the potential to support many tasks involved in the policy cycle, there are also risks that must be weighed. Conducting a risk assessment and cost-benefit analysis are important.
  • The integration of AI into evidence review workflows requires careful management, with strict adherence to responsible use and quality standards. This is vital within the context of policymaking as it impacts on people’s lives, so the stakes are higher and the risks are greater. The use of AI can be encouraged where it demonstrably improves the quality of the review process and outputs. At the same time, humans must always be fully responsible for verifying output and making judgement decisions. Full transparency about the use of AI and justification for its use are essential. Compliance with the law is vital, and intellectual property rights must be respected.
  • This is a landscape that is changing constantly as AI evolves and matures as a technology and more is known about its performance, so continuous monitoring and adjustment are recommended.

Introduction

This introductory literature search is designed to provide input to SAPEA’s Work Package 7, The use of artificial intelligence (AI) in the Scientific Advice Mechanism. It has been conducted by Cardiff University. Our report focuses on how AI might be used in policymaking, in science advice and in scientific processes related to evidence review. It does not cover other areas of SAPEA’s work, such as communications or administration.

Section 1: Timeline, definitions of AI
and AI’s potential in policymaking
and in science advice

Introduction

Section 1 starts with a timeline and definitions, then examines the potential opportunities and challenges of using AI in policymaking. It considers the practical implementation of AI in policymaking contexts, including European examples. It provides an overview of international legislation, principles and guidelines on the use of AI. It then looks at the available evidence on the use of AI for science advice.

Timeline and definitions

In 1950, the famous computer scientist Alan Turing posed the question of whether machines can think, proposing the famous Turing Test or imitation game1. In 1956, Professor John McCarthy introduced the term “artificial intelligence”. In 1959, the first AI laboratory was established at MIT and took the first steps in Natural Language Processing (NLP). A paper by Artur Samuel in 1959 introduced the term “machine learning” (ML). Publications on Artificial Neural Networks (ANNs) date from 1979 (Koltsakis, Klontzas & Karantanas, 2023).

There is no single accepted definition of AI that is widely recognised across all countries, contexts and organisations. AI is a broad field that covers a range of technologies, methodologies and applications; definitions depend on the AI’s area of focus and purpose (UNESCO & OECD, 2024).

The European Union High Level Expert Group on Artificial Intelligence (2019) defined AI technologies as:

“software (and possibly also hardware) systems designed by humans that, given a complex goal, act in the physical or digital dimension by perceiving their environment through data acquisition, interpreting the collected structured or unstructured data, reasoning on the knowledge, or processing the information derived from this data and deciding the best action(s) to take to achieve the given goal” (cited in Bohland, Rechkemmer & Rogers, 2024).

The OECD definition of an AI system is set out in its Recommendation on Artificial Intelligence and is in broad alignment with the EU, Japan and other OECD jurisdictions. According to the 2024 (updated) OECD Council definition:

“An AI system is a machine-based system that, for explicit or implicit objectives, infers, from the input it receives, how to generate outputs such as predictions, content, recommendations, or decisions that can influence physical or virtual environments. Different AI systems vary in their levels of autonomy and adaptiveness after deployment” (UNESCO & OECD, 2024).

The UNESCO Recommendation on the Ethics of AI notes that AI systems are designed to operate with varying degrees of autonomy (UNESCO & OECD, 2024). Levels of AI autonomy include:

  1. ‘Human-out-of-the-loop’ (machines work autonomously and make decisions, with humans setting objectives and constraints)
  2. ‘Human-over-the-loop’ (humans supervise and can intervene in decision-making)
  3. ‘Human-in-the-loop’ (humans are integrated into AI systems) (UNESCO & OECD, 2024).

AI is also distinguished between ‘narrow AI’ and ‘general AI’, with most systems today categorised as ‘narrow’. These are systems that perform specific tasks or operate within designated domains. For example, narrow AI excels in tasks like natural language processing (NLP) for interpreting text, detecting and classifying objects and speech recognition. By contrast, general AI can operate across a range of tasks. Machine learning is often the most utilised form of AI, and includes approaches such as unsupervised and supervised, reinforcement and deep learning (UNESCO & OECD, 2024).

The following diagram is designed to illustrate existing and emerging fields of AI, taken from the UK Government’s AI Playbook (2025):

A conceptual diagram illustrating the relationships between machine learning, deep learning, neural networks, generative AI, large language models, natural language processing, ethics and societal impact, and related subfields and technologies.

Description generated by AI
  1. Existing and emerging fields of AI and interdependencies (Source: UK Government, 2025)

AI in policymaking

The OECD (2025a) conducted research in 11 core functions of government across 200 AI use cases. The findings were that AI is most frequently used in internal operations and public service delivery, with lower use in government oversight and policymaking. Another OECD report (2025b) found that only a small number of EU Member States had established AI tools to support regulatory compliance and policymaking processes (OECD 2025b), with most of these initiatives at the early stages of development. The EU Coordinated Plan on AI 2 identifies AI as a crucial technology for enhancing public services, improving citizen-government interactions, enabling smarter analytical capabilities and increasing efficiency across the public sector, alongside support for democratic processes. The OECD cautions that governments’ failure to leverage AI signifies a missed opportunity, creating a widening gap between the public and private sectors (OECD, 2025a).

Potential opportunities of AI in policymaking

AI has the potential to help governments increase their productivity, resulting in more efficient operations. An obvious example is that AI could automate repetitive tasks (OECD, 2024a). It could also support in designing and delivering public policies and services that are more inclusive and responsive to the needs of citizens and specific groups; for example, AI could mine large datasets to gain more granular insights into user needs or anticipate trends. AI could strengthen the accountability of governments by enhancing their capacity for oversight, for example, by identifying risks and detecting fraud (OECD, 2024a).

Arora et al. (2024) suggest that AI has the potential to transform governance by enabling policymakers to analyse large volumes of data, identifying trends and gaining insights. The authors suggest that AI can streamline administrative procedures and enhance governance systems. It could also automate routine tasks, analyse complex datasets and provide real-time information, enabling quicker policy decisions. AI could process information from diverse sources, including public opinion, social media and citizen feedback, which could enable the development of policy that is more aligned with societal expectations. It could also be used for conducting complex policy simulations and scenario modelling, assessing the possible impact of different policy options (Arora et al., 2024).

AI has the potential to be useful throughout the policy cycle, as follows (see also Figure 2):

  1. Agenda setting: By analysing large datasets, AI could identify critical challenges and/or priorities (APEC, 2022). Governments could monitor emerging topics, detect social problems and initiate faster policy responses (UNESCO & OECD, 2024).
  2. Policy formulation: AI could provide evidence-based insights, estimate the impact of policies, and/or analyse the costs and benefits of policy options (APEC, 2022). AI could support consultation and engagement processes, analysing stakeholders’ perspectives (UNESCO & OECD, 2024).
  3. Decision-making: AI could improve processes within legislative bodies (APEC, 2022).
  4. Implementation: AI could lead to increases in the quality, speed and efficiency of policy implementation (APEC, 2022), as well as ongoing policy improvements (UNESCO & OECD, 2024).
  5. Evaluation: AI could provide faster and more accurate data on the impact of policies (APEC, 2022), with better insights and quicker policy adjustments when needed (UNESCO & OECD, 2024).


A circular infographic illustrating the policy cycle with three main stages labeled agenda setting and policy formulation, policy implementation, and policy research and evaluation, each accompanied by key benefits such as accuracy, legitimacy, accountability, cost savings, and improved decision making.

Description generated by AI
  1. Potential benefits of AI at each stage of the policy cycle
    (Pencheva, Esteve & Mikhaylov, 2020, reproduced in UNESCO & OECD, 2024).

According to the OECD, AI can help governments in three key opportunity areas: productivity, responsiveness and accountability. At each stage of the policy cycle, AI can bring complementary benefits. These include (see also Figure 3):

  • Automated, streamlined and tailored processes and services
  • Better decision-making, sense-making and forecasting
  • Enhanced accountability and anomaly detection
  • Opportunities for external stakeholders through AI as a public good for all (OECD, 2025a).
  1. AI at each stage of the policy cycle. Source: Based on Pencheva, Esteve and Mikhaylov, 2018;
    adapted to align with OECD terminology (OECD, 2025a)

Craglia, Hradec and Troussard (2020) discuss the opportunities and challenges of using big data and AI to modernise the entire policy cycle. They point to the increasing gap between the speed of the policy cycle and that of technological and social change. According to the authors, it should be possible to evaluate policy more effectively, based on regular and short feedback loops. It should also be feasible to measure policy impact on specific groups of people, based on individual characteristics like age, gender, residence and socioeconomic profile (Craglia et al., 2020).

Potential of AI in policy evaluation

The OECD (2025a) looks at the role AI can play in supporting policy evaluation, where use has been limited and slower than in other functions of government. The OECD describes the current state-of-play:

  • Evaluation design and implementation. AI can support the synthesis of existing evidence or provide summaries of previous evaluations. AI has the potential to enhance causal inference in policy evaluation or simulate various impact scenarios.
  • Management and communication. AI can help with administrative processes, as well as drafting, translating and communicating policy evaluation results. It can help with searching across large volumes of reports.
  • Risks and challenges. There is little research on the risks and challenges of using AI in policy evaluation. Associated risks include inadequate or skewed data; automation bias; lack of explainability and transparency. Some of these potential risks are significant given their possible impact; these include the use of poor data, so it is essential to ensure that data is of good quality and is representative (for example, across all relevant societal groups).
  • Evidence of impact. The impact of AI on actual practice of policy evaluation is still modest and difficult to measure (2025a).

The OECD cautions that many people perceive AI systems and their decisions as neutral and impartial, even though there is a risk of inaccuracies. Human operators may over-rely on AI (known as ‘automation bias’), reducing the role of human judgement and oversimplifying complexity. The lack of transparency of AI can hinder policymakers in understanding and justifying AI-driven insights. Given these shortcomings, the implementation of AI requires sound digital and numeracy skills (OECD, 2025a).

Over the longer term, AI has the potential to change the approach to policymaking, allowing policy evaluations to feed into decision-making at multiple stages. This may make it possible to shift to an approach where evaluative evidence is available to shape policies almost in real-time, creating a ‘Dynamic Public Policy Cycle’. To enable this to happen, governments need to invest in skills and develop a strong data infrastructure (OECD, 2025a).

Potential of AI in foresight activities

A 2025 paper looks at foresight, an activity that requires intuitive perception, creativity and imagination, where AI’s capabilities can often be limited. Forward-looking innovation is not always identifiable from training data, which is usually based on the past. The paper’s conclusion is that AI can support foresight exercises but a hybrid approach, combining AI and human expertise, is optimal (Petrakis, Vasilis & Kanzola, 2025).

A white paper by the World Economic Forum/OECD (2025) evaluates AI’s potential impact in conducting strategic foresight, based on a survey of 167 foresight experts from 55 countries. The results show that many foresight practitioners now use AI in their work. Survey respondents report that AI is useful for trend analysis, scenario development and the identification of emerging themes and issues. AI can save time, particularly on repetitive and labour-intensive tasks, as well as for processing large datasets to uncover trends and insights. However, the survey reveals diverging opinions about AI’s usefulness, accessibility and reliability. Many respondents express their concerns about the quality and trustworthiness of AI-generated content, that it is prone to hallucinations, operates without transparency and can produce biased results. AI also has limited capacity for inductive reasoning, as the technology relies on existing knowledge. The survey highlights differences in how AI is regarded by practitioners in the public and private sectors, academia and civil society. Recommendations include increasing AI literacy, as well as encouraging experimentation. In conclusion, AI tools can help humans with certain tasks, freeing up time for experts’ higher-level analysis, interpretation and critical thinking (WEF/OECD, 2025).

Challenges and risks of AI in policymaking

Alongside the potential benefits of AI, there are concerns about the risks of any AI deployment that is fragmented and ungoverned. These risks include bias, lack of transparency in AI system design, breaches in data privacy and security (OECD, 2024a). When using AI, it is vital to uphold human and civil rights, protect personal privacy, ensure algorithmic transparency and accountability, promote ‘explainability’ and avoid unfair and biased policy outcomes (OECD, 2024a). Risks include potential harms of misuse, poor design, bias and discrimination, lack of transparency, invasion of privacy, infringements of intellectual property rights (IPR) and poor-quality outcomes (UNESCO & OECD, 2024). Responsible and transparent use of AI is key to maintaining public trust (Arora et al., 2024).

Approaches to the implementation of AI

The OECD’s review of AI use cases in government indicates a high presence of early-stage initiatives, such as experiments and pilots. The OECD suggests this may indicate the following challenges for governments:

  • Difficulties in transitioning from experimentation to implementation
  • Skills gaps
  • Problems in obtaining and/or sharing quality data
  • A lack of actionable frameworks and guidance on AI usage, including for specific policy areas
  • Risk aversion
  • A need to demonstrate results and return on investment (ROI)
  • Inflexible or outdated legal and regulatory environments
  • High or uncertain costs of AI adoption and scaling
  • Outdated legacy IT systems (OECD, 2025a).

A brief by the Asia-Pacific Economic Cooperation forum states that AI should not be seen as a ‘silver bullet’ but rather as a tool for improving human and social welfare. Its use should therefore remain ‘human-centric’ (APEC, 2022). The brief poses three fundamental questions over the use of AI in policymaking:

  1. Is it appropriate to use AI in a particular situation, and what are its limitations?
  2. Who develops the AI, and what biases might influence algorithms and models?
  3. What data is provided and what is its quality? (APEC, 2022).

A further consideration is that proprietary AI may lack transparency, with limitations in the ability to review its performance or reproduce results (APEC, 2022).

The G7 toolkit (UNESCO & OECD, 2024) sets down five steps when considering the use of AI:

  1. Framing the problem to be addressed and evaluating the available data.
  2. Prototyping to enable early testing and refinement, ensuring solutions are viable and meet user needs, as well as addressing risks such as privacy, intellectual property rights (IPR) and cybersecurity.
  3. Piloting and scaling under a controlled, real-world environment, whilst allowing for evaluation of the system.
  4. Ensuring transparency and explainability, clarifying who is responsible for the AI system.
  5. Monitoring performance, accuracy and impact, as well as compliance with regulations (UNESCO & OECD, 2024).

In a report produced by the Bennett Institute for Public Policy at the University of Cambridge (Turobov, 2025), guidance is given on how to use LLMs effectively in the UK civil service. According to the report, successful LLM implementation rests on three core principles:

  1. Evidence, which should underpin implementation based on an assessment of capabilities, measurement of outcomes and regular evaluation
  2. Transparency, through clear documentation and traceable processes
  3. Continuous learning, through shared knowledge and experience (Turobov, 2025).

Different government functions may require contrasting approaches to LLM implementation. LLMs make probability calculations, meaning their outputs represent statistical predictions and not logical deductions; results can vary even with identical inputs, and confidence levels in outputs differ across tasks and contexts. Understanding this probabilistic feature is essential to the selection of appropriate tasks, assessing risk, establishing quality control and output validation. LLMs also work with fixed knowledge, cutoff dates and training parameters that define their boundaries of expertise; awareness of this limitation is crucial in government contexts, where current and accurate information is vital for decision-making. There are also resource considerations, requiring an assessment of the cost-benefit for each task. LLM implementation should involve:

  1. Development of a pilot programme
  2. Integration into the policy framework
  3. Risk mitigation strategies
  4. Progressive scaling (Turobov, 2025).

A structured approach would start by identifying a specific challenge that LLMs might help to address. It then involves an assessment of available resources and capabilities, potential risks and mitigation, and ongoing management. A comprehensive training programme would involve technical skills for interaction with LLMs, competencies in output evaluation, risk management, and best practice application. It is important to document and share best practices, such as prompt strategies, quality control methods, integration into workflows and problem-solving (Turobov, 2025).

A manual by Eurostat looks at LLMs and their relevance for statistical offices (European Commission, Eurostat, 2024), based on a literature review and best practices. It sets out a comprehensive list of criteria when considering the application of LLMs:

1. Strategic Alignment: Does the application align with the organisation’s current objectives and long-term vision?

2. Impact Potential: What is the potential of the application to enhance data quality, insights, or stakeholder engagement?

3. Resource Requirements: How intensive are the resource demands of the application in terms of finances, manpower, and technology?

4. Risk Profile: What are the potential risks associated with the application, including data security, ethical concerns, and error potential?

5. Stakeholder Value: How valuable is the application to external stakeholders, such as policymakers, the public, or partner organisations?

6. Operational Efficiency: Can the application streamline workflows, reduce manual labour, or improve overall operational efficiency?

7. Scalability: Does the application have the potential to handle increasing data volumes or expand in scope as the organisation grows?

8. Innovation Potential: Does the application offer a novel approach or solution that sets the organisation apart from peers?

9. Implementation Timeframe: How long will it take to fully integrate the application and start seeing tangible results?

10. Cost Efficiency: Beyond initial costs, what are the long-term savings or cost benefits associated with the application?

11. Ethical and Responsible Use: Does the application adhere to ethical standards, and does it promote responsible data handling and usage?

12. Feedback and Evaluation: Is there a mechanism in place to gather feedback and continuously evaluate the application’s effectiveness?

(Source: European Commission, 2024).

The AI Playbook for the UK Government (UK Government, 2025) defines 10 principles to guide the use of AI in government and public sector organisations:

Principle 1: You know what AI is and what its limitations are

Principle 2: You use AI lawfully, ethically and responsibly

Principle 3: You know how to use AI securely

Principle 4: You have meaningful human control at the right stage

Principle 5: You understand how to manage the AI life cycle

Principle 6: You use the right tool for the job

Principle 7: You are open and collaborative

Principle 8: You work with commercial colleagues from the start

Principle 9: You have the skills and expertise needed to implement and use AI

Principle 10: You use these principles alongside your organisation’s policies
and have the right assurance in place

(Source: UK Government, 2025).

The playbook (UK Government, 2025) emphasises that use cases should be led by organisational and user needs, and not what the technology can do. Examples might be problems or difficulties that can be solved by AI, or where AI offers significant advantages over existing techniques. Evaluating AI’s potential impact could involve cost-benefit analysis to consider where there could be efficiency, increased accuracy or cost reduction. It is also important to assess whether the necessary skills and infrastructure are in place.

Risk assessment

According to the OECD (2025a), there is no such thing as risk-free AI adoption and risks need to be mitigated. Risks include biased algorithms, misuse of AI, lack of transparency, explainability and public understanding, as well as overreliance on AI. Governments must manage ethical risks, operational risks, exclusion risks, public resistance risks and also risks of inaction. Risk assessments are important; there should be safeguards and human oversight, particularly in situations where decisions have a significant policy and societal impact (APEC, 2022). Policymakers should be transparent about their use of AI, its role in decision-making, and whether decisions can be reversed (APEC, 2022). According to the OECD, it requires a ‘whole-of-government’ approach, the involvement of stakeholders (including citizens) in AI’s design and deployment, robust data governance and accountability, along with the means to scrutinise AI systems and models (OECD, 2024a).

Risk assessment should consider the task’s complexity and potential impact; different tasks carry different levels of risk. For example:

  • Level 1 could involve foundation tasks and/or basic information processing with minimal risk, such as document summarisation. Control would require basic review and verification.
  • Level 2 could involve analytical support, including pattern identification, trend analysis and policy research. Control would require regular validation and peer review.
  • Level 3 could involve policy development and would require closer oversight. Control would require a more rigorous review process at each stage.
  • Level 4 could involve critical decisions, for example, policy implementation guidance. Control would involve comprehensive oversight, with extensive safeguards
    (Turobov, 2025).

Some specific situations may exclude LLM use completely, for example, possible critical decision-making contexts such as crisis briefings, emergency response, or situations directly affecting individual rights. Human judgement and direct accountability are needed when stakes are high, time-critical, or outcomes directly impact public safety or individual rights (Turobov, 2025). Examples of use cases that should be avoided are those that involve decisions that are significant and impactful, or where there is a risk to health, safety, fundamental rights or the environment (UK Government, 2025).

Cost-benefit assessment

The Eurostat manual recommends a cost-benefit consideration of implementing LLMs. Costs can include acquiring and training an LLM, its maintenance, establishing the necessary infrastructure such as hardware and cloud services, staff training, oversight and governance mechanisms. Consideration of potential benefits includes the efficiency of data processing at speed and scale, improved public interaction, multilingual capabilities, and greater consistency in data analysis and reporting (European Commission, Eurostat, 2024).

Prompt engineering

Understanding prompt engineering allows civil servants to achieve more consistent, reliable results. Developing effective prompts involves:

  1. Purpose: Identifying specific task requirements, clarifying desired outcomes, considering governance constraints
  2. Structure: Matching prompt type to task, considering complexity level, ensuring compliance with protocols and requirements
  3. Testing and refinement: Validating outputs, checking for compliance, then further iteration based on results (Turobov, 2025).

Validation and documentation of outputs

Effective validation of AI-supported outputs involves:

  1. A quality assessment for accuracy and completeness, alignment with policy, practical applicability, and robustness of the evidence base
  2. Expert validation of technical accuracy, policy implications, implementation feasibility, and risk assessment (Turobov, 2025).

LLM-generated content should be identified, with metadata and clear attribution statements. Documentation should be maintained to ensure transparency and ensure continuous improvement (Turobov, 2025).

Competences

The JRC has developed a competency framework to guide civil servants in adopting and managing AI, based on literature reviews, expert workshops, interviews and case studies. The framework classifies competences into three main dimensions:

  1. Technical competences, such as knowledge and skills related to data management, machine learning and system implementation
  2. Managerial competences around areas such as project ownership and leadership, knowledge brokering and decision-making on AI-related initiatives
  3. Policy, legal and ethical competences covering areas like awareness of ethical issues and sustainability, auditing and compliance, consultation with experts on ethical matters (Medaglia, Mikalef & Tangi, 2024).

Within each dimension, the report identifies several cross-cutting competences:

  1. Attitudinal competences, which are defined by mindsets or attitudes that contribute towards the effective use of AI
  2. Operational competences, which refer to practical application of the knowledge required to use AI
  3. Literacy competences, which relate to the knowledge required to use and work with AI (Medaglia et al., 2024)

Policies, legislation and ethical frameworks

The EU AI Act3 is a European regulation approved by the European Parliament in 2024. The regulation establishes obligations, based on potential risks and level of impact. The Act identifies different levels of risks for governments’ use of AI. Four risk levels are defined:

  1. Unacceptable risk, prohibiting AI use
  2. High risk, where AI use is regulated
  3. Limited risk, where developers and deployers must ensure that end-users are aware that they are interacting with AI
  4. Minimal risk, where there is no regulation, but a code of conduct is suggested.

Implementation of the AI Act is supervised by bodies such as the European AI Office4, European Artificial Intelligence Board5 and the European Centre for Algorithmic Trust (ECAT)6.

Policymakers are employing both ‘soft’ (standards and principles) and ‘hard’ law (legislation, regulations and directives) to ensure the design and deployment of AI are responsible, ethical and trustworthy. International standards and principles include the following:

The EC’s Ethics Guidelines for Trustworthy AI were developed by the High-Level Expert Group on Artificial Intelligence and are based on three key components:

  1. Lawfulness: AI should comply with all applicable laws and regulations.
  2. Ethical Alignment: AI should adhere to ethical principles and values.
  3. Robustness: AI should be robust from both a technical and social perspective, minimising unintentional harm (cited in European Commission, Eurostat, 2024).

The EC’s Assessment List for Trustworthy Artificial Intelligence (ALTAI) for Self-Assessment was also developed by the High-Level Expert Group on AI and comprises the following principles:

  • Human Agency and Oversight: AI systems should support and respect human decision-making and autonomy, enabling an equitable society and maintaining fundamental rights, underpinned by adequate human oversight.
  • Technical Robustness and Safety: AI systems developers should take a preventative approach to risks, ensuring they behave reliably and as intended, minimising and preventing unintentional harm, and maintaining resilience in changing environments or against adversarial interactions.
  • Privacy and Data Governance: Privacy should be upheld as a fundamental right by means of adequate data governance, ensuring the quality, integrity and appropriate data processing in a way that safeguards privacy and aligns with the deployment domain.
  • Transparency: There should be traceability of decisions and processes, explainability of the system’s functioning and decisions, and open communication about the system’s limitations,
  • Diversity, Non-discrimination, and Fairness: Inclusion and diversity should be promoted throughout the AI system’s lifecycle, ensuring the system is free from biases and discrimination, designed to be user-centric and accessible to all individuals.
  • Societal and Environmental Wellbeing: The broader societal and environmental impacts of AI systems should be considered, addressing sustainability and global concerns while carefully monitoring effects on social relationships and individual well-being.
  • Accountability: Mechanisms should ensure responsibility and accountability in the development, deployment and use of AI systems, incorporating transparent risk management and provision for third-party auditing to address any adverse impacts (cited in European Commission, Eurostat, 2024).

The OECD Recommendation on Artificial Intelligence (the “OECD AI Principles”) was adopted in 2019 and included the first intergovernmental set of principles on AI. It was further updated in 2024 to take account of policy and technology developments (OECD, 2024b). It includes five values-based principles and five recommendations for policymakers (see Figure 4):

  1. OECD AI principles overview (Source: OECD.AI
    Principles overview, https://oecd.ai/en/ai-principles)

In a report, the OECD (2025) sets out its framework for trustworthy AI in government (see Figure 5
and key below):

A circular infographic illustrating trustworthy AI in government with sections on engagement, enablers, guardrails, and their related components such as governance, ethics, monitoring, and partnerships.

Description generated by AI


  1. OECD framework for trustworthy AI in government (OECD, 2025a)
  • “Enablers” include establishing key governance mechanisms and processes, understanding data’s role as the foundation for AI, building digital infrastructure, fostering skills and talent, investing purposefully, using public procurement effectively and expanding AI’s potential through partnerships.
  • “Guardrails” can be binding and non-binding policy levers, transparency processes and accountability mechanisms.
  • “Engagement” with stakeholders can take the form of citizen assemblies, engaging with civil servants, involving users in AI development and collaborating across borders (OECD, 2025a).

Adopted in 2021 by UNESCO, the UNESCO Recommendation on the Ethics of AI was the first global standard on AI ethics (UNESCO, 2021). It aims to ensure AI systems are designed, developed and used in ways that respect human rights and freedoms, foster just and interconnected societies, ensure diversity and inclusiveness and promote sustainable development. The Recommendation lays out 11 key areas for policy action required by Member States. It is complemented by a Readiness Assessment Methodology (RAM) to help countries assess their readiness to develop and deploy AI technologies responsibly and effectively (UNESCO, 2023, also cited in UNESCO & OECD, 2024). The UNESCO Ethical Impact Assessment (EIA) (UNESCO, 2023, also cited in UNESCO & OECD, 2024) is a tool designed to identify and assess AI systems’ benefits, concerns, risks and appropriate measures for the prevention, mitigation, remediation and monitoring of identified risks.

Developed by the G7, the International Code of Conduct for Organizations Developing Advanced AI Systems recommends that organisations take a risk-based approach (Ministry of Foreign Affairs of Japan, 2023).

The Council of Europe’s Framework Convention on Artificial Intelligence was the first legally binding treaty in the field, which “aims to ensure that activities within the lifecycle of artificial intelligence systems are fully consistent with human rights, democracy and the rule of law, while being conducive to technological progress and innovation” (Council of Europe, 2024). In 2024, the UN General Assembly adopted a resolution on the promotion of “safe, secure and trustworthy AI systems that will also benefit sustainable development for all” 7.

‘HHH’ is an example of a framework to assess how Helpful, Honest and Harmless a model is. Helpful AI is trained with users’ needs and values in mind. Honest AI provides accurate information, expresses uncertainty and the reason behind it, and is developed transparently. Harmless AI does not comply when prompted to perform a dangerous task; it is trained within frameworks that mitigate bias (European Commission, Eurostat, 2024).

Transparency, fairness and inclusion

Transparency of AI systems includes:

  • Technical transparency: Information about the technical operation of the system, such as the code and underlying datasets used for training
  • Process transparency: Information about the design, development and deployment decisions made
  • Outcome-based transparency: Clarifying to any user how the AI works and what influences its decision-making and outputs
  • Internal transparency: Retaining up-to-date records on technology and processes
  • Public transparency: Communication about the use of AI systems, made in an open and accessible way (UK Government, 2025).

Fairness is about ensuring that a system’s outputs are unprejudiced and do not amplify existing social, demographic or cultural disparities. This means ensuring that AI systems allocate resources and services fairly to all people, and that certain subgroups are not disproportionately adversely impacted or harmed (UK Government, 2025).

In a recent paper, Radeljić (2025) contends that an unquestioning adoption of AI in policymaking, along with a misplaced belief in its neutrality, undermines democratic governance. He sees AI as a sociotechnical artifact that is deeply embedded in ideological and geopolitical dynamics. Viewed through this lens, AI systems are regarded as far from neutral, but instead reflect and encode the values, assumptions and strategic interests of AI’s creators, whether large corporations or state actors. The paper puts forward the concept of a new “public intellectual” as a way to reimagine the relationship between technology, politics and public reason. This calls for an AI governance model that is more reflexive, participatory and accountable, in which the voices of the marginalised can be heard (Radeljić, 2025).

Copyright issues

A recent report by the European Intellectual Property Office (EIPO, 2025) is based on desk research and stakeholder interviews. In the EU, two legal instruments are especially relevant from a copyright perspective. One is the Copyright in the Single Market Directive (CDSM)8, which creates a legal framework for text and data mining (TDM). In the context of AI, this often involves the rights of copyright and database owners regarding training content. The CDSM allows for TDM by scientific research organisations (Article 3), whereas Article 4 allows TDM by any user, including commercial, subject to the rightsholders’ abilities to opt-out of the TDM exception. To use content for AI training where an opt-out reservation has been placed, authorisation is needed from the rights-holder, for example, through licences (EIPO, 2025).

As stated, the EU AI Act sets out a regulatory framework for AI in the EU, with specific obligations on the providers of general-purpose AI (GPAI) models. The AI Act addresses issues such as risk management, transparency, data governance, ethical considerations and compliance with fundamental rights (EIPO, 2025). GPAI system providers must publish sufficiently detailed summaries of the training data they use, so that copyright holders can enforce their rights, if needed. System deployers must also ensure that generative output is detectable in a machine-readable format (EIPO, 2025).

There are a rising number of legal disputes between rightsholders and system providers, particularly in the US. Several agreements on the use of copyrighted material have been reached in the context of AI, and direct licensing has the potential to bring in new revenue streams, including content for Retrieval Augmented Generation (RAG). At the same time, there are concerns about scientific research privileges (Article 3) being exploited for commercial purposes (EIPO, 2025). Datasets that are publicly available for training may include pirated content or display incorrect licence information. All of this may result in copyright liability being passed down the chain (EIPO, 2025). Some AI models may have ‘memorisation’, where outputs resemble or even replicate training outputs; this creates a legal issue over plagiarism and content ‘regurgitation’, where trained content is explicitly reproduced. There are measures to address memorisation, including tools that compare generated content with input sources, filters that prevent duplicative output, and prompt rewriting or filtering, along with ‘model unlearning’ and ‘model editing’, which enable AI developers to solve issues detected after the model’s deployment. Some system providers offer some form of legal indemnification (EIPO, 2025).

Accountability and responsibility

Accountability and responsibility mean that individuals and organisations can be held responsible for the effects of the AI systems they develop, deploy or use. There can be a chain of human responsibility across an AI project life cycle. In cases of harm or errors caused by AI, there would need to be recourse and feedback mechanisms for affected people. This would mean identifying the specific people involved in AI systems, which could include developers, policymakers, regulators, system operators and end-users. In each case, it is important to define roles and responsibilities, and align these with legal and ethical standards. Auditability means demonstrating adequate responsibility and trustworthiness by means of robust reporting and documentation protocols and traceability throughout an AI project’s lifecycle. Liability is about making sure parties involved are acting lawfully. An end user must take responsibility for an AI system’s outputs and potential consequences – outputs should be accurate, non-discriminatory and not violate the law. There must be oversight in place in situations with significant impact or high risk. Ultimately, responsibility for an output or decision made with AI rests with the public organisation. Contestability and redress are about how outputs and decisions can be challenged, and how individuals who are impacted can seek remedy. Societal wellbeing and public good means ensuring that AI is developed in a way that minimises and mitigates harms, as well as delivering good for society (UK Government, 2025).

The European Commission’s Internal Guidelines on Generative AI state the following:

  1. Staff must never share any information that is not already in the public domain, nor personal data, with an online available generative AI model.
  2. Staff should always critically assess any response produced by an online available generative AI model for potential biases and factually inaccurate information.
  3. Staff should always critically assess whether the outputs of an online available generative AI model are not violating intellectual property rights, in particular copyright of third parties.
  4. Staff shall never directly replicate the output of a generative AI model in public documents, such as the creation of Commission texts, notably legally binding ones.
  5. Staff should never rely on online available generative AI models for critical and time-sensitive processes (cited in European Commission, 2024).

Examples of AI for policymaking in Europe

A chapter in the Research Handbook on Public Management and Artificial Intelligence looks at the JRC’s role in collecting information on the use of AI in the public sector, based on input from the platforms AI Watch9 and Innovative Public Services10 (Tangi, Ulrich, Schade & Manzoni, 2024). It reports a growing number of cases where AI is being used, particularly machine learning and NLP. Most of the examples are internally oriented and aimed at improving efficiency within the public organisations themselves (Tangi et al., 2024). The EU has also invested in R&D projects around AI for policymaking, such as the AI4PublicPolicy project11 (Papadakis et al., 2024).

A more recent report (European Commission, 2025) provides a broad description of the adoption of generative AI in the European public sector, based on data from the Public Sector Tech Watch (PSTW) observatory and evaluating guidelines and procedures established by EU Member States, together with interviews with five public administration managers. The report states that the public sector is adopting GenAI solutions, particularly in general public services. One emerging trend is for the development of national language models. However, many projects are still in the planning, development or piloting stages. There are challenges related to implementation processes and effective public-private collaboration. There are other challenges around human oversight; accountability; the importance of data protection; and governance, safety, fairness and transparency (European Commission, 2025).

A JRC team describes the AI tool called GPT@JRC (De Longueville et al., 2025). Launched as a pilot in 2023, it leverages the JRC’s expertise with LLMs and includes Retrieval Augmented Generation (RAG)12
that enables the system to draw on relevant information from a knowledge base, which can include sources like scientific papers. In the article, it was reported that the JRC is developing an agency component, which would enable the system to make use of specialised LLM tools and other software components. The GPT@JRC pilot was extended and rolled out to other EC DGs and EU institutions, and by November 2024 had more than 12,000 users. In addition to individual user access, there is also API13
access by project teams, who can request it for more advanced uses. Use cases are classified according to their technical complexity, the AI expertise of the requesting team and the resource investment required. Level 1 ranges from more generic tasks (like text enhancement) to more specific ones (like coding assistance, data analysis and literature review). The article suggests that efficiency gains and quality improvements vary, depending on the nature of the tasks and user understanding of AI. At the basic Level 1, the breadth of application has been wide but there have been limitations, such as an in-depth understanding of specific concepts and issues. Level 2 includes more skilled users, such as data scientists and IT engineers. At this level, the tool is used for tasks like data processing and exploratory research. Although the uses are fewer, quality and impact tend to be higher. It suggests that certain processes can be faster, cheaper and more accurate by combining AI with human expertise, for example, for public consultation analysis and systematic literature reviews. Level 3 involves the development of specialised GenAI systems that can tackle complex tasks. Level 3 use cases envision a JRC “virtual scientific assistant” that can perform systematic literature reviews, produce digests of research on a given topic, or engage in brainstorming on research topics. See Figure 6 for a mapping of use cases, based on levels of user and technological capabilities. The article suggests that measuring impact is required via a mix of quantitative and qualitative indicators. These could include efficiency through time saving and/or productivity gain, improved output quality and/or higher stakeholder satisfaction. The JRC’s findings so far are that LLMs are not able to replicate the cognitive abilities of humans, especially in complex domains. The aim going forward is to produce agentic systems, in which LLMs have access to the right knowledge and can choose which computer programmes to utilise for which purpose. The article concludes that each organisation needs its own use cases and robust experimentation within specific contexts (De Longueville et al., 2025).

A two-axis chart maps AI applications by user AI literacy from low to high on the horizontal axis and AI-IQ from low to high on the vertical axis, showing various AI uses such as administrative chatbots, document summarization, external communication assistants, and science-for-policy advice.

Description generated by AI
  1. Mapping selected GPT@JRC use cases according to the JRC GenAI Compass
    with people and technology as key dimensions (Source: De Longueville et al., 2025)

The potential use of AI in scientific advice

This section starts by looking at the potential of AI in science, with a focus on two recent reports. It then examines information on AI in scientific advice.

AI’s impact on science

There have been a number of reports on how AI is impacting on science, and here we only look at two major reports published in 2025. A report from the International Science Council (2025) looks at how AI is reshaping scientific enterprise, not only scientific discovery but also scientific reasoning, influencing what counts as evidence, how explanations are formed and who participates in the construction of knowledge. The ISC describes a range of forms of AI and how they impact science:

  • Descriptive AI extracts and structures data patterns, helping researchers to explore and understand complex datasets.
  • Predictive AI anticipates outcomes, based on historical trends.
  • Generative AI produces novel content and scientific hypotheses. It can simulate experiments, model complex systems and help researchers explore ideas and research directions.
  • Optimisation AI identifies efficient solutions within complex constraints.
  • Prescriptive AI recommends concrete courses of action.
  • Privacy-aware AI enables analysis of sensitive information. Federated learning is a way to train AI models collaboratively without sharing sensitive data.
  • Causal AI seeks to uncover underlying cause-effect relationships.
  • Explainable AI enhances transparency of models.
  • Reinforcement learning enables an AI model to learn through interaction with its environment, for example, automating experimental design and control protocols.
  • Meta-scientific AI supports the scientific process by generating hypotheses, designing experiments, analysing data or connecting insights across different fields.
  • Agentic AI can operate with a level of autonomy, by planning and taking actions towards scientific goals (ISC, 2025).


The ISC report suggests that the convergence of meta-scientific and agentic capabilities means that AI can act as a collaborator with researchers. At the same time, greater AI autonomy would call for increased transparency, reproducibility and accountability. This would require ethical foresight, with the design of AI systems that allow for human interpretability, can support inclusive and representative datasets and provide governance frameworks that audit, validate and contest AI-generated knowledge. The impact of AI should be to “deepen science trustworthiness, widen its accessibility and extend its capacity for epistemic inclusiveness” (ISC, 2025).

A recent report by the JRC (Purificato et al., 2025) also asserts that AI is reshaping science, affecting how knowledge is generated, experiments are designed and results are shared. Like the ISC, the JRC report states that AI has transformative potential across every stage of the scientific process, including: (1) asking questions and formulating hypotheses; (2) designing and conducting experiments; (3) collecting and analysing data; (4) interpreting results and drawing conclusions; (5) publishing and communicating findings. At the same time, the impact of AI depends on how it is deployed and governed. Recommendations made by the JRC include: (1) support for open science principles to foster innovation and ensure reproducibility and trustworthiness of AI-driven research; (2) a new skillset for researchers, with hybrid roles in teams that combine domain knowledge with knowledge of AI and data science methods. The report also flags the risk of epistemic drift in terms of what scientific knowledge is and how it is produced. For example, AI may reinforce existing research paradigms or create an environment where knowledge production becomes detached from human oversight. AI may fabricate information, distorting scientific understanding. All these issues require the promotion of AI literacy, critical thinking and multidisciplinary collaboration (Purificato et al., 2025).

The potential for AI and science advice

Namdarian (2025) bases her article on a systematic literature review and data analysis. She suggests that AI has the potential to completely change our understanding of evidence and its role in policy formulation. This may require a thorough re-evaluation of the nature of evidence and evidence-based policymaking, as well as methodologies for evidence collection and analysis. She recommends ethical frameworks and guidelines for the integration of AI into evidence-based policymaking. Policymakers may need new competencies in data science, AI interpretation and digital literacy. There could be a need to build trust and demonstrate the reliability of AI-enabled decisions.

In a commentary, Cong-Lem (2025) considers the impact of AI on evidence-informed policy and practice in the field of education. He suggests that a new model would place critical thinking at the heart of evidence use, requiring not only cognitive skill but a professional stance that involves “interpretive judgement, epistemic reflexivity and ethical responsibility” and that “fosters context-sensitive, deliberative engagement with evidence” (Cong-Lem, 2025). Text generated by AI may appear authoritative but may lack provenance, can reflect algorithmic bias and may incorporate incorrect or outdated information. Users of AI need to be evaluators and interpreters of information. Such critical thinking involves questioning assumptions, interrogating sources, considering whose knowledge is being represented and why. It requires reviewing and revising prompts, cross-checking information and analysing the limitations of the AI system. At policy level, it means not only questioning AI-generated advice but also embedding safeguards into the decision-making process. This means that AI literacy must go beyond technical proficiency; verification and citation mechanisms are needed for AI-generated outputs (Cong-Lem, 2025).

Tyler et al. (2023), in a commentary in Nature, suggest that AI tools could increase the capacity of science advisers, as well as improving science advice to make it more agile, rigorous and targeted. However, it would require guidelines, as well as careful design and responsible use of AI. The authors present two examples to illustrate potential, one on synthesising evidence and the other on drafting briefing papers. They suggest machine learning can automate the early stages of a systematic review, such as search, screening and data extraction, including foreign language material. However, assessing data quality and drawing conclusions from the evidence requires human judgement. The paper also suggests that AI tools could be used for generating policy briefs and options. The authors highlight three main issues they think need to be addressed in this regard. Firstly, adopting more consistency in how research methods and results are presented would help AI tools make sense of research outputs. Secondly, experts need to agree on standards for research quality that can be used in AI tools. Thirdly, there is a need for open access and greater interoperability of databases. The authors flag the “black box” nature of AI, which is problematic for science advice. They suggest that science advisers should be involved in the selection of training data and processes. Algorithms should be fully transparent and explainable, to ensure accountability. Rigorous, unbiased systems are essential for evidence synthesis. Disinformation is a risk, requiring oversight and an understanding of training data and processes. A further issue is the protection of sensitive or classified information; this requires clear guidelines on what type of data can be used in external AI tools and where internal AI systems are needed. Science advisers will also need to be trained as responsible AI users. In a published response to the commentary, Canali and Barone-Adesi (2025) express their concern about the black-box nature of AI and its impact on the opacity of science advice processes. In their view, governments could claim to be “following AI” when decisions may prove unpopular.

Published case studies

In a scoping review, Meeske and others look at language technology and AI in policymaking on food security (Meeske, Cruijssen, van der Lee, & Krahmer, 2026). In the last decades, there has been a significant increase in the amount of unstructured data from sources such as social media, research publications and news articles. NLP is an opportunity to collect and structure this data, helping policymakers to navigate complex information landscapes and enabling more effective policymaking. Across 60 reviewed studies, six application areas were identified14. One finding was that there was uneven coverage of dimensions of food security. Another finding was around access to high-quality data, where there is variability between geographic regions, some of which may be overlooked. The majority of NLP for food security policymaking projects remain in research labs, rather than implemented in practice. Meeske et al. (2026) found that only 25% of the studies they reviewed explicitly discussed ethics, bias or fairness. Of those that did, the most frequently mentioned concern was the potential underrepresentation of marginalised populations in the data. Another was the limitation of data to English-language content, along with the dominance of the global NLP market by the United States and China. Of the reviewed studies, 77% of authors were affiliated with Western institutions. One study highlighted the environmental impact of running LLMs, as climate change affects global food security. Limitations of the scoping review included the cut-off date (May 2024), the limitation to English and the data types considered. Future recommendations include establishing stakeholder partnerships; enhancing data quality and fostering a “data culture”; piloting innovations, scaling up and ensuring sustainability of NLP projects. Fostering communities of practice, training programmes and knowledge networks are also seen as important, whilst being guided by ethical and responsible AI principles (Meeske et al., 2026).

Table 1: Automation potential of AI for the peer review process (Kankanhalli, 2024, p. 79).

A paper by Voutyrakou and Skordoulis (2025) looks into gender bias in AI-generated policymaking, based on a literature review and four experiments conducted across diverse policy-making contexts to assess whether AI-generated recommendations include or overlook gender. The findings were that unless prompts explicitly mention gender, AI systems tend to overlook gender-specific needs. Rather than being neutral, such tools replicate dominant social norms present in the training data.

An article in Nature (Ziegler, Lothian, O’Neill, Anderson, & Ota, 2025) looks into whether AI language models can help or hinder equity in marine policymaking, based on a case study of an AI chatbot. Although the case study suggested some promising opportunities, it also raises concerns that LLMs could deepen existing inequalities by maintaining or introducing biases that overrepresent Western views, for example.

Zhao et al. (2026) look into the issue of plastics pollution, and the need for an AI-driven science-policy interface that can translate scientific evidence into the knowledge required for policy. Current applications of LLMs in plastics’ Life Cycle Assessment (LCA) lack the framework required to systematically extract, reconcile, and validate the contextual elements required. A systematic understanding of the Lifecycle Environmental Impacts (LCEIs) of plastics needs to pull fragmented study-level findings into a structured, reconciled, and searchable evidence base. However, science-policy practice still lacks reproducible workflows for AI-assisted scientific information synthesis, and the study highlights the urgency and necessity of developing AI toolkits designed specifically for this. The authors propose an AI-driven science-policy interface framework for plastics pollution governance, although it could be applied to other environmental fields. Despite its capabilities, the framework also had limitations. For example, the accuracy of knowledge extraction is constrained by the completeness and clarity of the source literature. The performance of the LLM is also bounded by its knowledge and limited ability for domain-specific inference, when compared with expert interpretation. Instead, it does provide an evidence map that highlights trends, consensus and critical knowledge gaps. There are other challenges; AI trained on existing literature inherits biases and inconsistencies from the underlying body of literature and data. To ensure responsible and ethical integration of AI into the science-policy process, transparent model documentation, traceable data provenance and human oversight need to be embedded throughout the workflow. AI should serve as a collaborative analytical instrument, rather than replace human judgement (Zhao et al., 2026).

In a study by Salekpay et al. (2024), responses provided by ChatGPT-4 on climate policy were compared against those received from almost 800 experts worldwide. ChatGPT was presented with the same questions as the experts, as well as follow-up questions aimed at clarifying responses and identifying the sources. The answers received from the experts and from ChatGPT were compared quantitatively and qualitatively. The authors observed significant differences in the importance that ChatGPT and experts gave to the four criteria selected for evaluating climate policy instruments (effectiveness, efficiency, equity, and feasibility). ChatGPT deemed all four criteria as “extremely important” (5 out of 5), whereas there were more variations among the experts. Salekpay et al. (2024) also compared the evaluation of various policy instruments by ChatGPT and experts against the four criteria. Again, there were substantial differences between the evaluations, with ChatGPT rating some instruments higher and some lower than the experts. Finally, Salekpay et al. (2024) compared how ChatGPT and experts rated the importance of policy instruments. There was some agreement on the evaluations, but some major differences, too (for example, ChatGPT rated the importance of information provision as significantly lower than did the experts).

Another study compared briefing notes created by AI with those by a human (Safaei and Longo, 2024). In this study, three types of briefing notes were created: (1) artificial intelligence policy analyst (AIPA) notes, exclusively created by AI (2) policy analyst (PA) notes, created by a human using non-AI methods, and (3) intelligence augmented policy analysis (IAPA) notes, where a human edited AIPA outputs. The notes were produced on two topics: (1) the effectiveness of carbon tax in reducing greenhouse emissions and (2) prevention policies and practices regarding COVID-19. The algorithm15 was trained on briefing notes from the Government of British Columbia and Situation Reports from the World Health Organization (WHO)16. A single human – a student in a graduate programme in public policy – was involved in the creation of the briefing notes. The notes were then independently evaluated by retired senior public servants. Four were informed that some of the notes might have been developed using AI and were also asked to identify which they thought those were. The other four were not informed that AI had been used in any of the notes. A cut-off of 50% was used to determine acceptability of the briefing notes. Neither of the briefing notes created by AI (AIPA notes) approached this cut-off value: on average (across the eight experts), the carbon tax note was scored at 40% and the COVID-19 note at 34%. The notes created by AI and edited by a human (IAPA) scored somewhat better across the eight assessments, but the COVID-19 note, scored at 40%, still failed to reach the acceptability threshold, and the carbon tax one was deemed just about acceptable at 56% (Safaei & Longo, 2024). Although the notes created by the human participant received over 50% in most cases, the ratings were still rather low (Safaei & Longo, 2024). A limitation of the study is that the notes were written by a graduate student rather than a professional who in a real-world situation would be a more appropriate comparator.

Conclusions

A number of reports suggest that AI could have a transformative impact on the way that science is conducted, not only in terms of scientific discovery but also the construction of knowledge and how evidence is regarded. The evidence base on science advice is currently limited, with a lack of ‘real-world’ applications. It suggests that there may be potential for AI to support science advisers in terms of their capacity, helping on tasks such as evidence aggregation and synthesis. However, there are challenges. AI-generated output can appear authoritative but may be biased or inaccurate. Therefore, verification methods and skills like critical thinking are essential. The limited evidence from actual practical examples shows contrasting levels of performance between AI and human experts. AI may be able to help with information aggregation but may exhibit bias towards certain types of knowledge or towards certain groups. This would require greater transparency in terms of the provenance of information and calls for human oversight of processes.

Section 2: Examining the potential use of AI to support evidence review in the SAM

Introduction

In Section 2, we examine several key tasks that are undertaken as part of the SAM’s evidence review process. These are peer review, literature searching and synthesis, and literature screening. In all cases, we have conducted a search of the literature and published case studies.

Overview of opportunities and challenges

Below, we summarise some of the general literature relating to opportunities and challenges in evidence review.

A paper by Siemens et al. (2025) looks at the opportunities, challenges and risks of using AI for systematic reviews. They include:

  • Searching for studies. LLMs can be used to generate Boolean search queries and assist with the development of search strategies. However, LLMs can also create misleading controlled vocabulary terms and caution is therefore needed.
  • Screening of studies. Examples show mixed results, so more evidence is needed.
  • Translating reports. Evidence needs no longer be restricted to the language of the reviewer authors, but human judgement of complex texts remains essential.
  • Extracting study data. LLMs can assist with retrieving study information, such as methods and results. Early results are promising, but it is too early to generalise results.
  • Assessing the risk of bias. Results are inconsistent so far. Using an LLM for double-checking (e.g. replacing a second reviewer) may be a useful starting point. It may be necessary to report human and machine scores separately to maintain transparency.
  • Generating text. LLMs may be able to generate sections of a review with a standard format, such as the description of a screening process. LLMs have potential to improve the quality of academic writing.
  • Validation of LLM outputs. The application of LLMs in most review tasks has not been sufficiently validated, results are variable and validation remains challenging. Developing outcomes and standardised metrics is critical for comparison and replication of methods. LLM outputs are often difficult to reproduce; different LLMs show variations and training data may be biased.
  • Reference standards. Validation studies comparing the performance of LLMs and human reviewers need to define reference standards to provide a realistic benchmark.
  • Prompt engineering. Prompts are emerging as a specialised task. Effective prompt engineering can significantly enhance the performance of LLMs. Research is needed to understand how different types of prompts impact output.
  • Governance. LLMs have broad societal and environmental risks that require public debate and governance.
  • Over-reliance. There is a risk of over-reliance on AI, leading researchers to neglect critical thinking and contextual understanding that a trained human reviewer would provide. LLM outputs always need to be critically analysed and interpreted. Over-reliance can also lead to losing important research skills, such as being able to search and assess the literature.
  • Guidance. Consensus on what constitutes responsible use of LLMs in evidence synthesis has been missing until the drafting of RAISE (see below for details), which should be a living document that is adapted as the technology evolves.
  • Intellectual property. LLMs pose a risk of improper use of data, bias, unintentional plagiarism and possible copyright infringement.
  • Data privacy. Researchers must ensure compliance with data privacy regulations. LLM providers must be transparent about the processing and storage of submitted requests.
  • Environmental impact. Researchers and policymakers should recognise the large carbon footprint of LLMs and promote their sustainable use (Siemens et al., 2025).

Future measures should include capacity building to equip researchers with the knowledge and skills to use LLMs responsibly and understand their limitations; engagement with stakeholders, including policymakers and developers; and debate around societal concerns such as data privacy and environmental impact (Siemens et al., 2025).

A systematic review based on 19 studies in health sciences (Clark et al., 2025) concluded that the current evidence does not support the use of GenAI in evidence synthesis without human involvement or oversight. For most other tasks other than searching, GenAI may have a role in assisting humans. Only one study reported time savings with searching, although it did not lead to reliable outputs. The accuracy of outputs from GenAI tools frequently did not meet the standards achieved by human researchers, particularly in searching and screening tasks (Clark et al., 2025).

A study by Rose et al. (2025) used AI (but not GenAI, which was not available at the time) to analyse a sample of reviews in health and welfare. It found that reviews using AI/ML used more resources but were completed slightly faster. It concluded that outcomes are uncertain and further studies are needed.

Potential of AI in peer review

Peer review is a core part of SAPEA’s work and an area of interest for the potential use of AI.

Literature review

Potential opportunities of AI in peer review

The literature describes how the current system of scientific peer review faces numerous challenges. Firstly, peer review and research evaluation exercises consume considerable resources of time and staffing (OECD, 2023). There is an increasing number of manuscripts for review, a lack of peer reviewers, and concerns over “effectiveness, fairness and efficiency” (Bauchner & Rivara, 2024). There is a wider debate about how research evaluation and assessment might be done in a more open and inclusive way (GYA, IAP and ISC Scoping Group, 2023).

A 2023 study for ITHAKA suggests that publishers are now developing AI to bring efficiencies to peer review (Rieger & Schonfeld, 2023). In a cross-sectional analysis of AI integration in peer review, policies from 439 high impact factor journals across a range of disciplines were examined between March and August 2025. During this time, 103 had updated their AI policies for peer review (involving publishers such as Karger, Springer Nature, Wiley and IEEE), bringing the total number of these journals with AI policies related to peer review to 83% (Wang & Gong, 2026). In another study that selected the top 100 medical journals to examine guidance on the use of AI in peer review, it was found that 78% of the journals provided guidance on the use of AI in peer review. Of those that provided guidance, 59% prohibited using AI, while 32% allowed its use if confidentiality was maintained and authorship rights respected (Li et al., 2024).

A 2024 study for ITHAKA looks at the role of generative AI in scholarly publishing (Bergstrom & Ruediger, 2024). Experts’ views on the role of AI in peer review varied; they mentioned advantages for ‘pre-review’ feedback, identifying peer reviewers, detection of research misconduct, copy editing and speeding up the review process. An OECD report suggested that AI could save time, for example, on tasks like pre-review screening to detect early problems in papers (OECD, 2023). Other potential uses of AI included conducting initial checks of submissions for issues like relevance to the journal’s scope, formatting, adherence to the citation style, matching references with in-text citation, and detecting plagiarism or data manipulation (Doskaliuk, Zimba, Yessirkepov, Klishch, & Yatsyshyn, 2025; Ebadi, Nejadghanbar, Salman, & Khosravi, 2025; Giray, 2024; Hoyt, Limon, & Chang, 2025; Kadri et al., 2024; Kankanhalli, 2024; Kousha & Thelwall, 2024). AI could flag grammatical errors, spelling mistakes, phrasing and inconsistencies in terminology, references, and data reporting (Doskaliuk et al., 2025). AI could also be used to scan for statistical errors or unusual reporting patterns that might indicate problems (Purificato et al., 2025). AI could be used to help reviewers format and structure their feedback and ensure they use an appropriate language tone (Doskaliuk et al., 2025; Ebadi et al., 2025; Giray, 2024; Lee et al., 2025). Kankanhalli (2024) provided a summary of the potential applications of AI tools in different peer review tasks (see Table 1):

Task

AI automation potential

Example of AI tool

Format check: checking that the manuscript follows publication outlet format rule of structure, styles, references, and metadata

High

Penelope.ai

Plagiarism detection: identifying the extent and nature
of copying from other sources without source attribution

Low

iThenticate, ZeroGPT

Language quality: assessing audience-appropriate
readability, cohesion, logic

Medium

UNSILO

Manuscript-reviewer matching: finding suitable reviewers
for a manuscript using reviewer profiling

High

TPMS

Scope/relevance: assessing fit with the scope
of the publication outlet

Medium

UNSILO, GPT-4

Soundness/rigor: checking that study methodology
and analysis are rigorous and robust

Medium-Low

Enago, StatCheck, StatReviewer

Novelty: newness or departure from the existing body
of knowledge

Low

ReviewAdviser

Significance: importance of the phenomenon
that the manuscript is focusing on

Low

ReviewAdviser

Writing and integrating reviews: reviewer and editor roles
of writing and integrating reports

Medium

GPT-4

Reproducibility check: author code and data analysis checks
for reproducibility of findings

High

GPT-4

  1. Automation potential of AI for the peer review process (Kankanhalli, 2024, p. 79).

Kadri et al. (2024) outline automated tools and technologies commonly used for peer review, with their benefits and challenges (see Table 2):

Category

Tool/Technology

Benefits

Limitations/Challenges

Linguistic rules

Grammar, spelling,
and punctuation check

Improves manuscript language quality

May not detect all nuanced language issues, especially in scientific writing

Machine learning algorithms

Plagiarism detection
and grammar check software

Enhances the integrity of peer review and discourages unethical practices

False positives/negatives, accuracy issues with obscure sources

Data analysis and visualisation

Data visualisation tools
(e.g. tableau, MS PowerBI)

Enhances data clarity and simplifies data-driven discussions

Requires reviewer familiarity with the tool

Natural language processing

Content analysis

Speeds up content analysis and ensures consistency in language use

Accuracy can be affected by nuanced language or context

Reference and data management

Reference management software (e.g. EndNote, Mendeley)

It improves citation accuracy and minimises
the risk of missing or incorrect references

Requires access to comprehensive reference databases

Statistical software

Data validation
(e.g. R, python)

Ensures statistical robustness reduces the risk of data manipulation

Requires reviewer proficiency in statistical analysis

  1. Commonly used tools and technologies for peer review
    (Kadri, Dorri, Osaiweran, Garyali, & Petkovic, 2024, p. 400).

In a recent paper, Sun (2025) suggests that LLMs demonstrate significant promise in tasks such as generating automatic feedback, simulating reviews, identifying key evaluation aspects, and providing human-like recommendations and review scores. 

Potential risks of AI in peer review

Peer review quality. AI may fail to address specific contexts and unique challenges, resulting in feedback that is too general and therefore of limited use (Ben Saad et al., 2025). AI lacks in-depth expertise in subjects and so may miss methodological flaws or theoretical inconsistencies and/or fail to interpret complex arguments (Doskaliuk et al., 2025; Hoyt et al., 2025; Salman et al., 2025). Human reviewers over-relying on AI may result in missed errors or poor assessments of manuscripts (Doskaliuk et al., 2025). AI tools are also limited in their ability to evaluate the novelty and significance of research (Doskaliuk et al., 2025; Ebadi et al., 2025; Hoyt et al., 2025). Data used to train AI models may become outdated, particularly in fast-evolving fields (Hoyt et al., 2025). AI is vulnerable to security challenges like invisible text injections, where instructions to accept a manuscript are written in white font on white backgrounds, visible to LLMs but not humans (Choi et al., 2026).

Biases. AI may risk propagating biases and be open to misuse (GYA, IAP & ISC Scoping Group, 2023). A report from the OECD (2023) concurs that the use of AI raises ethical and institutional challenges. If training data contains biases (for example, related to gender, race, geographic region, research topic, or publication trends), they will be transferred onto the model and AI may end up perpetuating them (Doskaliuk et al., 2025; Giray, 2024; Hoyt et al., 2025; Kadri et al., 2024; Kankanhalli, 2024; Lee et al., 2025). Algorithmic biases can result in systematic prejudices that undermine an objective evaluation of manuscripts (Ben Saad et al., 2025).

Privacy, confidentiality, data protection, and IP. Peer review using AI should take into account data privacy, confidentiality, and intellectual property concerns (Science Europe, n.d.). Data submitted to AI tools may be used for training or other purposes, raising concerns about privacy and data security, especially in the case of sensitive or confidential information (Ebadi et al., 2025; Giray, 2024; Hoyt et al., 2025; Kadri et al., 2024; Kankanhalli, 2024; Lee et al., 2025). This risk may be mitigated by local hosting of information uploaded to AI (Hoyt et al., 2025)).

Transparency of AI. AI tools often perform as ‘black boxes’, making it difficult to understand the rationale behind their decisions and recommendations (Ebadi et al., 2025; Kadri et al., 2024; Kankanhalli, 2024).

Transparency and policies about the use of AI. A failure to declare the use of AI in peer review could mislead authors and editors and undermine the ethics of scholarly communication (Ben Saad et al., 2025). There needs to be transparency about the use of AI; it should be explicitly communicated, including addressing potential concerns that may arise (Ebadi et al., 2025).

Accountability. It is important to define who is responsible for any decisions or recommendations and to provide clear guidelines to authors, publishers, reviewers, and engineers who develop AI systems (Kadri et al., 2024).

Deskilling. Some literature suggests that the use of AI is associated with reduced critical thinking because overreliance results in ‘cognitive offloading’. In the case of peer review, this has significance because the task of human reviewers is to use their best judgment to evaluate manuscripts (Hoyt et al., 2025).

Sun (2025) highlights several limitations of applying LLMs in peer review. These include lack of accuracy and depth to provide constructive feedback and contextually nuanced critiques, as well as outputs that often diverge from those of human reviewers. Such shortcomings support the view that LLMs should be used as support tools rather than replacements for human expertise in the peer review process. Future innovation would need to focus on improving the performance and reasoning abilities of LLMs in task-specific contexts, evaluating the alignment of LLM outputs and human reviewers, advancing the integration of LLMs into human peer review workflows, and exploring the potential of LLM agents.

Published case studies

There are a number of published case studies that examine the use of AI in peer review, and new ones are being published all the time.

Farber (2025) compared the performance of human reviewers versus Claude-3, based on evaluation of ten unpublished papers by independent researchers. The results of the analysis are shown below. Claude-3 showed consistently lower scores than human reviewers, especially on literature engagement.

  1. A bar chart compares the mean quality scores of human reviewers and an AI system, showing human reviewers consistently scoring higher with visible error bars.

Description generated by AI Comparative analysis of human and AI review quality (Farber, 2025, p. 878).

In another study, five human reviewers and three AI models (Microsoft Co-pilot, ChatGPT-3.5, and Google Gemini 1.0) reviewed eight manuscripts each, resulting in a total of 64 reviews that were anonymised and submitted to two experts, who ranked each submission according to different criteria related to the quality of feedback and the quality of recommendations (Lim et al., 2025). As can be seen in the table (below), the human reviews scored higher than AI on each dimension of the quality of recommendations, but lower on the quality of feedback.

Effectiveness in providing feedback and recommendation

Characteristic

Overall, N = 128

AI, N = 48

Humans, N = 80

P-value

Feedback

Relevance

2.75 (0.92)

3.50 (0.76)

2.30 (0.68)

<0.001

Completeness

2.70 (0.95)

3.54 (0.72)

2.19 (0.67)

<0.001

Accuracy

2.70 (0.92)

3.45 (0.79)

2.25 (0.68)

<0.001

Error identification

2.55 (1.00)

3.22 (1.00)

2.14 (0.77)

<0.001

Constructive

2.73 (0.92)

3.52 (0.72)

2.25 (0.67)

<0.001

Overall

2.68 (0.92)

3.44 (0.77)

2.23 (0.68)

<0.001

Recommendation

Relevant

3.18 (0.85)

2.83 (0.91)

3.38 (0.75)

0.003

Complete

3.17 (0.83)

2.88 (0.95)

3.35 (0.69)

0.009

Accurate

3.16 (0.85)

2.84 (0.92)

3.35 (0.75)

0.013

Overall

3.17 (0.83)

2.85 (0.92)

3.36 (0.71)

0.008

  1. Effectiveness and efficiency of generative AI tools compared to human reviewers
    (Lim, Tan, Hoe, & Koh, 2025, p. 4).

A study assessed comments from human peer-reviewers and GPT-4 for 3,096 submitted manuscripts in 15 of the Nature family of journals and 1,709 manuscripts submitted to the International Conference on Learning Representations (ICLR). It found an average overlap of comments of 30.1% in the Nature journals and 35.3% in the ICLR submissions (Liang et al. 2023, cited in Bauchner & Rivara, 2024).

In another study, 22 manuscripts submitted to the Journal of Oral and Maxillofacial Surgery were used to compare reviews between human reviewers and four LLMs17 (Joachim, Dodson, & Laviv, 2025). Human reviewers rejected 15 (68.2%) of the manuscripts, while the AI models rejected between 0 and 2 (9.1%). Of the 22 manuscripts, human reviewers recommended 2 (9.1%) for acceptance with minor revisions, while for the AI models this figure ranged between 8 (36.4%) and all 22 manuscripts (Joachim et al., 2025).

In a study, 20 accepted research papers were reviewed, using an LLM review system based on ChatGPT-4. The generated reviews were then analysed to assess the semantic similarity and quality between LLM-generated and human reviews. The results show that LLMs could generate quality reviews of papers that match human reviews in semantic aspects. However, it remains challenging for researchers to evaluate the critical thinking process of LLM peer review systems. The paper recommends further research, based on more types of papers. It also mentions that LLM reviews may be overly positive in certain situations. The authors suggest that LLM reviews could be used as a baseline for peer reviews, assisting human reviewers (Zhao & Mahmoud, 2024).

A study by Zhu et al. (2025) used Claude to generate peer review reports, recommendations for rejection, requests for additional citations and refute requests to cite literature for 20 manuscripts in the field of cancer biology. AI detection tools18 were used to assess whether the reviews were identifiable as LLM-generated. Two experts evaluated all the LLM-generated outputs. The results showed that LLM-generated reviews had some consistency with human reviews but lacked depth and detailed critique. The LLM was able to generate convincing rejection comments and plausible citation requests, although these included requests for completely unrelated references. The AI detection tools struggled to identify LLM-generated reviews (in the case of GPTzero, over 80% were classified as human-written). The conclusion was that LLM-generated review comments do not reach the level of high-quality human review.

A discussion paper for the IZA Institute of Labor Economics (Pataranutaporn, Powdthavee & Maes, 2025) analysed 27,090 evaluations of 9,030 unique submissions in the field of Economics, using an LLM19.
The base data contained a variation by author (gender), institutional affiliation (by rank) and journal (top-, mid- and bottom-tier), as well as a set of AI-generated papers that mimicked top-tier quality. The results found that the LLM was very effective at distinguishing submissions by paper quality, although it struggled to recognise the AI-generated papers. The LLM did exhibit some small but important bias towards prominent institutions, male authors and renowned economists, underlining the bias demonstrated by human reviewers in previous published studies (cited in the paper). The paper’s conclusion was that LLMs could offer efficiency gains by streamlining editorial workflows, especially at the early stages of desk rejection. However, AI algorithms would need to recognise and mitigate bias. Hybrid models would integrate human judgement with AI-assisted evaluations.

AI-assisted peer review. The authors prototyped a system called MetaWriter, trained on five years of open peer review data to support meta-reviewing (which summarises a set of peer reviews and offers a final decision). The study involved 32 participants, each writing meta-reviews for two papers, one with and one without MetaWriter. It found that MetaWriter significantly shortened the time taken and improved coverage of the meta-reviews, as rated by experts. However, the participants raised concerns around trust, over-reliance, and agency, with some expressing concerns about over-reliance and the potential for bias and abuse of the system (Sun, Tao, Hu & Dow, 2024).

A study looked at using AI to support peer reviewers with suggestions. The experiment compared two different designs on 31 participants, one involving an inline interface and the other a list of suggestions, embedded within a writing assistant tool. The system was trained on a corpus of nearly 12,000 peer reviews in German business studies. Those using the inline interface wrote longer reviews on average, and those with the list experienced more ease in using the tool (Neshaei, Rietsche, Su & Wambsganss, 2024).

  1. Line graphs show a sharp increase in the frequency of the words commendable, innovative, meticulous, intricate, notable, and versatile from 2023 to 2024.

Description generated by AI “Shift in Adjective Frequency in ICLR 2024 Peer Reviews” (Liang et al., 2024, p. 1)

An online experiment, involving 137 participants, examined the impact of appending an automatically generated positive summary to peer reviews of a writing task, alongside an overall evaluation score (high to low). It found that a positive summary increased authors’ acceptance of the feedback, whereas a low score increased authors’ revision efforts. The study concluded that AI assistance could provide an opportunity to create a more constructive and friendly peer review process (Yang, Uhde, Yamashita & Kuzuoka, 2025).

Detecting undeclared use of AI in peer reviews: Liang et al. (2024) found a significant increase in the use of certain adjectives that tend to be disproportionately produced by AI in the ICLR20 2024 conference peer reviews, compared to the previous years. Analysing peer review submissions to a number of machine learning (ML) conferences after the release of ChatGPT to the public, Liang et al. (2024) found that there was evidence that a significant proportion of peer review text submitted to some of the conferences had been substantially modified (beyond grammar and spell checking) by ChatGPT; the estimated usage of ChatGPT in reviews peaked within three days of review deadlines; and that reviewers who did not respond to author rebuttals had a higher estimated usage of ChatGPT21.

Figure 9. Taxonomy of metrics (Thomas et al., 2025b, p. 20).

Potential of AI for literature searching

A key task in the SAM is in finding literature relevant to the scoping questions posed. The literature we found on this topic entirely focused on case studies.

Published case studies

We found several papers relating to the use of AI in literature searching, as follows.

The Research Information Services at Canada’s Drug Agency looked at 3 AI tools for information retrieval22 (Featherstone et al. 2025). Performance was compared with the usual information retrieval practices used at the Agency and the study was based on a sample of 7 projects within a range of health-related topics. The results showed that AI search tools have inconsistent and variable performance. Although AI tools for searching and screening sometimes achieved a direct time saving compared to standard searching, they did so at greater cost to search sensitivity23 and/or NNR24). For some topics, it may be worth an expert operator spending time to build effective prompts as this is likely to improve the performance of the tools. The overall recommendation was that it is worth supplementing current methods with AI search tools when it is likely that they will improve sensitivity (recall), or retrieve unique records, without increasing unreasonably the NNR and/or cost to screen results. The study warns that AI search tools can give users a false sense of confidence and skillful prompt engineering is required when developing search strategies. Experience in evidence review is needed to validate AI-generated search strategies and results. Understanding the individual strengths and weaknesses of individual AI tools is crucial (Featherstone et al., 2025).

A study in Conservation Science evaluated the performance of ten LLMs across six different database retrieval strategies against human experts in answering multiple-choice exam questions using the Conservation Evidence25 database (Iyer et al., 2025). It found that LLM performance was comparable to human experts and far better than random guessing over 45 restricted multiple-choice questions, both in terms of correctly answering them and retrieving the document used to generate them. Newer models of LLMs performed better than older models. In contrast, in response to unfiltered questions, the LLM’s level of specific knowledge varied across topic areas. The findings suggest that LLMs could enable expert-level use of evidence syntheses and databases in different disciplines. However, general LLMs used ‘out-of-the-box’ are likely to perform poorly and misinform decision-makers. The authors hypothesise that LLMs’ performance on more complex, nuanced questions and/or using free-text open-ended answers would be lower, but investigation would be needed into how and why, and whether mitigation strategies could be built in. There is also the question of whether advanced prompt engineering could improve the performance of LLMs. The paper flags a number of risks, including de-skilling, for example, less critical thinking about the evidence, its limitations, source and validity. There could also be a lack of accountability in the use of evidence, including bias. It is important that users understand the limitations of the evidence base. Nonetheless, the paper foresees RAG systems capable of expert-level, evidence-based advice using databases and syntheses. At the same time, evaluation of LLM-based decision support systems is essential to ensure that decision-makers do not receive biased, erroneous or misleading information (Iyer et al. 2025).

A study by Lau and Golder (2025) compared four evidence syntheses, using Elicit Pro. It found the sensitivity (recall) of Elicit to be poor, but it did identify some studies not included in the original searches and had a higher level of precision. Across the four case studies, an average of over 80% of the included studies not found by Elicit were indexed in Semantic Scholar (the literature source for Elicit), which points to Elicit’s poor searching skills rather than lack of access to sources. Elicit also has an issue with transparency, not sharing the underlying Semantic Scholar query and version history, thus limiting reproducibility. The conclusion was that although Elicit did not search with high enough sensitivity to replace traditional literature searching, it could be useful for preliminary searches and as a useful adjunct. As Elicit continues to improve, further evaluations should be conducted.

Spillias et al. (2024) evaluated the usefulness of an LLM, based on a scoping review of the literature on community-based fisheries management. The findings suggest that generalised AI tools can assist in brainstorming for relevant search terms and designing the structure of a search string, although AI was unable to provide an actual search string without substantial prompting. For literature screening, AI can provide a reliable classification of relevant articles, although for complex screening there were inconsistent results, requiring additional repeat calls to the AI tool. Overall, the results highlight the need for collaboration between humans and AI systems and the continued need for human researchers in identifying relevant literature. AI has the capability to support the search and retrieval of relevant articles, improving the efficiency of the review process. The question is whether AI systems can be designed that enable better search and retrieval for specific topics, without the need for manual searching and screening.

Another study used 9 test questions in current marketing research, which were then searched on traditional and AI-based research platforms (Tomczyk, Brüggemann, Mergner, & Petrescu, 2024). The AI tools gave a considerably higher number of results for each research question, but the quality of the results were frequently lower. On the other hand, the AI tools, and particularly Elicit, were good at finding articles that were not discovered via conventional means. This makes some AI tools a useful complement to traditional search methods.

A study by Gwon et al. (2024) aimed to explore the potential of ChatGPT as a literature search tool for systematic reviews and clinical decision support systems in health care settings. It found that ChatGPT did not conduct a systematic search and could not be used for such a purpose. Bing AI had more accurate answers and provided website references, in contrast to ChatGPT. Like ChatGPT, it gave more answers when given more specific questions. Another study examined the benefits of ChatGPT for literature search from an end-user perspective, comparing it with conventional search methods (Yip, Sun, Bassuk & Mahajan, 2025). It also tested whether ChatGPT’s user support tools could improve performance. Basic ChatGPT functions had limitations in consistency, accuracy and relevancy of results. The user-support tools showed improvements, but there were still limitations. The conclusion was that ChatGPT does not have the level of reliability needed for it to be adopted for literature searching.

Potential of AI for literature screening

Literature review

The first appraisal phase of a systematic review requires at least two independent reviewers to screen a large number of abstracts. Although this step is crucial to the quality and validity of the review, it is also highly time-consuming, taking on average 33 days (Nordmann & Fischer, 2026; Nordmann, Schaller, Sauter & Fischer, 2026). Moreover, only a small proportion (typically 1% to 2.9%) of abstracts screened are deemed relevant, and the process is prone to human error (Nordmann et al., 2026).

To address these challenges, AI has been explored as a supporting tool in this context. Text mining, an AI technique, has supported automation in systematic reviews since 2006, but processes have been refined and improved since then (Nordmann et al., 2026). The most popular machine learning technique for screening is active learning. Using this method, AI models learn from human decisions about abstract eligibility to predict the inclusion likelihood of an unclassified abstract. It then ranks these abstracts by predicted relevance and continuously refines its predictions as additional human screening decisions are incorporated (Zhan et al., 2025).

Several studies estimate that machine learning can halve the screening workload, while detecting 95% of relevant articles. However, AI tools for screening have three main limitations:

  1. Existing tools lack the versatility needed to screen articles with a diverse range of research fields
  2. Validation is mostly lacking in detail in many machine learning tools for systematic reviews
  3. Existing tools apply stopping rules to determine when screening should end, but these rules are often arbitrary and uncertain, often relying on fixed thresholds (e.g., stopping after 100 irrelevant articles have been screened consecutively) (Zhan et al., 2025).

Published case studies

The National Collaborating Centre for Methods and Tools (NCCMT) integrated 4 AI features into the screening process for literature, looking at 35 rapid reviews on 20 topics (Dobbins, Traynor, Clark & Neil-Sztramko, 2024). The results were compared against those produced manually. AI screening correctly excluded up to 80% of irrelevant search results across reviews and involved less screening time. It concludes that AI holds promise as a way to improve screening efficiency.

A study by Hu et al. (2025) examined the performance of an agentic-AI assisted tool for screening and classifying evidence for the Global Repository of Epidemiological Parameters (grEPI)26. It evaluated 2000 citations from a systematic review on measles, using four LLMs27. It found that the agent was able to incrementally improve accuracy, regardless of the LLM. Performance was significantly impacted by screening thresholds and thresholds for optimised workload reduction strategies (i.e. the trade-off between speed and recall/sensitivity). The conclusion was that the agent showed promise in improving the efficiency and effectiveness of evidence synthesis in public health, although further development and refinement of human-in-the-loop systems would be essential going forward (i.e. interaction between human reviewers and the agentic tool). The study highlights the importance of prompt language in the performance of LLMs. The results show that extra care is needed on the clarity of the screening questions and the inclusion/exclusion criteria, with iterative fine-tuning needed. Although workload reduction and speed of review are key motivators in using AI, a further consideration is the time to develop or adapt an AI-assisted system and the cost of using LLMs within that system.

Rogers et al. (2025) examined the use of AI to update the Health Evidence28 database, which is resource-intensive to maintain, given the growth of published literature. It looked at the ability of AI-assisted screening to correctly predict non-relevant references. Over three years, the number of references requiring manual screening was reduced by 70%, reducing the time spent manually screening by an estimated 382 hours. The conclusion was that AI-assisted screening can be an important tool to supplement manual screening. However, the training, testing and analysis of the AI-assisted screening required significant human resources (not quantified), and the overall cost-benefit would depend on the volume of references that require relevance screening. Thus, the approach may be useful for systematic reviews with large search results and systematic reviews that are regularly updated. Another factor is that AI models are trained on historical data, so model performance may degrade over time and needs to be monitored. An advantage of AI is that it does not need the frequent retraining needed for new staff performing manual screening. Future work could explore the use of explainable AI to better understand results and integrate new insights into improving the AI model (Rogers et al., 2025).

Ruan et al. (2025) evaluated 5 AI tools29 in literature screening, using a random sample of 1,000 publications. The tools assessed performed well, and reduced screening time. All the tools took between 2 and 6 seconds per article (i.e. between 33 and 100 minutes for all the articles), in contrast to human screening which took 2 weeks. However, the higher error rate (False Positive Fraction, FPF) means that they are not yet suitable as standalone solutions; rather, the tools could serve as effective aids as part of a hybrid approach that integrates human expertise with AI to optimise efficiency and accuracy.

A study looked at the role of AI in applying eligibility criteria for evidence-screening in environmental science (Zuo et al., 2025). The AI model demonstrated substantial agreement at title/abstract review and moderate agreement at full text review with expert reviewers. AI showed potential for structured screening assistance. Combining AI with domain knowledge could provide a step towards evaluating the feasibility of AI-assisted screening in a field with diverse, large volume and interdisciplinary studies, such as environmental science.

Based on 124 studies and an in-depth analysis of 40, a paper by Adel and Alani (2025) evaluates the capabilities and limitations of generative AI, specifically ChatGPT, in conducting systematic literature reviews. The findings show that ChatGPT can enhance efficiency, with reported workload reductions, but accuracy varies widely by task and context. It suggests that AI can match or possibly exceed human performance in simple screening but underperforms in nuanced screening. Whilst ChatGPT typically accelerated tasks, humans did better at accuracy and contextual understanding, particularly in full-text evaluation and nuanced methodological assessments. The value of AI-human collaboration is emphasised, where human oversight ensures methodological rigour and integrity.

Conclusions

Literature screening appears to be a promising area for the use of AI, as it can save considerable time. AI can be accurate and precise on simple screening with well-defined criteria, but less so as complexity increases. There are also considerations such as the resource requirement needed for training, testing, and evaluation. Going forward, it would seem that human oversight will remain essential, with effective prompt engineering required.

Potential of AI for search, data extraction
and literature synthesis

Literature review

We found limited evidence on search, data extraction and synthesis. Zhang et al., (2024) flag some of the risks, such as hallucinations, which can either be ‘intrinsic’ (meaning that the output contradicts the source) or ‘extrinsic’ (where the output cannot be either supported or contradicted by the source). They suggest that copyright is a major constraining factor, limiting access to full text for training. Whilst many automatic evaluation metrics have been developed for assessing AI-generated summaries, they do not correlate strongly with human expert evaluations on factual consistency, relevance, coherence and fluency. The authors call for more research to develop evaluation metrics that correlate with human judgements. They conclude that domain-specific LLMs may perform better than generalist LLMs (Zhang et al., 2024).

Guidelines and frameworks for the use of AI
in systematic review

A number of bodies involved in systematic review are working together on guidelines and frameworks. Currently, the most robust guidelines available for use of AI in evidence synthesis are Responsible Use of AI in Evidence SynthEsis (RAISE), which is a joint endeavour between evidence synthesis organisations such as the EPPI Centre, Cochrane Collaboration, Campbell, JBI, and many others (Thomas et al., 2025a,b,c). Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence (CEE) have formed a joint AI Methods Group (Flemyng et al., 2025). The Group supports the aims of RAISE, which states the need to work together to ensure that AI does not compromise the principles of research integrity. It suggests that decisions involving trade-offs are required, related to capacity, resource availability, urgency, relevance and scope. For example, biases may include using English-only or open-access only material, which may or may not be important. Moreover, evaluation studies are essential to determine whether an AI tool performs at an adequate level. There is a lack of consensus on what constitutes trustworthy AI for evidence synthesis; trust in AI and trustworthy AI are co-dependent. Trustworthy AI is seen as a transparent, open-access system, whereas trust is a perception by users. Considering which tasks AI should or should not undertake can help address concerns around trust, along with integrating human interaction through human-in-the-loop procedures, such as combining manual screening with machine-learning prioritisation.

Gartlehner et al. (2025) suggest that there is likely to be an increase in AI-generated outputs, which may appear well-written and methodologically sound, but lack the required critical evaluation for use in decision-making. It may also become increasingly difficult to detect these. The transition to greater use of AI therefore needs to be carefully managed, with strict adherence to rigorous conduct and reporting standards. In a position statement, the Cochrane Rapid Reviews Methods Group outlines its stance on the use of AI in rapid reviews. AI should not be used to fully automate a rapid review; however, the use of AI is encouraged when it serves to improve review quality. For example, AI might be helpful in providing an additional check or suggesting additional sources. Humans remain fully responsible for verifying AI outputs and making final decisions, such as evidence interpretations or methodological judgements. Its recommendations are as follows:

  • Do not rely on AI to fully automate any step of a rapid review. Human oversight is essential to maintain methodological rigour and accountability.
  • Use AI for quality assurance where it complements human judgment and helps mitigate risks, for example, flagging missed studies, spotting inconsistencies, or detecting extraction errors.
  • Be transparent about AI use and justify its use. Specify in the protocol which tools will be used, the tasks they will perform, and how their outputs will be verified. If using a generative large language model, document the model version and prompts applied.
  • Demonstrate that the selected AI tool upholds the methodological rigour and integrity of the rapid review.
  • Maintain author responsibility. Authors remain fully accountable for the accuracy, interpretation, and overall validity of the rapid review. AI tools cannot be listed as authors.
  • Respect copyright and intellectual property. Ensure that AI tools are used in compliance with licensing terms.
  • Adhere to RAISE principles and the Cochrane position statement on AI use in evidence synthesis to ensure that AI applications are ethical, transparent, and fit for purpose (Gartlehner et al., 2025)

Below, we provide a summary of the RAISE guidelines of most importance to the evidence synthesis work within SAPEA. At the time of writing (January 2026), RAISE is on version 2.2, and it is a set of regularly updated documents, reflecting the fast developments of the field.

The RAISE guidelines define responsible AI as “a framework for developing and using AI that does not compromise the principles of research integrity, which include honesty, rigour, transparency and open communication, care and respect, and accountability. It also involves ensuring ethical, moral and legal values are upheld” (Thomas et al., 2025a, p.9). The guideline authors acknowledge that due to its use in making policy and practice decisions, evidence synthesis “is a challenging area for automation, due to its need for higher accuracy and transparency than other AI use cases, as its outputs can influence decisions that affect people’s lives” (Thomas et al., 2025b, p.3).

The guidelines recognise eight distinct roles involved in evidence synthesis – methodologists, evidence synthesists, publishers of evidence synthesis, users of evidence synthesis, trainers of evidence synthesis methods, organisations producing evidence synthesis, funders of evidence synthesis, and AI development teams – which they encourage to work cooperatively to develop and improve AI tools and integrate them in evidence synthesis workflows in a responsible way. It is recognised that the same individual may fill more than one of these roles (Thomas et al., 2025a).

Below are some of the main considerations outlined in the RAISE guidelines.

Accountability and transparency. Authors of evidence synthesis are responsible for the content they produce, the methods they use, and for the findings of evidence synthesis, including the decision to use AI tools, the ways in which they are used, and their impact on the work. Authors should declare if they have used AI if it makes or suggests any judgements that may potentially influence the work. AI tools should not receive credit for authoring evidence synthesis because they cannot satisfy criteria for authorship (Thomas et al., 2025a).

Choice of AI tools: capability and performance. Relevant evaluations should be referenced when justifying the use of a particular automation tool (Thomas et al., 2025c). For generative AI tools, the RAISE guidelines suggest that issues of stability, robustness, and the potential for hallucinations are considered (Thomas et al., 2025b). ‘Stability’ refers to the level of random variability. This can be assessed, for example, in a classification task by putting the same data through an LLM multiple times and evaluating how often predicted categories change. ‘Robustness’ refers to a system’s ability to “handle unexpected inputs in a predictable manner”, and assessing this will sometimes require human judgment and interpretation, in addition to metrics, to ensure nuance is captured (Thomas et al., 2025b, p.19). The presence of hallucinations can be assessed by directly asking the LLM about content that is not present in the input data or by asking nonsensical questions to check if the system states that such information is not available. Verification of sources cited by LLMs is also necessary: both that they exist and that they contain the information the LLM claims they contain. In terms of evaluating tools’ performance, both generative AI and other kinds of AI, the RAISE guidelines propose the following taxonomy of metrics, provided in Figure 9 below, as appeared in Thomas et al. (2025b):

A taxonomy chart categorizing common performance metrics for AI tools in evidence synthesis into accuracy, reliability, efficiency, usability, and risk assessment with detailed sub-metrics listed under each category.

Description generated by AI
  1. Taxonomy of metrics (Thomas et al., 2025b, p. 20).

This taxonomy is not meant to be prescriptive, but rather to help with the selection of the most appropriate ways to evaluate AI tools. The RAISE guidelines outline the following types or evaluations:

  • Performance evaluation that addresses how well AI performs
  • Process evaluation that addresses how well AI works in realistic settings
  • Human evaluation that addresses how the use of AI affects the human processes involved in evidence synthesis (Thomas et al., 2025b).

Choice of AI tools: regulatory issues. The RAISE guidelines state that “issues related to plagiarism, provenance, copyright, intellectual property, jurisdiction, licensing, confidentiality, compliance, and privacy responsibilities, including data protection laws” need to be considered (Thomas et al., 2025a, pp.13-14). Use of AI must adhere to the laws in the relevant jurisdictions. The guidelines suggest that AI systems should be chosen on the basis of having clear and transparent information about their processes and methods, ideally existing in the public domain; that they should have the option to disable data collection for training and turn off or delete history or interaction tracking; and that they should not gain any rights to users’ content (Thomas et al., 2025c). The RAISE guidelines cite Wiley’s advice to be wary of: “‘perpetual’ rights to your content; ‘royalty-free’ usage rights; ‘transferable’ rights to third parties; broad statements about using content ‘for training’ or ‘for any purpose’; classifying the AI-generated content as ‘open’, ‘free to use’, or subject to a permissive license or ‘Creative Commons’ license (which may open use of your work up to the rest of the world), or subject to a restrictive license; any limitations on your ability to use your content or the output; or unclear policies about data retention and deletion” (cited in (Thomas et al., 2025c, p.21). Another consideration is that there may be restrictions on the ‘export’ of services in some parts of the world, so, for example, a USA-based AI tool may not be used on a particular territory (Thomas et al., 2025c). There may also be considerations such as terms and conditions of an organisation’s insurance which would preclude it from using AI (Thomas et al., 2025c).

Limitations and biases. Evidence synthesis authors should consider and discuss any biases and other limitations of AI tools and their potential impact on the work (Thomas et al., 2025a).

Open science. Where possible and practical, the inputs (such as prompt development), outputs, datasets, or code used to develop, validate, and implement AI tools should be made publicly available (Thomas et al., 2025a).

Conflict of interests. Any financial or non-financial interests in the AI tool and any funding sources should be declared (Thomas et al., 2025a).

Environmental impacts. AI tools can use a large amount of energy and water and use of such resources can lead to significant carbon emissions, so it is important to consider the environmental impacts of using AI, where possible prioritise tools that report their carbon footprint transparently and make an effort to minimise it, and balance the environmental impacts against potential benefits including reducing research waste through increasing efficiency (Thomas et al., 2025c). The guidelines refer to the Green Algorithms project30, which provides calculators and resources that can be used to estimate and reduce the carbon footprint of research projects.

Use of AI in the work of others. Evidence synthesis teams should consider that some of the published research papers that may be of relevance to their review work may have been fully or partially generated by AI (Thomas et al., 2025a).

Conclusions

The integration of AI into workflows relating to evidence review requires careful management, with strict adherence to responsible use and quality standards. This is even more vital within the context of policymaking, given that it affects people’s lives and therefore the stakes are higher and the risks are greater. The use of AI may be encouraged where it demonstrably improves the quality of reviews. At the same time, humans must always be fully responsible for verifying output and making judgement decisions. Transparency about the use of AI and justification for its use are essential. Compliance with the law is vital, and intellectual property rights must be respected.

References

Adel, A., & Alani, N. (2025). Can generative AI reliably synthesise literature? Exploring hallucination issues in ChatGPT. AI & Society, 1-14. https://doi.org/10.1007/s00146-025-02406-7

APEC (Asia-Pacific Economic Cooperation. (2022). Artificial intelligence in economic policymaking. Retrieved from https://www.apec.org/publications/2022/11/artificial-intelligence-in-economic-policymaking

Arora, A. et al. (2024). Towards intelligent governance: The role of AI in policymaking and decision support for e-governance. In So In, C., Londhe, N.D., Bhatt, N., & Kitsing, M. (Eds.), Information systems for Intelligent Systems. ISBM 2023. Smart Innovation, Systems and Technologies, 379. Springer, Singapore. https://doi.org/10.1007/978-981-99-8612-5_19

Bauchner, H., & Rivara, F. P. (2024). Use of artificial intelligence and the future of peer review. Health Affairs Scholar, 2(5), qxae058. https://dx.doi.org/10.1093/haschl/qxae058

Ben Saad, H., Dergaa, I., Ghouili, H., Ceylan, H. I., Chamari, K., & Dhahbi, W. (2025). The assisted technology dilemma: A reflection on AI chatbots use and risks while reshaping the peer review process in scientific research. AI & Society. https://doi.org/10.1007/s00146-025-02299-6

Bergstrom, T. & Ruediger, D. (2024). A third transformation? Generative AI and scholarly publishing. Retrieved from https://sr.ithaka.org/wp-content/uploads/2024/10/SR-Brief-Generative-AI-and-Scholarly-Publishing-103024.pdf

Canali, S., & Barone-Adesi, F. (2023). Can AI deliver advice that is judgement-free for science policy?. Nature624(7991), 252-252. Retrieved from https://www.nature.com/articles/d41586-023-03949-9

Choi, B., Jun, T. J., Sung, J. W., Park, I. W., Lee, J. M., Cho, S. I., ... & Suh, J. (2026). Invisible text injection and peer review by AI models. JAMA Network Open9(1), e2552099. https://doi.org/10.1001/jamanetworkopen.2025.52099

CIGI (Centre for International Governance Innovation. (2023). Building trust in AI: A landscape analysis of government AI programs. Retrieved from https://www.cigionline.org/publications/building-trust-in-ai-a-landscape-analysis-of-government-ai-programs/

Clark, J., Barton, B., Albarqouni, L., Byambasuren, O., Jowsey, T., Keogh, J., ... & Jones, M. (2025). Generative artificial intelligence use in evidence synthesis: A systematic review. Research Synthesis Methods, 1-19. https://doi.org/10.1017/rsm.2025.16

Cong-Lem, N. (2025). Rethinking evidence-informed policy and practice in the age of generative artificial intelligence. London Review of Education23(1), 1-8. https://doi.org/10.14324/LRE.23.1.16

Council of Europe. (2024). Framework convention on Artificial Intelligence. Retrieved from https://www.coe.int/en/web/artificial-intelligence/the-framework-convention-on-artificial-intelligence

Craglia, M., Hradec, J., & Troussard, X. (2020). The big data and artificial intelligence: Opportunities and challenges to modernise the policy cycle. In Science for policy handbook (pp. 96-103). https://doi.org/10.1016/B978-0-12-822596-7.00009-7

De Longueville, B., Sanchez, I., Kazakova, S., Luoni, S., Zaro, F., Daskalaki, K., & Inchingolo, M. (2025). The proof is in the eating: Lessons learnt from one year of generative AI adoption in a science-for-policy organisation. AI, 6(6), 128. https://doi.org/10.3390/ai6060128

Dobbins, M., Traynor, R., Clark, E., & Neil-Sztramko, S. (2024). Using artificial intelligence to support and streamline rapid systematic evidence reviews. European Journal of Public Health34(Supplement_3), ckae144-1053

Doskaliuk, B., Zimba, O., Yessirkepov, M., Klishch, I., & Yatsyshyn, R. (2025). Artificial intelligence in peer review: Enhancing efficiency while preserving integrity. Journal of Korean Medical Science, 40(7), e92. https://dx.doi.org/10.3346/jkms.2025.40.e92

Ebadi, S., Nejadghanbar, H., Salman, A. R., & Khosravi, H. (2025). Exploring the impact of generative AI on peer review: Insights from journal reviewers. Journal of Academic Ethics. https://doi.org/10.1007/s10805-025-09604-4

European Commission, Directorate-General for Digital Services. (2025). Analysis of the generative AI landscape in the European public sector, Publications Office of the European Union. https://data.europa.eu/doi/10.2799/0409819

European Commission, Eurostat (2024). An introduction to large language models and their relevance for statistical offices – 2024 edition. Publications Office of the European Union. https://data.europa.eu/doi/١٠.٢٧٨٥/٧١٦٢١٧

European Commission, Directorate-General for Communications Networks, Content and Technology & High-Level Expert Group on Artificial Intelligence. (2019). Ethics guidelines for trustworthy AI. Publications Office. https://data.europa.eu/doi/10.2759/346720

European Union Intellectual Property Office. (2025). The development of generative artificial intelligence from a copyright perspective. European Union Intellectual Property Office.  https://data.europa.eu/doi/10.2814/3893780

Farber, S. (2025). Comparing human and AI expertise in the academic peer review process: Towards a hybrid approach. Higher Education Research & Development, 44(4), 871-885. https://doi.org/10.1080/07294360.2024.2445575

Featherstone, R., Walter, M., MacDougall, D., Morenz, E., Bailey, S., Butcher, R., ... & Kaunelis, D. (2025). Artificial intelligence search tools for evidence synthesis: Comparative analysis and implementation recommendations. Cochrane Evidence Synthesis and Methods3(5), e70045. https://doi.org/10.1002/cesm.70045

Flemyng, E., Noel‐Storr, A., Macura, B., Gartlehner, G., Thomas, J., Meerpohl, J. J., ... & Grainger, M. (2025). Position statement on artificial intelligence (AI) use in evidence synthesis across Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence. Cochrane Database of Systematic Reviews, (10). https://doi.org/10.1186/s13750-025-00374-5

Gartlehner, G., Nussbaumer‐Streit, B., Hamel, C., Garritty, C., Griebler, U., King, V. J., ... & Kamel, C. (2025). Responsible integration of artificial intelligence in rapid reviews: A position statement from the Cochrane Rapid Reviews Methods Group. Cochrane Evidence Synthesis and Methods3(6), e70063. https://doi.org/10.1002/cesm.70063

Giray, L. (2024). Benefits and challenges of using AI for peer review: A study on researchers’ perceptions. Serials Librarian, 85(5), 144-154. https://doi.org/10.1080/0361526X.2024.2428377

Gwon, Y. N., Kim, J. H., Chung, H. S., Jung, E. J., Chun, J., Lee, S., & Shim, S. R. (2024). The use of generative AI for scientific literature searches for systematic reviews: ChatGPT and Microsoft Bing AI performance evaluation. JMIR Medical Informatics12, e51187. https://medinform.jmir.org/2024/1/e51187

GYA, IAP & ISC Scoping Group. (2023). The future of research evaluation: A synthesis of current debates and developments, discussion paper. Retrieved from https://www.interacademies.org/sites/default/files/2023-05/2023-05-11%2BEvaluation%2B-%2BWEB.pdf

Hoyt, R., Limon, A., & Chang, A. (2025). Generative AI and scientific manuscript peer review. Intelligence-Based Medicine, 11. https://doi.org/10.1016/j.ibmed.2025.100246

Hu, B., Tomini, E., Corrin, T., Pussegoda, K., Sandner, E., Henriques, A., ... & Waddell, L. (2025). Enhancing evidence synthesis efficiency: Leveraging large language models and agentic workflows for optimized literature screening. Cochrane Evidence Synthesis and Methods3(6), e70042. https://doi.org/10.1002/cesm.70042

International Science Council (2025). Types of AI and their use in science. https://doi.org/10.24948/2025.09

Iyer, R., Christie, A. P., Madhavapeddy, A., Reynolds, S., Sutherland, W., & Jaffer, S. (2025). Careful design of Large Language Model pipelines enables expert-level retrieval of evidence-based information from syntheses and databases. PLoS One20(5), e0323563. https://doi.org/10.1371/journal.pone.0323563

Joachim, M. V., Dodson, T. B., & Laviv, A. (2025). How artificial intelligence differs from humans in peer review. Journal of Oral and Maxillofacial Surgery: Official Journal of the American Association of Oral and Maxillofacial Surgeons. https://dx.doi.org/10.1016/j.joms.2025.03.015

Kadri, S. M., Dorri, N., Osaiweran, M., Garyali, P., & Petkovic, M. (2024). Scientific peer review in an era of artificial intelligence. In P. B. Joshi, P. P. Churi, & M. Pandey (Eds.), Scientific publishing ecosystem: An author-editor-reviewer axis (pp. 397-413). Singapore: Springer Nature. https://doi.org/10.1007/978-981-97-4060-4_23

Kankanhalli, A. (2024). Peer review in the age of generative AI. Journal of the Association for Information Systems, 25(1). https://doi.org/10.17705/1jais.00865

Koltsakis, E., Klontzas, M.E. & Karantanas, A.H. (2023). What is artificial intelligence: History and basic definitions. In Klontzas, M.E., Fanni, S.C., & Neri, E. (Eds.), Introduction to Artificial Intelligence. Springer-Verlag. https://doi.org/10.1007/978-3-031-25928-9_1

Kousha, K., & Thelwall, M. (2024). Artificial intelligence to support publishing and peer review: A summary and review. Learned Publishing, 37(1), 4-12. https://doi.org/10.1002/leap.1570

Lau, O., & Golder, S. (2025). Comparison of elicit AI and traditional literature searching in evidence syntheses using four case studies. Cochrane Evidence Synthesis and Methods3(6), e70050. https://doi.org/10.1002/cesm.70050

Lee, J., Lee, J., & Yoo, J.-J. (2025). The role of large language models in the peer-review process: Opportunities and challenges for medical journal reviewers and editors. Journal of Educational Evaluation for Health Professions, 22(101490061), 4. https://dx.doi.org/10.3352/jeehp.2025.22.4

Li, Z.-Q., Xu, H.-L., Cao, H.-J., Liu, Z.-L., Fei, Y.-T., & Liu, J.-P. (2024). Use of Artificial Intelligence in peer review among top 100 medical journals. JAMA Network Open, 7(12), e2448609. https://dx.doi.org/10.1001/jamanetworkopen.2024.48609

Liang, W., Izzo, Z., Zhang, Y., Lepp, H., Cao, H., Zhao, X., . . . Berkenkamp, F. (2024). Monitoring AI-modified content at scale: A case study on the impact of ChatGPT on AI conference peer reviews. Paper presented at the Proceedings of Machine Learning Research. https://doi.org/10.48550/arXiv.2403.07183

Lim, G. H., Tan, M. L., Hoe, V. C. W., & Koh, D. (2025). Generative AI in peer review process for occupational health. Occupational Medicine (Oxford, England). https://dx.doi.org/10.1093/occmed/kqaf051

Meeske, M., Cruijssen, F., van der Lee, C., & Krahmer, E. (2026). The role of language technology and artificial intelligence in food security policymaking. Discover Sustainability. 7, 34. https://doi.org/10.1007/s43621-025-02209-2

Medaglia, R., Mikalef, P., & Tangi, L. (2024). Competences and governance practices for artificial intelligence in the public sector. Publications Office of the European Union.

Ministry of Foreign Affairs of Japan. (2023). Hiroshima process international code of conduct for organizations developing advanced AI systems. Retrieved from https://www.mofa.go.jp/files/100573473.pdf

Namdarian, L. (2025). Towards AI-enabled evidence-based policymaking in science and technology: A conceptual framework. Technology Analysis & Strategic Management, 1-16. https://doi.org/10.1080/09537325.2025.2486003

Neshaei, S.P., Rietsche, R., Su, X. Wambsganss, T. (2024). Enhancing peer review with AI-powered suggestion generation assistance: Investigating the design dynamics. In 29th International Conference on Intelligent User Interfaces (IUI ’24), March 18–21, 2024, Greenville, SC, USA. New York: ACM. https://doi.org/10.1145/3640543.3645169

Nordmann, K., & Fischer, F. (2026). Harnessing ChatGPT for abstract screening in health-related scoping reviews: The role of structured eligibility criteria. BMC Health Services Research, 26(1), 109. https://doi.org/10.1186/s12913-025-13901-4

Nordmann, K., Schaller, M., Sauter, S., & Fischer, F. (2026). Capability of chatbots powered by large language models to support the screening process of scoping reviews: A feasibility study. JAMIA Open9(1), ooaf098. https://doi.org/10.1093/jamiaopen/ooaf098

OECD. (2023). Artificial intelligence in science: Challenges, opportunities and the future of research. OECD Publishing. Retrieved from https://doi.org/10.1787/a8d820bd-en

OECD. (2024a), Governing with Artificial Intelligence: Are governments ready? OECD Artificial Intelligence Papers, No. 20. OECD Publishing. https://doi.org/10.1787/26324bc2-en

OECD. (2025a), Governing with Artificial Intelligence: The state of play and way forward in core government functions. OECD Publishing. https://doi.org/10.1787/795de142-en

OECD. (n.d.). OECD AI principles overview. Retrieved from https://oecd.ai/en/ai-principles

OECD. (2025b). Progress in implementing the European Union Coordinated Plan on Artificial Intelligence (Volume 1): Member States’ actions. OECD Publishing. Paris, https://doi.org/10.1787/533c355d-en

OECD. (2024b). Revised recommendation of the Council on Artificial Intelligence. C/MIN(2024)16/FINA. Retrieved from https://one.oecd.org/document/C/MIN(2024)16/FINAL/en/pdf

Papadakis, T., Christou, I. T., Ipektsidis, C., Soldatos, J., & Amicone, A. (2024). Explainable and transparent artificial intelligence for public policymaking. Data & Policy6, e10. https://doi.org/10.1017/dap.2024.3

Pataranutaporn, P., Powdthavee, N. & Maes, P. (2025). Can AI solve the peer review crisis? A large-scale experiment on LLM’s performance and biases in evaluating economics papers. Retrieved from https://docs.iza.org/dp17659.pdf

Petrakis, P. E., Vasilis, G., & Kanzola, A. M. (2025). Utilizing artificial intelligence and new technologies for foresight. In The political economy of long-term planning: Using strategic analysis and foresight to support policymaking (pp. ٢٦٩-٢٧٨). Cham: Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-86437-7_12

Purificato, E., Bili, D., Jungnickel, R., Ruiz Serra, V., Fabiani, J. et al. (2025). The role of artificial intelligence in scientific research - A science for policy, European perspective. Publications Office of the European Union. https://data.europa.eu/doi/10.2760/7217497

Radeljić, B. (2025). Artificial intelligence and the algorithmic discursive sphere: Policymaking dilemmas and the rise of a new public intellectual. Global Society, 1-30. https://doi.org/10.1080/13600826.2025.2592708

Rieger, O.Y. & Schonfeld, R.C. (2023). Common scholarly communication infrastructure landscape review. ITHAKA. Retrieved from https://sr.ithaka.org/wp-content/uploads/2023/04/SR-Report-Common-Scholarly-Communication-Infrastructure-Landscape-Review042423.pdf

Rogers, K., Miller, A., Girgis, A., Clark, E. C., Neil-Sztramko, S. E., & Dobbins, M. (2025). Leveraging AI to optimize maintenance of health evidence and offer a one-stop shop for quality-appraised evidence syntheses on the effectiveness of public health interventions: Quality improvement project. Journal of Medical Internet Research27, e69700. https://www.jmir.org/2025/1/e69700

Rose, C. J., Meneses‐Echavez, J. F., Muller, A. E., Berg, R. C., Borge, T. C., Jardim, P. S. J., & Cooper, C. (2025). Artificial intelligence and machine learning to improve evidence synthesis production efficiency: An observational study of resource use and time‐to‐completion. Cochrane Evidence Synthesis and Methods3(3), e70030. https://doi.org/10.1002/cesm.70030

Ruan, M., Fan, J., Liu, M., Meng, Z., Zhang, X., & Zhang, C. (2025). Artificial intelligence for the science of evidence synthesis: How good are AI-powered tools for automatic literature screening? BMC Medical Research Methodology25(1), 199. https://doi.org/10.1186/s12874-025-02644-9

Safaei, M., & Longo, J. (2024). The end of the policy analyst? Testing the capability of artificial intelligence to generate plausible, persuasive, and useful policy analysis. Digital Government: Research and Practice, 5(1), 1-35. https://doi.org/10.1145/3604570

Salekpay, F., van den Bergh, J., & Savin, I. (2024). Comparing advice on climate policy between academic experts and ChatGPT. Ecological Economics226, 108352. https://doi.org/10.1016/j.ecolecon.2024.108352

Siemens, W., von Elm, E., Binder, H., Böhringer, D., Eisele-Metzger, A., Gartlehner, G., ... & Meerpohl, J. J. (2025). Opportunities, challenges and risks of using artificial intelligence for evidence synthesis. BMJ Evidence-Based Medicine. https://doi.org/10.1136/bmjebm-2024-113320

Salman, H. A., Ahmad, M. A., Ibrahim, R., & Mahmood, J. (2025). Systematic analysis of generative AI tools integration in academic research and peer review. Online Journal of Communication and Media Technologies, 15(1), e202502. https://doi.org/10.30935/ojcmt/15832

Science Europe. (n.d.). A European strategy for AI in science: Science Europe input to the European Commission’s Call for Evidence. Retrieved from https://www.scienceeurope.org/media/zcmn5hcg/science_europe_input_strategy_on_ai_in_science.pdf

Spillias, S., Tuohy, P., Andreotta, M., Annand-Jones, R., Boschetti, F., Cvitanovic, C., ... & Trebilco, R. (2024). Human-AI collaboration to identify literature for evidence synthesis. Cell Reports Sustainability1(7). https://doi.org/10.1016/j.crsus.2024.100132

Sun, L., Tao, S., Hu, J. & Dow, S. (2024). MetaWriter: Exploring the potential and perils of AI writing support in scientific peer review. Proc. ACM Hum.-Comput. Interact. 8, CSCW1, Article 94 (April 2024). https://doi.org/10.1145/3637371

Sun, Z. (2025). Large language models in peer review: Challenges and opportunities. Scientometrics, 1-44. https://doi.org/10.1007/s11192-025-05440-w

Tangi, L., Ulrich, P., Schade, S., & Manzoni, M. (2024). Taking stock and looking ahead: Developing a science for policy research agenda on the use and uptake of AI in public sector organisations in the EU. In Charalabidis, Y., Medaglia, R., & van Noordt, C. (Eds.), Research handbook on public management and artificial intelligence (pp. 208-225). Edward Elgar Publishing.

Thomas J, Flemyng E, Noel-Storr, A. et al. (2025a). Responsible use of AI in evidence SynthEsis (RAISE): recommendations for practice (version 2.2; updated 7 November 2025). In: Open Science Framework [https://osf.io/], Washington DC: Center for Open Science. http://doi.org/10.17605/OSF.IO/FWAUD

Thomas J, Flemyng E, Noel-Storr, A. et al. (2025b). Responsible use of AI in evidence SynthEsis (RAISE): Building and evaluating AI evidence synthesis tools (version 2.2; updated 7 November 2025). In: Open Science Framework [https://osf.io/], Washington DC: Center for Open Science. http://doi.org/10.17605/OSF.IO/FWAUD

Thomas J, Flemyng E, Noel-Storr, A. et al. (2025c). Responsible use of AI in evidence SynthEsis (RAISE): Selecting and using AI evidence synthesis tools (version 2.2; updated 7 November 2025). In: Open Science Framework [https://osf.io/], Washington DC: Center for Open Science. https://doi.org/10.17605/OSF.IO/FWAUD

Tomczyk, P., Brüggemann, P., Mergner, N., & Petrescu, M. (2024). Exploring AI’s role in literature searching: Traditional methods versus AI-based tools in analyzing topical e-commerce themes. In Digital Marketing & eCommerce Conference (pp. ١٤١-١٤٨). Springer, Cham. https://doi.org/١٠.١٠٠٧/٩٧٨-٣-٠٣١-٦٢١٣٥-٢_١٥

Turobov, A. (2025). Using Large Language Models responsibly in the UK civil service: A guide to implementation. Bennett Institute for Public Policy, University of Cambridge

Tyler, C., Akerlof, K. L., Allegra, A., Arnold, Z., Canino, H., Doornenbal, M. A., ... & Sutherland, W. J. (2023). AI tools as science policy advisers? The potential and the pitfalls. Nature622(7981), 27-30. https://doi.org/10.1038/d41586-023-02999-3

UK Government. (2025). Artificial Intelligence playbook for the UK Government. Retrieved from https://www.gov.uk/government/publications/ai-playbook-for-the-uk-government/artificial-intelligence-playbook-for-the-uk-government-html

UNESCO & OECD. (2024). G7 toolkit for artificial intelligence in the public sector: Report prepared for the 2024 Italian G7 presidency and the G7 digital and tech working group. Retrieved from https://unesdoc.unesco.org/ark:/48223/pf0000391566

UNESCO. (2022). Recommendation on the ethics of Artificial Intelligence. Retrieved from https://one.oecd.org/document/C/MIN(2024)16/FINAL/en/pdf

UNESCO. (2023). Ethical impact assessment: A tool of the Recommendation on the Ethics of Artificial Intelligence. https://doi.org/10.54678/YTSA7796

UNESCO. (2023). Readiness assessment methodology. Retrieved from https://unesdoc.unesco.org/ark:/48223/pf0000385198

Valizadeh, A., Moassefi, M., Nakhostin-Ansari, A., Hosseini Asl, S. H., Saghab Torbati, M., Aghajani, R., . . . Faghani, S. (2022). Abstract screening using the automated tool Rayyan: results of effectiveness in three diagnostic test accuracy systematic reviews. BMC Medical Research Methodology, 22(1). https://doi.org/10.1186/s12874-022-01631-8

Voutyrakou, D.A. & Skordoulis, C. (2025) Algorithmic governance: Gender bias in AI-generated policymaking? Human-Centric Intelligent Systems, 5, pp. 385–417. https://doi.org/10.1007/s44230-025-00109-2

Wang, Z., & Gong, M. (2026). A cross‐disciplinary analysis of AI policies in academic peer review. Learned Publishing, 39(1), e2035. https://doi.org/10.1002/leap.2035

World Economic Forum/OECD (2025), AI in strategic foresight: Reshaping anticipatory governance. The World Economic Forum. https://doi.org/10.1787/aa573076-en.

Yang, C.L., Uhde, A., Yamashita, N. & Kuzuoka. H. (2025). Understanding and supporting peer Review using AI-reframed positive summary. CHI Conference on Human Factors in Computing Systems (CHI ’25), April 26–May 01, 2025, Yokohama, Japan. New York: ACM. https://doi.org/10.1145/3706598.3713219

Yip, R., Sun, Y. J., Bassuk, A. G., & Mahajan, V. B. (2025). Artificial intelligence’s contribution to biomedical literature search: revolutionizing or complicating? PLOS Digital Health4(5), e0000849

Zhan, J., Suvada, K., Xu, M., Tian, W., Cara, K. C., Wallace, T. C., & Ali, M. K. (2025). Accelerating the pace and accuracy of systematic reviews using AI: A validation study. Systematic Reviews. 15 (24). https://doi.org/10.1186/s13643-025-02997-8

Zhang, G., Jin, Q., McInerney, D. J., Chen, Y., Wang, F., Cole, C. L., ... & Peng, Y. (2024). Leveraging generative AI for clinical evidence synthesis needs to ensure trustworthiness. Journal of Biomedical Informatics153, 104640. https://doi.org/10.1016/j.jbi.2024.104640

Zhao, K., Li, Y., Peng, X., Wang, C., Tew, Y. S., Fu, H., ... & Hu, S. (2026). Artificial intelligence-driven framework for science-policy interface on global plastic life cycle environmental impacts. Nexus3(1).

Zhao, W. & Mahmoud, Q.H. (2024). Evaluating the efficacy of large language models in automating academic peer reviews. 2024 International Conference on Machine Learning and Applications (ICMLA), Miami, pp. 1208-1213. https://doi.org/10.1109/ICMLA61862.2024.00187

Zhu, L., Lai, Y., Jiarui, X., Weiming, M., Lihaoyun, H., Chang, Q.,…& Peng, L. (2025). Evaluating the potential risks of employing large language models in peer review. Clinical and Translational Discovery 2025, 5,e70067. https://doi.org/10.1002/ctd2.70067

Ziegler, M., Lothian, S., O’Neill, B., Anderson, R. & Ota, Y. (2025). AI language models could both help and harm equity in marine policymaking. npj Ocean Sustainability 4, 32. https://doi.org/10.1038/s44183-025-00132-7

Zuo, C., Yang, X., Errickson, J., Li, J., Hong, Y., & Wang, R. (2025). AI-assisted evidence screening method for systematic reviews in environmental research: integrating ChatGPT with domain knowledge. Environmental Evidence, 14(1), 5. https://doi.org/10.1186/s13750-025-00358-5

Annex: Search strategy

This taxonomy is not meant to be prescriptive, but rather to help with the selection of the most appropriate ways to evaluate AI tools. The RAISE guidelines outline the following types or evaluations:

  • Performance evaluation that addresses how well AI performs
  • Process evaluation that addresses how well AI works in realistic settings
  • Human evaluation that addresses how the use of AI affects the human processes involved in evidence synthesis (Thomas et al., 2025b).

Choice of AI tools: regulatory issues. The RAISE guidelines state that “issues related to plagiarism, provenance, copyright, intellectual property, jurisdiction, licensing, confidentiality, compliance, and privacy responsibilities, including data protection laws” need to be considered (Thomas et al., 2025a, pp.13-14). Use of AI must adhere to the laws in the relevant jurisdictions. The guidelines suggest that AI systems should be chosen on the basis of having clear and transparent information about their processes and methods, ideally existing in the public domain; that they should have the option to disable data collection for training and turn off or delete history or interaction tracking; and that they should not gain any rights to users’ content (Thomas et al., 2025c). The RAISE guidelines cite Wiley’s advice to be wary of: “‘perpetual’ rights to your content; ‘royalty-free’ usage rights; ‘transferable’ rights to third parties; broad statements about using content ‘for training’ or ‘for any purpose’; classifying the AI-generated content as ‘open’, ‘free to use’, or subject to a permissive license or ‘Creative Commons’ license (which may open use of your work up to the rest of the world), or subject to a restrictive license; any limitations on your ability to use your content or the output; or unclear policies about data retention and deletion” (cited in (Thomas et al., 2025c, p.21). Another consideration is that there may be restrictions on the ‘export’ of services in some parts of the world, so, for example, a USA-based AI tool may not be used on a particular territory (Thomas et al., 2025c). There may also be considerations such as terms and conditions of an organisation’s insurance which would preclude it from using AI (Thomas et al., 2025c).

Limitations and biases. Evidence synthesis authors should consider and discuss any biases and other limitations of AI tools and their potential impact on the work (Thomas et al., 2025a).

Open science. Where possible and practical, the inputs (such as prompt development), outputs, datasets, or code used to develop, validate, and implement AI tools should be made publicly available (Thomas et al., 2025a).

Conflict of interests. Any financial or non-financial interests in the AI tool and any funding sources should be declared (Thomas et al., 2025a).

Environmental impacts. AI tools can use a large amount of energy and water and use of such resources can lead to significant carbon emissions, so it is important to consider the environmental impacts of using AI, where possible prioritise tools that report their carbon footprint transparently and make an effort to minimise it, and balance the environmental impacts against potential benefits including reducing research waste through increasing efficiency (Thomas et al., 2025c). The guidelines refer to the Green Algorithms project-1, which provides calculators and resources that can be used to estimate and reduce the carbon footprint of research projects.

Use of AI in the work of others. Evidence synthesis teams should consider that some of the published research papers that may be of relevance to their review work may have been fully or partially generated by AI (Thomas et al., 2025a).

Conclusions

The integration of AI into workflows relating to evidence review requires careful management, with strict adherence to responsible use and quality standards. This is even more vital within the context of policymaking, given that it affects people’s lives and therefore the stakes are higher and the risks are greater. The use of AI may be encouraged where it demonstrably improves the quality of reviews. At the same time, humans must always be fully responsible for verifying output and making judgement decisions. Transparency about the use of AI and justification for its use are essential. Compliance with the law is vital, and intellectual property rights must be respected.

References

Adel, A., & Alani, N. (2025). Can generative AI reliably synthesise literature? Exploring hallucination issues in ChatGPT. AI & Society, 1-14. https://doi.org/10.1007/s00146-025-02406-7

APEC (Asia-Pacific Economic Cooperation. (2022). Artificial intelligence in economic policymaking. Retrieved from https://www.apec.org/publications/2022/11/artificial-intelligence-in-economic-policymaking

Arora, A. et al. (2024). Towards intelligent governance: The role of AI in policymaking and decision support for e-governance. In So In, C., Londhe, N.D., Bhatt, N., & Kitsing, M. (Eds.), Information systems for Intelligent Systems. ISBM 2023. Smart Innovation, Systems and Technologies, 379. Springer, Singapore. https://doi.org/10.1007/978-981-99-8612-5_19

Bauchner, H., & Rivara, F. P. (2024). Use of artificial intelligence and the future of peer review. Health Affairs Scholar, 2(5), qxae058. https://dx.doi.org/10.1093/haschl/qxae058

Ben Saad, H., Dergaa, I., Ghouili, H., Ceylan, H. I., Chamari, K., & Dhahbi, W. (2025). The assisted technology dilemma: A reflection on AI chatbots use and risks while reshaping the peer review process in scientific research. AI & Society. https://doi.org/10.1007/s00146-025-02299-6

Bergstrom, T. & Ruediger, D. (2024). A third transformation? Generative AI and scholarly publishing. Retrieved from https://sr.ithaka.org/wp-content/uploads/2024/10/SR-Brief-Generative-AI-and-Scholarly-Publishing-103024.pdf

Canali, S., & Barone-Adesi, F. (2023). Can AI deliver advice that is judgement-free for science policy?. Nature, 624(7991), 252-252. Retrieved from https://www.nature.com/articles/d41586-023-03949-9

Choi, B., Jun, T. J., Sung, J. W., Park, I. W., Lee, J. M., Cho, S. I., ... & Suh, J. (2026). Invisible text injection and peer review by AI models. JAMA Network Open, 9(1), e2552099. https://doi.org/10.1001/jamanetworkopen.2025.52099

CIGI (Centre for International Governance Innovation. (2023). Building trust in AI: A landscape analysis of government AI programs. Retrieved from https://www.cigionline.org/publications/building-trust-in-ai-a-landscape-analysis-of-government-ai-programs/

Clark, J., Barton, B., Albarqouni, L., Byambasuren, O., Jowsey, T., Keogh, J., ... & Jones, M. (2025). Generative artificial intelligence use in evidence synthesis: A systematic review. Research Synthesis Methods, 1-19. https://doi.org/10.1017/rsm.2025.16

Cong-Lem, N. (2025). Rethinking evidence-informed policy and practice in the age of generative artificial intelligence. London Review of Education23(1), 1-8. https://doi.org/10.14324/LRE.23.1.16

Council of Europe. (2024). Framework convention on Artificial Intelligence. Retrieved from https://www.coe.int/en/web/artificial-intelligence/the-framework-convention-on-artificial-intelligence

Craglia, M., Hradec, J., & Troussard, X. (2020). The big data and artificial intelligence: Opportunities and challenges to modernise the policy cycle. In Science for policy handbook (pp. 96-103). https://doi.org/10.1016/B978-0-12-822596-7.00009-7

De Longueville, B., Sanchez, I., Kazakova, S., Luoni, S., Zaro, F., Daskalaki, K., & Inchingolo, M. (2025). The proof is in the eating: Lessons learnt from one year of generative AI adoption in a science-for-policy organisation. AI, 6(6), 128. https://doi.org/10.3390/ai6060128

Dobbins, M., Traynor, R., Clark, E., & Neil-Sztramko, S. (2024). Using artificial intelligence to support and streamline rapid systematic evidence reviews. European Journal of Public Health, 34(Supplement_3), ckae144-1053

Doskaliuk, B., Zimba, O., Yessirkepov, M., Klishch, I., & Yatsyshyn, R. (2025). Artificial intelligence in peer review: Enhancing efficiency while preserving integrity. Journal of Korean Medical Science, 40(7), e92. https://dx.doi.org/10.3346/jkms.2025.40.e92

Ebadi, S., Nejadghanbar, H., Salman, A. R., & Khosravi, H. (2025). Exploring the impact of generative AI on peer review: Insights from journal reviewers. Journal of Academic Ethics. https://doi.org/10.1007/s10805-025-09604-4

European Commission, Directorate-General for Digital Services. (2025). Analysis of the generative AI landscape in the European public sector, Publications Office of the European Union. https://data.europa.eu/doi/10.2799/0409819

European Commission, Eurostat (2024). An introduction to large language models and their relevance for statistical offices – 2024 edition. Publications Office of the European Union. https://data.europa.eu/doi/١٠.٢٧٨٥/٧١٦٢١٧

European Commission, Directorate-General for Communications Networks, Content and Technology & High-Level Expert Group on Artificial Intelligence. (2019). Ethics guidelines for trustworthy AI. Publications Office. https://data.europa.eu/doi/10.2759/346720

European Union Intellectual Property Office. (2025). The development of generative artificial intelligence from a copyright perspective. European Union Intellectual Property Office.  https://data.europa.eu/doi/10.2814/3893780

Farber, S. (2025). Comparing human and AI expertise in the academic peer review process: Towards a hybrid approach. Higher Education Research & Development, 44(4), 871-885. https://doi.org/10.1080/07294360.2024.2445575

Featherstone, R., Walter, M., MacDougall, D., Morenz, E., Bailey, S., Butcher, R., ... & Kaunelis, D. (2025). Artificial intelligence search tools for evidence synthesis: Comparative analysis and implementation recommendations. Cochrane Evidence Synthesis and Methods3(5), e70045. https://doi.org/10.1002/cesm.70045

Flemyng, E., Noel‐Storr, A., Macura, B., Gartlehner, G., Thomas, J., Meerpohl, J. J., ... & Grainger, M. (2025). Position statement on artificial intelligence (AI) use in evidence synthesis across Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence. Cochrane Database of Systematic Reviews, (10). https://doi.org/10.1186/s13750-025-00374-5

Gartlehner, G., Nussbaumer‐Streit, B., Hamel, C., Garritty, C., Griebler, U., King, V. J., ... & Kamel, C. (2025). Responsible integration of artificial intelligence in rapid reviews: A position statement from the Cochrane Rapid Reviews Methods Group. Cochrane Evidence Synthesis and Methods3(6), e70063. https://doi.org/10.1002/cesm.70063

Giray, L. (2024). Benefits and challenges of using AI for peer review: A study on researchers’ perceptions. Serials Librarian, 85(5), 144-154. https://doi.org/10.1080/0361526X.2024.2428377

Gwon, Y. N., Kim, J. H., Chung, H. S., Jung, E. J., Chun, J., Lee, S., & Shim, S. R. (2024). The use of generative AI for scientific literature searches for systematic reviews: ChatGPT and Microsoft Bing AI performance evaluation. JMIR Medical Informatics12, e51187. https://medinform.jmir.org/2024/1/e51187

GYA, IAP & ISC Scoping Group. (2023). The future of research evaluation: A synthesis of current debates and developments, discussion paper. Retrieved from https://www.interacademies.org/sites/default/files/2023-05/2023-05-11%2BEvaluation%2B-%2BWEB.pdf

Hoyt, R., Limon, A., & Chang, A. (2025). Generative AI and scientific manuscript peer review. Intelligence-Based Medicine, 11. https://doi.org/10.1016/j.ibmed.2025.100246

Hu, B., Tomini, E., Corrin, T., Pussegoda, K., Sandner, E., Henriques, A., ... & Waddell, L. (2025). Enhancing evidence synthesis efficiency: Leveraging large language models and agentic workflows for optimized literature screening. Cochrane Evidence Synthesis and Methods3(6), e70042. https://doi.org/10.1002/cesm.70042

International Science Council (2025). Types of AI and their use in science. https://doi.org/10.24948/2025.09

Iyer, R., Christie, A. P., Madhavapeddy, A., Reynolds, S., Sutherland, W., & Jaffer, S. (2025). Careful design of Large Language Model pipelines enables expert-level retrieval of evidence-based information from syntheses and databases. PLoS One20(5), e0323563. https://doi.org/10.1371/journal.pone.0323563

Joachim, M. V., Dodson, T. B., & Laviv, A. (2025). How artificial intelligence differs from humans in peer review. Journal of Oral and Maxillofacial Surgery: Official Journal of the American Association of Oral and Maxillofacial Surgeons. https://dx.doi.org/10.1016/j.joms.2025.03.015

Kadri, S. M., Dorri, N., Osaiweran, M., Garyali, P., & Petkovic, M. (2024). Scientific peer review in an era of artificial intelligence. In P. B. Joshi, P. P. Churi, & M. Pandey (Eds.), Scientific publishing ecosystem: An author-editor-reviewer axis (pp. 397-413). Singapore: Springer Nature. https://doi.org/10.1007/978-981-97-4060-4_23

Kankanhalli, A. (2024). Peer review in the age of generative AI. Journal of the Association for Information Systems, 25(1). https://doi.org/10.17705/1jais.00865

Koltsakis, E., Klontzas, M.E. & Karantanas, A.H. (2023). What is artificial intelligence: History and basic definitions. In Klontzas, M.E., Fanni, S.C., & Neri, E. (Eds.), Introduction to Artificial Intelligence. Springer-Verlag. https://doi.org/10.1007/978-3-031-25928-9_1

Kousha, K., & Thelwall, M. (2024). Artificial intelligence to support publishing and peer review: A summary and review. Learned Publishing, 37(1), 4-12. https://doi.org/10.1002/leap.1570

Lau, O., & Golder, S. (2025). Comparison of elicit AI and traditional literature searching in evidence syntheses using four case studies. Cochrane Evidence Synthesis and Methods3(6), e70050. https://doi.org/10.1002/cesm.70050

Lee, J., Lee, J., & Yoo, J.-J. (2025). The role of large language models in the peer-review process: Opportunities and challenges for medical journal reviewers and editors. Journal of Educational Evaluation for Health Professions, 22(101490061), 4. https://dx.doi.org/10.3352/jeehp.2025.22.4

Li, Z.-Q., Xu, H.-L., Cao, H.-J., Liu, Z.-L., Fei, Y.-T., & Liu, J.-P. (2024). Use of Artificial Intelligence in peer review among top 100 medical journals. JAMA Network Open, 7(12), e2448609. https://dx.doi.org/10.1001/jamanetworkopen.2024.48609

Liang, W., Izzo, Z., Zhang, Y., Lepp, H., Cao, H., Zhao, X., . . . Berkenkamp, F. (2024). Monitoring AI-modified content at scale: A case study on the impact of ChatGPT on AI conference peer reviews. Paper presented at the Proceedings of Machine Learning Research. https://doi.org/10.48550/arXiv.2403.07183

Lim, G. H., Tan, M. L., Hoe, V. C. W., & Koh, D. (2025). Generative AI in peer review process for occupational health. Occupational Medicine (Oxford, England). https://dx.doi.org/10.1093/occmed/kqaf051

Meeske, M., Cruijssen, F., van der Lee, C., & Krahmer, E. (2026). The role of language technology and artificial intelligence in food security policymaking. Discover Sustainability. 7, 34. https://doi.org/10.1007/s43621-025-02209-2

Medaglia, R., Mikalef, P., & Tangi, L. (2024). Competences and governance practices for artificial intelligence in the public sector. Publications Office of the European Union.

Ministry of Foreign Affairs of Japan. (2023). Hiroshima process international code of conduct for organizations developing advanced AI systems. Retrieved from https://www.mofa.go.jp/files/100573473.pdf

Namdarian, L. (2025). Towards AI-enabled evidence-based policymaking in science and technology: A conceptual framework. Technology Analysis & Strategic Management, 1-16. https://doi.org/10.1080/09537325.2025.2486003

Neshaei, S.P., Rietsche, R., Su, X. Wambsganss, T. (2024). Enhancing peer review with AI-powered suggestion generation assistance: Investigating the design dynamics. In 29th International Conference on Intelligent User Interfaces (IUI ’24), March 18–21, 2024, Greenville, SC, USA. New York: ACM. https://doi.org/10.1145/3640543.3645169

Nordmann, K., & Fischer, F. (2026). Harnessing ChatGPT for abstract screening in health-related scoping reviews: The role of structured eligibility criteria. BMC Health Services Research, 26(1), 109. https://doi.org/10.1186/s12913-025-13901-4

Nordmann, K., Schaller, M., Sauter, S., & Fischer, F. (2026). Capability of chatbots powered by large language models to support the screening process of scoping reviews: A feasibility study. JAMIA Open9(1), ooaf098. https://doi.org/10.1093/jamiaopen/ooaf098

OECD. (2023). Artificial intelligence in science: Challenges, opportunities and the future of research. OECD Publishing. Retrieved from https://doi.org/10.1787/a8d820bd-en

OECD. (2024a), Governing with Artificial Intelligence: Are governments ready? OECD Artificial Intelligence Papers, No. 20. OECD Publishing. https://doi.org/10.1787/26324bc2-en

OECD. (2025a), Governing with Artificial Intelligence: The state of play and way forward in core government functions. OECD Publishing. https://doi.org/10.1787/795de142-en

OECD. (n.d.). OECD AI principles overview. Retrieved from https://oecd.ai/en/ai-principles

OECD. (2025b). Progress in implementing the European Union Coordinated Plan on Artificial Intelligence (Volume 1): Member States’ actions. OECD Publishing. Paris, https://doi.org/10.1787/533c355d-en

OECD. (2024b). Revised recommendation of the Council on Artificial Intelligence. C/MIN(2024)16/FINA. Retrieved from https://one.oecd.org/document/C/MIN(2024)16/FINAL/en/pdf

Papadakis, T., Christou, I. T., Ipektsidis, C., Soldatos, J., & Amicone, A. (2024). Explainable and transparent artificial intelligence for public policymaking. Data & Policy, 6, e10. https://doi.org/10.1017/dap.2024.3

Pataranutaporn, P., Powdthavee, N. & Maes, P. (2025). Can AI solve the peer review crisis? A large-scale experiment on LLM’s performance and biases in evaluating economics papers. Retrieved from https://docs.iza.org/dp17659.pdf

Petrakis, P. E., Vasilis, G., & Kanzola, A. M. (2025). Utilizing artificial intelligence and new technologies for foresight. In The political economy of long-term planning: Using strategic analysis and foresight to support policymaking (pp. ٢٦٩-٢٧٨). Cham: Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-86437-7_12

Purificato, E., Bili, D., Jungnickel, R., Ruiz Serra, V., Fabiani, J. et al. (2025). The role of artificial intelligence in scientific research - A science for policy, European perspective. Publications Office of the European Union. https://data.europa.eu/doi/10.2760/7217497

Radeljić, B. (2025). Artificial intelligence and the algorithmic discursive sphere: Policymaking dilemmas and the rise of a new public intellectual. Global Society, 1-30. https://doi.org/10.1080/13600826.2025.2592708

Rieger, O.Y. & Schonfeld, R.C. (2023). Common scholarly communication infrastructure landscape review. ITHAKA. Retrieved from https://sr.ithaka.org/wp-content/uploads/2023/04/SR-Report-Common-Scholarly-Communication-Infrastructure-Landscape-Review042423.pdf

Rogers, K., Miller, A., Girgis, A., Clark, E. C., Neil-Sztramko, S. E., & Dobbins, M. (2025). Leveraging AI to optimize maintenance of health evidence and offer a one-stop shop for quality-appraised evidence syntheses on the effectiveness of public health interventions: Quality improvement project. Journal of Medical Internet Research, 27, e69700. https://www.jmir.org/2025/1/e69700

Rose, C. J., Meneses‐Echavez, J. F., Muller, A. E., Berg, R. C., Borge, T. C., Jardim, P. S. J., & Cooper, C. (2025). Artificial intelligence and machine learning to improve evidence synthesis production efficiency: An observational study of resource use and time‐to‐completion. Cochrane Evidence Synthesis and Methods, 3(3), e70030. https://doi.org/10.1002/cesm.70030

Ruan, M., Fan, J., Liu, M., Meng, Z., Zhang, X., & Zhang, C. (2025). Artificial intelligence for the science of evidence synthesis: How good are AI-powered tools for automatic literature screening? BMC Medical Research Methodology, 25(1), 199. https://doi.org/10.1186/s12874-025-02644-9

Safaei, M., & Longo, J. (2024). The end of the policy analyst? Testing the capability of artificial intelligence to generate plausible, persuasive, and useful policy analysis. Digital Government: Research and Practice, 5(1), 1-35. https://doi.org/10.1145/3604570

Salekpay, F., van den Bergh, J., & Savin, I. (2024). Comparing advice on climate policy between academic experts and ChatGPT. Ecological Economics, 226, 108352. https://doi.org/10.1016/j.ecolecon.2024.108352

Siemens, W., von Elm, E., Binder, H., Böhringer, D., Eisele-Metzger, A., Gartlehner, G., ... & Meerpohl, J. J. (2025). Opportunities, challenges and risks of using artificial intelligence for evidence synthesis. BMJ Evidence-Based Medicine. https://doi.org/10.1136/bmjebm-2024-113320

Salman, H. A., Ahmad, M. A., Ibrahim, R., & Mahmood, J. (2025). Systematic analysis of generative AI tools integration in academic research and peer review. Online Journal of Communication and Media Technologies, 15(1), e202502. https://doi.org/10.30935/ojcmt/15832

Science Europe. (n.d.). A European strategy for AI in science: Science Europe input to the European Commission’s Call for Evidence. Retrieved from https://www.scienceeurope.org/media/zcmn5hcg/science_europe_input_strategy_on_ai_in_science.pdf

Spillias, S., Tuohy, P., Andreotta, M., Annand-Jones, R., Boschetti, F., Cvitanovic, C., ... & Trebilco, R. (2024). Human-AI collaboration to identify literature for evidence synthesis. Cell Reports Sustainability1(7). https://doi.org/10.1016/j.crsus.2024.100132

Sun, L., Tao, S., Hu, J. & Dow, S. (2024). MetaWriter: Exploring the potential and perils of AI writing support in scientific peer review. Proc. ACM Hum.-Comput. Interact. 8, CSCW1, Article 94 (April 2024). https://doi.org/10.1145/3637371

Sun, Z. (2025). Large language models in peer review: Challenges and opportunities. Scientometrics, 1-44. https://doi.org/10.1007/s11192-025-05440-w

Tangi, L., Ulrich, P., Schade, S., & Manzoni, M. (2024). Taking stock and looking ahead: Developing a science for policy research agenda on the use and uptake of AI in public sector organisations in the EU. In Charalabidis, Y., Medaglia, R., & van Noordt, C. (Eds.), Research handbook on public management and artificial intelligence (pp. 208-225). Edward Elgar Publishing.

Thomas J, Flemyng E, Noel-Storr, A. et al. (2025a). Responsible use of AI in evidence SynthEsis (RAISE): recommendations for practice (version 2.2; updated 7 November 2025). In: Open Science Framework [https://osf.io/], Washington DC: Center for Open Science. http://doi.org/10.17605/OSF.IO/FWAUD

Thomas J, Flemyng E, Noel-Storr, A. et al. (2025b). Responsible use of AI in evidence SynthEsis (RAISE): Building and evaluating AI evidence synthesis tools (version 2.2; updated 7 November 2025). In: Open Science Framework [https://osf.io/], Washington DC: Center for Open Science. http://doi.org/10.17605/OSF.IO/FWAUD

Thomas J, Flemyng E, Noel-Storr, A. et al. (2025c). Responsible use of AI in evidence SynthEsis (RAISE): Selecting and using AI evidence synthesis tools (version 2.2; updated 7 November 2025). In: Open Science Framework [https://osf.io/], Washington DC: Center for Open Science. https://doi.org/10.17605/OSF.IO/FWAUD

Tomczyk, P., Brüggemann, P., Mergner, N., & Petrescu, M. (2024). Exploring AI’s role in literature searching: Traditional methods versus AI-based tools in analyzing topical e-commerce themes. In Digital Marketing & eCommerce Conference (pp. ١٤١-١٤٨). Springer, Cham. https://doi.org/١٠.١٠٠٧/٩٧٨-٣-٠٣١-٦٢١٣٥-٢_١٥

Turobov, A. (2025). Using Large Language Models responsibly in the UK civil service: A guide to implementation. Bennett Institute for Public Policy, University of Cambridge

Tyler, C., Akerlof, K. L., Allegra, A., Arnold, Z., Canino, H., Doornenbal, M. A., ... & Sutherland, W. J. (2023). AI tools as science policy advisers? The potential and the pitfalls. Nature, 622(7981), 27-30. https://doi.org/10.1038/d41586-023-02999-3

UK Government. (2025). Artificial Intelligence playbook for the UK Government. Retrieved from https://www.gov.uk/government/publications/ai-playbook-for-the-uk-government/artificial-intelligence-playbook-for-the-uk-government-html

UNESCO & OECD. (2024). G7 toolkit for artificial intelligence in the public sector: Report prepared for the 2024 Italian G7 presidency and the G7 digital and tech working group. Retrieved from https://unesdoc.unesco.org/ark:/48223/pf0000391566

UNESCO. (2022). Recommendation on the ethics of Artificial Intelligence. Retrieved from https://one.oecd.org/document/C/MIN(2024)16/FINAL/en/pdf

UNESCO. (2023). Ethical impact assessment: A tool of the Recommendation on the Ethics of Artificial Intelligence. https://doi.org/10.54678/YTSA7796

UNESCO. (2023). Readiness assessment methodology. Retrieved from https://unesdoc.unesco.org/ark:/48223/pf0000385198

Valizadeh, A., Moassefi, M., Nakhostin-Ansari, A., Hosseini Asl, S. H., Saghab Torbati, M., Aghajani, R., . . . Faghani, S. (2022). Abstract screening using the automated tool Rayyan: results of effectiveness in three diagnostic test accuracy systematic reviews. BMC Medical Research Methodology, 22(1). https://doi.org/10.1186/s12874-022-01631-8

Voutyrakou, D.A. & Skordoulis, C. (2025) Algorithmic governance: Gender bias in AI-generated policymaking? Human-Centric Intelligent Systems, 5, pp. 385–417. https://doi.org/10.1007/s44230-025-00109-2

Wang, Z., & Gong, M. (2026). A cross‐disciplinary analysis of AI policies in academic peer review. Learned Publishing, 39(1), e2035. https://doi.org/10.1002/leap.2035

World Economic Forum/OECD (2025), AI in strategic foresight: Reshaping anticipatory governance. The World Economic Forum. https://doi.org/10.1787/aa573076-en.

Yang, C.L., Uhde, A., Yamashita, N. & Kuzuoka. H. (2025). Understanding and supporting peer Review using AI-reframed positive summary. CHI Conference on Human Factors in Computing Systems (CHI ’25), April 26–May 01, 2025, Yokohama, Japan. New York: ACM. https://doi.org/10.1145/3706598.3713219

Yip, R., Sun, Y. J., Bassuk, A. G., & Mahajan, V. B. (2025). Artificial intelligence’s contribution to biomedical literature search: revolutionizing or complicating? PLOS Digital Health4(5), e0000849

Zhan, J., Suvada, K., Xu, M., Tian, W., Cara, K. C., Wallace, T. C., & Ali, M. K. (2025). Accelerating the pace and accuracy of systematic reviews using AI: A validation study. Systematic Reviews. 15 (24). https://doi.org/10.1186/s13643-025-02997-8

Zhang, G., Jin, Q., McInerney, D. J., Chen, Y., Wang, F., Cole, C. L., ... & Peng, Y. (2024). Leveraging generative AI for clinical evidence synthesis needs to ensure trustworthiness. Journal of Biomedical Informatics153, 104640. https://doi.org/10.1016/j.jbi.2024.104640

Zhao, K., Li, Y., Peng, X., Wang, C., Tew, Y. S., Fu, H., ... & Hu, S. (2026). Artificial intelligence-driven framework for science-policy interface on global plastic life cycle environmental impacts. Nexus, 3(1).

Zhao, W. & Mahmoud, Q.H. (2024). Evaluating the efficacy of large language models in automating academic peer reviews. 2024 International Conference on Machine Learning and Applications (ICMLA), Miami, pp. 1208-1213. https://doi.org/10.1109/ICMLA61862.2024.00187

Zhu, L., Lai, Y., Jiarui, X., Weiming, M., Lihaoyun, H., Chang, Q.,…& Peng, L. (2025). Evaluating the potential risks of employing large language models in peer review. Clinical and Translational Discovery 2025, 5,e70067. https://doi.org/10.1002/ctd2.70067

Ziegler, M., Lothian, S., O’Neill, B., Anderson, R. & Ota, Y. (2025). AI language models could both help and harm equity in marine policymaking. npj Ocean Sustainability 4, 32. https://doi.org/10.1038/s44183-025-00132-7

Zuo, C., Yang, X., Errickson, J., Li, J., Hong, Y., & Wang, R. (2025). AI-assisted evidence screening method for systematic reviews in environmental research: integrating ChatGPT with domain knowledge. Environmental Evidence, 14(1), 5. https://doi.org/10.1186/s13750-025-00358-5

Annex: Search strategy

Source 

Search interface if relevant  (e.g OVID/Proquest) 

Search date 

Searcher 

Search strategy1 

Papers retrieved2  

Notes  

Scopus

N/A

11/04/2025

MK

( TITLE-ABS-KEY ( ai OR “artificial intelligence” OR “machine learning” OR “deep learning” OR “large language model*” ) ) AND ( TITLE-ABS-KEY ( ( advic* OR advis* ) W/5 ( policy OR policies OR scien* OR politician* OR government* ) ) ) AND PUBYEAR > 2021 AND PUBYEAR < 2026 AND ( EXCLUDE ( SRCTYPE , “p” ) OR EXCLUDE ( SRCTYPE , “b” ) OR EXCLUDE ( SRCTYPE , “k” ) )

94

All downloaded and deduplicated against the WoS and ProQuest searches conducted the same day.

Web of Science

N/A

11/04/2025

MK

1 TS=(ai OR “artificial intelligence” OR “machine learning” OR “deep learning” OR “large language model*”)

2 TS=((advic* OR advis*) NEAR/5 (policy OR policies OR scien* OR politician* OR government*))

3 #1 AND #2 Timespan: 2022-01-01 to 2025-12-31 (Publication Date)

75

All downloaded and deduplicated against the Scopus and ProQuest searches conducted the same day.

Social Science Premium Collection

ProQuest

11/04/2025

MK

noft(ai OR “artificial intelligence” OR “machine learning” OR “deep learning” OR “large language model*”) AND noft((advic* OR advis*) NEAR/5 (policy OR policies OR scien* OR politician* OR government*))

Additional limits - Date: After 01 January 2022

39

All downloaded and deduplicated against the Scopus and WoS searches conducted the same day, resulting in 140 unique records. 8 selected at title and abstract for further consideration. 6 were included.

WoS 

N/A 

18/07/2025 

MK 

1 TI=(ai OR “artificial intelligence” OR “machine learning” OR “deep learning” OR “large language model*”) 

2 TI=(“peer review*”) 

3 #2 AND #1 

4 #2 AND #1 | Timespan: 2024-01-01 to 2025-12-31 (Publication Date) 

65 

All downloaded and deduplicated against the Scopus and Medline searches 

Scopus 

N/A 

18/07/2025 

MK 

(TITLE (( peer review*) ) AND (TITLE (ai OR (artificial intelligence) OR (machine learning) OR (deep learning) OR (large language model*)) AND PUBYEAR > 2023 AND PUBYEAR < 2026 

90 

All downloaded and deduplicated against the WoS and Medline searches 

MEDLINE(R) ALL 

Ovid 

18/07/2025 

MK 

1. (ai or “artificial intelligence” or “machine learning” or “deep learning” or “large language model*”).ti.133271 

2.”peer review*”.ti.6549 

3.1 and 2 80 

4.limit 3 to english language 80 

5.limit 4 to yr=”2024 -Current” 46 

46 

All downloaded and deduplicated against the Scopus and WoS searches. Auto deduplication on Rayyan at 98% similarity used first, which removed 37 articles. The rest of the duplicates were resolved manually, resulting in 103 unique articles. 

Overton 

 

22/07/25 

LE 

(AI OR “artificial intelligence” AND “peer review”. Limit to post-2020 

 

15 reports/blogs downloaded for review

LibSearch 

 

12/8/25 

LE 

Title: synthes* OR summar* OR writ* OR condens* AND title: academ* OR research* OR scient* OR scholar* AND title: AI OR “artificial intelligence” OR “machine learning” OR LLM* OR “ChatGPT” OR CoPilot OR Gemini 

1,118 

204 articles marked for further review, based on metadata (title, source, date) 

 

Scopus 

N/A 

22/08/2025 

MK 

TITLE ( ( expert* OR reviewer* OR researcher* OR academic* OR scientist* OR specialist* ) W/3 ( find* OR identif* OR select* OR choos* OR choice* OR pick* OR assign* OR allocat* ) ) AND TITLE ( ai OR ( artificial intelligence ) OR ( machine learning ) OR ( deep learning ) OR ( large LANGUAGE model* ) ) AND PUBYEAR > 2023 AND PUBYEAR < 2026 

26 

All exported and manually deduplicated against the WoS search for literature on expert identification in Rayyan, resulting in 31 unique articles to screen 

WoS 

N/A 

22/08/2025 

MK 

1 TI=((expert* OR reviewer* OR researcher* OR academic* OR scientist* OR specialist*) NEAR/3 (find* OR identif* OR select* OR choos* OR choice* OR pick* OR assign* OR allocat*)) 

2 TI=(ai OR (artificial intelligence) OR (machine learning) OR (deep learning) OR ( large language model*)) 

3 #1 AND #2 

4 #1 AND #2 | Timespan: 2024-01-01 to 2025-12-31 (Publication Date) 

16 

All exported and manually deduplicated against the Scopus search for literature on expert identification in Rayyan, resulting in 31 unique articles to screen 

LibSearch 

 

14/1/26 

LE 

Title: AI OR “artificial intelligence” OR LLM* OR “large language” OR “deep learning” OR “machine learning” AND Title: policymak* OR “policy mak*”8 

 

62 

8 downloaded for review 

LibSearch 

 

14/1/26 

LE 

Title: AI OR “artificial intelligence” OR LLM* OR “large language” OR “deep learning” OR “machine learning” AND title: evidence AND title: review* OR synthes* 

288 

44 downloaded for review 

LibSearch 

 

14/1/26 

LE 

Title: AI OR “artificial intelligence” OR LLM* OR “large language” OR “deep learning” OR “machine learning” AND title: “science for policy” OR “science advice” OR “scientific advice” 

 

6 downloaded for review 

LibSearch

22/1/26

LE

Title: literature AND (synthes* OR summar*) AND Title: AI OR “artificial intelligence” OR “machine learning” OR chatgpt* OR LLM* OR “copilot”. Last 2 years only.

3 downloaded for review

LibSearch

22/1/26

LE

Title: literature AND search* AND Title: AI OR “artificial intelligence” OR “machine learning” OR chatgpt* OR LLM* OR “copilot”. Last 2 years only

4 downloaded for review

Table of contents