That was the central question of a recent webinar organised by the Scientific Advice Mechanism, drawing on two freshly published reports: a Rapid Evidence Review Report by SAPEA and a Statement by the Group of Chief Scientific Advisors (GCSA). The webinar panel brought together members and contributors to the SAPEA working group, two members of the Group of Chief Scientific Advisors (GCSA), and a French firefighting officer, offering an interesting blend of academic expertise and practical experience.
The Goldilocks dilemma
Professor Tina Comes (German Aerospace Center & TU Delft), who chaired the SAPEA working group, explained that, when the working group started, the dominant narrative in practitioner and conference circles was essentially: How do we make people trust AI so they will use it?, as if unlocking trust would automatically unlock AI’s vast potential. “There are risks of over-trusting AI if we use it like an autopilot,” Comes warned, “blindly accepting something that AI is stating, especially if it sounds very credible and plausible or if we are under a lot of stress.” She called this the “Goldilocks dilemma”: too little trust means useful tools go unused, but too much trust can be dangerous.
Comes also stressed that there is no single “AI.” The field spans everything from ChatGPT to specialised flood-prediction platforms like Google’s Flood Hub, and each application demands its own assessment. Rather than evaluating whether a specific tool is good or bad right now, the working group focused on principles and guidelines, given how rapidly the technology is evolving. She insisted, “AI has to serve us and should not become an objective in and of itself.”
EU AI Act Guardrails
Professor Andrej Zwitter (University of Klagenfurt), a member of the working group, walked the audience through the European regulatory landscape. For crisis management, the first question is whether the AI will serve a military or national security purpose. If so, the AI Act does not apply, while it mostly does for civil protection mechanisms.
The EU AI Act classifies AI systems by risk level, from minimal (spam filters) to prohibited (social scoring). Many AI tools for crisis management can be classified as ‘high’ risk, which carries a heightened challenge to their deployment, including risk management and requirements on data quality, transparency, explainability and human oversight.
In crisis situations, one of the critical tasks is triage: deciding who or what gets priority when dispatching emergency responders such as firefighters or medical teams. Under the EU AI Act, any AI system used to make or support these prioritisation decisions is explicitly classified as high risk. This means it must meet strict requirements: risk management throughout its entire life cycle, high-quality data, thorough technical documentation, transparency, explainability, and human oversight.
He also pointed to a complication: general-purpose AI tools like ChatGPT, not designed for crisis use, may still be deployed in emergencies, and they carry their own obligations. If powerful enough, they fall under even stricter systemic risk rules, meaning they are subject to stricter rules designed to ensure these powerful models are transparent, well-documented, and properly disclosed to regulators and users.
While the AI Act provides a useful framework, Professor Zwitter noted that not all AI used in crisis situations fit neatly into its categories; some may fall under limited risk, and others may escape its scope entirely due to national security or military exemptions. Finally, he pointed out that existing soft law mechanisms (such as guidelines and recommendations) strive towards the same goals as those the European regulators are aiming for. These, he noted, should be always applied in parallel.
Lessons from the COVID-19 pandemic
Associate Professor Olya Kudina (Delft University of Technology) examined how AI was deployed during the COVID-19 pandemic. Contact tracing apps, typically built on corporate infrastructure from Google and Apple, created an immediate conflict of values. People were asked to share personal data for the common good, but many were uncomfortable doing so through platforms owned by tech giants, especially when governments failed to clearly communicate how data was being used. In the Netherlands, Kudina noted, the technology was technically ready and reasonably effective, but adoption was crippled by the government’s delayed and inconsistent communication. By the time official endorsement arrived, “it was almost the end of the pandemic, and people really had a strong mistrust of the technology.”
The diagnostics story was equally cautionary. By early 2020, more than 700 AI models had been rapidly developed to diagnose COVID-19 from chest scans, an extraordinary burst of innovation. But Kudina argued that “it’s sometimes smart to resist, this time, the urgency of adopting technological innovation” when robust validation processes and representative datasets are lacking. Technology, she said, should not be treated as a “technological fix for a complex societal problem.” Quarantine enforcement, used in countries including Ukraine and Poland, combined GPS geo-fencing with facial recognition, requiring quarantined individuals to submit selfies to prove they were staying isolated. These systems frequently produced false positives, triggering monitoring teams to check on people who had not actually violated their quarantine. Kudina concluded that the crisis accelerates adoption, but sometimes at the expense of scrutiny. Governance strategies need to be in place before an emergency strikes, not improvised in the heat of the moment.
Crisis response is fundamentally about relationships.
Commandant Quentin Brot, Head of Innovation and Foresight at the National Fire Officers Academy, offered valuable insights drawn from his experience as a firefighter. Brot drew a sharp line between AI applications that work and those that do not, or worse, that are actively harmful. Early wildfire detection, predicting operational needs in a given area, and logistics planning can be efficient and even useful. But the operational chatbot, a tool designed to give firefighters real-time advice and guidelines during operations, is, in his view, “quite useless or even dangerous.”
His objections were bluntly practical. First, ergonomics: when you are fighting a fire, you do not have time to type on a screen. Second, legal risk: making such a chatbot effective requires feeding it private and sensitive data. Third, and most psychologically insidious: “As a decision maker, I tend to trust the machine more than my intelligence, and it could be dangerous because sometimes you are right, but when you trust the machine, you tend to change the way you are thinking.” And finally, the human dimension. Crisis response, Brot said, is fundamentally about relationships: about motivating people, giving them courage, showing empathy. “If you put an AI chatbot between you and your people, I think it’s not efficient.” He noted that similar conclusions had been drawn by the US Army regarding AI tools for field officers.
“As a decision maker, I tend to trust the machine more than my intelligence, and it could be dangerous because sometimes you are right, but when you trust the machine, you tend to change the way you are thinking.”
A case for sovereign European AI tools
Professor Rémy Slama, speaking on behalf of the Group of Chief Scientific Advisors (GCSA), delivered what may have been the panel’s most cautious message. The Advisors identified sovereignty as a critical issue. Managing crises requires handling extremely sensitive data, personal information, details about critical infrastructures, and “it’s essential that this information does not circulate outside the EU and in particular in rogue states,” Slama said, making the case for sovereign European AI tools.
He also introduced the concept of opportunity cost: not just asking what AI can do, but carefully assessing “what we would lose if we switched to a strong reliance on AI in crisis management.” Could it lead to decreased human expertise? Could it create an over-emphasis on data analysis at the expense of data collection? In a crisis, data quality is typically poor, and Slama invoked the statistical maxim: “Garbage in implies garbage out.” Whether using classical methods or AI, unreliable data produces unreliable results.
On the question of AI in decision-making, the Advisors went further than what the SAPEA report presented. The evidence review noted that AI should not be used for “morally engaging decisions.” But the Advisors argued that it is difficult to know in advance which decisions are morally engaging. “Is closing a school a morally engaging decision?”, Slama asked. “The safe option is not to use AI tools currently for any decision, and to keep humans at the centre of decision-making.”
Discussions about anticipatory AI and accountability
Asked whether anticipatory AI could have helped European governments respond more quickly to the onset of COVID-19, the panellists were sceptical, but for different reasons. Kudina argued that trust is “not a single and monolithic property but a multi-dimensional practice” combining technological, social, and institutional factors. The Dutch contact tracing app, she reiterated, failed not because of technology but because of governance. Slama pointed to the data problem: if a crisis originates in “a rather closed country” then “whatever tools you may have do not allow you to make relevant modelling of an epidemic.” He urged stronger investment in infrastructure, coordination, and harmonised surveys outside the context of acute crisis. Comes argued that the issue during COVID was not really about AI or data at all. Governments could see the exponential curves emerging; the data was there. The problem was a failure of trust in predictive modelling and science more broadly, not a failure of AI.
When asked whether the real trust question is about AI itself or about the people behind it, Zwitter pointed to a compounding problem. Not only are many AI models still “considered black boxes and hard to disentangle,” but the organisations providing the training capabilities for large models are primarily in the corporate sector, mostly outside the EU, specifically in the US, and now to some degree in China.” Comes added that with large language models, even the developers may not be able to predict what their models will do, which “makes all questions on accountability hard.” On the unresolved question of accountability, Zwitter acknowledged that the AI Act has brought some clarity, particularly around the obligations of providers and deployers of high-risk AI systems. But significant gaps remained, As Zwitter bluntly put it, “a lot of those that develop AI tools will shy away from taking responsibility for the results and the decisions taken by others with that tool.”
The webinar left a clear impression: while the use of AI in crisis management looks set to grow across Europe, the experts’ collective message was one of vigilant caution. The technology is advancing faster than the governance frameworks, validation processes, data preparedness, and shared guardrails or benchmarks needed to deploy it responsibly, leaving governments and organisations increasingly dependent on a small number of providers and raising concerns about digital sovereignty. In a crisis, the instinct to reach for any available tool is strong, but as these researchers and practitioners made clear, reaching too quickly can itself become a source of danger.
You can watch the full webinar here, and the full report and statement are available on the SAM website.

