SaferAI – Malcolm Murray – AI risk management
Investor discussion: AI risk management
Summary notes from an investor expert discussion with Malcolm Murray the research lead at SaferAI on AI risk managment.
Malcolm leads research at SaferAI which is a non-profit producing research, tools, and assessments on AI risk management. SaferAI produces one of the only assessments of AI model developers on risk management – SaferAI Ratings. The session highlighted that AI risk management is falling behind model capabilities, but investors can help by pushing companies to adopt existing best practices.
- The gap between predicted AI risks and current real-world harm is still wide, but it will probably close soon. Cyber could be first for a major AI incident (as highlighted by Mythos recently) and labour market disruptions are steadily increasing.
- Current AI risk management practices are lacking in key areas such as novel risk identification (for example, manipulation and multi-agent risks) and board oversight. Frontier safety frameworks have a high degree of discretion and approaches are weakening over time.
- On the positive side, if companies adopted best practices across SaferAI’s 65 sub-criteria there would be a significant improvement from a median score of 18% to 51%. Hence, investors can play an important role in encouraging companies to adopt these.
Summary of investor discussion
Malcolm Murray: Malcolm leads research at SaferAI which is a non-profit producing research, tools, and assessments on AI risk management. SaferAI produces one of the only AI risk management assessments of leading AI model developers – SaferAI Ratings. Malcolm also led the work on AI risk management for the recent International AI Safety Report 2026.
Links
- SaferAI paper on ratings
- SaferAI ratings website
- Reports from The Centre for Long-Term Resilience on AI risk governance 1) Why frontier AI safety frameworks need to include risk governance (Feb 2025), 2) Transforming risk governance at frontier AI companies (July 2024).
5 key findings
- Frontier Safety Frameworks fall short of what is achievable: If companies adopted the existing best practice in each of the 65 sub-criteria there could be significant improvement to a score of 51% vs the median of 18%.
- Missing risk management aspects: 1) Novel risk identification is missing. The AI industry zeroed in on a few risks (cyber and CBRN), and places little focus on identifying completely new risks, such as persuasion and manipulation (Google is starting to cover this), or multi-agent risks. 2) Limited assessment of risk tolerance levels and hence what amount of risk it is appropriate to take.
- Little independent oversight: Limited – board involvement, dedicated board committee looking at AI risk, internal audit or external audit. Limited use of three lines of defence as exist for risk management in other sectors.
- High levels of discretion: A lot of the language in frontier safety frameworks leaves room for discretion.
- Weakening over time: Some weakening over time in risk management practices. As risk benchmarks are surpassed and no major incidents happen, companies move the goalposts.
Q&A
Role of investors: Pushing companies to adopt existing best practices. SaferAI’s research can provide the analysis on which best practices to ask companies to adopt.
AI incidents: There is still a gap between predictions of AI incidents and real-world harm.
- Cyber could be the canary in the coal mine. Cyber risks were highlighted in the recent threat assessment report and incident published by Anthropic (Anthropic: Disrupting AI espionage).
Labour market impacts are starting to increase.
Analysis of public vs private information: Will need both, for example companies have more capable models internally that they don’t release, but a lot can be done by analysing public information. Will need analysis of private information such as third-party assurance.
Disclosure as part of the EU Code of Practice: There are three levels 1) public framework (assessed by SaferAI), 2) more in-depth framework only provided to EU AI Office to analyse, 3) commitments in in Code of Practice that don’t go in the framework e.g. model cards.
Third-party AI audits: Investors could play a role in asking companies to conduct these, alongside insurance and regulations. There is currently a lack of providers of audits, but a range of organisations that could step into the role.
Model evals: There is increasing concern on the gap between model evaluations before deployment and what happens in the real-world. The answer is more onerous and longer models evals that more closely reflect the real-world. But then the model evals become expensive to run and the current model eval organisations are non-profits.
Best practice in risk governance: Recommended reading – two reports from The Centre for Long-Term Resilience 1) Why frontier AI safety frameworks need to include risk governance (Feb 2025), 2) Transforming risk governance at frontier AI companies (July 2024).
Regulation: Key drivers for companies to adopt best practices are the EU Code of Practice and SB53.
