Reeves-McLaren N · 2026 · ACS omega
Paper
Artificial intelligence models for materials discovery are only as reliable as the data on which they are trained, yet systematic audits across the chemical sciences reveal quantifiable error rates in published experimental and computational data that often exceed the predictive accuracy of these models. Domain experts distinguish real from AI-generated characterization data at accuracy levels indistinguishable from random chance. This Viewpoint argues that the established culture of rigorous data collection, calibration, and metadata documentation that is standard at large experimental facilities represents a model for data integrity that the wider community must now adopt. Protecting and investing in these facilities, and in the expert staff who operate them, is not merely a matter of maintaining experimental capability; it is essential to ensuring that the data driving AI-enabled discovery can be trusted.
Analysis
This paper argues that the reliability of AI models for materials discovery is fundamentally limited by the quality of training data, advocating for the adoption of rigorous data collection and documentation standards from experimental facilities to ensure trustworthy AI-driven discovery.
Discovery
Mohamed Hamidi; Mohamed Loutou
Paul Hagemann; Simon Müller; Janine George; Philipp Benner
Mihail Kolev
Prudvi Saisaran Ponduru
Wen Qian
Cheng Li; Yuehui Xian; Yumei Zhou; Xiangdong Ding; Jun Sun; Dezhen Xue
Source record