Artificial intelligence (AI) tools, particularly those that use natural language processing (NLP), are becoming increasingly prominent in healthcare delivery. For these systems to operate safely and reliably in clinical settings, they require high-quality training datasets drawn from multiple sources, including electronic health record (EHR) data, that closely reflect real-world conditions. Although a range of publicly available AI training datasets exists, these resources are not housed in a centralized repository. Establishing such a repository could improve the management and organization of de-identified, interoperable AI training datasets and streamline access for researchers and AI tool developers.
This series of reports provides a framework in three areas: (1) requirements for creating AI-ready de-identified or synthetic datasets using EHR data for comparative effectiveness and patient-centered outcomes research; (2) the data governance structures needed to ensure usability while protecting the privacy and safety of patient data housed in a publicly available repository of AI model training datasets; and (3) a pediatric ADHD annotation schema designed to create AI-training datasets by extracting patient-centered outcomes from EHR data.