The authors lead by emphasising the importance of thorough validation in relation to patient safety and effective application in health care, progressing to detailing the challenges in achieving that. They highlight the complex landscape LLMs face due to substantial natural language generation task variation and the versatility in outputs from LLMs, in comparison to quantitative artificial intelligence (AI) prediction algorithms. They note this makes it much harder to predict the clinical impact and unintended consequences of LLMs.
The authors continue by proposing the validation process follows a three-tier approach, with each tier increasing in specificity graduating to clinical impact validation. The article provides suggestions and prompts for developers and researchers to consider.
Enjoyed a book, blog, video or podcast recently that you think other CHIAs might also like? Tell us what and why by emailing [email protected].
