Opportunity
AI scribes and other large language model (LLM)-based clinical documentation tools are increasingly used in clinical practice without robust assessment. Current research primarily focuses on a narrow set of evaluation criteria, such as:
- Qualitative clinician experience
- Clinical efficiency
- Reduction of documentation burden
This project aims to develop an evaluation framework for LLM-based clinical documentation tools that reaches beyond current benchmarks. The framework will define a full spectrum of evaluation criteria, with the objective of guiding comprehensive assessment of the quality and impacts of AI Scribes.
Project Objectives
The project team will conduct a review of published research on AI scribe evaluation and
- Undertake research with clinical stakeholders, through structured interviews, to identify key evaluation dimensions.
- Develop a structured evaluation framework for LLM-based clinical documentation tools based on these findings.
- Establish protocols for applying the framework in future assessments.
Framework development will be grounded in evidence from user experiences.
One key evaluation dimension is the veracity (accuracy and quality) of automatically generated documentation compared to original consultations. The project aims to create a realistic dataset for benchmarking tools on this dimension. The evaluation framework will be applied to one or more LLM-based clinical documentation tools to assess their performance regarding documentation veracity.


