Can Language Models Improve the Creation of Patient Assessment Tools?

“`html

Can Language Models Improve the Creation of Patient Assessment Tools?

The creation of tools to collect patient testimonies relies on essential steps where language plays a central role. These tools, designed to capture individuals’ lived experiences, require careful attention to every word used. Researchers must first precisely define the concepts to be studied, then draft clear and appropriate questions. Each question is then refined through interviews with patients to ensure it is understood as intended and aligns with their reality.

Advanced language models, capable of understanding and producing text in sophisticated ways, now offer new possibilities to support these steps. They do not replace human judgment or traditional validation methods, but they can facilitate certain tasks. For example, they help analyze large amounts of existing data to identify what has already been well measured and what has been measured inconsistently or inadequately. In a recent project on children’s eating behaviors, researchers reviewed over 80 existing tools and 1,400 questions. Language models could have accelerated this step by summarizing these questions, identifying overlaps between concepts, and proposing preliminary definitions for human review.

Another practical use is generating candidate questions. Models can propose varied formulations, adapt the language level, or suggest contextual examples. In the same project on eating behaviors, researchers used these tools to fill conceptual gaps. For example, instead of focusing on very specific details like meals at precise times of the week, the models helped draft broader questions. These addressed general habits, such as how often children eat alone or with their loved ones, or how nutritious foods are integrated into family routines.

Language models are also useful for revising questions based on patient feedback. Cognitive interviews often reveal comprehension issues or formulations perceived as too vague, too broad, or stigmatizing. Models can then suggest clearer or less emotionally charged alternatives. In the mentioned project, the term healthy was replaced with nutritious after families expressed that the latter was less judgmental and less likely to induce guilt.

Another interesting application is checking semantic consistency. Researchers can submit a set of questions to a model without indicating which concept they relate to, then ask the model to guess the common theme. If the model mistakenly associates questions about emotional eating with purely practical behaviors, this signals a clarity or conceptual distinction issue. This method does not validate the quality of the questions, but it helps detect ambiguities before empirical testing.

Finally, these tools facilitate adapting questions to different contexts. A question about recognizing hunger in an infant will not be phrased the same way as for a teenager. Models help generate versions tailored to age, reading level, or cultural context while preserving the original meaning. However, each adaptation must be carefully reviewed by experts to avoid unintentionally altering the meaning or scope of the question.

Despite these advantages, using language models comes with risks. They can invent concepts, confuse similar ideas, or reproduce biases present in their training data. For example, an automatically generated revision might broaden a question about parents’ eating practices to general parenting behaviors, thereby losing precision. Additionally, the use of sensitive data, such as patient narratives, raises privacy and ethical compliance concerns. Researchers must therefore ensure this information is protected and use secure environments.

Transparency and reproducibility also pose challenges. The results produced by these models can vary depending on the version used or the timing of the query. To address this, it is advisable to document the model versions, prompting strategies, and decisions made during the process. Another solution is to use multiple different models to compare their interpretations. If the results diverge, this may indicate a clarity or semantic structure issue requiring human review.

In practice, integrating these tools follows a structured process. Researchers begin by defining the concepts to be studied through a literature review and expert opinions. Language models then step in to summarize existing tools, identify overlaps, and propose preliminary definitions. Once the concepts are clearly defined by humans, the models generate candidate questions or alternative formulations. These proposals are then subjected to thorough examinations: expert review, readability testing, cognitive interviews with patients, and semantic consistency checks. Only after these qualitative steps are the questions tested quantitatively to assess their psychometric properties, such as reliability or sensitivity.

In the eating behaviors project, this approach accelerated certain steps without compromising rigor. The models helped explore the structure of concepts, draft questions for underrepresented areas, and propose revisions based on family feedback. However, they were never used to validate the relevance of concepts or the quality of questions. This validation still relies on proven methods: theoretical review, expert opinions, patient interviews, and empirical testing.

The opportunity offered by these tools, therefore, does not lie in their ability to replace traditional methods but to enhance them. They allow for broader exploration of concepts, faster drafting of proposals, and more systematic comparison of alternatives. In a field where translating lived experience into valid measurements is essential, this assistance is invaluable. The challenge is to use it in a way that expands methodological possibilities without weakening the evidence standards on which the quality of these tools depends.

“`


Legal Attributions

Study Citation

DOI: https://doi.org/10.1186/s41687-026-01112-2

Title: Leveraging large language models in patient-reported outcome measure development: practical opportunities, cautions, and a human-in-the-loop roadmap

Journal: Journal of Patient-Reported Outcomes

Publisher: Springer Science and Business Media LLC

Authors: Constance Mara; Tiffany Rybak

Speed Reader

Ready
500