Cross-Evaluation of LLMs on Consumer Health Question Answering
An empirical benchmark of large language models on consumer-based (informal, non-specialist) medical question answering using the MedRedQA dataset of AskDocs questions with verified-expert answers, employing a model-as-judge cross-evaluation scheme in which every model scores every model's answers (including its own) against the expert response to reduce single-judge bias. It ranks models by alignment with expert answers and characterizes LLM strengths and limits on real-world consumer health queries.
2501.00208
This paper empirically evaluates several large language models on consumer-based medical question answering using MedRedQA, a dataset of informal medical questions from the AskDocs subreddit paired w…