Evaluation & Benchmarks
Datasets and evaluation protocols built with the domain experts who have to trust the result. Rather than scoring models on proxies, this work curates expert-annotated ground truth (clinicians reviewing health claims, rubrics written by practising obstetricians) and asks whether language models can stand in for that judgement. Usually the answer is partly, and the interesting part is where the gap sits.
2 papers
-
PATHFinder Agent for Tailored Prenatal Care
Interactive Health 2026
-
RELIANCE: Curating and Evaluating Reproductive Health Information on Social Media
KDD 2026 · Datasets and Benchmarks Track