Evaluation & Benchmarks

Datasets and evaluation protocols built with the domain experts who have to trust the result. Rather than scoring models on proxies, this work curates expert-annotated ground truth (clinicians reviewing health claims, rubrics written by practising obstetricians) and asks whether language models can stand in for that judgement. Usually the answer is partly, and the interesting part is where the gap sits.

2 papers

  1. PATHFinder Agent for Tailored Prenatal Care

    Vaibhav Balloli, Carissa Samuel, Samia Abdelnabi, Alex Peahl, Elizabeth Bondi-Kelly

    Interactive Health 2026

  2. RELIANCE: Curating and Evaluating Reproductive Health Information on Social Media

    Vaibhav Balloli, Laura Peyton Ellis, Vishala Mishra, Alice M Chi, Alex Friedman Peahl, Elizabeth Bondi-Kelly

    KDD 2026 · Datasets and Benchmarks Track

← All publications