Researchers at the Indraprastha Institute of Information Technology Delhi have published MH-INDIC, a framework for evaluating how well language models handle maternal-health conversations in urban and semi-urban north India.
The method starts with people rather than with the models. A 26-item survey was administered to 102 pregnant and postpartum women, and their answers were used to define ten dimensions of maternal-health reasoning, the kind shaped by household and social norms rather than by clinical fact. Ten models were then measured against that human baseline. The finding the paper leads with is a split between two things usually reported as one. At the population level, the models approximate the women’s distribution of answers reasonably well. Across demographic and household profiles, they vary far less than the women did. Aggregate cultural alignment, in other words, can conceal a near-total insensitivity to the individual asking.
The authors also test whether prompting closes the gap: conditioning generated dialogues on human-grounded profiles produced better profile alignment and dialogue-quality ratings than zero-shot or self-conditioned prompting.
Why it matters: health-model evaluation mostly measures factual correctness, safety and fluency, none of which would catch this. The sample is small and regional, and the paper makes no clinical claim.
Source: arXiv:2610.11586, announced 9 October 2026.
