By Endless Health and The Agency Fund
As health organizations deploy AI tools in real-world settings, it has become clear that models are no longer the primary bottleneck – context and implementation are. While model errors in health have serious consequences, recent evaluations suggest that LLMs perform strongly on technical benchmarks.
The key question now is how these tools perform in specific contexts. Will they work safely and reliably in a crowded primary care clinic in Nairobi? Will they help nurses in rural India triage patients more effectively over WhatsApp, or enable maternal health workers to better support new mothers?
Standardized AI evaluations rarely capture these realities. Most organizations rely on fragmented and non-repeatable approaches, with limited agreement on how to define criteria for quality and safety and few structured ways to learn from peers. This makes performance claims difficult to compare and even harder to interpret.
We see this across the social sector. Organizations are deploying AI but lack shared, practical ways to evaluate what “good” looks like. In response, we’ve developed the Playbook for AI Evaluation, the experiment engine Evidential, and the Agency Fund Accelerator for nonprofits building AI tools to learn from each other and generate ecosystem-wide insights.
Building on this work, The Agency Fund and Endless Health recently launched the Health AI Benchmarking Initiative – a cohort-based program for organizations deploying AI tools tailored to the health sector.
The Health AI Benchmarking Initiative
Rooted in real-world health delivery, the initiative aims to develop grounded benchmarks, practical evaluation approaches, and clearer standards for quality and safety in frontline AI health systems. Instead of ranking models on abstract metrics, it focuses on answering context-specific questions. How should we define accuracy in triage workflows? What does it mean for a system to avoid harm? How should escalation reliability be measured?
The goal is to design benchmarks that hold up in real-world use. By working collaboratively, we hope to make AI evaluation in health more transparent, comparable, and actionable. Over a 120-day sprint, participating organizations will receive a $20,000 catalytic grant to focus on four areas:
Curating real-world data drawn from active health service workflows.
Developing evaluation frameworks with clearly defined quality dimensions (e.g., medical accuracy, harm avoidance, contextual appropriateness, escalation reliability) and measurable metrics.
Running benchmarks to generate documented performance results.
Documenting and sharing evaluation methodology, data preparation processes, assumptions, and known limitations or biases.
For the initial cohort, we are partnering with 10 organizations whose operational experience and early AI adoption span primary care, maternal health, digital triage, and community-based services: Armman (India), E-Health Africa (West and Central Africa), Jacaranda Health (Kenya), Intelehealth (India, Kyrgyzstan), Intron (Sub-Saharan Africa), Medtronic Labs (Sub-Saharan Africa and South Asia), Myna Mahila (India), Pinky Promise (India), Penda Health (Kenya), and SameSame (South Africa).
Building shared standards
The initiative’s broader goal is to strengthen how the health sector evaluates AI tools. Currently, a strong benchmark score does not necessarily mean that a tool will perform safely and reliably in clinical settings. For evaluation results to be meaningful, they need to reflect real contexts, workflows, and constraints.
To help address this, the initiative will prioritize:
Sharing benchmarks and approaches. Evaluation frameworks and benchmark results from the cohort will be released under open licenses and shared via platforms such as Hugging Face, making approaches easier to understand, reuse, and build on.
Reducing duplication through shared learning. By comparing approaches across the cohort, we aim to surface common methodological gaps, differences in safety definitions, workflow constraints that shape evaluation choices, and tooling limitations that slow down testing.
While this initiative is a solid start, it only addresses the first level in our AI Evaluation Framework: model performance. Other levels of evaluation, including for product adoption and user experience, require additional resources and expertise. Still, making Level 1 evaluation methods visible across our community could help ensure that differences in benchmark results reflect real differences in model behavior – not simply differences in how the evaluations were designed.
Learnings from the cohort will inform funder strategies in AI for development and contribute to an upcoming high-level stakeholder gathering, co-convened by The Agency Fund, Clinton Health Access Initiative, and Endless Health at the Rockefeller Foundation’s Bellagio Center.
Looking ahead
This initiative is one part of a broader effort to improve how the social sector evaluates AI in practice. In addition to the Playbook and Evidential highlighted above, we’re also developing and will soon publish Product Cards for GenAI programs run by grantee organizations: standardized, easy-to-read profiles that show how specific tools perform across contexts.
As more health organizations deploy AI in live service environments, evaluation approaches need to keep pace. When evaluation is disconnected from clinical realities, it becomes difficult to interpret what results mean for safety or service quality. Shared approaches for evaluations of model performance, product adoption, and user experience can help make these trade-offs visible and support more informed decisions about what and when to deploy AI in healthcare.



