Using AI in Evaluation Practice: What Works Across the Evaluation Lifecycle
Mesa redonda | En línea
-
Organizado por:
CLEAR South Asia
Sobre el evento
Artificial intelligence is becoming the newest addition to the evaluator’s toolkit. As it takes on everyday evaluation tasks such as data processing, analysis, and reporting, it is changing the nature of evaluation practice and evidence use. In doing so, AI is placing evaluations at a crucial crossroads between using these tools to strengthen rigor, or adopting them in ways that prioritise speed but leave much of the clarity and credibility wanting.
This session, moderated by J-PAL South Asia, will examine how artificial intelligence is being applied across different stages of the lifecycle of evaluation, from instrument design and data collection to analysis and reporting. Alongside current applications, the discussion will focus on practical considerations for using these emerging technologies responsibly, including how to keep the evaluator in the loop, how outputs are reviewed and validated, and how AI is used alongside established methods to ensure meaningful impact.
The session will feature practitioners and technical experts from multilateral organisations, along with implementing partners who have applied these tools in practice. Through these examples, participants will engage with where AI has added value, and where it has raised challenges, especially in verifying outputs and maintaining confidence in evaluation findings. The discussion will aim to move beyond adoption of AI, towards clearer choices on where AI should be relied upon, where it should be treated with caution, and how its use can strengthen, rather than blur, the foundations of credible evaluation.
This session, moderated by J-PAL South Asia, will examine how artificial intelligence is being applied across different stages of the lifecycle of evaluation, from instrument design and data collection to analysis and reporting. Alongside current applications, the discussion will focus on practical considerations for using these emerging technologies responsibly, including how to keep the evaluator in the loop, how outputs are reviewed and validated, and how AI is used alongside established methods to ensure meaningful impact.
The session will feature practitioners and technical experts from multilateral organisations, along with implementing partners who have applied these tools in practice. Through these examples, participants will engage with where AI has added value, and where it has raised challenges, especially in verifying outputs and maintaining confidence in evaluation findings. The discussion will aim to move beyond adoption of AI, towards clearer choices on where AI should be relied upon, where it should be treated with caution, and how its use can strengthen, rather than blur, the foundations of credible evaluation.
Presentador/a
| Nombre | Título | Biografía |
|---|---|---|
| Neeta Goel | Senior Evaluation Specialist, Independent Evaluation Department, Asian Development Bank | |
| Maaike Bijker | Chief of Evidence (Data, Research, and Evaluation), UNICEF India | |
| Paramjyoti Chattopadhyay | Head of Research, Assessment, Monitoring (RAM) and Evaluation, World Food Programme India | |
| Tithee Mukhopadhyay | Deputy Executive Director, J-PAL South Asia |
Moderador/a
| Nombre | Título | Biografía |
|---|---|---|
| Megha Pradhan | Director, CLEAR South Asia |
Resumen
The discussions underscored that AI is already reshaping the evaluation lifecycle, shifting the focus from whether to use AI to how to use it responsibly. As routine tasks become increasingly automated, the role of evaluators becomes even more critical, with greater emphasis on interpretation, contextual judgement, ethical reasoning, causal inference, and accountability for findings. A key takeaway was that the evaluation community must engage with AI on two fronts simultaneously: harnessing it as a tool to strengthen evaluation practice, while also building the capacity to assess AI as a development intervention in its own right.
The conversation highlighted a number of areas requiring greater attention. These include the development of shared benchmarking standards for AI-assisted evaluation to balance innovation with accountability, stronger norms for disclosing the use of AI in commissioned evaluations, and the creation of shared prompt banks and use-case libraries to reduce inconsistencies across teams. The need to close the evaluator skills gap was also emphasised, enabling practitioners to recognise the limitations of AI and critically assess its outputs. Participants further underscored the importance of embedding evaluation considerations into AI product development from the outset to ensure that marginalised populations are not excluded from representation.