Stanford Graduate School of Business

World Development Report 2026: How Can AI Improve the Delivery of Public Services?: summary

YouTube summary28 sectionsWatch on YouTube ↗

This is an AI-generated summary of the YouTube video "World Development Report 2026: How Can AI Improve the Delivery of Public Services?" (Stanford Graduate School of Business), made with Samuraize and published by Samuraize. It condenses the YouTube video into 28 titled sections you can read in a couple of minutes, each linking to the moment in the video it covers.

1
Filed under💻 Technology0 comments🍱 Add to trayReport
Study this
Export

World Development Report 2026: How Can AI Improve the Delivery of Public Services?

Stanford Graduate School of Business

切

Introducing the Panel on AI 0:00

The session opens after lunch with an introduction to a panel on how artificial intelligence can improve public service delivery. The two speakers are Anya Salman, a co-author of the World Development Report working in the World Bank's development economics research group, and Josh Blumenstock, a professor at UC Berkeley. Each speaker is given twenty minutes before an open panel discussion follows.

切

Capacity Gaps in Poor Countries 1:01

Anya Salman begins by describing the severe capacity gaps across health, education, and public administration in low and middle income countries. She cites a striking figure from the report: low and low-middle income countries together have fewer than fifty thousand radiologists for a combined population of about 3.7 billion people. Her talk draws on a government survey, an analysis of procurement contracts, and two evidence reviews, one on AI tutoring and one on AI in frontline health care, to assess where AI can realistically expand access to services.

切

Front End Versus Back End AI 3:01

She divides government AI use into two spaces. The front end covers apps and tools used directly by citizens or local service providers, constrained by the last mile problem of connectivity, devices, language, and literacy. Here governments mostly procure private-sector solutions with light adaptation like prompt engineering. The back end involves centralized agencies with large datasets building custom machine learning models, for things like weather forecasting or tax enforcement, where the real challenges are data infrastructure, governance, and internal technical capacity.

切

What Governments Report Using 5:01

Survey data shows many governments already use chatbots, direct citizen-facing apps, and frontline worker support tools, with lower-income countries lagging somewhat behind. Procurement records over the past ten to fifteen years show AI use skewing toward prediction and optimization rather than generative AI, alongside applications like biometric identification and intelligent transport systems.

切

Two Ways AI Helps Services 6:00

Front end AI use falls into two categories. One is improving the quality of existing services, such as using large language models to handle administrative burdens for school principals so they can focus on instructional leadership, or supporting infrastructure maintenance and clinical decision making. The other is expanding access by automating specialist tasks entirely, either delivering services directly to citizens or through task shifting, where a lower-level worker performs a higher-level task with AI assistance.

切

Evidence From Chile and Tutoring 7:31

A chatbot called Consilium Bot in Chile gave personalized college advice to students and increased college attendance by up to 44 percent, with about 20 percent more students enrolling in higher-ranked programs. High-intensity tutoring, long known to work but limited by the need for roughly one tutor per three students, is now being scaled through AI, though concerns remain about whether it might reduce student effort or erode learning in unexpected ways.

切

Scanning the Health AI Evidence 9:31

Looking specifically at frontline healthcare, the team scanned nearly 300,000 papers, narrowed to 10,000 on AI and health, and ultimately found only 18 studies that actually measured impact on patient outcomes or cost effectiveness in low resource settings. Much existing research focuses on applications irrelevant to primary care in places like Nigeria, such as fully automated colonoscopy guidance meant for tertiary hospitals in rich countries. Among the 18 studies, eight concerned diabetic retinopathy and cataract screening and two concerned X-ray based tuberculosis screening, with only two using large language models for clinical decision support.

切

Tuberculosis Screening Cost Effectiveness 11:32

AI-assisted tuberculosis screening from chest X-rays is a mature application, important because roughly a million people die needlessly from tuberculosis each year, and the World Health Organization has approved vendors with excellent detection performance. The approach is cost effective where X-ray machines already exist and compares well against human readers, but in a Nigerian trial comparing it to simple symptom-based sputum testing, the extra detection of asymptomatic cases did not justify the added expense of devices and AI.

切

Limits of AI Clinical Advice 12:30

Large language models used for clinical decision support are good at flagging inconsistencies in care plans, such as a wrong antibiotic or a missed malaria test before prescribing malaria medication. But they fail at spotting missing conditions or diseases altogether, which is a major source of real patient harm, because the model can only react to what health workers submit rather than catch what is absent. The broader challenge is identifying high value applications before the evidence has fully caught up, especially in high-risk fields like health and education.

切

Barriers to Frontline Adoption 14:01

Deploying AI tools to citizens and frontline workers requires devices, connectivity, subscriptions, and training, which either raises the cost per user or excludes people who cannot access the tools on their own. Government officials surveyed point to exactly these barriers: limited citizen skills, undertrained frontline workers, and insufficient AI support in local languages, with these concerns notably more severe in lower income economies.

切

Back End Patterns and Forecasting 15:01

Procurement data shows that even low-income countries often build or customize their own AI tools internally rather than relying solely on commercial products, and tax authorities are notably ahead in adoption because the financial incentives are strong. Prediction remains the dominant back-end function, covering predictive analytics, remote sensing, geospatial intelligence, and environmental monitoring. One promising area highlighted is improved forecasting, from weather to floods to pandemics, with the Google Flood Hub website mentioned as an example.

切

Back-office AI applications and risks 16:34

Flood forecasting in rivers can now be done with AI using only past data rather than expensive sensor networks, a free tool already covering 80 countries. In Sierra Leone, a custom supply chain tool improved delivery of essential medicines to rural areas. AI also helps monitoring and oversight, such as spotting illegal logging or fraudulent firms, though this only works if governments also invest in enforcement and auditing. These back-end uses need less last-mile infrastructure, but they demand strong data systems, secure access, clear standards, and attention to risks like misuse or cyberattacks, especially since many governments struggle to maintain systems they deploy.

切

Recommendations and survey findings 19:30

Surveys show data quality and budget are top government concerns, and nearly 60 percent of low-income countries report having no AI policies or guidelines at all. Eighty-one percent of lower-income respondents report little or no performance monitoring of AI investments. Recommended priorities include benchmarking models before rollout, tracking their impact on both beneficiaries and non-beneficiaries, procuring systems with interoperability and data rights retained, building accountability frameworks with clear red lines, and matching deployment to each government's actual capacity, since maturity varies widely even within income groups.

切

Togo's resource allocation problem 24:00

Minister Cina Lawson of Togo faced a stark COVID-era problem: 10 million dollars in benefits against 6 million potential beneficiaries, enough to cover only 150,000 people. She approached it as a resource allocation problem rather than an AI problem, and a pragmatic, non-elegant approach worked. Globally, over half the population receives some social protection, costing about 15 trillion dollars yearly, yet many governments lack current data on beneficiary needs. Research shows targeting accuracy drops roughly 10 percentage points per year since data was collected, and most countries update registries only every four or five years.

切

Satellite imagery for poverty mapping 28:02

In Nigeria, satellite imagery combined with household survey data let AI models estimate wealth and poverty at high resolution across urban wards, producing far more detailed maps than survey data alone while losing little accuracy. This helped the government target benefits to the poorest wards. In Togo, this approach did not fit politically, since geographic targeting would have concentrated all transfers in just 12 or 13 cantons, risking accusations of favoritism, pushing the team toward using mobile money data to identify poor individuals within communities instead.

切

Phone data outperforms geographic targeting in Togo 32:31

Working with Togo's government, a machine learning model that predicted poverty from mobile phone usage patterns proved far more accurate at identifying poor people than the geographic targeting approach previously used in Nigeria. But the clean performance charts hide a key bias, since the program distributed aid through mobile money, so anyone without a phone was automatically excluded, and phone ownership gaps are especially large between men and women and across age groups. Much of the exclusion and lost impact happens before performance is even measured, echoing broader points about missing complements that limit basic technology, not just AI.

切

Interpretability, privacy, and fairness trade-offs 35:01

When Togo's minister asked for a simpler, explainable rule, such as basing eligibility on monthly phone spending, accuracy dropped sharply compared to the full model. Similar trade-offs appear with privacy, since adding differential privacy protections to the data reduces targeting accuracy along a measurable frontier, letting policymakers choose their preferred balance. In Bangladesh, phone-based AI targeting was more accurate at finding poor households, yet communities trusted and preferred decentralized, community-based targeting, raising the question of whether accuracy is even the right metric.

切

Weighing cost against targeting benefits 38:00

Traditional proxy means test surveys cost about four dollars per household on average, which is not free, while digital data screening can cost almost nothing at the margin. Looking at roughly a hundred World Bank social protection programs, cost becomes decisive mainly for programs screening very large populations with modest budgets. Under a welfare-based comparison, about 10 of 95 such programs are better served by AI-driven targeting than by conventional surveys.

切

Three broader lessons for AI deployment 40:02

AI systems are most useful for filling information gaps, as seen with satellite data in Nigeria and phone data in Togo. Strong technical performance often fails to translate into better policy outcomes, so evaluation must cover the whole deployment stack, including training, incentives, and rollout, not just the model itself. Maximizing accuracy frequently conflicts with other goals that better reflect a policymaker's true objectives, and these lessons apply well beyond social protection to clinics, schools, and other public services.

切

Panel reflects on privacy and evaluation 42:05

Moderating the panel, Ana noted that both talks emphasized predictive, back-end AI rather than the generative, front-end tools dominating public debate, and stressed that impact depends on analog and digital complements. She then asked about privacy and data security, about what evaluation metrics need to change to capture cost effectiveness, and about what must shift in government systems to scale pilots successfully.

切

Rethinking what privacy means locally 44:01

Josh said survey data on AI concerns shows strong privacy worries in lower and middle income countries, but his direct experience in Togo suggests privacy often isn't the top concern for partners facing more immediate needs. Fieldwork by a PhD student found Togolese beneficiaries' privacy worries differed sharply from Silicon Valley style concerns, focusing less on governments or corporations and more on neighbors in their own village learning more about them through the data collected.

切

Governments can still set data protections 46:31

Building on this, the point was made that people's lack of concern about government or corporate data use may simply reflect limited exposure to those risks, not that the risks are absent. Governments can still retain control over data, set standards, and apply privacy-protecting principles to their own systems, since prediction tasks can often be done without disclosing data in harmful ways. The discussion then turned to cost effectiveness, noting that development research across public health, education, and other fields rarely measures it seriously, partly because tallying costs can feel like a less interesting task for academic researchers.

切

The Missing Cost Data Problem 48:30

A researcher points out how hard it is to know what it actually cost to achieve a result, like detecting a certain number of new tuberculosis cases, because pilot costs differ wildly from costs at scale. A meta analysis of health applications found almost no studies that reported both patient benefits and costs together, leaving policymakers with little useful guidance from the academic literature.

切

Three Levels of Evaluating AI 50:00

Josh describes frameworks, including ones from JPAL and Agency Fund, that separate evaluation into layers: model performance, proxy outcomes like clicks and engagement, and costly outcomes like behavioral change and welfare, where development economists should focus. He warns against applying old RCT style thinking, like the bednet analogy, to AI, since AI involves an evolving stack of deployment choices rather than a fixed product.

切

Taxation, Corruption, and Enforcement Capacity 52:30

Asked about AI reducing corruption in tax enforcement, the response is that the bigger issue is usually too little enforcement capacity rather than misused discretion, citing environmental land surveys as an example. AI can also remove biased discretion, as shown in credit approval research where loan officers under-favored firms run by unbanked women, a bias algorithmic scoring helped eliminate.

切

Medicine Shortages and Demand Prediction 55:01

Rural clinics, like those referenced in Sierra Leone, often stock out of essential medicines and vaccines because requesting resupply from central warehouses can take weeks. Simple predictive tools using limited available data to forecast near-term medicine demand could have high impact in such decentralized resource systems.

切

Which AI Companies Are Chosen 56:31

Asked about competition between US and Chinese AI firms, Josh says governments rarely choose freely among providers; constraints like resource limits, regulation, or a company offering free access usually force their hand, pushing some toward open source and others toward whichever provider is available, with few countries like India having real choice.

切

Too Few Studies to Generalize 57:00

Evidence on what deployment factors matter most, such as electricity, connectivity, or local language models, remains thin, drawn from only a few dozen studies. Contrasting labor studies, one on Philippine call center workers showing low ability workers benefit most, another on Kenyan entrepreneurs showing high ability workers benefit most, aren't contradictory but show how little data exists to generalize who truly benefits from AI.

AI-generated summary. It can be wrong or incomplete - check anything that matters against the original.

Summarize your own YouTube video

Paste a YouTube link, article, PDF, ebook or slide deck and get a summary like this in seconds. Free to try, no sign-up needed.

⚔️ Try the YouTube summarizer

Discussion

Sign in to join the discussion. Sign in

More from the Bento Box

Browse the Bento Box →

We use Microsoft Clarity and Google Analytics to see what breaks and where visitors come from. They set cookies and send data to the US. Product events are counted without cookies either way. Cookie details

Samuraize.ai - World Development Report 2026: How Can AI Improve the Delivery of Public Services?