When AI Stops Being a Project: Turning Technology into Real Value for Patients and Providers
Stanford Online
A Father's Day Shopping Assist 0:12
Sundeep Dudlani, CEO of Optimum Insight, opens by describing a Father's Day trip to the mall with his daughters, ages 23 and 19. The store had few options and nothing fit right, so he photographed an item and uploaded it to Gemini, which identified the color, size, fit, and stock availability, found it online at the right brand, and placed the order. Having been a retail consumer his whole life, he was struck by how much of the actual work the AI completed on his behalf, not just answering a question but finishing a task end to end.
From Answering Questions to Doing Work 2:00
The hosts connect this to a new OpenAI paper showing that AI use has shifted from simple question-answering, the kind of thing most of the world's billion ChatGPT users treat as a Google replacement, toward agents actually completing work across finance, recruiting, and legal tasks. The paper found that about 25 percent of the tasks agents performed were estimated by humans to take over eight hours to do manually. Healthcare, they note, hasn't caught up to this shift. Most health systems have only done light question-and-answer work, and the idea of an agent independently carrying out an eight-hour task is still hard to imagine, partly because it's difficult to even name a continuous eight-hour task to hand off.
Governance Gaps and Agent Hype 5:00
Dudlani says that even inside UnitedHealth Group, with about 22,000 engineers experimenting with large language models, long-form agentic work is still out of reach because of missing governance, guardrails, security, and imagination. He describes sitting through a business review where every feature was labeled an "agent," even a simple API call, when a true agent requires actual agency: the ability to reason, decide, and act. He warns that stacking agents in silos won't reimagine healthcare; real transformation requires looking past isolated fixes to the full end-to-end process, whether that's getting a claim paid or getting a patient to the right care.
Cost, Lock-In, and Model Anxiety 9:00
The conversation turns to enterprise worries about agentic AI: the token costs of giving a model real agency, and the risk of being locked into one frontier model or provider only to see it change, get restricted, or turn out to have a security issue. Dudlani shares that UnitedHealth built an internal system called United AI Studio, which now runs 117 language models, about 20 of them internally built open-weight small models, plus a token gateway to control spending. The organization consumes between 10 and 25 billion tokens a day, and Dudlani's own personal token spending limit is $1,000, beyond which the enterprise CTO tells him he needs to move things to production rather than keep experimenting. About 92 registered agents work across domains like payer, provider, and policy, all monitored for security and data responsibility.
The GLM Moment and Shifting Costs 13:30
The discussion shifts to GLM 5.2, an open-source, lower-cost model out of China that matched prior frontier-level coding benchmarks and unsettled assumptions about needing to stick with established providers. A separate study found that most US companies routing through the platform OpenRouter are now increasingly using Chinese open-source models. One host notes that GLM appears to hit frontier-level results only by using five to ten times more tokens than something like a leading closed model, meaning it's cheaper per contract but not necessarily more efficient, and that current public benchmarks don't reveal an upper limit on performance gains from just adding more tokens. The point is made that companies with strong internal knowledge of their own use cases can judge within days of a new release whether it's actually better for their needs, regardless of the marketing noise.
Generalist Models Beat Specialized Clinical AI 17:30
A widely discussed study compared general-purpose large language models, including Claude Opus 4.6, GPT-5.2, and Gemini 3.1, against specialized clinical AI tools, finding the generalist models performed better than tools like OpenEvidence and other expert medical AI systems. Dudlani connects this to Optimum Health's own scale, serving 20 million patients a year and employing 10,000 physicians, explaining that when deciding whether to build or buy AI capability for claims, patient intake, or care pathways, the temptation is always to build, since training an open-weight model is no longer especially hard. Instead, he says the company focuses on building a harness that lets the underlying model be swapped out as better ones emerge, rather than betting on any one specialized model, since general models keep improving and clinical reasoning is best treated as one embeddable piece of a larger workflow.
Value Sits Above the Model 22:30
Both guests agree the real debate isn't which model wins a benchmark, since the underlying models are becoming commoditized, but the value built on top of them, meaning business context, specific use cases, and the harness that ties them together. They point out that peer-reviewed papers test models that are already 18 months old by the time they're published, so the industry is often arguing over outdated comparisons while the real-world question is which model creates value for a given task, not which one wins an abstract horse race. Dudlani adds perspective by noting the industry still processes 9 billion faxes a year, and that some teams inside UnitedHealth still rely on old optical character recognition tools embedded in workflows years ago simply because it remains the cheapest way to handle that fax volume, illustrating how far behind the frontier much of healthcare's actual infrastructure remains.
Moving 100 Times Faster 25:30
Asked about pacing AI adoption, Dudlani recalls a philosophy from his earlier work at Mars, a $50 billion company, built around the idea that a large company could move 100 times faster by closing the distance between problem solvers and the actual point of contact with users, patients, or providers, rather than letting layers of abstraction slow decisions down. He describes setting up six "tiger teams" this week to tackle six processes end to end, insisting they be fluent with tools like Claude 4.8 and Codex. He shares a recent moment in a strategy meeting where a competitor's strikingly polished app design left the team discouraged, until someone fed a login screen into Claude Code and, within about 60 seconds, generated an HTML file inspired by that design but built on their own context.
Monthly AI accountability meetings 29:00
At United Health Group, every CEO of every business, along with CTOs, CFOs, and chief medical officers, meets for six hours a month to review AI use cases, AI token usage, outcomes, and NPS results before and after AI deployment. The speaker admits these are draining sessions but says the regular cadence injects a healthy paranoia into the organization, pushing peers to show results within thirty days. This creates real tension, since investment has to produce a return and affordability still has to be solved.
AI as a leadership issue 31:02
The speaker urges clients to replicate this rhythm, and some health system partners now run similar tracking sessions with CFOs and leadership every other week. The point is that AI cannot be treated as a delegated problem handed off to IT or an innovation team; it has to be a CEO level discussion if an organization wants real pace. Historically, early technology adopters in health systems have not been the ones who won, which is why this level of leadership engagement is unusual, though it is starting to spread.
AI as a way of working 33:01
One speaker predicts that within two years, AI will either remain a side project or it will effectively be the business itself. Unlike a static platform shift such as moving to the cloud, AI reshapes work itself, moving both top down and bottoms up, touching every part of an organization. The moment a team's light bulb turns on, someone close to the user can build something on the fly that instantly addresses a real need, and leaders should step back and let that happen. For United Health Group, this same inside out logic is turning internal fixes into external products like Optum, Crimson, and Digital Prism.
AI-generated summary. It can be wrong or incomplete - check anything that matters against the original.

