Stanford Online

Webinar: What AI Can and Cannot Do: Intelligence Augmentation in Practice with Michael Bernstein: summary

YouTube summary12 sectionsWatch on YouTube ↗

This is an AI-generated summary of the YouTube video "Webinar: What AI Can and Cannot Do: Intelligence Augmentation in Practice with Michael Bernstein" (Stanford Online), made with Samuraize and published by Samuraize. It condenses the YouTube video into 12 titled sections you can read in a couple of minutes, each linking to the moment in the video it covers.

1
Filed under💻 Technology0 comments🍱 Add to trayReport
Study this
Export

Webinar: What AI Can and Cannot Do: Intelligence Augmentation in Practice with Michael Bernstein

Stanford Online

切

Introduction and Speaker Background 0:00

This session, hosted by Stanford Online with Global Alumni, introduces Michael Bernstein, a professor of computer science at Stanford, a Bass University Fellow, and a senior fellow at the Stanford Institute for Human-Centered Artificial Intelligence. He is described as a bestselling author whose research on generative AI simulations is the most cited in the history of the UIST conference. His work has appeared in the New York Times, TED, and MIT Technology Review, and he holds degrees from Stanford and MIT. The talk is framed as excerpts from an online course that has resonated with company boards, CEOs, and CTOs, aiming to help viewers think through AI strategy.

切

Three Curves of AI Progress 2:30

Michael opens by tracing three trends over the same period of time. The pace of AI model releases feels like it is accelerating rapidly, with new models appearing almost weekly. The number of AI-based products and services launched follows a similarly steep curve, as organizations try to figure out what problems AI can solve. But the number of genuinely successful AI products is far more modest, and he wants to explain why so many efforts turn into embarrassing, wasted misfires, splitting his answer into what is feasible and what is worthwhile.

切

Sharp-Edged Versus Rough-Edged Problems 5:30

Because AI capability shifts faster than typical business planning cycles, Michael offers a durable framework rather than a snapshot forecast, built on the distinction between rough-edged and sharp-edged problems. Rough-edged problems have many acceptable solutions, like writing marketing copy, summarizing a lecture, or generating an image, where getting most of the way there is still useful. Sharp-edged problems have only one correct answer, like a bug-free software deployment or a correct hospital-readmission prediction, where being close still means being wrong. He notes agentic AI systems, which choose and execute tool commands like writing code, reading files, or sending emails, can contain both rough and sharp elements, and that coding agents can even handle sharp problems when errors are automatically verifiable. Design work, he argues, is inherently rough-edged despite often being treated as sharp. The real reason AI struggles more with sharp-edged problems isn't lower technical performance, it's that human tolerance for error is much thinner there, creating an adoption threshold rather than a gradual improvement curve, as illustrated by frustrating automated phone routing systems and the twenty-five years it took spam filters to earn real trust.

切

Sharp-edged problems compound in agent chains 16:30

Spam filters and voice assistants like Siri are classic sharp-edged problems, where the answer is simply right or wrong with no partial credit, and improving accuracy is what finally makes them reliable enough to trust. The danger grows when AI agents chain several sharp-edged steps together, like a game of telephone: an inventory agent must find the right part, decide whether to reorder, pick the right quantity, and ship to the right warehouse. If any single step misfires, the whole workflow fails, which is why coding agents tend to succeed mainly when their outcomes are automatically verifiable. Real failures show this pattern too, such as lawyers submitting legal citations to cases that do not exist, and a Princeton study on AI agent reliability found that recent capability gains have only produced small improvements in consistency.

切

Rough edge always outpaces sharp edge 23:31

Customer support bots illustrate the split well: sharp-edged tasks like order tracking or arranging returns already work reliably, while rough-edged tasks like drafting suggested responses or handling open-ended concierge questions tolerate more variation and still prove useful even when imperfect. The set of problems AI can solve in a rough-edged, good-enough way will always be larger than the set it can solve in a sharp-edged, near-perfect way, since any system accurate enough for 99 percent reliability is automatically good enough for 80 percent. As models improve, a task moves from unsolved, to solved roughly, to eventually solved sharply, and both the rough and sharp circles of solvable problems expand with each new model generation. You can watch which side of that boundary your own use case sits on to predict what the next model release will unlock, and you can also convert a sharp-edged idea, like an autonomous hospital readmission predictor, into a rough-edged one, like a report flagging risk factors for a human decision maker, to get value sooner.

切

The human-AI handoff matters most 29:01

Even a strong AI gets abandoned if it does not fit how people actually work, as shown by robots that use slower gestures and eye contact instead of the most efficient possible motion, because that human-plus-AI combination works better in practice. A self-driving car simulator study found that even when the car correctly detected an emergency and handed control back to the person in time, the driver still nearly crashed twice within thirty seconds, showing that failures often happen at the seam, the handoff point between AI and person, rather than in either party alone. The lesson, as researcher Eytan Adar puts it, is not to let your user interface promise more than your AI can deliver. This gap between promise and reality is what produces the trough of disillusionment, where tech-first products overpromise a sharp-edged experience while only delivering a rough-edged one. The better framing is intelligence augmentation, remembered as IA, or AI spelled backward, which aims to make people smarter and more capable rather than to replace them.

切

Augmentation Meets Resistance 33:30

Most AI use today augments people rather than replacing them, letting them iterate on rough-edged problems where a human stays in the loop. This idea traces back to Doug Engelbart, the Turing Award winner who in the 1960s pioneered augmenting human intellect and invented tools like the computer mouse and interactive word processing. Replacement efforts often backfire: research from Stanford's Angele Christin found that journalists and legal professionals push back, drag their feet, or openly criticize data tools when they feel threatened with replacement. A Carnegie Mellon study found doctors only adopted an AI tool once it was made to feel unremarkable, tucked to the side as a quiet, optional suggestion rather than placed front and center.

切

When Complementarity Works and Fails 38:00

The goal of intelligence augmentation is complementarity, where human plus AI outperforms either alone. Examples include Stanford researcher Erik Brynjolfsson's finding that call center workers do better with AI advice, a large medical trial showing improved breast cancer detection when AI helps prioritize mammograms, and BCG consultants improving on creative tasks with AI help. But the same BCG report found performance dropped 23 percent on a different task when people took AI feedback at face value, and studies show software engineers often feel faster with coding tools while actually working slower. An MIT meta-analysis pooling many studies found that, on average, human-plus-AI performance was worse than human or AI alone, with the biggest losses in sharp-edged decision-making tasks like predicting hospital readmission, and the biggest wins in rough-edged content creation tasks.

切

Overreliance and Algorithm Aversion 40:00

People often overrely on AI, checking out mentally in a pattern called cognitive surrender, especially since longer AI explanations tend to feel more convincing even when wrong. This overreliance eventually produces a bad mistake, which pushes people toward algorithm aversion, trusting AI less than they would trust a human who made the identical error, since trust in algorithms is more brittle than trust in people. During the question period, one attendee asked whether you need expertise to judge sharp-edged AI results; the answer was that once a task is reliable enough, like simple coding operations, deep expertise becomes unnecessary, but below that reliability threshold expertise remains essential, since AI-generated content can bias judgment the way pre-highlighted passages in a used textbook do, even when you know the highlights are wrong.

切

Choosing the Right Success Metrics 50:30

When asked about metrics for evaluating success, Michael Bernstein warns against defaulting to replacement-style measures, like doing ten times a day something you now do once. Those numbers only capture optimizations worth maybe ten percent, not the creation of whole new industries. He argues that if you are truly building something intelligence augmenting, you need different metrics, such as whether products are shipping better, performing better in the market, or crashing less often. Executives and boards often push for easy ROI numbers, but those are usually the wrong ones, so he pushes teams to ask who or what is being augmented and how they would know it is working.

切

Risk Tolerance in Customer Facing AI 52:30

On a question about risk tolerance in customer-facing situations, Bernstein points to the classic frustration of yelling operator, operator at an automated airline phone system as evidence of thin trust margins and algorithm aversion. He stresses that AI errors are not random but clustered, consistently failing on certain kinds of tasks, so deploying AI is a strategic business decision about how much catastrophic or bad-experience risk is acceptable for the gains offered. He notes professionals like doctors and lawyers proceed responsibly, balancing rewards against risk, and suggests more risk tolerance is reasonable for consumer contexts, while safety-critical, one-shot situations demand much greater caution.

切

Wrappers, UX, and Getting Sherlocked 55:31

Asked about building wrappers around AI models to improve user experience, Bernstein says many current startups are essentially wrappers around frontier models, which is promising short term but risky long term. Good user experience is a real differentiator, but it exposes companies to being sherlocked, a term he compares to Apple absorbing a great third-party Mac feature into its own operating system, as when Anthropic or others launch design tools that swallow smaller wrapper products. He expects the next ten years to be driven by experience improvements fueling AI adoption, and says the strongest position combines both a modeling capability and a strong product and UX capability, though that combination is harder to achieve.

AI-generated summary. It can be wrong or incomplete - check anything that matters against the original.

Summarize your own YouTube video

Paste a YouTube link, article, PDF, ebook or slide deck and get a summary like this in seconds. Free to try, no sign-up needed.

⚔️ Try the YouTube summarizer

Discussion

Sign in to join the discussion. Sign in

More from the Bento Box

Browse the Bento Box →

We use Microsoft Clarity and Google Analytics to see what breaks and where visitors come from. They set cookies and send data to the US. Product events are counted without cookies either way. Cookie details

Webinar: What AI Can and Cannot Do: Intelligence Augmentation in Practice with Michael Bernstein — Summary — Samuraize