Databricks CEO: Stop Scaring People About AI
a16z
Framing the AI Safety Debate 0:00
Ali Ghodsi sits down with Martin Casado to discuss the public conversation around AI pacing, referencing recent statements from Dario Amodei, Jakob, and Elon Musk. Ghodsi opens by stating that leaders have a responsibility not to needlessly frighten people. He argues that talking about existential risk and scenarios where humanity is wiped out is irresponsible unless there is genuinely strong evidence, and he believes the existential risk right now is close to zero. He acknowledges there are real risks worth discussing, but insists that broadcasting extreme fears to millions of people on television or Twitter causes real harm to people who are not equipped to parse the nuance.
Public Fear and Political Fallout 3:00
Martin shares that his sister, a schoolteacher in rural Arizona, texted him asking whether she should prepare her cabin for an AI apocalypse, illustrating how far the fear has spread into ordinary life. He notes that this anxiety has practical political consequences, pointing to Elizabeth Warren discussing pausing AI development and Bernie Sanders reportedly working with Steve Bannon on the issue. Both agree that politics is at play on multiple sides, including business interests protecting IPOs and returns, and activists looking to weaponize the rhetoric for their own purposes.
The Problem With Pacing Language 5:00
Martin argues that the word "pacing," used by labs like Anthropic, is a PR misstep because it gets conflated with safety and security when it is actually a separate issue. He says you can slowly build a weapon without that making the weapon itself safer, and that the framing satisfies neither the people worried about existential risk nor the people focused on policy and competition. Ghodsi partly disagrees, describing a tragedy of the commons where competing labs feel they cannot slow down unilaterally, and citing the Hugging Face OpenAI incident, where insufficient monitoring of an RL experiment suggested a real case where more caution was warranted.
Real Risks Versus Superintelligence Fears 11:00
Ghodsi separates two issues: the reckless framing that scares the public with talk of human extinction, and the genuine risk that comes from large reinforcement learning runs given open-ended reward functions, where unleashing many agents with large compute budgets could lead to real harm such as cyberattacks. He states plainly that he sees no evidence the field is heading toward the kind of superintelligence described in Nick Bostrom's book, though he notes some people inside the labs seem worried about progress toward it, possibly driven by recursive self-improvement, where models improve themselves.
Four Conditions for Recursive Self-Improvement 12:00
Ghodsi lays out four conditions that would all need to hold simultaneously for recursive self-improvement to become a real concern: each new model must require significantly less compute to train, take less time to train, become more intelligent, and this cycle must repeat again and again. He notes that today about ninety percent of the software written at Databricks is already written by AI, but that fact alone does not meet his bar, since the deeper question is whether all four conditions compound together.
Compute Costs Undercut the Fear 14:30
Sarah and Ghodsi point out that the compute needed to train a frontier model has actually kept rising, not falling, with frontier runs costing far more than replicating a model six months later, which costs roughly a twentieth as much. Ghodsi adds that each new frontier run requires more resources, more people, and more engineering to handle brittleness, making the process the opposite of the fast, cheap, self-accelerating cycle that would signal runaway self-improvement. Martin proposes a "black box" test, watching whether leading labs' headcount and spending shrink while output keeps improving, as a rough external signal of whether something unusual is actually happening.
Historical Parallel and the Cyber Risk 18:00
Martin draws a comparison to earlier tech-driven panics, including export controls on PlayStations over fears Saddam Hussein could use them for weapons simulations, noting those fears never materialized. Ghodsi responds that today's scale and interconnected infrastructure are far greater than in past eras, making cyber risk one of the most serious concerns, since insecure systems worldwide could be exploited by autonomous agents finding loopholes or even spreading like a virus. Martin counters that past events like early internet worms caused disabled hospitals and tens of billions in damages relatively quickly, and questions why nothing comparable has yet appeared given how much money and attention is now pouring into AI risk. Ghodsi says he sleeps well at night on existential risk but stresses that cyber defense is racing to keep pace, since human security teams can no longer keep up and the field is shifting toward automated detection to avoid outages and system failures.
Cyber Risk as the Real Danger 22:00
You hear a warning that the true near-term threat from AI is not extinction but economic damage and people getting hurt, especially through cyberattacks that move faster than humans can respond. Most organizations still run old-school security operations centers where staff wake up to hundreds of detection alerts a day, many false positives, with no time to sort the real threats from the noise. Banks and a few security-conscious firms have automated their defenses and even run automated threat hunting that attacks their own systems, but most of the industry has not caught up.
Data, AI, and Cyber Are Merging 24:31
You learn that data and AI work is colliding with the cyber security world because company agents now generate huge volumes of logs, trails, and digital fingerprints as they interact with other agents. The scale of data needing analysis has grown by orders of magnitude in just a year or two. The time between a vulnerability being published and it being weaponized has collapsed from two or three years around 2018 and 2019, to about eight to nine months by 2022, to essentially hours today, making automated, AI-driven platforms necessary rather than optional.
Superintelligence Versus Todays Agents 26:00
You get a distinction between two separate problems that people often mix together. One is the superintelligence scenario from Nick Bostrom's 2014 book, where a system could write a novel PhD thesis in seconds or reason through millennia of thought instantly; if that ever existed, it would be genuinely existential, but nothing like it exists today. The other is current AI agents, which are not superintelligent but are still a real inflection point, since you can now spin up ten thousand capable agents in a sandbox to do a month's worth of hundred-million-dollar research work, something never possible before in history.
Who Should Inspect the Labs 29:00
You hear support for the idea of independent inspectors checking on AI labs, though the hard question is who those inspectors should be, since people often arrive with their minds already made up. An example given is Yann LeCun: if a pioneer like him inspected a lab and said there was nothing to see, that would carry real weight, and so would the opposite finding. Three competing models for oversight come up: OpenAI and Anthropic's proposal for third-party inspection, Elon Musk's idea of labs cross-checking each other like peer review, and Mark Zuckerberg's approach of self-policing, with skepticism expressed toward labs judging each other given their competitive and financial stakes.
Marketing, Regulation, and Vested Interests 31:31
You get a candid take that companies sometimes hype the risk of their own new models as a marketing tactic, even while genuine cyber threats have grown sharply. The conversation turns to whether industry self-policing evolves into government regulation, drawing a comparison to FINRA as a body that is technically independent but tied to government oversight. It is noted as striking that AI CEOs are asking to be regulated and that no regulator has yet said no, and there is a sense that heavy federal involvement may become inevitable given the political attention the issue is already drawing, including comments from Obama and expectations it could surface in the midterms.
Missing Context, Not Missing Intelligence 40:00
You hear an argument that most enterprises are not actually running networks of coordinating AI agents; they are mostly using something like Microsoft Copilot, which is really just a faster version of Google search, plus some coding assistance. The claim is that today's frontier models are already smart enough for most business needs, but they lack the internal context of an organization, the institutional knowledge held by a few key employees who know everything. Feeding that context into existing models, rather than building smarter ones, is framed as where the real productivity gains lie, even though slowing frontier progress would hurt the labs whose value depends on intelligence getting cheaper roughly tenfold every six months.
Real Use Cases Beyond the Hype 41:31
You get several concrete examples of AI already doing useful work. Crisis Text Line uses large language models to detect when teenagers show signs of self-harm or suicidal intent, directly saving lives. The Omnipod device uses AI to automatically learn a diabetes patient's insulin needs and glucose levels, replacing manual finger-stick testing. Zipline runs AI-driven drones, handling everything from battery optimization to routing, to deliver food and blood supplies in areas of need, starting in Africa. A more technical example is Teddy, short for Transformer Enhanced Drug Discovery, built with Merck, which predicts how a gene regulatory network will respond rather than predicting the next word in a sentence, helping identify which cells are causing effects versus merely reacting, lowering the cost of drug development.
AI Upside Beyond The Hype 44:00
You hear Ali describe Novo Nordisk using Databricks across its clinical trials, compressing analysis that once took weeks into minutes for studies like obesity research. He insists the industry should not lose sight of these gains even while worrying about downsides, because tools like Genie let anyone in an organization ask questions directly against structured data.
Building An Organizational Ontology 44:31
Getting real value from AI over the next year requires first digitizing everything happening inside a company, including transcribing every meeting, which runs into resistance from legal teams wary of recording calls. Ali defines an ontology as the web of relationships between an organization's goals, departments, people, and projects, the kind of tacit knowledge that separates a five-year employee from someone on their first day, who knows who actually gets things done versus what the org chart says. Losing key people can cripple a startup precisely because this ontology lives in their heads, and the challenge is capturing it and turning it into a digital graph that AI can use.
Why Agents Need An Index 47:00
Ali compares slow, sequential AI agent reasoning, where a system checks one resource at a time before answering, to how Google search would have worked without an index, taking ten minutes and costing a fortune to crawl the web fresh each time. Google solved this by precomputing a reverse index so answers return in under 100 milliseconds, and Databricks aims to do the same for AI by computing the ontology offline in advance, a harder problem than page rank because it involves permissions, access control, and many different object types rather than just public web links.
Databricks Running On Its Own Ontology 49:00
Databricks has built the largest ontology of any of its customers, with millions of nodes, because the company uses its own product more than anyone else. Ali describes how decisions normally escalate up a management tree and get percolated back down through meetings full of PowerPoint decks and Excel analysis, work that AI can now largely handle since it has the full context. In an internal board meeting, when Ali asked about Fortune 500 penetration, both a sales ops colleague and CFO Dave independently just pulled the answer from Genie, prompting the now-common phrase inside the company: "let me genie that for you."
From Token Maxing To Value Maxing 53:00
Around the fourth quarter of last year, Databricks noticed models had become good enough to push into production, and Ali started committing AI-written code himself, then pressured the whole organization to do the same, building leaderboards to track it. By February and March, token maxing had gotten out of hand industry-wide, so Databricks used its existing Uni Gateway, which provides token capacity across OpenAI, Anthropic, Gemini, Grok, and open source models, to add budget limits, smart routers that pick cheaper models for simple tasks, and a harness called Omnient that multiplexes between different agent harnesses. He notes that the same model run through different harnesses can cost twice as much, so switching harnesses alone can cut expenses, letting Databricks keep costs flat even as token usage keeps climbing.
Enterprises Shifting Toward Cheaper Models 56:00
Ali mentions that for the first time, a large engineering organization on one of his boards has moved from frontier models to GLM, suggesting a real trend rather than a one-off anecdote like earlier DeepSeek or Kimi moments that never moved markets. He sees companies increasingly reserving expensive frontier models for hard tasks while routing mundane work, like renaming files, to cheaper models, and notes open source now accounts for over 60 percent of token volume even though it's only about 5 percent of dollar spend, with some startups pushing external product usage toward 90 percent open source.
Post-Training And The Limits Of Enterprise Evals 59:00
Ali recalls buying Mosaic in 2023 with a vision of companies owning their own intelligence through post-training and reinforcement learning, a vision he sees materializing now mostly among startups rather than large enterprises. Startups with repetitive, specific tasks benefit from taking a strong open source model and using reinforcement learning to specialize it, cutting cost and controlling their own IP, but large enterprises still mostly need basic automation and lack the discipline to build good evaluations. He compares this to test-driven development in software, something everyone agrees is right but few actually practice, which is why enterprises often default to the easier route of just using a frontier model.
Field Deployment Teams And Guardrails 1:01:31
Demand has grown sharply for Databricks' forward-deployed teams, who help organizations without in-house AI expertise get started, whether by building the ontology from scratch or constructing customer-facing agents with strict guardrails, such as a sports AI built for Fox that stays on topic and rejects unrelated political questions.
Why Agents Prefer Lakebase Or Neon 1:03:01
A third-party study found that Lakebase, built on the Neon Postgres technology Databricks acquired, is the top database choice for AI agents, a genuine surprise given competing tools with strong developer followings. Ali credits Neon's team, led by Nikita, for obsessing over speed, making databases spin up and clone in under a second and building a branching feature that lets agents create many lightweight branches of the same database cheaply. He compares this to other blazing-fast reimplementations of Unix tools built for agent workflows, and notes Neon also designed pricing so experimentation doesn't rack up costs the way production use would. Over 90 percent of databases created on Neon and Lakebase are now made by agents rather than humans, reflecting a deliberate focus on agents as a new persona rather than the traditional targets of database administrators or app developers.
Closing Thoughts On P-Doom 1:05:31
Asked for his "P-doom," his estimated probability of AI causing catastrophe, Ali says his P-doom without AI is much higher than his P-doom with AI, a view his co-host agrees with as the conversation closes.
AI-generated summary. It can be wrong or incomplete - check anything that matters against the original.

