The Ads Business Model Will Die & Lessons from Working with Elon at Twitter | Parag Agrawal
20VC with Harry Stebbings
Parallel: Google for Agents 0:00
Parag Agrawal describes Parallel as a search engine built for AI agents rather than humans, since agents need to search the web to complete almost any task they are given. The founding insight was that agents will use the web a thousand times more than humans do, a scale shift that means existing search technology cannot survive and both new technology and new business models are needed. Handling a thousand-fold increase in search volume would require far too much compute if built on today's approach, so the underlying system needs to become roughly ten to a hundred times more efficient.
How Agent Search Differs From Human Search 3:31
Humans type short, underspecified keyword queries, wait about half a second to a second, and skim ten blue links, but agents behave very differently. An agent might state a full sentence describing exactly what it needs, and it may be either extremely impatient, such as a voice agent needing an answer in 100 milliseconds, or willing to wait much longer as a background process focused only on getting the best answer. Because inputs, outputs, and time constraints all differ from human search, and outputs come back as tokens or files rather than links, the amount of compute worth spending varies enormously depending on the agent and the model behind it.
Allocating Compute Between Search and Model 7:30
Web search works by narrowing a few trillion documents down to about a thousand tokens a model can actually read, using cheap retrieval first and then progressively more expensive ranking on smaller and smaller sets of documents. The right amount of compute to spend on this narrowing depends on how expensive the model reading the result is: a cheap model can tolerate noisier input, while an expensive model justifies spending more on search first to avoid wasting its time. Parallel lets a program or the agent itself specify priorities like accuracy, latency, or cost through its API, and offers different modes such as a very fast, cheap mode called Turbo built for voice agents, and a slower, more thorough mode called Advanced for expensive background agents that can afford to think longer.
Where Agent Search Is Actually Used 10:02
Customers span many kinds of knowledge work, including coding, law, insurance underwriting, sales, and science, rather than being dominated by any one category. Coding is a large share of overall AI inference, but web search is only invoked in about five percent of coding prompts since most coding relies on internal codebases. Law, by contrast, is very web-search-heavy because it depends on case law and verifiable facts about people and companies, and insurance underwriting has a similar profile.
Why Smarter Models Still Need Search 12:30
Better and cheaper models are good for Parallel's business because they let agents do more, and agents are the customers Parallel builds for. Smarter models do not necessarily reduce search demand, because model intelligence and memorized parametric knowledge are separate things. Models tend to recall well-known facts about famous people or events but cannot memorize everything from training data, since their internal memory is a lossy compression that captures patterns rather than exact facts, and this effect gets stronger as models are shrunk down while keeping their reasoning ability.
Bigger Frontier Models and Open Models 16:01
Parag expects frontier models to keep growing larger while, at the same time, smaller models become able to match any fixed level of past performance within a matter of months. Larger frontier models will keep improving because some use cases are valuable enough to justify paying for incremental quality, and there appears to be no visible end yet to this scaling trend, including letting models think longer by spending more compute per problem. He declines to predict a fixed split of token activity or dollars between open and frontier models, saying it depends heavily on path-dependent factors, though he expresses hope that a strong American-built open model will emerge and ideally face real competition.
Why routing layers matter now 18:00
Parag explains that so-called routing, choosing which model and which GPU vendor serves a request, has real value today because GPU and token supply are unpredictable. Fast-growing startups often need more capacity than they forecasted, so having a flexible router that can switch models or providers when one runs short is genuinely useful. Whether this stays valuable depends on how GPU and token supply and demand evolve.
Data as the next valuable asset 20:01
He argues that unique data or insight will become far more valuable once markets figure out how to price it, especially at inference time rather than just training time. He uses the example of a venture capitalist who pays for a seat on Pitchbook, but whose AI agents cannot easily use that same subscription, forcing clunky workarounds. If companies could price data access for agents properly, it would unlock a large market; otherwise, data providers and agents will be stuck in an inefficient standoff.
Advertising model breaks down with agents 27:00
He warns that agents replacing human browsing will gut advertising, since ads only work when a human actually looks at the screen; Amazon's ad business, now bigger than its e-commerce business, illustrates the stakes. His answer is a kind of AdSense for agents, paying content owners a variable amount whenever an agent draws value from their content, calculated to match what a human visitor's attention would have earned through ads. He rejects the comparison to music streaming, arguing that because agents will use the web a thousand times more than humans do, the overall pie grows large enough for new business models to work, rather than shrinking the way music revenue initially did.
Margins, content payouts, and revenue scale 31:30
Asked about margin sustainability, he says the business focuses obsessively on quality, cost, and latency, achieving the same quality as human-built systems at roughly one 20th to one 150th the compute cost, with future margins depending mostly on competition rather than payouts to content owners. He describes paying content owners based on their marginal contribution to an agent's answer quality, comparing it to a cost curve where removing one source's data would only slightly lower quality, and that saved cost gets redirected to the content owner. On overall revenue, he frames the business as an adjacency to inference, estimating that 5 to 20 percent of all GPU spending on running agents will need to go toward a web search stack, so growth should roughly track the broader inference market's rapid expansion.
Path to a $100 Billion Business 37:00
Parag walks through the math behind reaching a hundred billion dollar business, starting from a 33 percent share of a market that could be worth six to seven billion dollars today. He says that figure, growing at the rate of inference demand, could scale to a hundred billion within a few years, though he stresses this depends on strong execution and a few things falling the right way. He identifies the main risks as agents failing to deliver real profit and value, society overspending on GPUs as a result, and the open question of whether his team builds the best technology, partners with enough content providers, and earns customer trust. He also notes real uncertainty around Fireworks and open models, since if open models dominate, revenue could shift heavily toward them instead.
Web Search Pricing Is Broken 39:33
Discussing commoditization and benchmarking, Parag argues the market has misjudged how much to spend on web search. He says once a model can handle sixty to eighty percent of a personal agent's work through many searches, spending eighty to ninety percent of the budget on web search would be absurd, meaning prices need to drop by an order of magnitude or more. He traces today's ten dollars per thousand searches back to old Google ads pricing logic, and claims his own product delivers the same quality for about one dollar per thousand, with another tenfold price drop possible within three years, even as usage grows far faster, echoing Jevons paradox.
Push Search and Agent Guardrails 44:59
He describes an internal alerting tool that flags newly registered European companies founded by people under twenty five, using it as an example of turning web search from something agents pull on demand into something the web pushes to them automatically through constant crawling, cutting compute costs sharply. Asked about guardrails, Parag distinguishes models mid reinforcement learning, which have caused most reported hacking incidents, from properly aligned released models, calling alignment an unsolved but partly effective effort. He argues hacks should be treated as embarrassments revealing insufficient safeguards, not badges of honor, and says he worries AI could be net positive yet still get diffused and used badly by a world that stays too concentrated.
Trust in AI agents is early 56:32
Parag compares giving an AI agent access to your bank account, login details, or email today to how strange it once sounded to put a credit card online or to find love online. He says the technology for agents to spend money or manage accounts already exists, but social acceptance lags far behind, and most people are nowhere near trusting an agent with that kind of access, even if bubble-dwelling tech insiders already do.
Wealth gaps and vertical integration 59:00
Asked about widening wealth inequality, he says some disparity is fine but worries the world is trending toward too much, though he hopes real corrective forces will kick in. On strategy, he acknowledges the pull toward vertical integration, as seen with Meta owning compute, chips, and the application layer, but argues that fully closing yourself off risks boxing you out, so his own company deliberately stops at the API layer to stay horizontal and reach a wide range of agents.
Elon, quickfire answers, and closing thoughts 1:01:32
On working with Elon Musk, Parag admires his urgency and ability to compress time, arguing that unreasonably high expectations often push people to realize they're capable of more than they assumed. In a rapid quickfire round he explains he never personally invests due to his wife's VC compliance rules, praises Vinod Khosla's technical intuition, and says he's grown to value sales and marketing far more than he once did. He distinguishes Perplexity, which he sees as a vertically integrated competitor rather than a search rival, from XAI, which he calls a direct competitor, and names the moment Josh and Todd flew out to meet him as his best VC meeting ever. Looking ahead, he says he's most excited for the chaos of rapid change over the next decade.
AI-generated summary. It can be wrong or incomplete - check anything that matters against the original.

