Y Combinator

What If We Stopped Using GPUs? | YC Paper Club: summary

YouTube summary19 sectionsWatch on YouTube ↗

This is an AI-generated summary of the YouTube video "What If We Stopped Using GPUs? | YC Paper Club" (Y Combinator), made with Samuraize and published by Samuraize. It condenses the YouTube video into 19 titled sections you can read in a couple of minutes, each linking to the moment in the video it covers.

1
Filed under💻 Technology0 comments🍱 Add to trayReport
Study this
Export

What If We Stopped Using GPUs? | YC Paper Club

Y Combinator

切

An Origin Question About Aliens 0:07

The host opens by describing a dinner conversation where a friend asked what single question you would pose to a superintelligent alien race. Instead of asking about power, the host says the more interesting question is how the aliens compute their flops, their floating point operations. This frames the whole night, which grew out of years spent thinking about hardware and software adapting to each other, often called hardware-software co-adaptation.

切

GPUs Chasing Transformers 1:00

Drawing on experience since 2012, the host traces how GPU design followed the dominant workload of the moment. Through the CNN era, roughly up to 2019 or 2020, Nvidia under Jensen Huang cared mostly about flops per joule, since convolutional networks were not limited by memory bandwidth. Crypto, especially Ethereum mining, was a bigger business driver than AI at the time. After GPT-2 and GPT-3 showed that transformers needed far more memory capacity and bandwidth, Nvidia shifted with chips like Ampere, prioritizing memory over raw compute efficiency, a trend the host ties to attention's quadratic scaling. Progress in gigaflops per joule has stalled over the last two years, which the host says demands new approaches beyond backpropagation, the standard training method, since the human brain runs on about 20 watts, like a light bulb, while doing far more.

切

Why the Brain Resists Backprop 4:03

The host argues the brain cannot be doing backpropagation because neurons fire forward only, and reversing that would require implausible mechanisms like the weight transport problem, where matching weights would have to arise through a different backward pathway. The brain instead seems to use forward-only, Hebbian-like learning supported by heavy inhibition between repeated structures called cortical columns, letting each small assembly learn independently.

切

Testing Gradient-Free Optimizers 8:32

The host describes PhD work testing non-gradient learning rules on a one-billion-parameter LSTM doing next-token prediction, finding that SPSA, a finite-difference method using paired perturbations, worked best and could even escape local minima that trip up backprop.

切

Sharding Small Expert Models 11:02

The host describes an approach called SOMA, clustering a dataset and training many small LSTM experts across GPUs worldwide, then routing between them at test time, noting this only works well when forward passes become cheap, as they would with optical, neural, or analog computing.

切

Introducing the Three Speakers 14:01

The host introduces Ilker, an EPFL PhD now at Stanford working on optical AI computing, Alok, a semiconductor PhD turned venture capitalist running a 1.5 billion dollar fund who will cover neuromorphics, and Sean, a neuromorphics-trained CEO of a company called Parasma.

切

Light Already Moves Our Data 15:00

Ilker begins by noting that optical fiber already carries nearly all data between AI data center racks and 99 percent of intercontinental transfer, because photons have about 10,000 times lower loss and 10,000 times greater bandwidth than electrons, raising the question of whether light could also be used for computing itself.

切

Why Optics and Electronics Are Converging 16:31

You might wonder why we don't just compute with photons if light carries information so efficiently over distance. The speaker explains that electronics and optics are meeting in the middle: digital processors once needed very high precision, like 32 or 64 bits, but AI workloads now tolerate much lower precision, while optical communication has gotten better at supporting higher bit depths. This creates a natural overlap. Light also has a built in advantage for computing because photons do not interact with each other, so many optical beams can share the same space without interference, unlike electrical signals which need separate physical paths.

切

Energy Savings and the Real Costs 19:31

In electronic matrix multiplication, energy use tends to grow with the square of the number of inputs and outputs, since every operation means charging and discharging wires. Optical systems can instead modulate input data directly onto light and let a physical weight mask passively shape it, with the result summed automatically at the detector, so energy use grows only linearly with the number of units, not quadratically. Yet optical computing isn't everywhere, and looking at companies like Lightmatter, which built an opto-electronic board rivaling GPUs, shows that photonics itself is only a small slice of the total energy budget. Most of the cost comes from converting digital data into analog light and back again, from the energy needed to keep optical devices calibrated, and from the difficulty of performing nonlinear operations optically, since passive optics naturally handle only linear math. The speaker's team tried using intense short light pulses in multimode fibers to create optical nonlinearities, but this added real complexity.

切

An Optical Diffusion Model for Images 22:02

The speaker describes a project from EPFL, done with Google, that used optical propagation to perform image generation through diffusion models, which normally work by starting from random noise and gradually transforming it into a realistic image through many repeated passes of a trained neural network, sometimes a thousand times. Instead of running one general purpose neural network over and over, they built a passive, application specific optical device where the physics of light propagation itself performs each denoising step, using layers that act like transparencies shaping the light's phase. Because reprogramming optical weights is hard, they split the generation process into fixed parameter stages, for example dividing 1000 steps into 100 step blocks, so a small number of physical devices could be reused without changing their weights. Tested on small datasets like MNIST and Fashion MNIST, the system behaved like a real diffusion model, producing increasingly realistic images, and its performance scaled with added parameters similarly to digital neural networks, showing a comparable power law relationship while promising better energy scaling. Compared to a similarly performing model on a commercial GPU, the optical system showed a significant energy advantage, assuming the weights were fixed and passively fabricated. Their actual prototype used a spatial light modulator paired with a mirror, letting them train weights by reading camera images and updating via back propagation, avoiding repeated fabrication cycles. The speaker closed by saying the next milestone is building an end to end optical system at a much larger, billion parameter scale, to see whether the energy and latency advantages hold up when scaled, followed by audience questions about analog to digital conversion bottlenecks, how backpropagation is handled using a digital twin, whether diffusion could be treated as a fully recurrent process to reduce digital interfacing, and what physical manufacturing challenges would arise at much larger scales, including the longer wavelength of light requiring larger component sizes and the ongoing difficulty of building nonlinearities into bigger optical systems.

切

Audience Questions on Photonic Limits 34:01

Attendees press the speaker on practical limits of optical computing. One asks whether photonic chips are constrained by the wavelength of visible light, since electrons are far smaller; the answer is that feature sizes in photonics do need to stay larger, though techniques like using hundreds of different wavelengths, known as optical comb multiplexing, let engineers pack in more computation without shrinking the hardware itself. Another question turns to optical storage, asking whether light-based memory could eliminate the need for analog-to-digital converters entirely. The speaker agrees this is a major bottleneck, noting that optics works well for archival storage like CDs and DVDs, but fast, cheap rewritable optical memory has never matched electronics, despite decades of attempts with holographic memories and phase-change materials.

切

Polarization as a Nonlinearity Source 37:30

A questioner suggests using polarized filters, which behave somewhat like trigonometric functions, to build nonlinearity into optical systems, similar to how Fourier methods construct signals from trigonometric building blocks. The speaker finds this promising, explaining that polarization is another degree of freedom in light beyond intensity and phase, and that manipulating it would show up as nonlinearity when measured by standard sensors. He admits uncertainty about how to stack this effect across many layers of a deep network, and says he isn't aware of significant research pursuing this specific idea.

切

Introducing Neuromorphic Computing 39:02

A new speaker, a venture investor with a background in semiconductors and photonics, introduces neuromorphic computing by praising the brain as a remarkable, multimodal, real-time computer that keeps learning for life without needing massive repeated examples, running on roughly 20 watts. He traces the field to Carver Mead at Caltech in the late 1980s, who coined the term after advances in neurobiology and electronics let researchers start mapping brain principles onto circuits. He outlines core brain traits such as intertwined memory and computation, communication through discrete events called spikes, and a mixed analog-digital style of operation, describing a spiking neuron as behaving like a leaky capacitor that fires only once accumulated charge crosses a threshold. He closes by noting that modern neural networks borrow only loosely from biology, since methods like backpropagation and transformer attention have no real counterpart in the brain, and that neuromorphic computing in 2026 remains mostly in early research, still searching for the right matches between biological inspiration and physical hardware.

切

Brain Inspired Hardware Directions 50:30

You can see several active research threads trying to borrow from biology's extreme energy efficiency, since the brain runs on about 20 watts while modern AI systems are far less efficient. One path co-mingles memory and computation, often done purely digitally by combining SRAM and logic, as seen in work from the startup Dmatrix and from IBM. Another path looks at edge devices that need to sense and react, exploring spiking networks, on-device learning, and event-driven computation similar to how the brain works, with Intel's research chip flying a self-piloting drone as an example. A third path starts from new physical substrates like photonics, resistive memories, memristors, and coupled oscillators, building models around their specific nonlinear behavior to co-design something genuinely new. The speaker calls this a thirty-thousand-foot view, noting neuromorphic computing has long been a source of inspiration without yet producing anything resembling the brain in practical use, much like neural nets themselves sat dormant for years before becoming useful.

切

Promising Startups And Packaging Limits 54:31

Asked which direction looks most promising, the speaker points to Dmatrix for co-mingling memory and compute digitally, since it is manufacturable and has a plausible path to market, while flagging Naveen Rao's company, working on coupled oscillators in pure CMOS, as a more ambitious bet on new physics that still respects the constraint of being manufacturable at scale. On packaging, a question about Dmatrix's approach of placing SRAM closer to compute to beat the decode bottleneck leads the speaker to note that the real problem is that interconnect capacitance is far greater than the capacitance of a transistor's gate, so most energy goes into moving bits across wires rather than switching them. The speaker's personal bet for solving this is optics.

切

Brain Limits And Silicon Choice 57:33

One questioner suggests that if neuromorphic computing is trying to recreate the brain, the simplest answer is just to use real neurons. The speaker counters that computers already beat brains in clock speed and storage capacity, so the open question is whether brain-like approaches can be scaled, a direction not yet seriously pursued. On hardware-software co-design, the speaker says the right order depends on whether you start from a physical insight, such as a well-understood optical computer, or from a novel model architecture needing a substrate. Asked about silicon versus carbon, the speaker notes chip materials have grown steadily more exotic, citing an unmentioned team working on diamond substrates, and argues the choice is less about silicon versus carbon than about electronics being our workhorse for computation. On noise, the speaker links it to thermodynamic computing and Boltzmann machines, noting reproducibility of noise, not noise itself, is the real obstacle, a point reinforced by Stanford neuroscientist Shaul Druckmann. The segment ends with an introduction to the Cortical Labs project teaching human brain cells to play Doom by encoding game variables like ammo and health into electrical stimuli and decoding neuron spikes into in-game actions.

切

From Pong to Doom 1:06:01

Cortical Labs had previously gotten brain cells to play Pong, a simple game where engineers could hand design the encoding and decoding because there were only a few possible outputs. Doom is far harder because it is a 3D environment with many enemies and a much larger action space, including strafing, turning, and attacking, which makes hand mapping impossible across only 59 channels. This pushed the team to conclude that for harder tasks, and eventually for GPT level intelligence, the encoding and decoding must be learned end to end rather than hardcoded.

切

Building the closed loop system 1:06:31

The team built a closed loop architecture using proximal policy optimization, a reinforcement learning algorithm, where Doom observations are encoded into stimulation for the neural culture, the resulting spikes are read by a decoder and turned into game actions, and feedback updates happen roughly every 2000 steps. The encoder network takes the game image through a CNN, combines it with scalar observations, and outputs the frequency and amplitude to stimulate each channel, later experiments also let it choose which channels to stimulate. Because gradients cannot pass through the physical electrode array, the team used stochastic, beta sampling based stimulation, which also changes the culture itself each time, making this different from standard reservoir computing. The decoder is a simple linear readout, and the team deliberately undersized it and zeroed its bias after discovering an oversized decoder could effectively play the game on its own by exploiting biases, a shortcut the speaker warns is common in brain organoid research, including recent fly brain connectome work, where silicon decoders do the real work instead of the biological tissue.

切

Feedback, surprise, and scaling questions 1:12:31

Good and bad feedback are delivered through synchronous stimulation for good outcomes and asynchronous stimulation for bad ones, an idea drawn from Karl Friston's free energy principle, where the culture is believed to want to avoid the disorder of asynchronous input. Because early gameplay is mostly bad, the team scaled feedback by surprise, using a critic to judge how unexpected an action's outcome was so that constant negative feedback would not overwhelm the system, and they also penalized entropy rather than encouraging it, unusual for typical reinforcement learning. In audience questions, the speaker clarified that the neurons feel nothing and have no pain receptors, that the system never settles into equilibrium but behaves as two dynamical systems pushing against entropy, that this lab stimulation differs greatly from real sensory signals, and that growing organoid eyes and wiring them to the brain edges toward recreating a human from first principles. Asked about scalability, the speaker said achieving intelligence in cells is understood, but distributing that intelligence to serve billions of people remains an open problem the team is still working to solve.

AI-generated summary. It can be wrong or incomplete - check anything that matters against the original.

Summarize your own YouTube video

Paste a YouTube link, article, PDF, ebook or slide deck and get a summary like this in seconds. Free to try, no sign-up needed.

⚔️ Try the YouTube summarizer

Discussion

Sign in to join the discussion. Sign in

More from the Bento Box

Browse the Bento Box →

We use Microsoft Clarity and Google Analytics to see what breaks and where visitors come from. They set cookies and send data to the US. Product events are counted without cookies either way. Cookie details