Skip to content

Semi Doped

Vikram Sekar and Austin Lyons
Semi Doped
Latest episode

50 episodes

  • Semi Doped

    Every Muse User Gets 2 CPUs and 8GB in Meta's Cloud. Your Phone Already Has That

    10/01/2026 | 50 mins.
    The push to run AI agents on edge devices will not be driven by model labs or hardware vendors, but by service providers like Meta whose ad-based business models create a powerful financial incentive to offload massive cloud compute costs onto the user's own powerful-but-idle hardware. Austin Lyons and Vik Sekar break down why the failed "AI PC" push has given way to a new, economically-driven strategy for on-device AI. They discuss how Qualcomm's acquisition of Modular provides the key software layer and why the ultimate vision is an "ecosystem of you" where the device acts as a context-aware orchestrator between local and cloud compute.
    This episode is presented by Crusoe. Try Crusoe's managed inference for open models with $5 in free credits: https://semidoped.com/crusoe
    "[Model Labs] business model is we make money when we do inference in the cloud. And yet we are saying, actually, there's a world where we would really like to just run this stuff locally... Meta's like, bro, I want to help you at the edge."
    — Austin Lyons, Chipstrat
    Key Takeaways:
    - The driver for edge AI is a business model conflict. Model labs like OpenAI sell cloud inference; service providers like Meta see it as a massive cost center for their ad-supported agents like Muse.
    - Your phone is an underutilized edge server. Its specs—like 8GB of RAM and 2 vCPUs of power—are comparable to the dedicated cloud VMs that services like Meta's Muse provision for each user.
    - The "AIPC" push failed because it was a hardware-first strategy without a use case. Today's edge AI push is a software-first strategy driven by the need to cut cloud OpEx for mass-market agents.
    - The role of the edge device is becoming a "context translation layer." It uses private data like location and biometrics to intelligently decide which AI tasks run locally versus in the cloud.
    - Qualcomm's acquisition of Modular is a key strategic move. It positions the open-source Mojo language as the "write once, run anywhere" software layer for deploying agents across diverse hardware.
    - Personal agents will lead to the "end of apps." Fixed UIs, which force a designer's paradigm on the user, will be replaced by interfaces generated on demand via natural language.
    - Google is missing its opportunity to lead in on-device AI. Despite vertical integration from TPUs to Pixel hardware and a leading ad business, it lacks the product execution of players like Meta.
    Chapters:
    0:00 Who Makes Local AI Happen?
    1:25 Making the Intro with AI
    4:43 Qualcomm's Agent Vision
    6:54 Why AIPCs Failed
    8:45 The Phone as a Portal
    13:52 Meta's Incentive for Edge AI
    24:05 The End of Apps & UIs?
    32:31 The Ideal Personal Agent
    34:25 The Edge as a Context Layer
    38:17 Qualcomm's Modular Acquisition
    43:16 The Ecosystem of You
    48:01 Google's Missed Opportunity
    Follow Chipstrat:
    Newsletter: https://www.chipstrat.com
    X: https://x.com/chipstrat
    Follow Vik:
    Newsletter: https://www.viksnewsletter.com/
    X: https://x.com/vikramskr
    Follow Semi Doped:
    Get more of Austin and Vik daily, free: https://daily.semidoped.com/
  • Semi Doped

    Flash Hit Its Physical Limit. ReRAM Is the Replacement: WeeBit Nano's Ilan Sever

    09/29/2026 | 56 mins.
    As embedded flash hits a physical and economic scaling wall below 22nm, a new on-chip memory is needed. Ilan Sever of WeeBit Nano joins to explain how Resistive RAM (ReRAM) is positioned to become the new standard by offering a low-cost, high-endurance alternative. Ilan details how ReRAM's simple back-end-of-line integration, automotive qualification, and future applications in AI in-memory compute create a compelling replacement for the incumbent technology.
    "In geometries below 28 or 22 nanometer, flash reached its physical limit, actually. So physically you cannot implement embedded flash anymore and the industry had been looking for alternative technologies for a few decades now."
    — Ilan Sever, WeeBit Nano
    Key Takeaways:
    - The inflection point for on-chip memory is physical. Embedded flash cannot scale below the 28/22nm node because its floating gate becomes too small to reliably hold charge.
    - The economic problem with eFlash is its 25-30% cost increase, a result of adding 8-10 extra masks to a standard CMOS process—a non-starter for cost-sensitive MCUs.
    - ReRAM's manufacturing advantage is its simplicity. It's built in the back-end-of-line (BEOL) metal layers, adding only two masks and not interfering with the sensitive transistors below.
    - The physics are different: ReRAM isn't storing charge, it's changing resistance by using voltage to form or dissolve a conductive filament of oxygen vacancies in a dielectric.
    - Automotive qualification is the key market validation. WeeBit's ReRAM is qualified for the AEC-Q100 standard and demonstrates 100,000 write cycles—10x that of typical eFlash.
    - The business model isn't just a recipe; it's a full IP module with the bit cell, analog circuits, digital logic, and programming algorithms, which is what SoC designers actually need.
    - Beyond replacing flash, ReRAM (as a memristor) is a foundational technology for future in-memory compute (IMC) AI accelerators that perform calculations inside the memory to save power.
    Chapters:
    0:00 The Overlooked Memory
    4:15 The eFlash Scaling Wall
    6:03 Four Uses of NVM
    23:19 The Cost of eFlash
    27:29 Introducing ReRAM
    32:25 How ReRAM Works
    37:12 Fab-Friendly Materials
    42:14 Automotive Qualification
    44:30 The Vertically Integrated IP Model
    51:13 Future: In-Memory Compute
    54:19 Innovation Story: Self-Trimming
    Follow Chipstrat:
    Newsletter: https://www.chipstrat.com
    X: https://x.com/chipstrat
    Follow Vik:
    Newsletter: https://www.viksnewsletter.com/
    X: https://x.com/vikramskr
    Follow Semi Doped:
    Get more of Austin and Vik daily, free: https://daily.semidoped.com/
  • Semi Doped

    Qualcomm's Durga Malladi: HBC vs HBM, Dragonfly AI 250, and Winning on TCO

    09/24/2026 | 20 mins.
    Qualcomm re-entered the data center inference market late — competitive GPUs, XPUs, and custom accelerators are already out there. Durga Malladi (EVP/GM, Technology Planning, Edge & Data Center) explains why they didn't build another me-too accelerator and instead placed a differentiating technical bet: High Bandwidth Compute (HBC).
    The difference between HBC and HBM: HBM is stacked DRAM that shuttles data across wide buses to a separate accelerator — very high bandwidth, but expensive and power-hungry. HBC bonds stacked DRAM directly on top of a custom logic die using wafer-on-wafer techniques, so a lot of the compute runs literally next to memory. The result is lower latency and lower power per bit — and 18x effective bandwidth on the AI 250 versus the AI 200 at the SAME 768 GB LPDDR capacity and the SAME 160 kW per rack.
    Why the late-mover bet works: Qualcomm's differentiation isn't just architectural. Their foundry and memory-vendor relationships let them credibly promise the multi-megawatt-to-gigawatt supply hyperscalers need. Memory vendors are developing their own HBC-equivalent variants as partners rather than pure competitors, and Qualcomm expects the two technologies to co-evolve. Silicon is back in the lab; proof points land in Q1/Q2 2027.
    Chapters:
    0:00 Introduction
    2:49 The Memory Wall Problem
    8:05 HBC Physical Architecture
    9:33 HBC vs. Custom HBM
    11:50 Silicon Proof Points Coming
    12:08 Managing Thermals
    13:34 Dragonfly AI 250
    14:48 Trillion-Parameter Model on One Card
    16:56 Tokenomics and TCO
    18:26 Why TCO is Key
    19:26 Manufacturing at Scale
    Follow Chipstrat:
    Newsletter: https://www.chipstrat.com
    X: https://x.com/chipstrat
    Follow Vik:
    Newsletter: https://www.viksnewsletter.com/
    X: https://x.com/vikramskr
    Follow Semi Doped:
    Get more of Austin and Vik daily, free: https://daily.semidoped.com/
  • Semi Doped

    Pacing AI Means More Compute, Not Less

    09/18/2026 | 36 mins.
    The debate over "pacing" AI for safety is a red herring, as the practical outcome of more safety work is not a slowdown but a net increase in demand for compute hardware. Austin Lyons and Vik Sekar unpack Dario Amodei's call to slow down frontier AI development, arguing that the true governors on AI are not voluntary agreements but physical data center bottlenecks and geopolitical realities. They conclude that the push for more safety and interpretability will ultimately drive more sales of both training (GPUs) and inference (XPUs) hardware.

    Today's sponsor is G2i — expert training data, evals, and RL environments for AI labs: https://fandf.co/4wZvKHS

    Chapters:
    0:00 Introduction: Pacing AI
    1:36 Recapping Dario Amodei's Essay
    7:42 Hot Take: Anthropic vs. OpenAI
    11:34 The John Deere Playbook
    15:22 The Problem with Interpretability
    17:53 Pulling the Plug
    19:03 The Real Bottleneck: Data Centers
    22:03 The Geopolitical Blind Spot
    23:41 Counter-narratives: Meta & Open Source
    27:26 The Investment Thesis: Pacing Increases Compute
    32:29 The Financial System Analogy
    36:05 AI Safety as the New Cybersecurity

    Follow Chipstrat:
    Newsletter: https://www.chipstrat.com
    X: https://x.com/chipstrat

    Follow Vik:
    Newsletter: https://www.viksnewsletter.com/
    X: https://x.com/vikramskr

    Follow Semi Doped:
    Get more of Austin and Vik daily, free: https://daily.semidoped.com/
  • Semi Doped

    TSMC Will Buy High NA After All: ASML's New 6x12-Inch Mask

    09/14/2026 | 50 mins.
    The emergence of user-friendly AI agents is creating massive new hardware demand, which in turn is forcing the semiconductor industry into a complex, multi-year transition to High NA EUV lithography and an entirely new 12-inch photomask standard. Austin Lyons and Vik Sekar unpack the physics and economics of ASML's move to High NA, explaining why it necessitates a shift from 6-inch to 6x12-inch masks and what this means for the timelines of TSMC, Intel, and Samsung.
    "This is where you have to get the whole supply chain coordinated around this. And this is why you really ultimately need ASML's biggest customers to stand up and say, 'We're going to buy this.'"
    — Austin Lyons, Chipstrat
    Key Takeaways:
    - ASML's High NA EUV (0.55 NA) solves resolution but creates a new problem: anamorphic optics (4x by 8x demagnification) cut the exposure field in half, effectively doubling the cost per wafer.
    - The industry's fix for High NA's halved output is a new 6x12-inch photomask standard — the 12-inch dimension compensates for the 8x demagnification, restoring the full 26mm x 33mm reticle size.
    - This shift to 12-inch masks is a full supply chain problem, which is why TSMC is waiting until 2030 for high-volume manufacturing while the ecosystem matures over the next 5-7 years.
    - Intel's aggressive first-mover strategy on High NA — with over 1 million wafers processed to date — is a direct reaction to its disastrous delay in adopting the previous generation of EUV.
    - Nearly 30% of ASML's revenue comes from recurring services (€2.8B of €9B in Q2), giving it a more stable financial profile than a typical equipment manufacturer.
    - The new wave of AI agents (Astra, Muse, Instinct) is moving from simple instruction-following to proactive partnership, creating hundreds of millions of new users for server CPUs and memory.
    - Integrating AI agents into ubiquitous text platforms like WhatsApp is key for mass adoption, as it solves the 'blank text box problem' for non-technical users who can simply text a request.
    Chapters:
    0:00 Episode Opening
    2:23 The Rise of AI Agents
    6:00 Instinct: The Autonomous Agent
    10:50 Productizing AI for Mass Adoption
    14:03 From AI Agents to ASML
    24:29 The Photomask Stitching Problem
    31:49 High NA's Anamorphic Optics Tradeoff
    36:39 The Cost of a Halved Reticle
    37:26 The 6x12-Inch Mask Solution
    39:08 A Full Supply Chain Problem
    41:17 TSMC, Samsung & Intel Timelines
    48:24 Intel's EUV History Lesson
    Follow Chipstrat:
    Newsletter: https://www.chipstrat.com
    X: https://x.com/chipstrat
    Follow Vik:
    Newsletter: https://www.viksnewsletter.com/
    X: https://x.com/vikramskr
    Follow Semi Doped:
    Get more of Austin and Vik daily, free: https://daily.semidoped.com/
More Technology podcasts
About Semi Doped
The business and technology of semiconductors. Alpha for engineers and investors alike.
Podcast website

Listen to Semi Doped, Dwarkesh Podcast and many other podcasts from around the world with the radio.net app

Get the free radio.net app

  • Stations and podcasts to bookmark
  • Stream via Wi-Fi or Bluetooth
  • Supports Carplay & Android Auto
  • Many other app features