
Top AI Podcasts for Engineers: 10 Shows Worth Your Time in 2026

AI podcasts are a dime a dozen these days, but they vary quite a bit in how much an engineer can take back from them and use in actual work. A lot of the shows will spend most of an episode reacting to the week's releases, whereas far fewer really go into questions deserving of an "add to queue": how inference systems are built, why agents fail, what changes in post-training, how evals are designed, or what it takes to run AI systems in production.
For engineers, the best podcasts on AI are the ones that repeatedly get into decisions, methods, and failures that are hard to pick up from headlines alone. That can mean proper time devoted to inference runtimes, a paper discussion on model behavior, or a detailed account of how coding agents are being built and used.
We checked what's current in September, 2026 and kept the main list to only the most worthwhile shows, all of them with a recently published episode.
So, without further ado, below are ten top AI podcasts for engineers, with a recent episode to start from and a clear indication of what each show spends its time on.
Top AI Podcasts for Engineers in 2026
This is a 'best for…' rather than your typical 'best of' listicle. Here's a brief overview table to narrow the list by the subject the podcast covers:
| If you want to follow... | Start with... | Why |
|---|---|---|
| AI engineering, agents, and inference | Latent Space | Frequent conversations with people building models, infrastructure, and agent systems |
| Applied AI and production architecture | Practical AI | Implementation-focused discussions at a manageable technical depth |
| ML research connected to production | TWIML AI Podcast | Research interviews with regular attention to inference, evaluation, retrieval, and deployment |
| Deep technical research discussions | Machine Learning Street Talk | Long episodes centered on papers, model behavior, and AI research questions |
| MLOps and production infrastructure | MLOps Community Podcast | Frequent coverage of inference cost, observability, agents, and production operations |
| Open models and post-training | Interconnects | Detailed coverage of RL, distillation, open models, and model development |
| Frontier research and AI labs | Dwarkesh Podcast | Long interviews on training, compute, continual learning, and automated AI research |
| Coding agents and AI-assisted software engineering | The Pragmatic Engineer | Connects current AI tooling with day-to-day software engineering |
| Agents, retrieval, and model capabilities | The Cognitive Revolution | Long-form coverage of agent systems, memory, evaluation, and frontier models |
| AI within the general software stack | Software Engineering Daily | Connects AI with databases, distributed systems, security, and infrastructure |
You probably don't need all ten. Pairing one research-heavy feed with one production-oriented show usually gives more variety than subscribing to several podcasts reacting to the same releases.
Latent Space: The AI Engineer Podcast
Focus: AI engineering, agents, inference, infrastructure, frontier models
Typical length: ~60–90 minutes
Publishing cadence: Several episodes per month
Start with: Humanity's Last Invention — Richard Socher of Recursive
Latent Space describes itself as a podcast "by and for AI Engineers." Its current feed covers agents, inference, model training, AI infrastructure, multimodality, and frontier research.
Guests are often directly involved in the systems being discussed. Richard Socher's September episode, for instance, digs into what he calls the "Eureka Machine" — a system meant to automate scientific invention itself — and gets specific about reward hacking, why he argues constitutional AI hasn't held up, and where he thinks current LLM architecture runs out of road.
Tune in if: you want technical conversations about building and running modern AI systems.
Skip if: you want short AI-news summaries or introductory machine learning content.
Practical AI
Focus: Applied AI, architecture, agents, MLOps
Typical length: ~45–60 minutes
Publishing cadence: Weekly
Start with: Computer-Use Agents and the Future of the Agentic Internet
Practical AI focuses on applied work across AI architecture, computer-use agents, multi-agent systems, and production operations. Its September 2026 feed includes episodes on agentic internet infrastructure and architecture for production AI.
The conversations usually stay centered on implementation decisions rather than open research questions — what changed after a production incident, which framework choice held up under load, how an integration actually got built. If you're building agent workflows yourself, that kind of detail is easier to act on than a research-only interview feed.
Tune in if: you want practical, implementation-level discussion of applied AI, architecture, agents, and production systems.
Skip if: you're after deep paper analysis or long research discussions.
The TWIML AI Podcast
Focus: ML research, inference, agents, RAG, applied AI
Typical length: ~45–75 minutes
Publishing cadence: Several episodes per month
Start with: Do AI Tokenomics Matter More Than Model Benchmarks?
The TWIML AI Podcast has published more than 775 episodes and continues to cover both research and production AI. Current topics include inference engineering, model economics, spatial AI, agent failures, multi-agent systems, and RAG.
Research interviews often connect the paper or model back to deployment, evaluation, infrastructure, or cost. The September episode with Stanford's Christopher Potts questions whether tokens are a meaningful unit of value at all as reasoning models spend more of them per answer, and gets into what that means for measuring return on AI spend beyond a benchmark leaderboard. A separate inference-engineering episode with Philip Kiely, from earlier in the year, covers batching, quantization, speculative decoding, KV-cache reuse, and serving runtimes in more implementation detail.
Tune in if: you want a steady mix of ML research and production-oriented AI discussion.
Skip if: you want a feed devoted to one narrow part of the AI stack.
Machine Learning Street Talk
Focus: AI research, cognitive science, model behavior, technical papers
Typical length: ~60–120+ minutes
Publishing cadence: Multiple episodes per month
Start with: How Replication Could Teach Machines What Good Science Looks Like — Edward Hughes
Machine Learning Street Talk (MLST) spends long episodes on individual research questions, papers, cognitive science, philosophy of mind, and model behavior. Tim Scarfe and recurring co-host Keith Duggar regularly push guests on assumptions and methodology instead of compressing the discussion into a short interview.
September 2026 alone includes episodes on physical AI, speech recognition, scientific replication, and AI development scenarios. The Edward Hughes conversation runs for roughly two hours and uses replication as a route into scientific judgment, creativity, and open-ended research.
Tune in if: you want long, technically demanding discussions of AI research and the ideas behind it.
Skip if: you want concise production advice or a quick way to follow AI news.
MLOps Community Podcast
Focus: Production AI, agents, infrastructure, MLOps, observability
Typical length: ~30–60 minutes
Publishing cadence: Several episodes per month
Start with: Why Cost Per Million Tokens Is a Useless KPI
The MLOps Community Podcast is heavily oriented toward running AI systems in production. Its current collection covers inference cost, infrastructure efficiency, observability, retrieval, agent operations, and the economics of scaling AI workloads.
The September conversation with Palo Alto Networks' Kuntal Patel and Abhinav Lad, joined by the Linux Foundation's Alex Salkever, argues that a single price-per-token number hides what agentic systems actually cost — retrieval, vector storage, and egress all add up separately, and agents consume tokens in bursts that don't follow typical forecasting curves.
Their fix was persona- and use-case-based cost tracking instead of one blended metric, plus AI gateways that catch runaway agent loops before they show up on a bill. An earlier September episode with Pinterest principal engineer Ambud Sharma covers similar ground from the infrastructure side — hardware procurement, model selection, and inference engines as one governance problem.
Tune in if: you want practical discussion of operating AI systems and agents in production.
Skip if: you're after model research, paper analysis, or theoretical ML.
Interconnects
Focus: Open models, post-training, RL, model development
Typical length: ~35–70 minutes
Publishing cadence: Irregular audio releases
Interconnects, hosted by Nathan Lambert, mixes audio essays with interviews on open models, post-training, reinforcement learning, distillation, and model development.
The audio feed publishes less often than the written publication. Its latest verified podcast episode at the time of writing this post is the July 22 open-model roundup with Florian Brand, which covers Chinese model labs, distillation, the economics of open models, and the state of the U.S. open-model ecosystem. Interconnects continued publishing written AI research coverage throughout September.
Tune in if: you want detailed coverage of open models, post-training, and model development.
Skip if: you want a predictable weekly audio feed or production-infrastructure discussions.
Dwarkesh Podcast
Focus: Frontier AI research, scaling, compute, continual learning, AI labs
Typical length: ~60–150 minutes
Publishing cadence: Several episodes per month
Start with: AI researchers debate how close we are to recursive self-improvement
Dwarkesh Podcast publishes deeply researched, long-form interviews with researchers and people working around frontier AI labs. Recent conversations cover recursive self-improvement, automated AI research, continual learning, compute, training data, and the limits of current model recipes.
A September 2026 discussion with John Schulman, Beren Millidge, and Charlie O'Neill spends more than 90 minutes on automated AI research, reinforcement learning, sample efficiency, continual learning, and possible bottlenecks to faster model progress.
Tune in if: you want long conversations about frontier models, training, compute, and current AI research.
Skip if: you want implementation advice, production architecture, or short episodes.
The Pragmatic Engineer
Focus: AI-assisted software engineering, coding agents, engineering practice
Typical length: ~60–120 minutes
Publishing cadence: Several episodes per month
Start with: Building Codex with Tibo Sottiaux
The Pragmatic Engineer covers software engineering more generally, with recurring 2026 episodes on coding agents and AI-assisted development. Recent guests include OpenAI's Tibo Sottiaux on Codex and Addy Osmani on AI engineering.
The Codex episode gets into the Rust-based CLI, the agent harness, model-provider support, internal use at OpenAI, and how AI coding changes review and software-development workflows. Sottiaux's framing of the harness as scaffolding the model still needs — guardrails and steerability that get stripped away as the model improves — connects to a broader argument in AI coding right now: that what keeps coding agents reliable over a long session is better memory, more than a bigger context window.
Tune in if: you want practical discussion of coding agents, AI-assisted engineering, and changes in software-development practice.
Skip if: you're after ML research, training methods, or inference systems.
The Cognitive Revolution
Focus: Agents, retrieval, model capabilities, frontier AI
Typical length: ~60–180 minutes
Publishing cadence: Multiple episodes per month
Start with: Write, Change, Recall, Forget: MongoDB's Pete Johnson on How Retrieval Drives Agent Performance
The Cognitive Revolution combines long-form interviews with recurring coverage of current AI releases and research. Its 2026 feed ranges across reinforcement learning, evaluation, agents, retrieval, infrastructure, and model capabilities.
The September conversation with MongoDB's Pete Johnson focuses directly on agent memory, covering vector and hybrid search, contextualized chunking, retrieval, writes, updates, and how stored context affects later agent behavior.
Tune in if: you want extended discussions of agents, retrieval, model behavior, and frontier AI.
Skip if: you want implementation tutorials or short production-focused episodes.
Software Engineering Daily
Focus: Software architecture, infrastructure, AI systems, agents
Typical length: ~45–80 minutes
Publishing cadence: Multiple episodes per week
Start with: Inside Google's Database Infrastructure for the AI Era
Software Engineering Daily publishes across the whole software stack, including regular episodes on AI agents, retrieval, databases, coding systems, infrastructure, and security.
The September conversation with Google's Sailesh Krishnamurthy digs into how databases are being pulled toward search — agents now write their own queries and propose their own schemas, which raises real governance questions about how that data gets structured and trusted.
Those are exactly the questions knowledge graph design has to answer as much as vector database design does. The show's earlier September episode on precomputed context for RAG covers similar ground from a retrieval-engine angle.
Tune in if: you want AI engineering discussed alongside the rest of modern software architecture.
Skip if: you want a feed devoted entirely to AI research or model development.
A Few More HonorA(I)ble Mentions
Here are a few other feeds worth keeping around if you want faster news, narrower technical subjects, or older engineering episodes.
The AI Daily Brief
The AI Daily Brief publishes near-daily episodes on model releases, policy, products, compute, and industry developments. September 2026 episodes generally run around 30–40 minutes, so it can fill the news role between longer technical interviews.
Tune in if: you want frequent AI news and analysis.
Skip if: you mainly want implementation details or research discussions.
No Priors
No Priors, hosted by Sarah Guo and Elad Gil, mixes AI research and infrastructure with founders, chips, enterprise software, and investment. Its 2026 feed includes conversations with Noam Brown, Satya Nadella, Cerebras CEO Andrew Feldman, Arm CEO Rene Haas, and others working across the AI stack.
Tune in if: you want technical AI discussed alongside infrastructure, startups, and business strategy.
Skip if: you want a consistently engineering-only feed.
Data Skeptic
Data Skeptic is a long-running biweekly podcast on data science, machine learning, AI, and statistics. Its 2026 recommender-systems season covers optimization goals, embeddings, explainability, fairness, privacy, and evaluation.
Episodes are often shorter than the main list, with many recent releases around 20–35 minutes.
Tune in if: you want compact, methodical discussions of machine learning and data-science topics.
Skip if: you're after frontier-model news or agent infrastructure.
How AI Is Built
How AI Is Built has a strong archive on production AI, retrieval, vector databases, embeddings, infrastructure, and coding agents. Its latest verified episode was published on September 11, 2025, so it isn't in the main current list.
Tune in if: you want an engineering-focused archive on putting AI systems into production.
Skip if: you specifically want a feed with new 2026 episodes.
AI Engineering Podcast
AI Engineering Podcast has an archive covering LLM observability, GPU infrastructure, MCP, agentic SaaS, AI sovereignty, and production architecture. Its latest verified episode was published on February 25, 2026, which puts it outside the freshness window used for the main ten.
Tune in if: you want a technical archive centered on building and operating AI applications.
Skip if: current publishing cadence is important to you.
How We Chose Our List of Top AI Podcasts
We prioritized current publishing activity, engineering relevance, technical specificity, and a distinct reason to subscribe. Each main feed had a recently published audio episode when we checked in September 2026, but podcast schedules change, so the dates and cadences above reflect only what's current at the time of writing.
We also checked recent episode subjects rather than relying on the podcast description alone. The main list favors shows that regularly cover areas such as agents, inference, evaluation, retrieval, model training, infrastructure, or AI-assisted software engineering.
Where several podcasts cover the same release or research area, we kept feeds that add technical detail, direct access to people doing the work, or a perspective that isn't already represented elsewhere in the list.
FAQ
Answers to the most common questions from this guide.
Which AI podcast is the most technical?
For research-heavy discussions, Machine Learning Street Talk spends the most time on papers, model behavior, and research methodology. Latent Space is more focused on AI engineering, infrastructure, agents, and inference, while Interconnects concentrates on open models, post-training, and reinforcement learning.
Which AI podcasts are best for keeping up with AI research?
TWIML AI Podcast, Machine Learning Street Talk, Interconnects, and Dwarkesh Podcast cover current research from different angles. TWIML connects research with production questions, MLST spends more time unpacking individual ideas, Interconnects follows open-model development and post-training, and Dwarkesh publishes long interviews with researchers working around frontier AI.
Are podcasts enough to learn AI engineering?
Podcasts are better for following current engineering practice, hearing how experienced practitioners reason through problems, and keeping up with research than for learning implementation fundamentals from scratch. Code, mathematics, system diagrams, and hands-on experimentation still need written material and actual building time.
What are the best short AI podcasts for keeping up with the news?
The AI Daily Brief is the clearest option in this article for frequent, relatively short updates on model releases, products, policy, compute, and industry developments. Data Skeptic is another shorter-format option, although its episodes are organized around technical themes rather than daily news.


