Find partners
The Data Engineering Show

The Data Engineering Show

Hosted by The Firebolt Data Bros

TechnologyInterviews guests

Episodes

63

Latest episode

Aug 2026

Language

EN

About the show

The Data Engineering Show is a podcast for data engineering and BI practitioners to go beyond theory. Learn from the biggest influencers in tech about their practical day-to-day data challenges and solutions in a casual and fun setting. SEASON 1 DATA BROS Eldad and Boaz Farkash shared the same stuffed toys growing up as well as a big passion for data. After founding Sisense and building it to become a high-growth analytics unicorn, they moved on to their next venture, Firebolt, a leading high-performance cloud data warehouse. SEASON 2 DATA BROS In season 2 Eldad adopted a brilliant new little brother, and with their shared love for query processing, the connection was immediate. After excelling in his MS, Computer Science degree, Benjamin Wagner joined Firebolt to lead its query processing team and is a rising star in the data space. For inquiries contact tamar@firebolt.io Website: https://www.firebolt.io

Listen to episodes

60 recent
August 18, 2026Episode 6017 min

How AI Is Reshaping Modern Data Teams and the Future of Platforms with Xavier Gumara Rigol

What if AI could transform how your entire data team works—without replacing them? In this episode, Benjamin Wagner explores with Xavier Gumara Rigol, Head of Data at Manychat, how natural language-to-SQL tools are reshaping data analyst and data scientist roles, why building a strong data platform is essential for AI-powered self-serve analytics, and the critical strategies for evaluating text-to-insight solutions in 2026. Whether you're leading a data organization or building your next analytics capability, discover how to balance build versus buy decisions and position your team for the AI-driven future of data engineering.

August 6, 2026Episode 5919 min

Building Modern Data Platforms Without Legacy Bottlenecks ft Andrew Jones

What if you could transform your data team from a bottleneck into a self-serve platform that empowers the entire organization? In this episode, host Benjamin Wagner sits down with Andrew Jones, Staff Data Engineer at LocalStack, to explore how to balance building new data infrastructure while supporting legacy systems, why reliability and data contracts are becoming non-negotiable, and how AI agents are accelerating platform development. Whether you're scaling a data organization or rethinking your data governance strategy, this conversation is packed with practical insights on managing different user personas, automating support requests, and staying open-minded as the data landscape evolves. Tune in to discover how to build resilient, self-serve data platforms that drive real business impact.

July 21, 2026Episode 5820 min

Why 99% of BI Tools Get Embedded Analytics Wrong And How Omni Fixed It ft. Chris Merrick

What if the future of analytics wasn't about building another middleware layer, but about getting closer to the actual business user? In this episode, Benjamin sits down with Chris Merrick, CTO and cofounder of Omni, to explore why semantic layers matter more than ever in an agentic world, how AI is reshaping embedded analytics and customer-facing data experiences, and the key strategies for keeping complex data models aligned across federated sources. Whether you're building analytics platforms, managing data infrastructure, or trying to make AI work at scale, this conversation is packed with practical insights on balancing governed analytics with exploratory AI, unifying disparate data sources, and capturing business intelligence beyond just the metrics in your warehouse. Tune in to discover how the next generation of analytics platforms will need to think about the entire business, not just the data.

June 16, 2026Episode 5719 min

AI for Data and Data for AI: The Dual Frontier of Modern Data Engineering with Pranav Motarwar

What if the data engineering skills you have today become obsolete in five years? In this episode, host Benjamin Wagner sits down with Pranav Motarwar, a data engineer who's witnessed the industry's transformation from traditional ETL to AI-powered pipelines, to explore how AI is fundamentally reshaping data engineering roles, why you need to master both "AI for data" and "data for AI" to stay relevant, and the emerging infrastructure required to handle multimodal data at scale. Whether you're a data engineer wondering about your career longevity or a builder curious about next-gen data stacks, this conversation unpacks the skills you'll need, the tools defining 2026, and why data engineers aren't disappearing - they're just evolving faster than ever.

May 7, 2026Episode 5618 min

AI Won't Replace Engineers, But This Framework Will Change How They Build with Rohit Girme

What if you could build AI features with confidence while moving at the pace of innovation? In this episode, Benjamin Wagner sits down with Rohit Girma, Staff Software Engineer at Airbnb, to explore how to evaluate generative AI in production, why breaking down complex problems into smaller chunks accelerates development, and the key strategies for scaling AI-powered products beyond zero-to-one. Whether you're shipping AI features or transforming your engineering workflow, this conversation offers practical insights on building reliable AI systems, leveraging LLMs as orchestration tools, and the future of software development. Tune in to discover why humans remain essential in the scaling phase and how your team can move faster without sacrificing quality.

April 28, 2026Episode 5522 min

The Framework Canva Uses for 200M+ Designers with Paul Tune

In this episode of The Data Engineering Show, Benjamin sits down with Paul Tune, Staff Research Scientist at Canva, to explore the advancement of machine learning at one of the world's leading design platforms. Learn how Canva is transitioning from traditional ML like recommendation engines for templates to cutting-edge agentic workflows that allow users and AI to collaborate on complex design tasks. Whether you're interested in the infrastructure behind distributed training or the nuances of post-training LLMs for aesthetic tasks, this deep dive offers a masterclass in scaling ML for millions of creative users.

April 8, 2026Episode 5422 min

Llama 2 & 3 Safety: Soumya Batra on Agentic AI Training

What if the expertise that built foundation models could reshape how you think about AI's future? In this episode, Benjamin sits down with Soumya Batra, founder and CEO of WisePort AI and former safety lead on Llama 2 and Llama 3 at Meta, to explore how foundation models evolved from traditional NLP, why post-training holds the highest leverage for safety and controllability, and what natively agentic AI means for the next frontier of AI development. Whether you're curious about the model training lifecycle or wondering what comes after large language models, this conversation unpacks the technical strategies and vision shaping tomorrow's AI systems.

March 24, 2026Episode 5318 min

The Data Fusion Secret & Why Custom Query Engines Fail with Nikita Lapkov

What if building a distributed SQL engine meant rethinking everything about how query execution works at scale? In this episode, Benjamin sits down with Nikita, Senior Software Engineer at Cloudflare, to explore how R2 SQL leverages object storage and distributed computing to power analytics across 300 global locations, why backward compatibility becomes critical when you can't control infrastructure rollouts, and the key strategies for handling joins and adaptive query execution in a stateless, point-to-point network architecture. Whether you're designing distributed systems or curious about how Cloudflare processes petabytes of data, this conversation reveals the real-world engineering challenges and innovations shaping the future of cloud data platforms.

March 10, 2026Episode 5224 min

How Zipline AI Turns Weeks of Engineering Into Minutes of SQL Queries ft. Nikhil Simha

What if you could deploy ML features and real-time data pipelines without building complex infrastructure from scratch? In this episode, host Benjamin sits down with Nikhil Simha, CTO at Zipline AI and co-author of Chronon AI, to explore how Chronon, an open-source system that generates data infrastructure from simple queries, is transforming feature engineering at companies like OpenAI and Airbnb. Learn why iteration speed matters for fraud detection, how to serve thousands of signals at a massive scale, and what the future of analytical databases looks like in an AI-first world. Whether you're scaling real-time ML systems or building customer-facing analytics, this conversation is packed with practical insights on bridging the gap between data scientists and ML engineers.

February 19, 2026Episode 5116 min

The Geo-Data Problem Nobody Talks About And How Voi Solved It ft. Magnus Dahlbäck

What if your data platform could power both critical business decisions and real-time product features at scale? In this episode, host Benjamin sits down with Magnus Dahlbäck, Senior Director of Data and Platform at Voi, to explore how a metrics-first approach and semantic layers transform data accessibility, why traditional ML and LLMs require different strategies for different problems, and how to balance FinOps costs while processing billions of IoT events daily. Whether you're building data infrastructure for a high-growth company or rethinking how your organization consumes data, this conversation is packed with practical strategies for unlocking data value and preparing your platform for AI. Tune in to discover how Voi ditched traditional BI tools and revolutionized their approach to enterprise analytics.

Is this your show?

Claim this listing to keep it up to date, reach guests who want to pitch you, and manage bookings with Guestify.

Claim this listing

More Technology podcasts