Data Engineering Podcast


This show goes behind the scenes for the tools, techniques, and difficulties associated with the discipline of data engineering. Databases, workflows, automation, and data manipulation are just some of the topics that you will find here.

Support the show!

07 October 2026

Why Metadata Still Matters: From Data Catalogs to AI-Ready Semantic Hubs - E518

Rewind 10 seconds
1X
Skip 30 seconds ahead
0:00/0:00

Share on social media:


Summary

In this episode, I spoke with Christian Bremeau and Simon Dynin about their decades-long work in metadata management and how that experience shaped the evolution of Meta Integration Technology and its direct-to-customer MetaKarta platform. We explored the recurring challenges of fragmented data ecosystems, why every tool inevitably creates its own metadata silo, and why organizations still need a vendor-neutral system to stitch lineage, semantics, and change history across databases, ETL/ELT pipelines, BI tools, and cloud platforms. Christian and Simon explained why deep lineage, versioning, configuration management, and deterministic understanding of transformations matter far more than a polished catalog UI when teams need to trust numbers, assess impact, or satisfy audits.

We also dug into the growing importance of metadata in the age of AI. Christian and Simon shared their view that metadata management is becoming foundational for both BI and agentic systems, especially as organizations struggle with semantic drift, tool sprawl, and non-deterministic AI behavior. They introduced Metacarta’s new Semantic Hub as a way to reconcile business logic across legacy and modern tools, compile shared semantics into downstream systems, and provide consistent answers for both human analysts and AI agents. Along the way, we discussed metadata ops, the organizational processes needed to make metadata programs succeed, and the practical situations where a cross-platform metadata system is—or isn’t—the right fit. 

Announcements 
  • Hello and welcome to the Data Engineering Podcast, the show about modern data management
  • Today’s episode is sponsored by Parallel - where agents find answers. Most engineers today closely follow new model releases, but don’t pay attention to their agent’s most important tool: web search. Parallel develops enterprise-grade infrastructure for agents to retrieve high-quality context from the web. Their core products are a suite of APIs for retrieving high quality information from the web with Pareto-optimal quality, cost, and speed. Whether you work on voice agents that need 200 millisecond latency, chat bots that balance speed, depth, and quality, or long-horizon agents to do thorough, overnight research for you, Parallel is a single platform for all your agentic research. Get started for free at dataengineeringpodcast.com/parallel
  • Your host is Tobias Macey and today I'm interviewing Christian Bremeau and Simon Dynin about his long history of dealing with metadata at Meta Integration Technology, and how they are adapting to the age of AI

Interview
 
  • Introduction
  • How did you get involved in the area of data management?
  • Can you start by sharing some of the highlights of your career in metadata and how the industry's focus on that subject has changed over the years?
  • What are the major sources of friction and fragmentation in metadata management in your experience?
  • There are numerous metadata management, metadata integration, data catalog, and other governance/discovery/semantic platforms in the ecosystem. How do you explain the differentiating factors of your system to people who are evaluating their options?
  • One of the constant challenges in data systems is the tension between data locality and data consistency. What is your approach to managing a canonical, reliable definition of a data asset, while making it available and accessible to all of the consuming systems?
  • There was a surge of interest in metadata systems in the 2018 - 2022 time-frame as the industry invested in cloud data warehouses, ELT, and all of the other "modern data stack" patterns. Now AI is driving a new surge of interest. What are the categorical differences between these periods in your experience?
    • What are the new stressors that you anticipate from the continued growth in LLM-mediated and agentic data interactions?
  • Can you describe the overall architecture of the MetaKarta platform and how it has evolved in the latest version?
  • For a team who is trying to tame their data estate, what does the adoption of MetaKarta look like?
    • How does it change their day-to-day work?
  • You're promoting the idea of "MetaDataOps" as a dedicated concern for an organization. What are the skills and activities of someone who is focused on that capability? 
  • What are the most interesting, innovative, or unexpected ways that you have seen MetaKarta used?
  • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Metadata and MetaKarta?
  • When is MetaKarta the wrong choice?
  • What do you have planned for the future of MetaKarta?

Contact Info

Parting Question
  • From your perspective, what is the biggest gap in the tooling or technology for data management today?

Links

The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

Share on social media:


Listen in your favorite app:



More options

Here are shows you might like

See show recommendations
AI Engineering Podcast
Tobias Macey
The Python Podcast.__init__
Tobias Macey

© 2025 Boundless Notions, LLC.
EPISODE SPONSORS Parallel
Parallel

Elevate your AI agents from simple chat bots to powerful, research-driven systems with Parallel. While model releases dominate the conversation, the most critical tool for any modern agent is reliable, high-quality web context. Parallel provides the enterprise-grade infrastructure needed to seamlessly retrieve accurate information from the web, offering a suite of APIs designed for Pareto-optimal combinations of quality, speed, and cost. Whether your application requires instant, 200-millisecond latency for voice agents, balanced speed and depth for chat bots, or deep, overnight research for long-horizon tasks, Parallel is your single platform for agentic context. Beyond best-in-class search, our full platform includes Web Monitors, Deep Research tools, and Batch Record Enrichment. Get started with complimentary MCP and $5 in free credits, and discover the best quality web search solution, no matter your latency or scale needs. Get started for free at Parallel.ai.