Summary
In this episode Ofer Mendelevitch shares what it really takes to evaluate RAG systems in production when your source data is incomplete, constantly changing, or difficult to validate against a clean ground truth. He explores how RAG has evolved from simple “chat with your PDF” demos into enterprise-grade retrieval systems that require robust ingestion pipelines, hybrid search, reranking, multimodal support, access controls, and refresh strategies for large and dynamic document collections. Ofer also explains why evaluation becomes one of the hardest parts of the stack, particularly when retrieval quality, generation quality, and system behavior can all shift as you change chunking, models, prompts, or ranking components.
He digs into practical approaches for offline and online evaluation, including synthetic question generation, sampling production traffic, and LLM-as-a-judge techniques such as UMBRELA and Auto Nuggetizer. Along the way, he discusses the rise of agentic RAG, context engineering, MCP-connected tools, text-to-SQL systems, and the growing role of coding agents in building and maintaining AI pipelines. Ofer closes with his perspective on where retrieval, memory, multimodal systems, and agent orchestration are headed next, along with a reminder that the biggest gaps in AI today are often not just technical, but educational and organizational as well.
Announcements
Contact Info
Parting Question
Links
The intro and outro music is from Hitman's Lovesong feat. Paola Graziano by The Freak Fandango Orchestra/CC BY-SA 3.0
In this episode Ofer Mendelevitch shares what it really takes to evaluate RAG systems in production when your source data is incomplete, constantly changing, or difficult to validate against a clean ground truth. He explores how RAG has evolved from simple “chat with your PDF” demos into enterprise-grade retrieval systems that require robust ingestion pipelines, hybrid search, reranking, multimodal support, access controls, and refresh strategies for large and dynamic document collections. Ofer also explains why evaluation becomes one of the hardest parts of the stack, particularly when retrieval quality, generation quality, and system behavior can all shift as you change chunking, models, prompts, or ranking components.
He digs into practical approaches for offline and online evaluation, including synthetic question generation, sampling production traffic, and LLM-as-a-judge techniques such as UMBRELA and Auto Nuggetizer. Along the way, he discusses the rise of agentic RAG, context engineering, MCP-connected tools, text-to-SQL systems, and the growing role of coding agents in building and maintaining AI pipelines. Ofer closes with his perspective on where retrieval, memory, multimodal systems, and agent orchestration are headed next, along with a reminder that the biggest gaps in AI today are often not just technical, but educational and organizational as well.
Announcements
- Hello and welcome to the AI Engineering Podcast, your guide to the fast-moving world of building scalable and maintainable AI systems
- Your host is Tobias Macey and today I'm interviewing Ofer Mendelevitch about how to evaluate your RAG systems when you have incomplete or changing data to grade against
- Introduction
- How did you get involved in machine learning?
- RAG was one of the first concepts to gain traction in the early explosion of LLMs and their application to real-world problems. How would you characterize its role in the current landscape?
- RAG is conceptually straight-forward, but can be quite complex to do well. What are some of the common pitfalls that teams encounter when bringing a retrieval-oriented LLM service into production?
- What are some of the predominant methods to avoid regressions or increased error rates as the data, code, and models change in a given deployment?
- There are numerous software projects and companies that offer various forms of evaluation/validation for RAG workloads. What are the factors that teams should be evaluating when selecting which one(s) to adopt?
- Especially when working with data that is subject to change (user-contributed content, changing products, news-based data, etc.), what are the techniques that teams can use to be sure that the generated responses based on that information is accurate?
- As models have become more powerful the idea of "agentic RAG" started circulating. How is that different from the first iteration of RAG systems?
- What new pressures and challenges does that prevent?
- What are the ways that it can improve or simplify a content-based generative system?
- How do you think about the differentiation of RAG and "context engineering"?
- What are the most interesting, innovative, or unexpected ways that you have seen teams address the production needs of RAG systems?
- What are the most interesting, unexpected, or challenging lessons that you have learned while working on RAG evaluation?
- How do you see the future of RAG, context engineering, and data-driven agents changing in the near to medium term?
Contact Info
Parting Question
- From your perspective, what are the biggest gaps in tooling, technology, or training for AI systems today?
Links
- Naive Bayes
- Hands on RAG for Production (affiliate link)
- RAG == Retrieval Augmented Generation
- Vectara
- Hybrid Search
- TF-IDF == Term-Frequency Inverse Document Frequency
- BM25
- UMBRELA
- AutoNuggetizer
- Agentic RAG
- Malloy
The intro and outro music is from Hitman's Lovesong feat. Paola Graziano by The Freak Fandango Orchestra/CC BY-SA 3.0