Episode 28
Stop Asking If AI Can Build It - Ask If You Can Trust It
Customers almost never ask, “Can you build this?” anymore. What they actually want to know is, “Can I trust it?”
In this episode of The AI Frontier Playbook, I sit down with Sapna Grover, Senior Director of Microsoft’s Cloud, Data and AI Center of Excellence. Drawing on nearly three decades in software, Sapna explains why the fundamentals of building great products have not changed, even though AI has introduced a probabilistic engineering and operating model.
We explore how teams define a good AI outcome, build golden datasets from real business scenarios, use LLMs as judges without removing human oversight, and trace failures back to their source. Sapna also explains why the most dangerous AI failures are often silent: infrastructure dashboards stay green while outdated data, model drift, rising escalation rates, or unnecessary model calls quietly erode customer trust and business value.
We close with the economics of AI, why the largest model is rarely the right default, and how Sapna uses personal agents to prepare executive briefings, draft her Friday update, and power a custom protein tracker.
You’ll Learn
- Why AI evaluation is becoming the unit testing discipline for AI systems
- How to define success criteria and build a useful golden dataset
- What tracing reveals when an AI system gives the wrong answer
- Why business metrics must sit beside latency, uptime, and availability
- How to recognize data drift, model drift, and silent AI failures
- Why model selection and inference costs require financial discipline
- How personal agents can automate repetitive work without replacing your voice
Listen and Watch
I think if I have to pick one, I'd say the biggest mistake is falling in love with technology.
Samuel Boulanger
Sapna is a senior director at Microsoft leading the Cloud, Data and AI Center of Excellence, with nearly three decades in software, including a scientist role at India's Defence Research and Development Organisation.
Sapna Grover
People often think that AI changes everything, but I actually think it changes less than what we imagine. You still have to start with the customer problem. The first question still always is, "What problem are you trying to solve?" and not, "Where can I use AI?"
The biggest misconception is that once you have deployed, you have built, you're done. I think that's where the real work starts. What's interesting is that the biggest AI failures aren't usually dramatic outages. They're the silent ones.
I actually think that we'll stop thinking about AI altogether. Not because it disappears, just because it'll become as invisible as internet or electricity.
Samuel Boulanger
What if the biggest mistake in AI isn't a bad model, but falling in love with the technology before you've defined the problem? What happens when your dashboard looks perfectly healthy while customers are losing trust? That's exactly why I sat down with Sapna Grover.
Sapna is a senior director at Microsoft leading the Cloud, Data and AI Center of Excellence, with nearly three decades in software, including a scientist role at India's Defence Research and Development Organisation. This episode is about what actually separates a working AI product from an impressive demo.
Before we jump in, I have a small favor to ask. If you're getting value from this show, hit subscribe. Dropping a comment or leaving a like is one of the best ways you can support me. It helps more people find these conversations and honestly, it means the world. Thank you. Now, let's jump in.
Sapna, welcome to the AI Frontier Playbook. It's very great to have you on today. Thank you so much for joining us.
Sapna Grover
Thanks for having me here, Sam.
Samuel Boulanger
Sapna, you're a senior director at Microsoft. You're leading the Cloud, Data and AI Center of Excellence. Before that, you spent close to three decades in software. You've been a scientist at India's Defence Research and Development Organisation. You've been in different leadership roles at Pegasystems. You've been in leadership roles in engineering at Microsoft across Windows and search. And now you're helping the world's largest companies actually put AI into production.
So that's exactly what I'd like to dig into today because there's been a lot of change. You've been building software the old way for almost 30 years, and now you're building it with a model instead of just code. So this is a big, big, big, big change.
The last couple of years leading implementation for customers at Microsoft, I think, give you this first place at understanding how it, it will change for the coming years and how it already changed the way we're developing software. So when you step back across all of those engagements, what's, what's the one thing that genuinely surprised you about how building with AI is different from everything that came before?
Sapna Grover
So when I look back, what surprised me the most is that the fundamentals of building great products haven't really changed, even though the technology underneath has changed dramatically. People often think that AI changes everything. But I actually think it changes less than what we imagine.
You still have to start with the customer problem. The first question still always is, "What problem are you trying to solve?" and not, "Where can I use AI?" Then you need to be clear on the business outcome. Customers don't really care about how many agents or how many copilots you have built. They care whether you're helping them reduce cost, grow revenue, improve their productivity or help them create better experiences for their own customers.
And all the engineering fundamentals still matter. The things like reliability, security, scalability, performance, great user experiences, they haven't become less important just because AI showed up.
What AI has really changed is the kind of engineering problems that we have to solve. And one of the biggest changes is that traditional software was deterministic. If you gave the same input, you'd expect the same output every single time. Two plus two is always four. AI doesn't really work that way. It's probabilistic. You ask the same question twice and you get two different answers, and sometimes both the answers are perfectly valid and right answers.
That changes the way you build software. Writing code is still an important part, but now you're also thinking about things like data quality, evaluations, observability, hallucinations, guardrails, governance, and even cost of every model call.
So what I've noticed is that in customer conversations, customers rarely ask, "Can you build this?" The questions are usually, "Can I trust it?" and, "How do I know when it goes wrong? How do I measure quality? How do I keep the cost in control? And how do I make sure that it is responsible?"
So I think these are the things that separate the AI product from another proof of con... concept. So probably that has been my biggest learning or surprise, that the principles of building products haven't really changed. AI has expanded what's possible, but it hasn't changed what matters.
What has changed is that now intelligence has become part of the software stack, and because it is probabilistic rather than deterministic, it demands a completely different engineering and operating model, if that makes sense.
Samuel Boulanger
I love it. You mentioned so many important points there that we will go deeper into in a second, but the first thing you started with was you start with a customer challenge. You start with a challenge.
I'm seeing so many gurus on, on LinkedIn or on Twitter that are building solutions where, where there's not even a problem to solve, and they're, they're starting with the technology instead of the human challenge or the process challenge. So this is, this is a very important point in the sense that we're still, even if we're, we're more agile and we're faster in building systems thanks to AI, we're still trying to solve basic challenges, right?
Sapna Grover
Right.
Samuel Boulanger
When you develop traditional software, the path is very clear. It's always design, build, test, deploy, and then you maintain. When you're starting from data instead and, and a model, what does that path actually look like for you? Because you've mentioned it, now data quality is more important than ever, and now you need to take into consideration cost. So the model you're using is important as well. So how does that change that traditional path, again, from like build to, to maintain?
Sapna Grover
Yeah, as you said, Sam, like if you think about traditional software, once you have agreed upon the requirements, the challenge is mostly execution. You design the architecture, write the code, test it, deploy it, and then maintain it. And even with agile, as you said, you're iterating on the features, fixing bugs and adding new functionality.
AI flips that around. You start with the business outcome that you're trying to achieve. And the next question isn't, "What code do we write?" The first question is, "Do we have the right data?"
So you're thinking about the data landscape. Do you really have the data that you need? Is the data structured and sitting in ERP or CRM systems, or is it unstructured data spread across your documents, emails, images, videos, meeting recordings or coming from a variety of devices that we all have?
Then you have to think about the characteristics of that data. How much do you have? How often does it change? Is it accurate, complete, governed? Because the right data strategy for a few thousand records is going to be very different from ingesting millions of real-time events every day.
So after you have understood the data, you decide how you will capture it, how you're going to store it, govern it and then make it available for the AI system. Then before even you worry about the right model, you have to decide how are you going to measure success. What's a good answer for your scenario? What's an acceptable answer?
As, as I said earlier, unlike traditional software, there is no single correct output. So you need an evaluation strategy that reflects the quality of the experience that you're trying to create, not whether the answer is technically correct. So only after that you start selecting the right model, experimenting, tuning your prompts, evaluating the results and iterating over it. And the cycle really doesn't stop.
I, I'd say that that's the biggest mindset shift. In traditional software, development often feels like a finish line, whereas in AI, development, sorry, deployment is really the beginning. From that point, you're constantly monitoring the performance, learning from the real-world usage and improving the system.
I sometimes use this analogy: building traditional software is like getting the interiors done for your house. Once, once the work is finished, you're mostly maintaining it for occasional changes or repairs, right? But building an AI product is more like running the business every day. You're looking at performance, responding to the new information, adapting to change around you and continuously making decisions to improve the outcome.
Samuel Boulanger
You, you've mentioned it earlier. In traditional software, two plus two is always equal to four. So you're basically building a unit test and then you, you, you test on what you already know should be working.
Now we've said it, if you ask a generative model the same question three times, you can get three different answers that can all be right. And I, I like the analogy of saying it's like asking a human being. If I'm asking a human a question today, I will get an answer. If I ask the same question tomorrow, I'll get the same kind of answer, but not with the exact same words. That's the same, that's the same for AI, right?
But how do you evaluate knowing that the answer might be right but not exactly the same? Do you... how do you actually evaluate whether the AI is right, and at scale, without a human sitting there checking every single output? Because now we, we took this example of asking three times the questions, but the truth is when I'm evaluating, I will assume you want to run hundreds, if not more, of the same questions to make sure it's always the... it's always pulling from the same data, right?
Sapna Grover
Yeah. So let me make it real with an example for you. So imagine you're building an AI customer support agent for an airline, and a customer comes and types, "My flight to London was cancelled. Can you help me?"
Now, in traditional software, you'd expect a predefined answer, and then you would say either the test is pass or fail. As you mentioned as well, with AI, that's not how it's going to work.
So one agent might come in response and apologize, explain the airline policy, and offer the customer to rebook another flight. Another agent might first check whether the customer is eligible for a refund before suggesting alternative flights. And the third advanced one might proactively look up the available seats on the next flight and ask the customer whether they would want to rebook.
Now all three responses are different. In fact, they might use very different wordings, different reasonings, different actions, and they can be excellent. So the question is not, "Did it produce the expected answer?" The question really becomes, "Did it produce the right outcome?"
Did it understand the customer intent? Did it retrieve the right information? Did it really follow the company policy? Was the information factually correct, and was the tone empathetic? Did it really complete the customer's task? And then most importantly, did it avoid making up policies or promising something that the airline cannot deliver?
So that's why AI evaluation is becoming the unit testing of AI systems. So how do we do it as a product leader? First responsibility is to define what really good means for this use case or for the business.
Now going back to the airline example, if I was the product leader for this use case, I'd sit with the customer service leaders, the operations team, and really understand what a good customer interaction would look like in this scenario.
And then once you have aligned with business in terms of definition of success, then you can translate it into evaluation criteria that software can do, like, as I mentioned, relevance. Did it retrieve the right booking? Did it get the right policy? Was it grounded? How grounded was the response? Was it factual? And how often is the agent able to do resolution without needing a human input? How long did it take? And how was the customer satisfaction, while keeping the operational cost under control?
So you define the rubric, and then you start measuring it systematically. So once you build that, the next... once you understand that, the next step is to build a golden data set. A golden data set is a collection of real business scenarios.
In the airline example, you'll include past conversations of cancelled flights, delayed baggage, missed connections, upgrade requests, refunds, loyalty programs, edge cases. And sometimes they might come from historic customer conversations, and sometimes they can be synthetically generated using AI as well so that you are covered with scenarios that don't, don't happen that often.
And then some bit of that evaluation can be automated. And it is interesting that you can use LLMs as a judge as well, where appropriate, compare the responses with the ground truth when it exists, and then you have to keep humans involved to review the complex cases and calibrate the automated evaluators.
Second practice, once... so this is, this is about evaluation. Second practice that is becoming increasingly essential is tracing. So now let's go back to the airline AI example. It tells a customer that they're eligible for a full refund when in reality they're not eligible, or maybe they're eligible for a travel voucher.
So without tracing, all you know is that AI gave a wrong answer. With tracing, you can see the entire decision, decision-making tree. Did it retrieve the wrong policy, or, or did it interpret the policy wrong, or did it call or receive incomplete information about the booking, or did the model reason incorrectly? So then you can get to the root cause of the real issue if you're able to trace it.
So that level of visibility is very powerful because then it gives you insight into where to improve and why it failed. What I'd say is that tracing really turns AI from a black box into a glass box.
And even after you deploy an AI system, the work isn't done. You're continuously monitoring the production metrics like customer satisfaction, task completion, the things that I mentioned which are relevant for your use case, be it relevance, groundedness, latency, cost, escalation rates, and whether the model performance is improving over time or drifting over time.
So it's not just... so AI evaluation is not just testing before launch. It is a capability that runs throughout the life cycle of the product. So all I would say is that if you can define what's good for your product, you'll be able to scale it.
Samuel Boulanger
So you define your success criteria, then you have your golden data set that can come from previous interactions with your customer, but you can generate some with an LLM. This is an interesting concept, right? You're using another model to generate those use cases. I will assume you need to... a human needs to validate the outcome, right, before you put it inside of your golden data set.
Sapna Grover
Golden data set is what the product leaders are supposed to come up with. And then you're using LLM not only to synthesize the data, you're using at times LLM to also act as a judge. So, but still with human in the loop, I'd say.
Samuel Boulanger
So once you have this golden data set, you're running evaluation. I will assume you, you run hundreds, if not thousands, of evaluations. You don't necessarily want a human to be reading all of those one by one. So that's where an LLM comes into play, the judge.
And then once it's done, you, you talked about tracing. So tracing is looking at the chain of thought of the model and why it, it takes one path or the other. Now I will assume that part should be done by a human, or are you still using LLM as the judge?
Sapna Grover
So that part is usually, comes into the... like during the testing phase, you want to make sure that it is taking the right path. And when things go wrong, wrong, that's when it becomes really, really handy. The visibility is powerful. You know exactly where AI failed and why it failed.
And once you know why, you can fix the right thing, whether it's improving the data or updating the prompts or refining the retrieval strategy or adding better guardrails. So you... that's, that's where I say tracing really helps in terms of fixing the issues later on as well.
Samuel Boulanger
Finally, it's the new way of debugging, right?
Sapna Grover
Correct.
Samuel Boulanger
Okay. Now once the systems are live, since it's... you might be changing your model along the way, right? You might... it's nondeterministic, so for, for, for different reasons, like you're changing the data set, it, it might start behaving differently.
So what does watching them actually look like day to day? Like what, what are you watching for? Are you looking at metrics? Are you looking at this tracing? And, and what's the worst thing that you've seen happen when, like, nobody was watching closely enough?
Sapna Grover
Yeah. Yeah. Yeah. So that's why I say the biggest misconception is that once you have deployed, you have built, you're done. Actually, I think that's where the real work starts. So it's hard to predict when a project will complete because when you put it on in production, that's where the real magic happens.
So let's take... go back to the example that we were talking about, the airline customer support agent. Imagine we have gone live. First few weeks are fantastic. Customer satisfaction goes up. Call volumes are down. Everybody is celebrating.
Then three months later, customer complaints start creeping in. Nothing dramatic. Just little things like customers are saying that the responses are not very useful or it gave outdated information. Now imagine if you were looking at your traditional IT dashboards. Everything is green. Server state is healthy. Response times are good. Nothing has crashed.
So what's really going on? And that's where AI observability and what you said, that we have to constantly be watching, comes in. Of course, you'll still monitor the traditional operational metrics like availability, latency, that we historically have been monitoring in production systems. But with AI, these system metrics are just half of the story.
You need to still... now you have to monitor whether AI is still producing good outcomes. Is it still giving accurate results? Is it, is it grounding those answers in the latest airline policies or not? Is it hallucinating more? Because people can prompt and ask, and it learns. Are customers asking completely new questions because the travel regulations have changed? Or has the cost per interaction doubled because now the agent is making unnecessary calls?
And the other challenge is that the world doesn't stand still. Maybe the airline introduces a new baggage policy. Maybe there's a new loyalty program. Maybe something outside the airline's purview, the visa requirements have changed for certain destinations. And suddenly the data that your AI agent is now seeing in production is very different from what it was evaluated on.
So the model hasn't necessarily become worse. It's the world around it that has changed, and that's when you start seeing things like data drift or model drift.
And what's interesting is that the biggest AI failures are... aren't usually dramatic outages. They are the silent ones. The retrieval system starts giving outdated refund policies, and maybe the quality drops by 10%, but nobody notices it because the system is still responding.
And maybe to fix something you switch to a newer model that gives better results, but it doubles your inference cost. Or maybe it has started escalating more conversations to a human agent, wiping out the productivity gains that you expected.
From an infrastructure perspective, everything might still look healthy, but from a business perspective, you're slowly losing customer trust and business value. That's why I'd have to repeat and say that combining these engineering metrics with business metrics is important for AI observability. The goal isn't anymore just to see whether the system is running or not. The goal is to know whether it is still creating business value.
Samuel Boulanger
You've been mentioning data drift and model drift. So are we talking about the data is changing and the model is changing as well?
Sapna Grover
Yeah. As I said, data can change because of multiple reasons, and often the underlying model, there are new versions of the model, or you change the model because you want better reasoning for one reason or the other to improve the quality, but without realizing you increase the cost by double.
Samuel Boulanger
Yeah, that's a good point. And you, you can totally change the behavior, right? Good example, if I've been using ChatGPT 4.1 for two years and now I'm changing to Opus 4.8, I don't think I, I can expect the same kind of result, right? You'll have to go through a whole new set of evals. And actually it's, it's almost a new project, right, when you're changing models that have a big difference in the architectures?
Sapna Grover
Yeah. So it is, is like if you have your business metrics right, you have defined your success properly from day one, I think then it is just a matter of observing the same, if you're drifting or not, if they are drifting or not. So not just going after the traditional metrics, but also the business metrics that are important.
Samuel Boulanger
So they all come back to, like, define the challenge, define what you want to do. I'm seeing so many customers starting with the tech and not starting with the process. So it's all... I'll drill down to it again, like define what success for you is, and then build on that. And like this is your North Star, right?
If I go back to your example with an airline giving refunds, like mistakenly giving refunds to customers, I will assume that since it's nondeterministic and it's AI, it can make very expensive mistakes. So what's the most expensive mistake you watched a company make with AI? What, what would you tell them to do differently?
Sapna Grover
Yeah, I, I think you, you said the answer yourself, Sam. I think if I have to pick one, I'd say the biggest mistake is falling in love with technology before falling in love with the business problem.
And let me take another example, changing the example. Like imagine a retail company decides that they want an AI shopping assistant because everybody else seems to be building one. So they spend, let's say, six months building. It works fantastic. It answers the product questions, recommends items and even explains refund policies.
They launch it, and six months later they discover the customers are still abandoning the shopping cart at exactly the same rate, at exactly the same point. And why? The reason is that because the real problem was never helping customers choose the product in this case, maybe. So I'm making this up. The real problem was a slow checkout process or expensive shipping.
So AI solved the problem that really wasn't a problem, right? That nobody had that problem. That's why it goes back to the basics, and start with the business question. What business outcome are we trying to improve? Who is going to own that outcome? How are we going to measure success? And how is it going to change the way people actually work?
So if you can't answer these questions, there is a good chance you're building an impressive demo rather than a valuable product.
And the second mistake I'd say, like you asked for one biggest mistake, but the second very common mistake I'd say is underestimating the economics of AI. With traditional software, once you have bought the license, your cost is predictable. AI is different. Every interaction has a cost. Every prompt, every context that you give, every document that is retrieved, every tool call, every reasoning step, every retry that we make, it all consumes compute.
Imagine if your AI agent handles millions of customer conversations in a month. And if every conversation is making, let's say, three unnecessary model calls instead of one, you've just multiplied your inference cost without creating any additional net new customer value.
That's why AI teams need the same financial discipline that was applied to cloud infrastructure. We, we need to understand which workflows are creating value. When you... when can you use a smaller or a cheaper model? You don't need the fanciest or best model for every possible use case. When to cache information rather than generating it again and again.
The final thing that I'd say is that organizations often think of AI as a technology transformation, but it's more of an organizational transformation as well. Technology is an easier part. The harder part is redesigning the workflow. If you just automate a bad process, you're not helping any business outcome, and old habits die hard.
So redesigning, earning the trust of all the users, defining new processes, building together, bringing together, sorry, business leaders, engineers and operations to solve the problem together.
Samuel Boulanger
You've mentioned the cost of models, which is really interesting to me because I'm seeing everyone wanting to default to the latest model, and I don't think most of the time that it's needed at all.
Like, not every process needs to use Fable 5 or GPT-5. I think it's 5.6, the latest release. You can still do a lot of interesting and powerful things with GPT-4.1 or even smaller models like Mistral or others.
So I, I, I think to your point, building your architecture ahead of time and maybe understanding which part of the process requires which level of complexity or, or power from a model is really important, right? Because there's a cost, but there's the time to inference as well, right?
If you're using a reasoning, a deep reasoning model versus a basic model, the time it will take to answer your customer is very, very different, and it can, it can affect the satisfaction of your, your end user, right?
Sapna Grover
That's right. Yeah. So financial discipline is extremely important. Otherwise the ROI is not going to be realized.
Samuel Boulanger
We're almost at the, the end of our time. I ask this question to every guest: how do you personally use AI on a day-to-day basis? Is there a tool or workflow that you've built for yourself that you're, you're particularly proud of?
Sapna Grover
Yeah. So I actually use AI, I'm pretty sure you do as well, dozens of times a day now, to a point where I really don't think about it anymore. It is, it has become a part of how I work.
Like most people in enterprises, I use it for tasks like drafting my emails, summarizing meetings, brainstorming ideas, reviewing documents, building presentations, doing some research and whatnot. But what has changed over the last year is that I've gone beyond just using AI assistants. I've started building AI agents that automate part of my own workflow.
For example, every morning I have an agent that prepares a personalized briefing for me. It pulls information from enterprise data sources, summarizes everything that happened overnight, updates from the headquarters, important announcements, conversations I may have missed or anything else that's relevant to me. So instead of spending the first hour of my day just catching up or searching for information, I spend the first hour actually acting on it.
And the other example is something that has become a bit of a habit, is the Friday ritual for me. So every week I send the Friday Feelings email to my organization, sharing updates, celebrating wins, reflecting on any interesting moments of the week.
So collecting that information used to take a couple of hours. So I built an agent that looks across my emails, team meetings, LinkedIn activity, any other sources I have access to. It identifies highlights and organizes them into themes and prepares a rough, well-structured first draft. So by Friday morning there's a draft sitting, waiting for me. I review it, add my own perspective and make it sound like me.
But the repetitive work is gone. I definitely don't want it to replace my voice, but it acts like an assistant and gives me more time to focus on the parts that only I can do.
Now, outside work, I have found AI equally valuable. So I'm a fitness enthusiast and a vegetarian. I'm particularly... I'm very particular about tracking my protein intake. So I've tried a number of nutrition apps, but they're either too generic or the features I wanted were hidden behind a subscription. So instead of looking for another app, I decided to build one.
So using the AI coding assistant, I created a protein tracker that is tailored for my lifestyle. It understands all the Indian foods that I eat. It knows my morning cappuccino has collagen in it. It remembers the recipes that I cook regularly, and recommendations aren't generic. They're based on my habits. So that's, that's the most exciting part of it, I'd say.
Samuel Boulanger
I'm really interested in this last one. Which data set are you using for, for tracking the amount of protein? Are you using the LLM itself, or are you connecting to a data set?
Sapna Grover
Yeah, that's, that's the interesting part. So if it is missing a recipe, it's just a prompt away, and my LLM would go and search the web and find the right amount of protein and whatnot for that recipe. So the data is also sourced from the web, and if any data I find is incorrect, I can check the label on the food choices and update it manually as well.
Samuel Boulanger
Yeah, I'm a fitness enthusiast as well. So this, this one's very interesting for me, and it's, it's, it's a good example of how AI is changing the landscape of how we live.
Like generally speaking, I had the same issue, like all the fitness apps I've been using so far, like there's always something that doesn't resonate with me, and now you can build your own. And I can just imagine how it will evolve over the next couple of years.
Sapna Grover
Yeah. And to me that's the most exciting part about AI, I'd say. So for years, software has forced us to adapt to the way the application worked. AI is reversing that. Now we can build software that adapts to the way we work. I think that's where I think the real opportunity lies. Not just making existing software smarter, but making, making things deeply personal to every individual.
Samuel Boulanger
So if we expand that to 10 years out, how do you think AI will change the way we, we live and work?
Sapna Grover
Yeah. Yeah. So if you ask me where AI is headed over the next decade, I actually think that we'll stop thinking about AI altogether. Not because it disappears, just because it will become as invisible as, let's say, internet or electricity.
We don't wake up in the morning and say that I'm going to use internet today. We just, we just get on with our lives, and I think AI is actually getting there. It will be part of every application we use, every business process and every device around us.
I also think that AI is going to become way more accessible. Today we were talking about the cost of AI. I think people think of it as capability over time. I, I'm certain, and I hope that the models are more effective, cost continues to go down, and AI will become something that is expected rather than an exception.
But the shift that I'm most excited about is AI becoming, as I said, deeply personal, and another thing is universally accessible. For, for most of history, expertise has been scarce. You have to find it. You travel, pay for it. Yeah, I think AI changes that equation. Expertise becomes abundant, affordable, available to everyone with a connected device.
So talking about India, take an example of an Indian farmer. Imagine walking through a field, noticing a crop that doesn't look healthy, and simply taking out your phone, and in your own native language you ask, "What's wrong with my crop?"
And AI identifies the disease, explains what's causing it, recommends the right treatment, checks the local weather forecast, maybe suggests the right time to irrigate, estimates how disease might affect the yield, and even tells you which nearby market is offering the best price.
So there it is not just answering a question. It's putting the knowledge of an agronomist, a weather expert, a market analyst in the hands of someone who may not have had access to them.
And I don't think that's limited to farming alone. Imagine a student in a small town having access to world-class tutors available 24/7 in their native language, or a small business owner getting guidance on cash flow, marketing, regulatory compliance, or a nurse in a rural clinic receiving guidance based on the latest medical practices.
So I think AI has the potential to narrow the gap in access to knowledge that has existed for decades. So I think, yeah, if I were to add, I think that it has the potential to do for knowledge what the mobile revolution did for communication. Mobile phones connected billions of people to each other. AI can connect billions of people to expertise.
Samuel Boulanger
I, I love your answer so much because everything you just mentioned is exactly what makes me passionate about AI, is how it can help so many people achieve so much more, and kind of making knowledge open source, I feel like, giving access to everyone from...
Sapna Grover
Yeah. Yeah.
Samuel Boulanger
So this, this was a very great conversation. Thank you so much for joining me today. I think if I, if I have to keep one thing on top of mind from our conversation, it will be start with a challenge. Start by defining your definition of success before anything else, right?
Sapna Grover
That's right. Thank you so much, Sam, for having me here. It was a pleasure talking to you.
Samuel Boulanger
It was a pleasure. Have a great day.
Sapna Grover
You have a good one. Bye.
Samuel Boulanger
Now, Sapna made one thing clear throughout this conversation. The technology underneath software has changed completely. But what separates a real AI product from an impressive demo hasn't moved an inch. It still comes down to knowing what business problem you're solving before you touch a model.
Three things I want you to walk away with from this episode.
First, before you write a line of code, sit down with the people who own the business outcome and define what a good answer actually looks like. Then build a golden data set from real conversations so you can evaluate against it systematically.
Second, watch for the failures that don't show up on a dashboard. Server uptime and response time can look perfectly healthy while data drift or model drift erodes accuracy. So track groundedness, escalation rates, and cost per interaction alongside your usual metrics.
Third, stop defaulting to the biggest, most expensive model for every task. Match the model to the complexity of the workflow. Cache what you can and treat inference costs with the same discipline you'll apply to any other infrastructure spend.
If you got any value from this one, subscribe to the AI Frontier Playbook wherever you listen, and sign up for the AI Frontier Playbook newsletter to stay sharp between episodes. Thank you so much for listening, and I'll see you in the next one. See you.
Related Episodes
You might also enjoy
Stay in the Loop
Never miss an episode.
Get new episodes and AI insights delivered to your inbox.