Zendesk serves around 100,000 brands including Airbnb and Shopify, handling billions of customer interactions per year.
Zendesk's AI chief says most AI failures aren't algorithmic — bad data, skills gaps, and fuzzy ROI kill more AI projects than bad models ever will.
The MongoDB Podcast
Zendesk's AI chief says most AI failures aren't algorithmic — bad data, skills gaps, and fuzzy ROI kill more AI projects than bad models ever will.
TL;DR
Akshaya Murthy, Director of AI Transformation at Zendesk, argues that most enterprises should integrate off-the-shelf AI rather than build from scratch — citing talent scarcity, speed-to-market, and rapidly closing capability gaps between commercial and open-source models [1] — Akshaya Murthy "Time to market, total cost of ownership, talent availability, and security compliance are the four axes that should drive your build-vs-buy…" 15:35 . He warns that "most AI failures are not algorithmic failures" but stem from messy data, skills gaps, and fuzzy ROI [2] — Akshaya Murthy "I put in my picture and said, create an action figure of me and it automatically named my action figure Ankur Patel." 45:00 . The single most actionable takeaway: start with a problem statement, not an AI mandate — then work backwards to data, model selection, and people [3] — Akshaya Murthy "Don't launch an AI initiative because AI is the thing to do. Start with a problem statement, validate that AI is part of the answer, then w…" 49:40 .
Shane McAlister, lead on MongoDB's developer relations team, interviews Akshaya Murthy, Director of AI Transformation at Zendesk, about the strategic and architectural decisions enterprises face when building AI products. Topics include the build-vs-integrate debate, AI washing, data quality, responsible AI, and the evolving expectations of enterprise customers.
Shane McAlister opens the MongoDB Podcast Live by framing the central tension that has dominated his two years of episodes: how should teams navigate the maze of building AI products? He introduces Akshaya Murthy — Director of AI Transformation at Zendesk — as the guide for today's journey through enterprise challenges, AI washing, and the architectural decisions that determine success or failure. The welcome is brief but purposeful, signalling that the conversation ahead will be both practical and opinionated.
Akshaya Murthy opens with a candid admission: his path to AI transformation is 'less orthodox'. It began two decades ago designing game AI, wrestling with philosophical questions like 'if the character doesn't know we exist, what does that mean from a creator perspective?' — a surprisingly deep foundation for thinking about machine intelligence. From there, he moved into enterprise software delivery, then spent a significant stretch at Big Four consulting firms working on transformations, enterprise redesign, and internal audit. That last role, he explains, gave him a rare bird's-eye view of how companies actually operate versus how they should. The thread connecting all of it is transformation: his current role at Zendesk is the natural culmination — applying the most powerful horizontal technology he's ever encountered to fundamentally change how a market-leading company operates, serves customers, and positions itself as an AI leader both inside and out.
Far from being a visionary role with a loose mandate, Akshaya describes AI transformation as a rigorous exercise in working backwards from business outcomes. The starting point is always the goal — improve employee experience, improve customer experience, make operations best in class — and AI is the means, not the end. He breaks the role into two complementary tracks: identifying tangible, high-return opportunities where AI can genuinely move the needle on KPIs like productivity lift, CSAT, revenue, and EBITDA; and advocating internally for adoption, because technology alone doesn't drive behaviour change. Credibility is earned incrementally, he argues, by promising specific outcomes and then delivering them — a discipline that separates real transformation from AI hype.
Shane's prompt to explain Zendesk for the uninitiated draws a clear, structured answer from Akshaya: think of it as the connective tissue between a business and its customers. The platform delivers on three dimensions — unified communications that collapse email, messaging, and phone into a single pane of glass; frictionless self-service powered by generative search and AI chatbots for instant resolution; and robust analytics that surface trends and patterns so organisations can improve continuously. With 100,000 brands from Airbnb to Shopify and billions of annual interactions, the scale is significant. [1] — Akshaya Murthy "100,000 Zendesk brand customers: Zendesk serves around 100,000 brands including Airbnb and Shopify, handling billions of customer interacti…" 07:47 The deeper point, though, is Zendesk's strategic pivot: it's no longer building better ticketing software, it's building for measurable outcomes. AI is the engine of that shift, and the conversation about what AI can do has changed from 'help me get to a resolution' to 'give me the resolution now'.
Akshaya walks through the three shifts that define Zendesk's AI-first redesign. The first is intelligent triage: what once required human review now happens in seconds, getting customers materially closer to resolution faster. The second is the move from macros — rigid, rule-based automations — to Copilots that provide dynamic, context-aware guidance to agents in real time. The third is workforce management, where load balancing and staffing optimisation become fully automated rather than siloed in a contact centre operation. The results speak loudly: in extreme cases, early adopters have cut email handle time by 92% and slashed customer question response time by 74%. [1] — Akshaya Murthy "92% email handle time reduction: Early Zendesk AI adopters cut email handle time by 92% using intelligent triage and routing powered by AI." 09:56 The through-line, Akshaya stresses, is a fundamental reframe: Zendesk has moved from building better ticketing software to delivering measurable business outcomes for customers.
The velocity of change in customer expectations is the theme of this segment. Akshaya identifies three seismic shifts: immediacy — SLAs are now measured in seconds, not hours; context — customers expect the service system to already know their history and what has been tried; and human-centricity — AI interactions must feel human, even when they aren't. The paradox he surfaces is striking: a 2025 survey shows 90% of companies say AI is central to customer loyalty, but the same data reveals that loyalty drops off a cliff when customers discover they're speaking to a bot. [1] — Akshaya Murthy "90% of companies say AI central to loyalty: A 2025 survey found 90% of companies say AI is central to customer loyalty, yet loyalty drops s…" 12:51 The implication for product designers is profound: the experience has to be seamlessly intelligent and warmly human at the same time. Customers, Akshaya concludes, simply don't care how it works — they want resolution, personalisation, and friction-free experience, and AI delivered at scale is the only mechanism that can meet that bar.
The build-vs-buy debate, Akshaya notes, is not new — he watched the same argument play out with SAP and Oracle ERP systems two decades ago, and enterprises eventually concluded that keeping up with the resulting tech debt was untenable. The same calculus is now playing out with AI, but under enormous time pressure as CFOs chase ROI proof for their investors. Akshaya structures the decision around four factors. Time to market and speed to impact: can you afford the months or years it takes to build from scratch when an off-the-shelf system can be online in one to two weeks? Total cost of ownership: the CapEx of a custom build almost always makes integration look attractive in the near term. Talent scarcity: the pool of AI engineers who can build systems at scale is tiny, and Stanford or Berkeley graduates in this space command $500K or more per year. [1] — Akshaya Murthy "AI experts: ~500K/year minimum salary: Hiring a single AI expert capable of building systems at scale costs at least $500K per year and is …" 17:11 Security and compliance: not every company has the expertise to build enterprise-grade security into a new product, and every AI vendor now ships with enterprise-grade privacy and GDPR compliance as table stakes. For most enterprises, integration wins on all four counts — and it frees up scarce development bandwidth to focus on building only what truly needs to be custom.
Akshaya paints a landscape where the case for building your own model is eroding almost by the day. The gap between GPT-4 and Llama's 70-billion-parameter open-source model is narrowing to the point where a self-hosted inference node can produce comparable results for many use cases. Security frameworks that once took years to mature are now shipped as standard in enterprise-grade off-the-shelf products — GDPR compliance, privacy controls, role-based access — because these are competitive table stakes. Model refreshes have gone from annual to weekly on commercial platforms, and daily on Hugging Face for open-source variants. And the per-token cost is compressing: OpenAI charges roughly $2 per million tokens, while running Llama yourself can bring that to fractions of a cent. [1] — Akshaya Murthy "OpenAI API cost: ~$2/million tokens: OpenAI API pricing is approximately $2 per million tokens, while open-source alternatives like Llama c…" 22:27 Most strikingly, the expertise required to work with these systems has inverted: in 2023 you needed PhDs with prompt engineering experience; today a product manager can achieve the same with a drag-and-drop interface. The only genuine exceptions to the 'integrate-first' rule are hyperscale operations where token costs make self-hosting cheaper, and mission-critical verticals like defence, banking, or healthcare where general-purpose training data is simply inadequate.
Akshaya identifies a structural tension in the current AI market: the industry is simultaneously watching per-token costs compress toward commodity pricing while also realising that the value delivered far exceeds what those tokens cost to produce. The result is a mispricing problem — vendors don't yet know how to charge for value, and buyers don't yet know how to quantify what they're getting. Zendesk's response has been to move toward resolution-based pricing: customers pay for solved problems, not for compute cycles consumed. That model change, Akshaya argues, offsets some of the cost pressure while aligning incentives more honestly with outcomes. His long-range view is that AI will follow the same arc as internet access — initially priced per megabyte or gigabyte, then compressed into flat-rate utility pricing. A 'value-add tax' will stabilise somewhere above pure commoditisation, but the gross cost trajectory is definitively downward. For enterprises making build-vs-buy decisions now, that means today's integration economics are still favourable, but the calculus will shift over time.
Shane raises the concern that stops many enterprise AI conversations in their tracks: are third-party AI vendors secretly training their models on sensitive customer data? Akshaya's answer is unambiguous — no — and he explains the three-layered protection that justifies that confidence. At the contractual level, every AI model vendor Zendesk works with is bound by zero-retention and zero-training clauses, making it legally prohibited to store or learn from Zendesk's data. At the product level, Zendesk's reputation as one of the most secure CX platforms on the market means security is culturally embedded, not an afterthought. And at the infrastructure level, enterprise AI APIs ship with role-based access controls, encryption in transit and at rest, as a baseline. Internally, Zendesk follows the NIST playbook to ensure no customer data is compromised. His takeaway for any company building an AI product: security and privacy must be first-principles decisions, not bolt-ons, because customer trust is the asset you can least afford to lose.
This chapter delivers what may be the episode's most practically useful insight: most enterprise AI failures are not caused by the technology. The algorithm works. What breaks AI projects is the plumbing, the people, and the inability to prove value. [1] — Akshaya Murthy "Most AI failures are not algorithmic failures. The algorithm works. It's the plumbing, the people, and proving out the value of the platfor…" 30:15 Akshaya opens with data: messy, siloed data is the original sin of enterprise AI. Without a single source of truth — a data lake house with clean pipelines and scrubbed PII — even the most sophisticated model will produce unreliable outputs. He coins a memorable inversion of the classic tech maxim: 'garbage in, hype out.' The second barrier is a skills gap that spans the full organisation, from MLOps engineers to sales reps and customer-facing employees. Legacy SaaS companies face a particular challenge because AI adoption requires company-wide buy-in from day one, not a gradual persuasion campaign. The third barrier is vague ROI: too many POCs promise 90% cost cuts without showing the mechanism. His solution to the people problem is hands-on training — one session beats 20 courses or 30 Slack announcements — plus clear, credible commitment from the C-suite. [2] — Akshaya Murthy "One session of hands-on training with AI beats 20 courses, beats 30 messaging through Slack." 33:10
AI washing, Akshaya explains, is what happens when 'AI' becomes a marketing suffix rather than a functional description — and he opens with a disarming example: a rice cooker bearing an 'AI Powered' sticker. In the enterprise context, the pattern is more insidious: a vendor adds a thin LLM plug-in to an existing product without re-architecting the solution or reimagining how value is delivered. The test for legitimacy comes down to four questions. Is the AI solving a specific, well-defined problem? Are concrete KPIs and ROIs attached to the solution, and is it clear how AI achieves them versus the prior approach? Is the claimed improvement consistent with realistic AI outcomes — Zendesk's internal POCs see 2x to 50x ROI, so a vendor promising 10% CSAT lift should raise eyebrows? [1] — Akshaya Murthy "2x–50x ROI on AI implementations: Zendesk's internal AI POC and platform implementations have delivered between 2x and 50x ROI." 35:43 And finally, is the vendor claiming a universal 'one size fits all' solution, which is almost always a red flag? Ticket summarisation might be one of the few truly general-purpose AI tasks, but beyond that, domain specificity matters enormously. If those four boxes aren't checked, you're looking at a rule-based engine wearing an AI costume.
Moving from strategy to implementation, Akshaya gives his clearest technical prescription of the episode. Clean data plumbing comes first: pipelines, a single source of truth, and PII scrubbing — because 'garbage in, hype out' and even the perfect transformer model will break on bad data. [1] — Akshaya Murthy "I say garbage in, hype out with AI. So bad data even breaks the perfect models. The transformer works, but your data is terrible." 38:50 Second, build a modular API-first model layer. Just as you wouldn't lock customers into a single ISP, you cannot lock them into a single LLM vendor. Containerised, microservice-sized model infrastructure lets you hot-swap models as the landscape evolves, giving both you and your customers optionality. Third, bring CI/CD discipline to model management: automated regression tests when you swap a model, one-click rollbacks when behaviour deviates from baselines — treating model changes with the same rigour as code releases. Fourth, invest in monitoring and drift detection so that six months of reinforcement learning doesn't silently push the model into a problem space it was never meant to solve. And throughout all of it, keep a human in the loop: the best learnings from model outputs should feed back into continuous improvement. Shane draws a parallel to MongoDB's MAP (MongoDB AI Applications Program) composable reference architecture, which follows the same hot-swap logic across LLMs and orchestration frameworks.
Shane asks whether the architectural considerations amount to 'responsible AI' — and Akshaya's answer is a layered yes, grounded in the NIST Risk Management Framework's four-function cycle: Govern, Map, Measure, and Manage. The key shift from traditional software, he argues, is that AI requires trustworthiness to be embedded at every phase of the model lifecycle, not just at the application layer. Unacceptable risks must be identified and eliminated early; high-risk use cases must be scoped out of the product footprint; bias and privacy protections must be baked in by design. He connects this to EU AI Act compliance, where the stakes for getting it wrong are both legal and reputational. The cautionary example is Google's early AI stumbles — high-profile bias incidents that damaged trust in ways that were slow to repair. His conclusion: any enterprise building an AI product should treat the NIST cycle as a minimum operating standard, not as optional overhead.
The abstract discussion of AI bias lands with sudden, personal force when Akshaya shares what happened at a Zendesk internal finance leadership event. The team was using ChatGPT to generate personalised action figures from staff photos. When Akshaya uploaded his own picture, the system automatically labelled the figure 'Ankur Patel' — a name he describes, wryly, as 'the stereotypical Indian name in the United States.' [1] — Akshaya Murthy "I put in my picture and said, create an action figure of me and it automatically named my action figure Ankur Patel." 45:00 The incident is both funny and pointed: it demonstrates, in a single data point, that the training data powering leading commercial models remains heavily skewed toward English-language, American cultural norms. Non-English languages and non-Western cultures occupy a far smaller slice of the training corpus, which means cultural nuance, dialectal variation, and naming conventions from outside the dominant training distribution are routinely flattened or stereotyped. Akshaya's assessment is measured but clear: bias is not a solved problem, it is a managed one, and controlling for it at the application layer — through guardrails, domain-specific fine-tuning, and careful output validation — is the current best practice until the underlying training pipelines improve.
Shane asks whether Zendesk pays attention to public model evaluation benchmarks when choosing which models to build products on. Akshaya's answer is nuanced: benchmarks serve a purpose — they signal general capability — but they are not a reliable proxy for enterprise task fit. The fact that a model can ace LSAT questions or solve complex mathematical problems says little about whether it can resolve a customer support query accurately and empathetically. His analogy is sharp: the top-scoring student in the class is not necessarily the brightest or most capable person in the room. At Zendesk, the practical response has been to evaluate both custom and commercial options in parallel and release both when appropriate — a custom triage model and a GPT-4o-powered feature launched on the same day. The lesson for enterprises is clear: use benchmark scores as a starting filter, not a final verdict. Test models against your own representative data and tasks, and maintain optionality by not over-committing to a single vendor or architecture. MongoDB draws a parallel: for 18 months the team has been working with code assistant providers to build MongoDB-specific evaluation sets, precisely because general benchmark performance doesn't predict performance on MongoDB-specific coding tasks.
As the conversation approaches its end, Shane asks Akshaya to distil his guidance for the C-suite into a single starting point. The answer is deceptively simple but cuts against a common failure mode: begin with the problem, not the technology. Too many AI initiatives launch because the board or the CEO has mandated that the company 'do something with AI', without first defining what specific problem needs solving or whether AI is genuinely part of the answer. Once a clear problem statement exists, the subsequent questions sequence naturally: is the data in place? What model best serves this particular purpose? Do we have the in-house talent, or do we need to buy? What will it cost, and does the projected outcome justify it? Shane also asks how Akshaya personally keeps up with the relentless pace of AI development — the answer is opportunistic rather than systematic, with Lex Fridman's podcast recommended for philosophical depth and the AI Engineer podcast for technical detail. The episode closes with a pointer to the Zendesk AI Hub for listeners who want to explore Zendesk's AI product updates, demos, and releases.
Shane closes by thanking Akshaya and the global audience that joined live from Denmark, Wisconsin, Burkina Faso, and Nairobi. Akshaya directs listeners to the Zendesk AI Hub as the single destination for AI product updates, demos, and release notes. The sign-off is warm but brisk — Akshaya is, quite literally, being asked to leave the room. It's a fittingly human end to a conversation about technology that aspires to be human-centric.
Chapter 3 · 05:35
Far from being a visionary role with a loose mandate, Akshaya describes AI transformation as a rigorous exercise in working backwards from business outcomes. The starting point is always the goal — improve employee experience, improve customer experience, make operations best in class — and AI is the means, not the end. He breaks the role into two complementary tracks: identifying tangible, high-return opportunities where AI can genuinely move the needle on KPIs like productivity lift, CSAT, revenue, and EBITDA; and advocating internally for adoption, because technology alone doesn't drive behaviour change. Credibility is earned incrementally, he argues, by promising specific outcomes and then delivering them — a discipline that separates real transformation from AI hype.
Zendesk serves around 100,000 brands including Airbnb and Shopify, handling billions of customer interactions per year.
Chapter 4 · 08:25
Shane's prompt to explain Zendesk for the uninitiated draws a clear, structured answer from Akshaya: think of it as the connective tissue between a business and its customers. The platform delivers on three dimensions — unified communications that collapse email, messaging, and phone into a single pane of glass; frictionless self-service powered by generative search and AI chatbots for instant resolution; and robust analytics that surface trends and patterns so organisations can improve continuously. With 100,000 brands from Airbnb to Shopify and billions of annual interactions, the scale is significant. [1] — Akshaya Murthy "100,000 Zendesk brand customers: Zendesk serves around 100,000 brands including Airbnb and Shopify, handling billions of customer interacti…" 07:47 The deeper point, though, is Zendesk's strategic pivot: it's no longer building better ticketing software, it's building for measurable outcomes. AI is the engine of that shift, and the conversation about what AI can do has changed from 'help me get to a resolution' to 'give me the resolution now'.
Zendesk's AI transformation produced three seismic product shifts: intelligent triage cuts resolution time to seconds, Copilots replace rule-based macros with live agent guidance, and workforce management becomes fully automated. In extreme cases, that's a 92% reduction in email handle time.
Early Zendesk AI adopters cut email handle time by 92% using intelligent triage and routing powered by AI.
Zendesk AI Copilot customers have cut time spent responding to customer questions by 74%.
Chapter 5 · 11:37
Akshaya walks through the three shifts that define Zendesk's AI-first redesign. The first is intelligent triage: what once required human review now happens in seconds, getting customers materially closer to resolution faster. The second is the move from macros — rigid, rule-based automations — to Copilots that provide dynamic, context-aware guidance to agents in real time. The third is workforce management, where load balancing and staffing optimisation become fully automated rather than siloed in a contact centre operation. The results speak loudly: in extreme cases, early adopters have cut email handle time by 92% and slashed customer question response time by 74%. [1] — Akshaya Murthy "92% email handle time reduction: Early Zendesk AI adopters cut email handle time by 92% using intelligent triage and routing powered by AI." 09:56 The through-line, Akshaya stresses, is a fundamental reframe: Zendesk has moved from building better ticketing software to delivering measurable business outcomes for customers.
Customer expectations have shifted from 'soon' to 'now'. They want instant resolution, full context of their history, and service that feels human — even if it's AI. Fail any one of those three and you've already lost them.
A 2025 survey found 90% of companies say AI is central to customer loyalty, yet loyalty drops sharply when customers discover they are talking to a bot.
Chapter 6 · 14:55
The velocity of change in customer expectations is the theme of this segment. Akshaya identifies three seismic shifts: immediacy — SLAs are now measured in seconds, not hours; context — customers expect the service system to already know their history and what has been tried; and human-centricity — AI interactions must feel human, even when they aren't. The paradox he surfaces is striking: a 2025 survey shows 90% of companies say AI is central to customer loyalty, but the same data reveals that loyalty drops off a cliff when customers discover they're speaking to a bot. [1] — Akshaya Murthy "90% of companies say AI central to loyalty: A 2025 survey found 90% of companies say AI is central to customer loyalty, yet loyalty drops s…" 12:51 The implication for product designers is profound: the experience has to be seamlessly intelligent and warmly human at the same time. Customers, Akshaya concludes, simply don't care how it works — they want resolution, personalisation, and friction-free experience, and AI delivered at scale is the only mechanism that can meet that bar.
Building AI from scratch is losing the argument in most enterprises. Integration wins on speed-to-market, TCO, and talent access — and you'd need months to years plus massive CapEx to match what you can get up and running in two weeks.
Time to market, total cost of ownership, talent availability, and security compliance are the four axes that should drive your build-vs-buy decision. Right now, on all four, integration usually wins.
Integrating an off-the-shelf AI system like Zendesk can be done in one to two weeks, whereas building from scratch can take months to years.
Chapter 7 · 16:40
The build-vs-buy debate, Akshaya notes, is not new — he watched the same argument play out with SAP and Oracle ERP systems two decades ago, and enterprises eventually concluded that keeping up with the resulting tech debt was untenable. The same calculus is now playing out with AI, but under enormous time pressure as CFOs chase ROI proof for their investors. Akshaya structures the decision around four factors. Time to market and speed to impact: can you afford the months or years it takes to build from scratch when an off-the-shelf system can be online in one to two weeks? Total cost of ownership: the CapEx of a custom build almost always makes integration look attractive in the near term. Talent scarcity: the pool of AI engineers who can build systems at scale is tiny, and Stanford or Berkeley graduates in this space command $500K or more per year. [1] — Akshaya Murthy "AI experts: ~500K/year minimum salary: Hiring a single AI expert capable of building systems at scale costs at least $500K per year and is …" 17:11 Security and compliance: not every company has the expertise to build enterprise-grade security into a new product, and every AI vendor now ships with enterprise-grade privacy and GDPR compliance as table stakes. For most enterprises, integration wins on all four counts — and it frees up scarce development bandwidth to focus on building only what truly needs to be custom.
Hiring a single AI expert capable of building systems at scale costs at least $500K per year and is rising as the talent wars heat up.
AI has become the number one investment line item in company priorities, overtaking cybersecurity.
AI model updates are happening on a weekly basis from commercial providers and on a daily basis on platforms like Hugging Face for open-source models.
Chapter 8 · 22:20
Akshaya paints a landscape where the case for building your own model is eroding almost by the day. The gap between GPT-4 and Llama's 70-billion-parameter open-source model is narrowing to the point where a self-hosted inference node can produce comparable results for many use cases. Security frameworks that once took years to mature are now shipped as standard in enterprise-grade off-the-shelf products — GDPR compliance, privacy controls, role-based access — because these are competitive table stakes. Model refreshes have gone from annual to weekly on commercial platforms, and daily on Hugging Face for open-source variants. And the per-token cost is compressing: OpenAI charges roughly $2 per million tokens, while running Llama yourself can bring that to fractions of a cent. [1] — Akshaya Murthy "OpenAI API cost: ~$2/million tokens: OpenAI API pricing is approximately $2 per million tokens, while open-source alternatives like Llama c…" 22:27 Most strikingly, the expertise required to work with these systems has inverted: in 2023 you needed PhDs with prompt engineering experience; today a product manager can achieve the same with a drag-and-drop interface. The only genuine exceptions to the 'integrate-first' rule are hyperscale operations where token costs make self-hosting cheaper, and mission-critical verticals like defence, banking, or healthcare where general-purpose training data is simply inadequate.
The gap between commercial LLMs and open-source alternatives is closing fast. OpenAI charges ~$2 per million tokens; Llama gets you there for fractions of a cent. The endgame looks like internet pricing — you'll pay for a utility, not per megabyte.
OpenAI API pricing is approximately $2 per million tokens, while open-source alternatives like Llama can reduce this to fractions of a cent.
At hyperscale — billions of tokens per hour — the per-token cost math flips and self-hosted models win. Mission-critical verticals like defence, banking, and healthcare also need bespoke models because general-purpose LLMs simply weren't trained on the right data.
Chapter 9 · 25:55
Akshaya identifies a structural tension in the current AI market: the industry is simultaneously watching per-token costs compress toward commodity pricing while also realising that the value delivered far exceeds what those tokens cost to produce. The result is a mispricing problem — vendors don't yet know how to charge for value, and buyers don't yet know how to quantify what they're getting. Zendesk's response has been to move toward resolution-based pricing: customers pay for solved problems, not for compute cycles consumed. That model change, Akshaya argues, offsets some of the cost pressure while aligning incentives more honestly with outcomes. His long-range view is that AI will follow the same arc as internet access — initially priced per megabyte or gigabyte, then compressed into flat-rate utility pricing. A 'value-add tax' will stabilise somewhere above pure commoditisation, but the gross cost trajectory is definitively downward. For enterprises making build-vs-buy decisions now, that means today's integration economics are still favourable, but the calculus will shift over time.
Chapter 11 · 30:10
This chapter delivers what may be the episode's most practically useful insight: most enterprise AI failures are not caused by the technology. The algorithm works. What breaks AI projects is the plumbing, the people, and the inability to prove value. [1] — Akshaya Murthy "Most AI failures are not algorithmic failures. The algorithm works. It's the plumbing, the people, and proving out the value of the platfor…" 30:15 Akshaya opens with data: messy, siloed data is the original sin of enterprise AI. Without a single source of truth — a data lake house with clean pipelines and scrubbed PII — even the most sophisticated model will produce unreliable outputs. He coins a memorable inversion of the classic tech maxim: 'garbage in, hype out.' The second barrier is a skills gap that spans the full organisation, from MLOps engineers to sales reps and customer-facing employees. Legacy SaaS companies face a particular challenge because AI adoption requires company-wide buy-in from day one, not a gradual persuasion campaign. The third barrier is vague ROI: too many POCs promise 90% cost cuts without showing the mechanism. His solution to the people problem is hands-on training — one session beats 20 courses or 30 Slack announcements — plus clear, credible commitment from the C-suite. [2] — Akshaya Murthy "One session of hands-on training with AI beats 20 courses, beats 30 messaging through Slack." 33:10
Most enterprise AI projects don't fail because the model is bad. They fail because the data is messy, the people aren't bought in, and nobody can prove the ROI. Fix those three things first.
The majority of enterprise AI failures come from poor data, people issues, and unclear ROI — not from the AI algorithm itself.
Having clean, centralised data is a prerequisite for any effective AI strategy; messy data will break even the best AI model.
Chapter 12 · 34:50
AI washing, Akshaya explains, is what happens when 'AI' becomes a marketing suffix rather than a functional description — and he opens with a disarming example: a rice cooker bearing an 'AI Powered' sticker. In the enterprise context, the pattern is more insidious: a vendor adds a thin LLM plug-in to an existing product without re-architecting the solution or reimagining how value is delivered. The test for legitimacy comes down to four questions. Is the AI solving a specific, well-defined problem? Are concrete KPIs and ROIs attached to the solution, and is it clear how AI achieves them versus the prior approach? Is the claimed improvement consistent with realistic AI outcomes — Zendesk's internal POCs see 2x to 50x ROI, so a vendor promising 10% CSAT lift should raise eyebrows? [1] — Akshaya Murthy "2x–50x ROI on AI implementations: Zendesk's internal AI POC and platform implementations have delivered between 2x and 50x ROI." 35:43 And finally, is the vendor claiming a universal 'one size fits all' solution, which is almost always a red flag? Ticket summarisation might be one of the few truly general-purpose AI tasks, but beyond that, domain specificity matters enormously. If those four boxes aren't checked, you're looking at a rule-based engine wearing an AI costume.
AI washing is a buzzword salad with no specifics. If a vendor can't show you the KPIs, the defined problem it solves, and a measurable ROI, they're masquerading a rule-based engine as AI. The rice cooker with an 'AI Powered' sticker is not a joke — it's a category.
Zendesk's internal AI POC and platform implementations have delivered between 2x and 50x ROI.
Chapter 13 · 38:30
Moving from strategy to implementation, Akshaya gives his clearest technical prescription of the episode. Clean data plumbing comes first: pipelines, a single source of truth, and PII scrubbing — because 'garbage in, hype out' and even the perfect transformer model will break on bad data. [1] — Akshaya Murthy "I say garbage in, hype out with AI. So bad data even breaks the perfect models. The transformer works, but your data is terrible." 38:50 Second, build a modular API-first model layer. Just as you wouldn't lock customers into a single ISP, you cannot lock them into a single LLM vendor. Containerised, microservice-sized model infrastructure lets you hot-swap models as the landscape evolves, giving both you and your customers optionality. Third, bring CI/CD discipline to model management: automated regression tests when you swap a model, one-click rollbacks when behaviour deviates from baselines — treating model changes with the same rigour as code releases. Fourth, invest in monitoring and drift detection so that six months of reinforcement learning doesn't silently push the model into a problem space it was never meant to solve. And throughout all of it, keep a human in the loop: the best learnings from model outputs should feed back into continuous improvement. Shane draws a parallel to MongoDB's MAP (MongoDB AI Applications Program) composable reference architecture, which follows the same hot-swap logic across LLMs and orchestration frameworks.
The architecture decisions that will make or break your AI product: clean data pipelines, a modular API-first model layer that lets you hot-swap LLMs, CI/CD for models themselves, and continuous monitoring with human review. These aren't nice-to-haves — they're the difference between a product and a liability.
Chapter 14 · 43:00
Shane asks whether the architectural considerations amount to 'responsible AI' — and Akshaya's answer is a layered yes, grounded in the NIST Risk Management Framework's four-function cycle: Govern, Map, Measure, and Manage. The key shift from traditional software, he argues, is that AI requires trustworthiness to be embedded at every phase of the model lifecycle, not just at the application layer. Unacceptable risks must be identified and eliminated early; high-risk use cases must be scoped out of the product footprint; bias and privacy protections must be baked in by design. He connects this to EU AI Act compliance, where the stakes for getting it wrong are both legal and reputational. The cautionary example is Google's early AI stumbles — high-profile bias incidents that damaged trust in ways that were slow to repair. His conclusion: any enterprise building an AI product should treat the NIST cycle as a minimum operating standard, not as optional overhead.
Responsible AI isn't a checkbox at the end of the build cycle. The NIST govern-map-measure-manage framework needs to run through every phase: pre-training, training, post-training, delivery, and application. Bias and privacy must be by design, not bolt-ons.
Akshaya Murthy asked ChatGPT to make an action figure from his own photo. It named the figure 'Ankur Patel' — a stereotypical Indian-American name. Bias isn't theoretical. It's baked into the training data and shows up in production.
Chapter 15 · 44:50
The abstract discussion of AI bias lands with sudden, personal force when Akshaya shares what happened at a Zendesk internal finance leadership event. The team was using ChatGPT to generate personalised action figures from staff photos. When Akshaya uploaded his own picture, the system automatically labelled the figure 'Ankur Patel' — a name he describes, wryly, as 'the stereotypical Indian name in the United States.' [1] — Akshaya Murthy "I put in my picture and said, create an action figure of me and it automatically named my action figure Ankur Patel." 45:00 The incident is both funny and pointed: it demonstrates, in a single data point, that the training data powering leading commercial models remains heavily skewed toward English-language, American cultural norms. Non-English languages and non-Western cultures occupy a far smaller slice of the training corpus, which means cultural nuance, dialectal variation, and naming conventions from outside the dominant training distribution are routinely flattened or stereotyped. Akshaya's assessment is measured but clear: bias is not a solved problem, it is a managed one, and controlling for it at the application layer — through guardrails, domain-specific fine-tuning, and careful output validation — is the current best practice until the underlying training pipelines improve.
Chapter 16 · 47:20
Shane asks whether Zendesk pays attention to public model evaluation benchmarks when choosing which models to build products on. Akshaya's answer is nuanced: benchmarks serve a purpose — they signal general capability — but they are not a reliable proxy for enterprise task fit. The fact that a model can ace LSAT questions or solve complex mathematical problems says little about whether it can resolve a customer support query accurately and empathetically. His analogy is sharp: the top-scoring student in the class is not necessarily the brightest or most capable person in the room. At Zendesk, the practical response has been to evaluate both custom and commercial options in parallel and release both when appropriate — a custom triage model and a GPT-4o-powered feature launched on the same day. The lesson for enterprises is clear: use benchmark scores as a starting filter, not a final verdict. Test models against your own representative data and tasks, and maintain optionality by not over-committing to a single vendor or architecture. MongoDB draws a parallel: for 18 months the team has been working with code assistant providers to build MongoDB-specific evaluation sets, precisely because general benchmark performance doesn't predict performance on MongoDB-specific coding tasks.
Chapter 17 · 49:40
As the conversation approaches its end, Shane asks Akshaya to distil his guidance for the C-suite into a single starting point. The answer is deceptively simple but cuts against a common failure mode: begin with the problem, not the technology. Too many AI initiatives launch because the board or the CEO has mandated that the company 'do something with AI', without first defining what specific problem needs solving or whether AI is genuinely part of the answer. Once a clear problem statement exists, the subsequent questions sequence naturally: is the data in place? What model best serves this particular purpose? Do we have the in-house talent, or do we need to buy? What will it cost, and does the projected outcome justify it? Shane also asks how Akshaya personally keeps up with the relentless pace of AI development — the answer is opportunistic rather than systematic, with Lex Fridman's podcast recommended for philosophical depth and the AI Engineer podcast for technical detail. The episode closes with a pointer to the Zendesk AI Hub for listeners who want to explore Zendesk's AI product updates, demos, and releases.
Don't launch an AI initiative because AI is the thing to do. Start with a problem statement, validate that AI is part of the answer, then work backwards to data, model selection, talent, and cost. Every AI project that skips this step is setting itself up to fail.
No indexed bits in this chapter.
This episode
Factual claims made this episode, and whether a source was named.
Zendesk AI early adopters have cut email handle time by 92% in extreme cases using AI-powered triage.
Zendesk AI Copilot early adopters reduced time spent responding to customer questions by 74%.
In 2025, 90% of companies say AI is central to customer loyalty, but customer loyalty drops sharply when customers discover they are talking to a bot.
Hiring an AI expert capable of building systems at scale costs a minimum of $500,000 per year in the United States.
AI has overtaken cybersecurity to become the number one investment line item in company priorities.
An off-the-shelf AI system like Zendesk can be brought online in one to two weeks, compared to months or years and significant CapEx to build from scratch.
OpenAI charges approximately $2 per million tokens for API access, while open-source Llama models can be run for fractions of a cent per token.
AI model refreshes on commercial platforms are occurring on a weekly basis, while on Hugging Face open-source updates arrive on a daily basis.
In 2023, building AI products required a team of PhDs with prompt engineering expertise; today a product manager can do the same with a drag-and-drop interface.
Zendesk has approximately 100,000 brand customers, including Airbnb and Shopify, and handles billions of customer interactions per year.
Zendesk's internal AI POC and platform implementations have achieved between 2x and 50x ROI.
ChatGPT generated an action figure from Akshaya Murthy's photo and automatically assigned it the name 'Ankur Patel', a stereotypical Indian-American name, demonstrating ongoing AI bias.
The gap in capability between GPT-4 and the 70-billion-parameter Llama open-source model is narrowing significantly.
This episode
Customer experience platform and the guest's employer; discussed as a case study for enterprise AI transformation and AI product development.
Host company and podcast producer; mentioned as a Zendesk customer and for its MAP (MongoDB AI Applications Program) composable architecture initiative.
Commercial LLM provider discussed in the context of API pricing (~$2/million tokens) and data privacy contractual obligations.
US standards body whose AI Risk Management Framework (Govern, Map, Measure, Manage) Zendesk follows internally for responsible AI governance.
Open-source AI platform cited as a source of daily model updates, enabling near-real-time model refresh cycles.
ERP software vendor used as a historical analogy for the current build-vs-buy debate in enterprise AI.
Cited as an example of a major Zendesk enterprise customer handling billions of interactions per year.
Cited alongside Airbnb as a major Zendesk customer among its 100,000 brand clients.
Mentioned as an example source pipeline for AI talent, with graduates costing $500K+ per year.
OpenAI's consumer AI chatbot cited as the moment AI entered mainstream public awareness (November 2022) and later used in an AI bias anecdote.
Open-source large language model cited as a low-cost alternative to commercial LLMs, with its 70B parameter version narrowing the gap with GPT-4.
OpenAI's large language model used as a benchmark comparison against open-source alternatives like Llama 70B.
OpenAI's multimodal model cited as an example of a commercially integrated feature Zendesk released alongside a custom triage model.
Recommended by Akshaya Murthy as a resource for keeping up with philosophical and technical AI developments.
Stats
We use essential and analytics cookies to run Vuci. To understand how the site is used: Privacy Policy.
Install Vuci on your phone
Add it to your home screen for a faster, app-like experience.
Install Vuci on your phone
Tap the Share button, then “Add to Home Screen”.
A new version is available
Reload to get the latest Vuci.