Don't Build Your Own AI (Unless You Have To)

Don't Build Your Own AI (Unless You Have To)

Zendesk's AI chief says most AI failures aren't algorithmic — bad data, skills gaps, and fuzzy ROI kill more AI projects than bad models ever will.

Mar 6, 2026 52:41 Difficulty: Intermediate Played

TL;DR

Akshaya Murthy, Director of AI Transformation at Zendesk, argues that most enterprises should integrate off-the-shelf AI rather than build from scratch — citing talent scarcity, speed-to-market, and rapidly closing capability gaps between commercial and open-source models. He warns that "most AI failures are not algorithmic failures" but stem from messy data, skills gaps, and fuzzy ROI. The single most actionable takeaway: start with a problem statement, not an AI mandate — then work backwards to data, model selection, and people.

#build vs buy AI #AI integration strategy #LLM commoditisation #AI washing #data quality for AI #AI bias #responsible AI #NIST AI framework #enterprise AI adoption #AI skills gap #modular AI architecture #customer experience AI #AI ROI #open source LLMs #AI change management #AI integration #build vs buy #enterprise AI #Zendesk #data strategy #LLM #RAG #customer experience #MLOps #CI/CD #NIST #open source models #AI transformation #skills gap #AI pricing

Shane McAlister, lead on MongoDB's developer relations team, interviews Akshaya Murthy, Director of AI Transformation at Zendesk, about the strategic and architectural decisions enterprises face when building AI products. Topics include the build-vs-integrate debate, AI washing, data quality, responsible AI, and the evolving expectations of enterprise customers.

Chapter list
  • Shane McAlister opens the MongoDB Podcast Live by framing the central tension that has dominated his two years of episodes: how should teams navigate the maze of building AI products? He introduces Akshaya Murthy — Director of AI Transformation at Zendesk — as the guide for today's journey through enterprise challenges, AI washing, and the architectural decisions that determine success or failure. The welcome is brief but purposeful, signalling that the conversation ahead will be both practical and opinionated.

  • Akshaya Murthy opens with a candid admission: his path to AI transformation is 'less orthodox'. It began two decades ago designing game AI, wrestling with philosophical questions like 'if the character doesn't know we exist, what does that mean from a creator perspective?' — a surprisingly deep foundation for thinking about machine intelligence. From there, he moved into enterprise software delivery, then spent a significant stretch at Big Four consulting firms working on transformations, enterprise redesign, and internal audit. That last role, he explains, gave him a rare bird's-eye view of how companies actually operate versus how they should. The thread connecting all of it is transformation: his current role at Zendesk is the natural culmination — applying the most powerful horizontal technology he's ever encountered to fundamentally change how a market-leading company operates, serves customers, and positions itself as an AI leader both inside and out.

  • Far from being a visionary role with a loose mandate, Akshaya describes AI transformation as a rigorous exercise in working backwards from business outcomes. The starting point is always the goal — improve employee experience, improve customer experience, make operations best in class — and AI is the means, not the end. He breaks the role into two complementary tracks: identifying tangible, high-return opportunities where AI can genuinely move the needle on KPIs like productivity lift, CSAT, revenue, and EBITDA; and advocating internally for adoption, because technology alone doesn't drive behaviour change. Credibility is earned incrementally, he argues, by promising specific outcomes and then delivering them — a discipline that separates real transformation from AI hype.

  • Shane's prompt to explain Zendesk for the uninitiated draws a clear, structured answer from Akshaya: think of it as the connective tissue between a business and its customers. The platform delivers on three dimensions — unified communications that collapse email, messaging, and phone into a single pane of glass; frictionless self-service powered by generative search and AI chatbots for instant resolution; and robust analytics that surface trends and patterns so organisations can improve continuously. With 100,000 brands from Airbnb to Shopify and billions of annual interactions, the scale is significant. The deeper point, though, is Zendesk's strategic pivot: it's no longer building better ticketing software, it's building for measurable outcomes. AI is the engine of that shift, and the conversation about what AI can do has changed from 'help me get to a resolution' to 'give me the resolution now'.

  • Akshaya walks through the three shifts that define Zendesk's AI-first redesign. The first is intelligent triage: what once required human review now happens in seconds, getting customers materially closer to resolution faster. The second is the move from macros — rigid, rule-based automations — to Copilots that provide dynamic, context-aware guidance to agents in real time. The third is workforce management, where load balancing and staffing optimisation become fully automated rather than siloed in a contact centre operation. The results speak loudly: in extreme cases, early adopters have cut email handle time by 92% and slashed customer question response time by 74%. The through-line, Akshaya stresses, is a fundamental reframe: Zendesk has moved from building better ticketing software to delivering measurable business outcomes for customers.

  • The velocity of change in customer expectations is the theme of this segment. Akshaya identifies three seismic shifts: immediacy — SLAs are now measured in seconds, not hours; context — customers expect the service system to already know their history and what has been tried; and human-centricity — AI interactions must feel human, even when they aren't. The paradox he surfaces is striking: a 2025 survey shows 90% of companies say AI is central to customer loyalty, but the same data reveals that loyalty drops off a cliff when customers discover they're speaking to a bot. The implication for product designers is profound: the experience has to be seamlessly intelligent and warmly human at the same time. Customers, Akshaya concludes, simply don't care how it works — they want resolution, personalisation, and friction-free experience, and AI delivered at scale is the only mechanism that can meet that bar.

  • The build-vs-buy debate, Akshaya notes, is not new — he watched the same argument play out with SAP and Oracle ERP systems two decades ago, and enterprises eventually concluded that keeping up with the resulting tech debt was untenable. The same calculus is now playing out with AI, but under enormous time pressure as CFOs chase ROI proof for their investors. Akshaya structures the decision around four factors. Time to market and speed to impact: can you afford the months or years it takes to build from scratch when an off-the-shelf system can be online in one to two weeks? Total cost of ownership: the CapEx of a custom build almost always makes integration look attractive in the near term. Talent scarcity: the pool of AI engineers who can build systems at scale is tiny, and Stanford or Berkeley graduates in this space command $500K or more per year. Security and compliance: not every company has the expertise to build enterprise-grade security into a new product, and every AI vendor now ships with enterprise-grade privacy and GDPR compliance as table stakes. For most enterprises, integration wins on all four counts — and it frees up scarce development bandwidth to focus on building only what truly needs to be custom.

  • Akshaya paints a landscape where the case for building your own model is eroding almost by the day. The gap between GPT-4 and Llama's 70-billion-parameter open-source model is narrowing to the point where a self-hosted inference node can produce comparable results for many use cases. Security frameworks that once took years to mature are now shipped as standard in enterprise-grade off-the-shelf products — GDPR compliance, privacy controls, role-based access — because these are competitive table stakes. Model refreshes have gone from annual to weekly on commercial platforms, and daily on Hugging Face for open-source variants. And the per-token cost is compressing: OpenAI charges roughly $2 per million tokens, while running Llama yourself can bring that to fractions of a cent. Most strikingly, the expertise required to work with these systems has inverted: in 2023 you needed PhDs with prompt engineering experience; today a product manager can achieve the same with a drag-and-drop interface. The only genuine exceptions to the 'integrate-first' rule are hyperscale operations where token costs make self-hosting cheaper, and mission-critical verticals like defence, banking, or healthcare where general-purpose training data is simply inadequate.

  • Akshaya identifies a structural tension in the current AI market: the industry is simultaneously watching per-token costs compress toward commodity pricing while also realising that the value delivered far exceeds what those tokens cost to produce. The result is a mispricing problem — vendors don't yet know how to charge for value, and buyers don't yet know how to quantify what they're getting. Zendesk's response has been to move toward resolution-based pricing: customers pay for solved problems, not for compute cycles consumed. That model change, Akshaya argues, offsets some of the cost pressure while aligning incentives more honestly with outcomes. His long-range view is that AI will follow the same arc as internet access — initially priced per megabyte or gigabyte, then compressed into flat-rate utility pricing. A 'value-add tax' will stabilise somewhere above pure commoditisation, but the gross cost trajectory is definitively downward. For enterprises making build-vs-buy decisions now, that means today's integration economics are still favourable, but the calculus will shift over time.

  • Shane raises the concern that stops many enterprise AI conversations in their tracks: are third-party AI vendors secretly training their models on sensitive customer data? Akshaya's answer is unambiguous — no — and he explains the three-layered protection that justifies that confidence. At the contractual level, every AI model vendor Zendesk works with is bound by zero-retention and zero-training clauses, making it legally prohibited to store or learn from Zendesk's data. At the product level, Zendesk's reputation as one of the most secure CX platforms on the market means security is culturally embedded, not an afterthought. And at the infrastructure level, enterprise AI APIs ship with role-based access controls, encryption in transit and at rest, as a baseline. Internally, Zendesk follows the NIST playbook to ensure no customer data is compromised. His takeaway for any company building an AI product: security and privacy must be first-principles decisions, not bolt-ons, because customer trust is the asset you can least afford to lose.

  • This chapter delivers what may be the episode's most practically useful insight: most enterprise AI failures are not caused by the technology. The algorithm works. What breaks AI projects is the plumbing, the people, and the inability to prove value. Akshaya opens with data: messy, siloed data is the original sin of enterprise AI. Without a single source of truth — a data lake house with clean pipelines and scrubbed PII — even the most sophisticated model will produce unreliable outputs. He coins a memorable inversion of the classic tech maxim: 'garbage in, hype out.' The second barrier is a skills gap that spans the full organisation, from MLOps engineers to sales reps and customer-facing employees. Legacy SaaS companies face a particular challenge because AI adoption requires company-wide buy-in from day one, not a gradual persuasion campaign. The third barrier is vague ROI: too many POCs promise 90% cost cuts without showing the mechanism. His solution to the people problem is hands-on training — one session beats 20 courses or 30 Slack announcements — plus clear, credible commitment from the C-suite.

  • AI washing, Akshaya explains, is what happens when 'AI' becomes a marketing suffix rather than a functional description — and he opens with a disarming example: a rice cooker bearing an 'AI Powered' sticker. In the enterprise context, the pattern is more insidious: a vendor adds a thin LLM plug-in to an existing product without re-architecting the solution or reimagining how value is delivered. The test for legitimacy comes down to four questions. Is the AI solving a specific, well-defined problem? Are concrete KPIs and ROIs attached to the solution, and is it clear how AI achieves them versus the prior approach? Is the claimed improvement consistent with realistic AI outcomes — Zendesk's internal POCs see 2x to 50x ROI, so a vendor promising 10% CSAT lift should raise eyebrows? And finally, is the vendor claiming a universal 'one size fits all' solution, which is almost always a red flag? Ticket summarisation might be one of the few truly general-purpose AI tasks, but beyond that, domain specificity matters enormously. If those four boxes aren't checked, you're looking at a rule-based engine wearing an AI costume.

  • Moving from strategy to implementation, Akshaya gives his clearest technical prescription of the episode. Clean data plumbing comes first: pipelines, a single source of truth, and PII scrubbing — because 'garbage in, hype out' and even the perfect transformer model will break on bad data. Second, build a modular API-first model layer. Just as you wouldn't lock customers into a single ISP, you cannot lock them into a single LLM vendor. Containerised, microservice-sized model infrastructure lets you hot-swap models as the landscape evolves, giving both you and your customers optionality. Third, bring CI/CD discipline to model management: automated regression tests when you swap a model, one-click rollbacks when behaviour deviates from baselines — treating model changes with the same rigour as code releases. Fourth, invest in monitoring and drift detection so that six months of reinforcement learning doesn't silently push the model into a problem space it was never meant to solve. And throughout all of it, keep a human in the loop: the best learnings from model outputs should feed back into continuous improvement. Shane draws a parallel to MongoDB's MAP (MongoDB AI Applications Program) composable reference architecture, which follows the same hot-swap logic across LLMs and orchestration frameworks.

  • Shane asks whether the architectural considerations amount to 'responsible AI' — and Akshaya's answer is a layered yes, grounded in the NIST Risk Management Framework's four-function cycle: Govern, Map, Measure, and Manage. The key shift from traditional software, he argues, is that AI requires trustworthiness to be embedded at every phase of the model lifecycle, not just at the application layer. Unacceptable risks must be identified and eliminated early; high-risk use cases must be scoped out of the product footprint; bias and privacy protections must be baked in by design. He connects this to EU AI Act compliance, where the stakes for getting it wrong are both legal and reputational. The cautionary example is Google's early AI stumbles — high-profile bias incidents that damaged trust in ways that were slow to repair. His conclusion: any enterprise building an AI product should treat the NIST cycle as a minimum operating standard, not as optional overhead.

  • The abstract discussion of AI bias lands with sudden, personal force when Akshaya shares what happened at a Zendesk internal finance leadership event. The team was using ChatGPT to generate personalised action figures from staff photos. When Akshaya uploaded his own picture, the system automatically labelled the figure 'Ankur Patel' — a name he describes, wryly, as 'the stereotypical Indian name in the United States.' The incident is both funny and pointed: it demonstrates, in a single data point, that the training data powering leading commercial models remains heavily skewed toward English-language, American cultural norms. Non-English languages and non-Western cultures occupy a far smaller slice of the training corpus, which means cultural nuance, dialectal variation, and naming conventions from outside the dominant training distribution are routinely flattened or stereotyped. Akshaya's assessment is measured but clear: bias is not a solved problem, it is a managed one, and controlling for it at the application layer — through guardrails, domain-specific fine-tuning, and careful output validation — is the current best practice until the underlying training pipelines improve.

  • Shane asks whether Zendesk pays attention to public model evaluation benchmarks when choosing which models to build products on. Akshaya's answer is nuanced: benchmarks serve a purpose — they signal general capability — but they are not a reliable proxy for enterprise task fit. The fact that a model can ace LSAT questions or solve complex mathematical problems says little about whether it can resolve a customer support query accurately and empathetically. His analogy is sharp: the top-scoring student in the class is not necessarily the brightest or most capable person in the room. At Zendesk, the practical response has been to evaluate both custom and commercial options in parallel and release both when appropriate — a custom triage model and a GPT-4o-powered feature launched on the same day. The lesson for enterprises is clear: use benchmark scores as a starting filter, not a final verdict. Test models against your own representative data and tasks, and maintain optionality by not over-committing to a single vendor or architecture. MongoDB draws a parallel: for 18 months the team has been working with code assistant providers to build MongoDB-specific evaluation sets, precisely because general benchmark performance doesn't predict performance on MongoDB-specific coding tasks.

  • As the conversation approaches its end, Shane asks Akshaya to distil his guidance for the C-suite into a single starting point. The answer is deceptively simple but cuts against a common failure mode: begin with the problem, not the technology. Too many AI initiatives launch because the board or the CEO has mandated that the company 'do something with AI', without first defining what specific problem needs solving or whether AI is genuinely part of the answer. Once a clear problem statement exists, the subsequent questions sequence naturally: is the data in place? What model best serves this particular purpose? Do we have the in-house talent, or do we need to buy? What will it cost, and does the projected outcome justify it? Shane also asks how Akshaya personally keeps up with the relentless pace of AI development — the answer is opportunistic rather than systematic, with Lex Fridman's podcast recommended for philosophical depth and the AI Engineer podcast for technical detail. The episode closes with a pointer to the Zendesk AI Hub for listeners who want to explore Zendesk's AI product updates, demos, and releases.

  • Shane closes by thanking Akshaya and the global audience that joined live from Denmark, Wisconsin, Burkina Faso, and Nairobi. Akshaya directs listeners to the Zendesk AI Hub as the single destination for AI product updates, demos, and release notes. The sign-off is warm but brisk — Akshaya is, quite literally, being asked to leave the room. It's a fittingly human end to a conversation about technology that aspires to be human-centric.

AI washing
Marketing that attaches 'AI' branding to a product without delivering genuine AI-driven value — analogous to greenwashing; the episode cites a rice cooker labelled 'AI Powered' as a consumer example.
RAG (Retrieval-Augmented Generation)
An AI architecture pattern where a model retrieves relevant documents from a database (e.g. via vector search) before generating a response, grounding outputs in specific data sources.
Fine-tuning
The process of further training a pre-trained AI model on a domain-specific dataset to specialise its behaviour, without rebuilding it from scratch; discussed as expensive but cheaper than full custom builds.
TCO (Total Cost of Ownership)
The full long-term cost of a technology investment including purchase, integration, maintenance, and support — used here to compare build-from-scratch vs integrate-off-the-shelf AI approaches.
LLM (Large Language Model)
A deep learning model trained on massive text datasets capable of generating, summarising, and reasoning about language; examples discussed include GPT-4, Llama, and GPT-4o.
Hyperscale
Operating at an extremely large scale of compute or usage — e.g. billions of tokens per hour — at which point the economics of self-hosted models can become more favourable than API-based pricing.
MLOps
Machine Learning Operations: the practices and tooling for deploying, monitoring, and maintaining ML models in production, analogous to DevOps for software.
CI/CD
Continuous Integration / Continuous Deployment: an automated software engineering practice for testing and releasing code changes; Akshaya Murthy advocates applying it to AI models so they can be rolled back like code.
NIST framework
A cybersecurity and risk management framework published by the US National Institute of Standards and Technology; its AI Risk Management Framework covers four functions: Govern, Map, Measure, and Manage.
GDPR
General Data Protection Regulation: EU law governing how organisations collect, store, and use personal data; referenced as a compliance requirement for enterprise AI vendors.
CSAT
Customer Satisfaction Score: a KPI measuring how satisfied customers are with a service interaction, used here as a metric for evaluating AI-driven customer support improvements.
EBITDA
Earnings Before Interest, Taxes, Depreciation, and Amortisation: a common measure of a company's core operating profitability, cited as one of the business outcomes AI transformation should move.
Drift detection
Monitoring a deployed AI model for statistical changes in its outputs or input data over time that indicate the model is becoming less accurate or relevant — a key element of responsible AI operations.
Token
The basic unit of text that LLMs process (roughly ¾ of a word); API pricing is charged per token, and Akshaya Murthy cites OpenAI's rate of ~$2 per million tokens as a key cost reference point.
Copilot (in AI context)
An AI assistant embedded in a workflow that provides real-time guidance, suggested next actions, or drafted responses to a human operator — contrasted here with older rule-based macro automations.
Modular API-first architecture
A software design approach where AI model components are exposed through standardised APIs and containerised microservices, enabling teams to swap in different models without rebuilding the whole system.
POC (Proof of Concept)
A small-scale pilot project used to test whether a technology or approach is feasible before committing to a full build; mentioned frequently in the context of validating AI ROI.
Hugging Face
An open-source AI platform hosting thousands of pre-trained models and datasets, where Akshaya Murthy notes new model releases arrive on a daily basis.
Bespoke
Custom-made or tailor-built for a specific purpose; used here to describe AI models built for specialised verticals like defence or healthcare that general-purpose LLMs cannot adequately serve.
ERP (Enterprise Resource Planning)
Integrated software platforms that manage core business processes; Akshaya Murthy uses the historical SAP/Oracle ERP buy-vs-build debate as a direct analogy for today's AI integration decisions.

Chapter 3 · 05:35

What AI Transformation Means in Practice at Zendesk

Far from being a visionary role with a loose mandate, Akshaya describes AI transformation as a rigorous exercise in working backwards from business outcomes. The starting point is always the goal — improve employee experience, improve customer experience, make operations best in class — and AI is the means, not the end. He breaks the role into two complementary tracks: identifying tangible, high-return opportunities where AI can genuinely move the needle on KPIs like productivity lift, CSAT, revenue, and EBITDA; and advocating internally for adoption, because technology alone doesn't drive behaviour change. Credibility is earned incrementally, he argues, by promising specific outcomes and then delivering them — a discipline that separates real transformation from AI hype.

Chapter 4 · 08:25

Zendesk Overview: Connective Tissue for Customer Relationships

Shane's prompt to explain Zendesk for the uninitiated draws a clear, structured answer from Akshaya: think of it as the connective tissue between a business and its customers. The platform delivers on three dimensions — unified communications that collapse email, messaging, and phone into a single pane of glass; frictionless self-service powered by generative search and AI chatbots for instant resolution; and robust analytics that surface trends and patterns so organisations can improve continuously. With 100,000 brands from Airbnb to Shopify and billions of annual interactions, the scale is significant. The deeper point, though, is Zendesk's strategic pivot: it's no longer building better ticketing software, it's building for measurable outcomes. AI is the engine of that shift, and the conversation about what AI can do has changed from 'help me get to a resolution' to 'give me the resolution now'.

Business
Zendesk's Three Big AI Shifts

Don't Build Your Own AI (Unless You Have To) · Mar 6, 2026 Business

Zendesk's AI transformation produced three seismic product shifts: intelligent triage cuts resolution time to seconds, Copilots replace rule-based macros with live agent guidance, and workforce management becomes fully automated. In extreme cases, that's a 92% reduction in email handle time.

Chapter 5 · 11:37

Zendesk's Three AI Product Shifts and Their Results

Akshaya walks through the three shifts that define Zendesk's AI-first redesign. The first is intelligent triage: what once required human review now happens in seconds, getting customers materially closer to resolution faster. The second is the move from macros — rigid, rule-based automations — to Copilots that provide dynamic, context-aware guidance to agents in real time. The third is workforce management, where load balancing and staffing optimisation become fully automated rather than siloed in a contact centre operation. The results speak loudly: in extreme cases, early adopters have cut email handle time by 92% and slashed customer question response time by 74%. The through-line, Akshaya stresses, is a fundamental reframe: Zendesk has moved from building better ticketing software to delivering measurable business outcomes for customers.

Chapter 6 · 14:55

The New Rules of Customer Expectations

The velocity of change in customer expectations is the theme of this segment. Akshaya identifies three seismic shifts: immediacy — SLAs are now measured in seconds, not hours; context — customers expect the service system to already know their history and what has been tried; and human-centricity — AI interactions must feel human, even when they aren't. The paradox he surfaces is striking: a 2025 survey shows 90% of companies say AI is central to customer loyalty, but the same data reveals that loyalty drops off a cliff when customers discover they're speaking to a bot. The implication for product designers is profound: the experience has to be seamlessly intelligent and warmly human at the same time. Customers, Akshaya concludes, simply don't care how it works — they want resolution, personalisation, and friction-free experience, and AI delivered at scale is the only mechanism that can meet that bar.

Chapter 7 · 16:40

Build vs Buy: The 4-Factor Framework for Enterprise AI

The build-vs-buy debate, Akshaya notes, is not new — he watched the same argument play out with SAP and Oracle ERP systems two decades ago, and enterprises eventually concluded that keeping up with the resulting tech debt was untenable. The same calculus is now playing out with AI, but under enormous time pressure as CFOs chase ROI proof for their investors. Akshaya structures the decision around four factors. Time to market and speed to impact: can you afford the months or years it takes to build from scratch when an off-the-shelf system can be online in one to two weeks? Total cost of ownership: the CapEx of a custom build almost always makes integration look attractive in the near term. Talent scarcity: the pool of AI engineers who can build systems at scale is tiny, and Stanford or Berkeley graduates in this space command $500K or more per year. Security and compliance: not every company has the expertise to build enterprise-grade security into a new product, and every AI vendor now ships with enterprise-grade privacy and GDPR compliance as table stakes. For most enterprises, integration wins on all four counts — and it frees up scarce development bandwidth to focus on building only what truly needs to be custom.

Chapter 8 · 22:20

Off-the-Shelf AI Maturity and the Commoditisation of LLMs

Akshaya paints a landscape where the case for building your own model is eroding almost by the day. The gap between GPT-4 and Llama's 70-billion-parameter open-source model is narrowing to the point where a self-hosted inference node can produce comparable results for many use cases. Security frameworks that once took years to mature are now shipped as standard in enterprise-grade off-the-shelf products — GDPR compliance, privacy controls, role-based access — because these are competitive table stakes. Model refreshes have gone from annual to weekly on commercial platforms, and daily on Hugging Face for open-source variants. And the per-token cost is compressing: OpenAI charges roughly $2 per million tokens, while running Llama yourself can bring that to fractions of a cent. Most strikingly, the expertise required to work with these systems has inverted: in 2023 you needed PhDs with prompt engineering experience; today a product manager can achieve the same with a drag-and-drop interface. The only genuine exceptions to the 'integrate-first' rule are hyperscale operations where token costs make self-hosting cheaper, and mission-critical verticals like defence, banking, or healthcare where general-purpose training data is simply inadequate.

Chapter 9 · 25:55

AI Pricing, Value-Based Models, and the Road to Utility

Akshaya identifies a structural tension in the current AI market: the industry is simultaneously watching per-token costs compress toward commodity pricing while also realising that the value delivered far exceeds what those tokens cost to produce. The result is a mispricing problem — vendors don't yet know how to charge for value, and buyers don't yet know how to quantify what they're getting. Zendesk's response has been to move toward resolution-based pricing: customers pay for solved problems, not for compute cycles consumed. That model change, Akshaya argues, offsets some of the cost pressure while aligning incentives more honestly with outcomes. His long-range view is that AI will follow the same arc as internet access — initially priced per megabyte or gigabyte, then compressed into flat-rate utility pricing. A 'value-add tax' will stabilise somewhere above pure commoditisation, but the gross cost trajectory is definitively downward. For enterprises making build-vs-buy decisions now, that means today's integration economics are still favourable, but the calculus will shift over time.

Chapter 11 · 30:10

The Three Real Barriers to Enterprise AI Adoption

This chapter delivers what may be the episode's most practically useful insight: most enterprise AI failures are not caused by the technology. The algorithm works. What breaks AI projects is the plumbing, the people, and the inability to prove value. Akshaya opens with data: messy, siloed data is the original sin of enterprise AI. Without a single source of truth — a data lake house with clean pipelines and scrubbed PII — even the most sophisticated model will produce unreliable outputs. He coins a memorable inversion of the classic tech maxim: 'garbage in, hype out.' The second barrier is a skills gap that spans the full organisation, from MLOps engineers to sales reps and customer-facing employees. Legacy SaaS companies face a particular challenge because AI adoption requires company-wide buy-in from day one, not a gradual persuasion campaign. The third barrier is vague ROI: too many POCs promise 90% cost cuts without showing the mechanism. His solution to the people problem is hands-on training — one session beats 20 courses or 30 Slack announcements — plus clear, credible commitment from the C-suite.

Chapter 12 · 34:50

AI Washing: How to Spot It and Why It Matters

AI washing, Akshaya explains, is what happens when 'AI' becomes a marketing suffix rather than a functional description — and he opens with a disarming example: a rice cooker bearing an 'AI Powered' sticker. In the enterprise context, the pattern is more insidious: a vendor adds a thin LLM plug-in to an existing product without re-architecting the solution or reimagining how value is delivered. The test for legitimacy comes down to four questions. Is the AI solving a specific, well-defined problem? Are concrete KPIs and ROIs attached to the solution, and is it clear how AI achieves them versus the prior approach? Is the claimed improvement consistent with realistic AI outcomes — Zendesk's internal POCs see 2x to 50x ROI, so a vendor promising 10% CSAT lift should raise eyebrows? And finally, is the vendor claiming a universal 'one size fits all' solution, which is almost always a red flag? Ticket summarisation might be one of the few truly general-purpose AI tasks, but beyond that, domain specificity matters enormously. If those four boxes aren't checked, you're looking at a rule-based engine wearing an AI costume.

Business
What AI Washing Actually Looks Like

Don't Build Your Own AI (Unless You Have To) · Mar 6, 2026 Business

AI washing is a buzzword salad with no specifics. If a vendor can't show you the KPIs, the defined problem it solves, and a measurable ROI, they're masquerading a rule-based engine as AI. The rice cooker with an 'AI Powered' sticker is not a joke — it's a category.

Chapter 13 · 38:30

Architectural Pillars for Production-Grade AI Products

Moving from strategy to implementation, Akshaya gives his clearest technical prescription of the episode. Clean data plumbing comes first: pipelines, a single source of truth, and PII scrubbing — because 'garbage in, hype out' and even the perfect transformer model will break on bad data. Second, build a modular API-first model layer. Just as you wouldn't lock customers into a single ISP, you cannot lock them into a single LLM vendor. Containerised, microservice-sized model infrastructure lets you hot-swap models as the landscape evolves, giving both you and your customers optionality. Third, bring CI/CD discipline to model management: automated regression tests when you swap a model, one-click rollbacks when behaviour deviates from baselines — treating model changes with the same rigour as code releases. Fourth, invest in monitoring and drift detection so that six months of reinforcement learning doesn't silently push the model into a problem space it was never meant to solve. And throughout all of it, keep a human in the loop: the best learnings from model outputs should feed back into continuous improvement. Shane draws a parallel to MongoDB's MAP (MongoDB AI Applications Program) composable reference architecture, which follows the same hot-swap logic across LLMs and orchestration frameworks.

Technology
Architectural Principles for Robust AI Products

Don't Build Your Own AI (Unless You Have To) · Mar 6, 2026 Technology

The architecture decisions that will make or break your AI product: clean data pipelines, a modular API-first model layer that lets you hot-swap LLMs, CI/CD for models themselves, and continuous monitoring with human review. These aren't nice-to-haves — they're the difference between a product and a liability.

Chapter 14 · 43:00

Responsible AI and the NIST Framework

Shane asks whether the architectural considerations amount to 'responsible AI' — and Akshaya's answer is a layered yes, grounded in the NIST Risk Management Framework's four-function cycle: Govern, Map, Measure, and Manage. The key shift from traditional software, he argues, is that AI requires trustworthiness to be embedded at every phase of the model lifecycle, not just at the application layer. Unacceptable risks must be identified and eliminated early; high-risk use cases must be scoped out of the product footprint; bias and privacy protections must be baked in by design. He connects this to EU AI Act compliance, where the stakes for getting it wrong are both legal and reputational. The cautionary example is Google's early AI stumbles — high-profile bias incidents that damaged trust in ways that were slow to repair. His conclusion: any enterprise building an AI product should treat the NIST cycle as a minimum operating standard, not as optional overhead.

Chapter 15 · 44:50

AI Bias Is Still Real: A Personal Demonstration

The abstract discussion of AI bias lands with sudden, personal force when Akshaya shares what happened at a Zendesk internal finance leadership event. The team was using ChatGPT to generate personalised action figures from staff photos. When Akshaya uploaded his own picture, the system automatically labelled the figure 'Ankur Patel' — a name he describes, wryly, as 'the stereotypical Indian name in the United States.' The incident is both funny and pointed: it demonstrates, in a single data point, that the training data powering leading commercial models remains heavily skewed toward English-language, American cultural norms. Non-English languages and non-Western cultures occupy a far smaller slice of the training corpus, which means cultural nuance, dialectal variation, and naming conventions from outside the dominant training distribution are routinely flattened or stereotyped. Akshaya's assessment is measured but clear: bias is not a solved problem, it is a managed one, and controlling for it at the application layer — through guardrails, domain-specific fine-tuning, and careful output validation — is the current best practice until the underlying training pipelines improve.

Chapter 16 · 47:20

Evaluating AI Models: Benchmarks vs. Task Fit

Shane asks whether Zendesk pays attention to public model evaluation benchmarks when choosing which models to build products on. Akshaya's answer is nuanced: benchmarks serve a purpose — they signal general capability — but they are not a reliable proxy for enterprise task fit. The fact that a model can ace LSAT questions or solve complex mathematical problems says little about whether it can resolve a customer support query accurately and empathetically. His analogy is sharp: the top-scoring student in the class is not necessarily the brightest or most capable person in the room. At Zendesk, the practical response has been to evaluate both custom and commercial options in parallel and release both when appropriate — a custom triage model and a GPT-4o-powered feature launched on the same day. The lesson for enterprises is clear: use benchmark scores as a starting filter, not a final verdict. Test models against your own representative data and tasks, and maintain optionality by not over-committing to a single vendor or architecture. MongoDB draws a parallel: for 18 months the team has been working with code assistant providers to build MongoDB-specific evaluation sets, precisely because general benchmark performance doesn't predict performance on MongoDB-specific coding tasks.

Chapter 17 · 49:40

Final Advice for C-Suite and Product Leaders

As the conversation approaches its end, Shane asks Akshaya to distil his guidance for the C-suite into a single starting point. The answer is deceptively simple but cuts against a common failure mode: begin with the problem, not the technology. Too many AI initiatives launch because the board or the CEO has mandated that the company 'do something with AI', without first defining what specific problem needs solving or whether AI is genuinely part of the answer. Once a clear problem statement exists, the subsequent questions sequence naturally: is the data in place? What model best serves this particular purpose? Do we have the in-house talent, or do we need to buy? What will it cost, and does the projected outcome justify it? Shane also asks how Akshaya personally keeps up with the relentless pace of AI development — the answer is opportunistic rather than systematic, with Lex Fridman's podcast recommended for philosophical depth and the AI Engineer podcast for technical detail. The episode closes with a pointer to the Zendesk AI Hub for listeners who want to explore Zendesk's AI product updates, demos, and releases.

No indexed bits in this chapter.

Show stoppers

Snapshots ()

Key Quotes ()

This episode

Claims & Sources

1 / 13 cited (8%)

Factual claims made this episode, and whether a source was named.

Zendesk AI early adopters have cut email handle time by 92% in extreme cases using AI-powered triage.

Akshaya Murthy no source cited

Zendesk AI Copilot early adopters reduced time spent responding to customer questions by 74%.

Akshaya Murthy no source cited

In 2025, 90% of companies say AI is central to customer loyalty, but customer loyalty drops sharply when customers discover they are talking to a bot.

Akshaya Murthy An unspecified 2025 industry survey

Hiring an AI expert capable of building systems at scale costs a minimum of $500,000 per year in the United States.

Akshaya Murthy no source cited

AI has overtaken cybersecurity to become the number one investment line item in company priorities.

Akshaya Murthy no source cited

An off-the-shelf AI system like Zendesk can be brought online in one to two weeks, compared to months or years and significant CapEx to build from scratch.

Akshaya Murthy no source cited

OpenAI charges approximately $2 per million tokens for API access, while open-source Llama models can be run for fractions of a cent per token.

Akshaya Murthy no source cited

AI model refreshes on commercial platforms are occurring on a weekly basis, while on Hugging Face open-source updates arrive on a daily basis.

Akshaya Murthy no source cited

In 2023, building AI products required a team of PhDs with prompt engineering expertise; today a product manager can do the same with a drag-and-drop interface.

Akshaya Murthy no source cited

Zendesk has approximately 100,000 brand customers, including Airbnb and Shopify, and handles billions of customer interactions per year.

Akshaya Murthy no source cited

Zendesk's internal AI POC and platform implementations have achieved between 2x and 50x ROI.

Akshaya Murthy no source cited

ChatGPT generated an action figure from Akshaya Murthy's photo and automatically assigned it the name 'Ankur Patel', a stereotypical Indian-American name, demonstrating ongoing AI bias.

Akshaya Murthy no source cited

The gap in capability between GPT-4 and the 70-billion-parameter Llama open-source model is narrowing significantly.

Akshaya Murthy no source cited

This episode

Cast

  • Track
  • Track
  • Track

Stats

Episode stats

Insight Overview

insights
chapters

Insight distribution

Sub-Categories

Speaker breakdown

Talk Time