Achieving 99% accuracy in internal testing does not guarantee the same performance in production because real users will surface edge cases and scenarios internal tests never covered.
Snapshot · The MongoDB Podcast
Achieving 99% accuracy in internal testing does not guarantee the same performance in production because real users will surface edge cases and scenarios internal tests never covered.
Where this was said
At 15:30 · chapter starts 10:40
This is where the conversation moves from biography to substance. Karthik systematically dismantles the assumption that AI product development resembles traditional software engineering. Before LLMs, observability was reactive — you built the product, shipped it, and monitored for failures. Errors had logical causes you could reason through. LLMs break this model entirely: the same prompt can yield different outputs at any time, and even an overnight model update can silently degrade a product the developer last tested hours before. [1] — Karthik Kalyanamaran "Before LLMs, software bugs were logical — a missing condition, a stack overflow. You could reason your way to the fix. LLMs are non-determi…" 11:50 Traditional unit tests are useless in this environment because they assume deterministic outputs. Karthik argues that the right mental model is an ongoing operational journey: develop with Langtrace tracing every API call, manually score traces to establish a baseline, ship to production when you have reasonable confidence (not 99% certainty — that's an illusion), then continue monitoring in production and iterating as real users expose edge cases your tests never anticipated. The tooling to support this loop is precisely what Langtrace is designed to provide.
Before LLMs, software bugs were logical — a missing condition, a stack overflow. You could reason your way to the fix. LLMs are non-deterministic: the same input can produce a different output every time, and traditional unit tests are useless against that. Observability is no longer reactive maintenance; it's a required part of the build process from day one.
Token count, cost per model call, time-to-first-token, and tokens-per-second for streaming — these are the core metrics Langtrace surfaces for every LLM invocation. For vector databases like MongoDB Atlas, Langtrace traces pipeline settings, retrieved results, and even embeddings, enabling replay analysis when something goes wrong in retrieval.
Sam's initial MVP was coded in approximately one week using ChatGPT voice mode and copy-pasting code, with no prior technical experience.
Sam argues Discord is 10x better than email for building relationships with younger users who rarely check their inbox.
Sam's monthly operating costs include Cursor ($200), AI image generation ($100), AI video generation ($200), hosting ($100), email marketing ($80), and AI compute ($300–$500).
Sam recommends copying days of Discord chat history into ChatGPT and prompting it to list recurring pain points as a fast, free market research technique.
Bhanu and his team built approximately 50 free tools to attract search traffic, each linked back to SiteGPT.
With AI coding tools like Cursor, Bhanu can now create a new free marketing tool in less than 5 minutes by referencing existing tools.
Bhanu filters Ahrefs keyword results to show only those with a keyword difficulty below 10, making them realistic ranking targets for any decent website.
Bhanu sets a minimum search volume of 1,000 monthly searches when selecting keywords to target with free tools.
PropGPT averaged 20 downloads per day right after launching on the App Store through influencer marketing.
We use essential and analytics cookies to run Vuci. To understand how the site is used: Privacy Policy.
Install Vuci on your phone
Add it to your home screen for a faster, app-like experience.
Install Vuci on your phone
Tap the Share button, then “Add to Home Screen”.
A new version is available
Reload to get the latest Vuci.