Language has billions of possible next tokens — the PUCT exploration heuristic assumes you'll visit the same node multiple times, but an LLM will almost never generate the exact same token sequence twice. The discrete action assumption breaks.
Language has billions of possible next tokens — the PUCT exploration heuristic assumes you'll visit the same node multiple times, but an LLM will almost never generate the exact same token sequence twice. The discrete action assumption breaks.
Where this was said
At 1:44:55 · chapter starts 1:41:09
Eric and Dwarkesh discuss why MCTS fails for LLM reasoning: language's vast token space violates the discrete finite-action assumption, and value estimation is much harder than in Go. [1] — Eric Jang "Language has billions of possible next tokens — the PUCT exploration heuristic assumes you'll visit the same node multiple times, but an LL…" 1:44:55
Scaling laws only emerge cleanly when your system already works, your data is good, and there are no bugs. Trying to extract scaling insights before you have a working baseline just gives you scaling laws on garbage.
Unlike naive RL which must figure out which of 100K+ tokens caused a win, MCTS provides a strictly better action target for every single move in every game.
A 10-layer neural network can compress what looks like an NP-class search problem into a single forward pass. This happened with Go, protein folding, and tensor decomposition. It might mean our understanding of computational hardness is fundamentally incomplete.
Eric Jang trained a strong Go bot for roughly $10K of donated compute from Prime Intellect, spending about $3K on the final training run.
AlphaGo Zero was trained on far more compute than any other AI model of its era — roughly 3e23 flops, comparable in order of magnitude to frontier LLMs.
Pre-training on 9x9 Go boards and then warm-starting a 19x19 model dramatically cuts the time needed to learn endgame value functions.
Game apps keep users engaged long enough for ads to pay off. Tool apps don't — so if you're building a utility, ads are almost always the wrong call and subscriptions are your only real lever.
Inside SEO Bot, a single button labelled 'Boost My Domain Rating' routes users directly to Listing Bot. That one interaction converts a user of one tool into a user of two — without any marketing cost.
Directory listings are a powerful but underrated growth channel — but only if your product is genuinely interesting enough to earn the click. AI products have a natural advantage here because they're easy to package in a compelling, clickable way.
The fastest path to Twitter growth isn't volume — it's forming sharp opinions about how the platform works and sharing them immediately. People cluster around those who understand the rules and say so out loud.
Sam had no coding knowledge, so he used ChatGPT voice mode to generate his entire codebase and copy-pasted it into Notepad. A friend later introduced him to Cursor, and he never looked back.
Copy days of Discord chat history, paste it into ChatGPT, and ask it to list recurring pain points. The ones that come up most often are your best product bets.
Sam's top advice: when prompting Cursor, tell it to architect code for 100,000 users from day one. The AI changes its approach, building scalable frameworks instead of brittle one-user code.
With AI coding tools like Cursor, Bhanu replicates an existing free tool for a new keyword in under 5 minutes. What used to be a multi-day build is now a lunch-break task.
Ahrefs, SiteGPT, Cal.com, PostHog, Datafast, Sibyl AI, Bento, Feather, Featurepace, Mintlify, Cloud Code, ChartMogul — Bhanu runs his entire business solo with these 12 tools.
We use essential and analytics cookies to run Vuci. To understand how the site is used: Privacy Policy.
Install Vuci on your phone
Add it to your home screen for a faster, app-like experience.
Install Vuci on your phone
Tap the Share button, then “Add to Home Screen”.
A new version is available
Reload to get the latest Vuci.