Office hours
Pete builds with ZeroWidth live on Twitch. Every stream is here, edited down, with chapters to jump straight to the part you want.
Oct 7, 2026 · 46 min · Prism · Watch on YouTube
Exploring a game where every NPC is an LLM
Pete uses Prism to explore a game where every character is a language model, checks whether the idea already exists, then writes and takes a survey with AI follow-up questions.
- 0:00Yesterday's recap
- 0:44A game where every NPC is an LLM
- 2:27What Prism is for
- 3:27Pushing an idea along an axis
- 7:19Generated images, and the flows behind Prism
- 8:25When your product outruns your testing
- 10:06Exploring without a plan
- 11:17What only works if the NPC is a model
- 15:48The ideas between two ideas
- 18:30Making it a web game
- 20:54When the model loses the plot
- 23:29Each guest is its own model
- 26:03Desk research: does it already exist?
- 28:49Not a new idea
- 29:26Drafting a survey with zv1
- 31:24AI follow-up questions
- 33:59Taking the survey
- 38:04Results and the printable report
- 39:25Research that calls you out
- 40:26Clustering the field into themes
- 42:00Sharing the field
- 43:17Quotas and automatic re-analysis
- 44:21Research on a schedule
- 45:07What's next
Oct 6, 2026 · 57 min · Watch on YouTube
Does your AI get better the more people use it?
Pete builds a texting-style agent in Workbench, measures it in Caliper, improves it with zv1 and adds memory, to show which kinds of better happen on their own and which ones you have to design.
- 0:00Does it get better the more you use it?
- 0:42Brand kits and on-brand decks in Napkin
- 3:39Three things people mean by better
- 4:29Building a texting-style agent in Workbench
- 10:27From anecdotes to evals
- 11:53Datasets and rubrics in Caliper
- 15:27Generating test cases with zv1
- 18:33Using the MCP server from Claude Code or Codex
- 19:43Why multi-turn evals matter
- 23:23When the judge misses tone
- 25:05A tone rubric: 88% to 74%
- 28:46Why the number matters
- 32:23Improving the agent with zv1
- 34:44Keep test cases out of your prompt
- 35:57Up 21 points
- 37:50Using an AI doesn't train it
- 38:58Shared memory or memory per person
- 40:39Profile data in the system prompt
- 42:20How the agent writes to memory
- 44:32Memory is retrieval
- 45:08Debugging a memory that won't stick
- 48:31Deciding what gets remembered
- 49:32So, did it get better?
- 50:07Evals for memory and personal preferences
- 53:15Prompt injection through a profile name
- 54:40None of this is machine learning
Try ZeroWidth on your own work. It's free to start.