Working methods for finding where AI belongs, building it, testing it, and measuring whether it worked.
How to design multi-agent systems
A sub-agent looks like a tool call from the outside and is a conversation on the inside. Here's what that means for how you draw the boundaries, how the orchestrator decides to call one, and what you have to evaluate at every level.
12 min
How to build a dataset for evals
How to choose the cases, write down the standard for each one, and decide how many you need.
20 min