Alex Lavaee and I did two masterclasses on this. This is the edited cut under 40 minutes so you can finish it.
Your coding agent can produce a diff. You still decide what runs, what counts as done, and what evidence is good enough to ship.
Alex walks through control loops and graphs, Atomic as a verifiable coding agent runtime, context as a budget, artifacts instead of dumping transcripts, durability and mid-run steering, and turning skills into inspectable workflows. Proof beats the model saying it is done.