An agent rebuilt a website. The pull request merged. The live site did not change. Nobody lied. The agent checked the wrong thing: it treated "merged" as "done". That small miss shows the real problem with AI agents today. They are fast at making work, and unreliable at knowing when the work is really finished. A good engineer solves this with three things: a way of working, a "done" that can pass or fail, and a memory of past mistakes. A new agent chat starts with none of them, so they have to be written down. pstack, by Lauren Tan, is that writing: skills, principles, and playbooks, tied together by /poteto-mode, with /figure-it-out for jobs no playbook covers. We follow our website story through it, one step at a time. Write "done" so it can fail. Check the real live site, not a stand-in. Keep a trail. Learn from the miss with /reflect. Then turn the lesson into structure, so the next run cannot repeat it. We finish by being honest about the cost and about when not to use it. The whole talk lands on one idea from Lauren's README: if you want to go fast, go deep first.
Each section has three parts. What to say is written to be read aloud. Why it matters is the point to land. Next is the bridge to the following slide.
The PR merged. The website didn't change.
What to say
Let me start with a small, true story.
On the 5th of October, an AI agent rebuilt the website for my studio, malohacoast.com. It did the work. It opened a pull request. The pull request was reviewed, and it was merged.
By every normal sign, the job was done.
Then I opened the real website in my browser. It was the old page. Exactly the same as before. Nothing had changed.
Why it matters
Most of us have felt this. The ticket is closed, but the customer still sees the bug. There is a gap between "the work says it is done" and "the world has actually changed". That gap is what this whole talk is about.
NextSo what went wrong? It was not a bug in the code.
What actually happened
What to say
Here is what happened, step by step.
The agent rebuilt the page and opened a pull request. It was merged. But the live site kept showing the old consulting page.
Why? The site is hosted on Cloudflare Pages. A host is the computer that hands your web page to every visitor. Some hosts are connected to your Git repository. When you merge to main, they notice, build the site, and publish the new version for you.
This one was not connected. The Pages project had no Git source at all. It was set up for manual uploads, where someone pushes the files to it by hand. So merging to main deployed nothing. The code changed in the repository, and the host never heard about it.
Think of a shop with a display window. In some shops, a person updates the window every time the stock changes. This shop had nobody doing that. We changed the stock in the back room, and the window stayed the same. In real terms: no Git source, so a merge triggers no deploy.
In the end we uploaded the new files by hand with Cloudflare's Wrangler tool, and the new page went live.
Now the key line. Nobody lied. The agent was not careless. It assumed that "merged to main" means "live". That was false here. It checked a stand-in, the merge, instead of the real thing, the live site. Engineers call a stand-in like that a proxy.
Why it matters
This is the pattern to remember. Most agent mistakes are not lies, and they are not bad code. They are checks of the wrong thing. If we want to trust agents, we need them to check the real thing.
NextThat is the story. Now let me tell you what this talk is.
pstack, from the problem up
What to say
This is a guide to pstack. pstack is a set of open-source agent skills made by Lauren Tan, who goes by poteto online. In her README she says she has worked with millions of lines of code at Meta, Netflix, and Cursor, and that she is on the React core team.
I am going to explain pstack from first principles. That means we start with the problem, not the tool. We build up one piece at a time, and each piece should follow from the one before. Then we look at how it works, and where it does not fit.
And we use the website story to hold it all together, because one real miss teaches more than a list of features.
Why it matters
A tool only makes sense once you feel the problem it solves. If I start with a list of commands, you will forget them by tomorrow. If you understand the problem, the commands start to look obvious.
NextSo let us name the problem clearly.
The problem pstack solves
What to say
Agents are fast at producing work. They are unreliable at knowing when it is done.
Writing code used to be the slow part of building software. Now an agent can write a lot of code in a few minutes. So the hard part has moved. The hard part now is one question: is this actually correct, and is it actually finished?
Lauren says it plainly in her README: "throughput without quality is not a goal i aspire to." Throughput just means how much you produce. She writes that there is a growing sense that AI writes too much sloppy code, and she agrees. Her aim is the opposite of more code. pstack, in her words, helps you write less, but higher quality code.
Why it matters
If you only measure how much an agent makes, you will get a lot of stuff, and some of it will be wrong in ways you cannot see. Our website was a perfect example. A lot of work, done fast, and the one thing that mattered, the live page, never changed.
NextAnd this problem does not stay small. It grows.
Why it compounds
What to say
Here is a line from the pstack guide: "Verification is the slowest step in most agent work, because it's the step that usually waits on a human."
Verification just means checking that the work is right. Think about who does that checking today. Usually you. The agent finishes, and then it waits for a person to look.
So the whole system moves at the speed of the person checking. Picture a bottle. It does not matter how fast you tip it. It pours at the speed of the narrow neck. Here the narrow neck is the human who has to check.
Now add more agents. Two agents, twice the checking. Ten agents, ten times the checking. You have not removed the slow step. You have made a longer queue in front of it.
The guide puts it this way. Make the agent able to do the checking, and you stop being the bottleneck. Skip it, and running more agents only gets you more unchecked work to review.
Why it matters
This is why checking sits at the center of pstack. Not because checking is polite, but because it decides how far you can scale. An agent that can check its own work can run without you. One that cannot will always come back to you.
NextSo what would it take for an agent to check its own work? Let us look at what a good human engineer brings.
First principles: what a good engineer brings that an agent doesn't
What to say
Let us strip it right back. When a strong engineer does good work on a team, what do they bring that an agent does not? Three things.
One. A way of working. How this team debugs, how it designs, how it ships.
Two. A checkable "done". Not "I think it is fine". Something that can pass or fail.
Three. A memory of misses. Last month's mistake changes this month's habit.
Now, a fresh agent chat starts with none of these. It is like a very skilled new hire on their first morning, every single time. The mechanism is simple: each new chat starts empty, and the agent only knows what it is given or what it reads. When the chat ends, the lessons end with it.
So if you want those three things in an agent, there is only one way. Write them down, somewhere the agent will read.
Why it matters
This gives us the shape of the answer before we have even seen the tool. Hold on to these three. Every piece of pstack we look at next is one of them, written down.
NextAnd that is exactly what pstack is.
So what is pstack?
What to say
In Lauren's article "How I Use Cursor", she says: "I've taken all the failure modes I've observed and turned them into skills."
A failure mode is a typical way that something goes wrong. For example: saying "done" when the code only compiled. Or checking the merge instead of the website. She collected the ones she kept seeing, and wrote each one down as a skill.
The practical facts are short. pstack is a Cursor plugin. You install it by typing /add-plugin pstack in a Cursor chat. It is MIT licensed, which means you are free to use it, change it, and share it. And the skills are written as markdown files, which is plain text with headings. You can open them and read every word the agent is told. Lauren even invites you to fork it and make it yours.
Why it matters
Because the skills are plain text, there is no hidden behavior to take on faith. You can read exactly what the agent will read. When the whole point is trust, that matters.
NextLet us open one of those files and see what is inside.
Building block 1: A skill is a markdown file the agent reads
What to say
The smallest piece of pstack is a skill. A skill is one file, called SKILL.md.
At the top there is a small header. It has a name, like figure-it-out. It has a description, which is the "when to use me" line, so the agent knows when this skill fits. Under the header come the instructions, written in normal sentences. That is all a skill is.
You can think of it as a recipe card. A title, a line saying what it is for, and then the steps.
There is one more line on this slide: disable-model-invocation: true. That flag means the agent will never start this skill by itself. It only runs when you type it, for example /figure-it-out. Some skills change how a whole run works, so a person should be the one who chooses them.
Now back to our site. No skill anywhere told the agent how this particular host deploys. So it filled the gap with a guess: merge means live.
Why it matters
An agent only knows what it reads. If a piece of knowledge is not in a file the agent reads, it is not in the run. A skill is the simplest form of the "write it down" idea from slide 6.
NextA skill is a full set of instructions. But sometimes you need something much shorter: a rule you can call by its name.
Building block 2: Principles, 24 short rules with names
What to say
The next building block is principles. pstack ships twenty-four of them. Each one is a single rule with a short, memorable name. Prove It Works. Subtract Before You Add. Fix Root Causes. Encode Lessons in Structure. Never Block on the Human.
Here is the clever part. /poteto-mode reads the index of these principles at the start of a multi-step task. So the agent already knows each rule. That means the name becomes a steering handle. The guide says you do not call principles directly. You use their names to steer.
It works like a word a team already shares. If someone shouts "fire drill", nobody needs a speech. Everyone knows the whole routine. In the same way, the name points at a full rule the agent has already read.
So instead of writing a long paragraph to correct the agent, you write one line. The guide's own example is: "apply prove it works. run the real import flow and show me the written records." For our site, one line would have done it: "apply prove it works. open the live site."
There is also a built-in honesty check. The guide says the agent has to say, in its reply, which decision the rule changed. If it names a principle but no decision changed, that is a sign it only name-dropped the rule.
Why it matters
Long corrections are slow, and the agent can still miss the point. A name is short, exact, and shared. Principles are the "way of working" from slide 6, broken into pieces you can point at.
NextPrinciples say how to think. But a lot of tasks come up again and again, and for those you want a full list of steps.
Building block 3: Playbooks, 23 recipes for recurring tasks
What to say
The third building block is playbooks. pstack comes with twenty-three. A playbook is an ordered list of steps for one common kind of task.
Take the bug fix playbook. It says: reproduce a defect, root-cause it, and fix with runtime evidence. In plain words: first make the bug happen on purpose, so you can see it. Then find the real reason, not just the symptom. Then fix it, and prove the fix with evidence from the running program, not just from reading the code.
There are playbooks for a new feature, a refactoring, a performance problem, a quick prototype, visual parity (making two screens match exactly), shipping, a long autonomous run, and more.
Each one is what a careful engineer would do, in the order they would do it. The order matters. It stops the agent jumping to "fix" before it has seen the bug.
Now our site. Our job was a redesign plus a deploy to a host that was set up by hand. None of the stock playbooks fit that. Hold that thought. We will come back to it.
Why it matters
A playbook bakes good order into the work. It is the "way of working" and the "checkable done" from slide 6, packed together for one kind of task.
NextSo now we have three building blocks. You do not want to pick them by hand every time. Something has to tie them together.
The router: /poteto-mode ties it together
What to say
You rarely use these pieces by hand. You type /poteto-mode and describe the goal in normal words.
Then four things happen. One, it reads your request. Two, it matches the task to a playbook. Three, it copies that playbook's steps, word for word, into a todo list you can see. Four, as each step comes up, it pulls in the skills and principles that step needs.
Here is the part I want you to notice. If it decides to skip a step, the step does not disappear. It stays in the list, marked skip: with a reason. So you can see what it chose not to do, and why.
A tip from the guide: say the goal, not the ceremony. Tell it what is wrong or what you want, and what "done" means. You do not need to list skills. The playbook already puts them in order.
Now our site. If our list had a step that said "verify on the live site", skipping it would have meant writing "skip" and a reason, right there on screen. That is much harder to miss than a check that simply never happened.
Why it matters
Visibility is the point. A skipped check you can see is a check you can question. A silent skip is exactly how our miss happened.
NextBut what happens when no playbook matches, like our site? The router has an answer for that too.
When no playbook fits: /figure-it-out designs one first
What to say
When the task is large, or matches no playbook, the guide shows /poteto-mode routing it to /figure-it-out. That is what we used for our site.
Its first output is not code. The skill file says: "The deliverable before any code is the workflow itself." In other words, when there is no recipe, it writes the recipe first. It has five phases.
A, Frame. Write down what "done" means as a falsifiable predicate. A predicate is a statement that is either true or false. Falsifiable means it can turn out false, and you would know. In this phase you also size the job, and choose how careful to be. The skill says to lean toward more care.
B, Design. Break the work into small pieces that can each land on their own. Do the riskiest unknown first. Build the check before the work, and capture a baseline from before the change, so the check reads "old value versus new value". Then write the list of phases down, because that list is what the human reviews.
C, Loop. Treat each piece as a small experiment. Make a guess, make the smallest change, measure it against the predicate on the real thing, then keep it if it helped or undo it if it did not. And one line I love: when something passes too easily, suspect the way you are looking before you trust the result.
D, Trail. Log every decision, with evidence.
E, Verify. At the end, check the whole thing against the Phase A predicate, on the real product. Not just on the test setup.
Why it matters
When there is no recipe, the worst move is to start cooking anyway. figure-it-out makes the plan first, and makes it something a human can read and challenge before any work begins.
NextLet us use it. Back to our site, and Phase A.
Back to our site, Phase A: Done, written so it can fail
What to say
Here is the frame we wrote for the next round of the site. Listen to the first words: "On live malohacoast.com and www, within about 10 seconds."
On live. Not in the pull request. On the real site.
And www. The site has two addresses, malohacoast.com and www.malohacoast.com. Both have to work, so both get checked.
Within about 10 seconds. That is a visitor's first impression. Then six clauses:
One. A visitor can name the three products. Two. Each product row has intentional media, meaning images or short clips chosen on purpose. Three. The page no longer reads as a plain text list. Four. Media stays small, and people who turn on reduced motion get a still image. Reduced motion is a setting some people use because moving things on screen make them feel unwell. Five. The hire link and the contact email work. Six. Zero client traces anywhere: not in video frames, file names, image descriptions, or the words on the page.
Compare that to "make the site nicer". You can never fail "make the site nicer". Every clause here can be checked, and every clause can fail. That is what makes it useful.
Why it matters
A "done" you cannot fail is a feeling. A "done" that can fail is a test. Only a test can be checked by an agent without you in the room. This is the "checkable done" from slide 6, made real.
NextWriting "done" well is half the job. The other half is how you check it.
The principle doing the work: Prove it works on the real artifact
What to say
The principle behind all this is Prove It Works. It says: check the real thing, not a stand-in. Not a proxy, not the agent's own report, not "it compiles".
The real thing has a name: the artifact. It is the actual thing you made. For us, the artifact is the live web page.
"It compiles" only tells you the code builds. It does not tell you it works. "It merged" only tells you the code is in main. It does not tell you a visitor can see it.
And the answer to a check has three values, not two. VERIFIED. NOT VERIFIED. INCONCLUSIVE. Inconclusive means "we could not tell". The skill is firm about it: "Inconclusive is not a pass. Don't hide a negative."
The principle also says the strongest proof is a script you can run again, not a one-time look.
Now the honest part about our run. Even after the fix went live, our first "live verified" was one text check. It looked for the product names on the page. It did not click the links. It did not check both addresses, or the 10 seconds, or the email. It never tested the predicate clause by clause. It happened to be right. That is lucky, not proven.
Why it matters
Most false "done" comes from a cheap check standing in for the real one. The third verdict stops "I could not check" from quietly turning into "it is fine".
NextA check gives you a verdict. But later, someone also needs to know why each choice was made. For that, you need a trail.
Leave a trail: show-me-your-work, one row per decision
What to say
While the agent works, the show-me-your-work skill keeps a decision log. It is one file, a TSV, which is a simple table stored as plain text.
Each row is one decision. What was done. Why. A pointer to the evidence. And the result.
The row on this slide is the example from the skill file itself. Decision: took screenshots of the old version before changing anything. Why: so we can compare old against new. Evidence: a script and a folder of screenshots. Result: saved 120 reference screenshots.
Notice the evidence is a pointer, not a story. A commit, a file path, a folder. Something you can open and check.
The log is append-only. If a decision was wrong, you do not edit it away. You add a new row that replaces it. It is like a ship's logbook. You never tear out a page. You write the correction underneath.
Before handing back, the skill also asks a reviewer running on a different AI model to read the trail and flag anything worth a closer look. Those flags go in an "Attention" section at the end of the reply.
The figure-it-out skill sums it up: "The trail plus the diff is what lets the human come back and trust the work."
Why it matters
You cannot watch every minute of an agent's work. A trail lets you check the decisions, without rereading the whole chat. It is the "memory" from slide 6, at the size of one run.
NextA trail remembers one run. To stop the same miss in the next run, you need to learn from it.
Learn from the miss: /reflect turns a run into skill edits
What to say
After our miss, we ran /reflect.
Here is how it works. Three reviewers read the session at the same time, each through a different lens. They are called Judgment, Tooling, and Divergent. Each one proposes lessons, and says where each lesson should go.
Then a fourth helper, the synthesizer, reads all three and sorts the proposals into three piles. Accepted. Rejected. Backlog, which means "good idea, later".
There is one more check before anything is applied. If a lesson would work better as a lint rule, a script, or an automatic check, it moves to Backlog, so it can be built as a real mechanism instead of more text. We will see why on slide 18.
And then the most important step. Nothing changes until a person approves it. The skill says: "Skill changes affect every future agent in the org. Do not auto-apply."
The skill also says that one-offs are not learnings. Only patterns are worth turning into rules.
Why it matters
This is the "memory of misses" from slide 6, turned into a process. The human gate matters because one bad lesson, once written into a skill, would spread to every future run.
NextSo what did reflect find for us?
What reflect found for us
What to say
The most useful lens for us was the Divergent one. It found three things.
First, it pointed at our own playbook. I mean our own engineering playbook here, the one we use for our projects, not one of pstack's twenty-three. Ours said "done is merged". But when "done" is a live website, merged is just a middle step. Our process had no place for "deployed and verified".
Second, it found the cheapest place to catch the problem. Our onboarding step asked for the repo and the host. It never asked how the host deploys. One quick look at the Cloudflare settings before starting would have shown there was no Git source. Cheap to ask on day one. Expensive to find out after the merge.
Third, it called out our check. "Live verified" was one lucky text check, not the predicate.
The other two reviewers agreed on the core point: before you call a merge "shipped", confirm how the host actually deploys.
And note the small print. These are proposals. They are waiting for approval. Nothing has been changed automatically.
Why it matters
Notice that none of these says "the agent should try harder". Each one points at a specific place in a written process. That is what makes it fixable.
NextFinding the lesson is not the end. The last step is the one people usually skip.
Close the loop: Encode the lesson in structure, not more text
What to say
The principle here is Encode Lessons in Structure. Its reason is one sentence: "Textual instructions are easy to miss."
A written reminder only works if the reader notices it, remembers it, and follows it. That is three chances to fail. A structure works without anyone's cooperation. The principle lists the kinds: a lint rule, a metadata flag, a runtime check, or a script.
Think of a sticky note on the door that says "remember to lock me", next to a door that locks itself when it closes. The sticky note is text. The self-locking door is structure. In software, that self-locking door is a check that runs on its own and fails loudly.
The principle also says to pick the strongest mechanism the situation allows, because agents copy whatever the surrounding code already does. And it has a line I like: the instruction is the symptom.
For us, the lesson went into the definition of done itself. Done now means live, verified clause by clause, on every hostname. Because it is in the predicate, which is what every run checks against, the next run cannot call the job finished without it.
And the same principle points one step further. The strongest version of that check is a script that tests every clause on both addresses, so nobody has to remember to run it.
Why it matters
This closes the loop we have been building. A clear problem. A written "done". A check on the real thing. A trail. A reflection. And finally a structure, so the lesson sticks.
NextThat is the whole mechanism. Now for the honest part. It is not free.
Trade-off: Rigor costs tokens
What to say
Tokens are the small pieces of text an AI model reads and writes. You pay for them. More model calls means more tokens.
pstack uses a lot of model calls on purpose. Helper agents, which are called subagents. Review panels. Parallel reviewers, like the three we saw in reflect. The setup guide says it directly: "pstack spends extra tokens on subagents and review panels. That's the price of the rigor."
The guide also gives ways to spend less. Run /setup-pstack again and choose a smaller reasoning budget, or cheaper models. Let some roles use the chat's own model. Use fewer reviewers on a panel. And the big one: save /poteto-mode for work that needs rigor. A small, obvious edit does not.
Why it matters
Rigor is a trade. It is worth paying for when a mistake would cost more than the checking. A website that looks shipped but is not is exactly the kind of mistake worth paying to catch.
NextAnd that leads straight to the question of when not to use it at all.
When not to use it
What to say
There are four times to leave pstack out.
One. A small, obvious edit. Fixing a typo does not need a review panel.
Two. When there is no checkable finish line. The guide says: "A duration is not a finish condition." If you tell an agent "work on this for four hours", it has nothing to check. You will wake up to four hours of motion instead of a result. Without a "done" that can pass or fail, there is nothing to verify against.
Three. A loop you have not earned trust in yet. A loop is an agent that keeps working on its own, maybe overnight. The guide warns that a loop that cannot check its own work "only makes unchecked work faster". It gives four things to have first. You have done the task once yourself, or watched an agent do it. The agent has the tools you would use. Every stage proves its work and can stop the line. And you have read a few transcripts and turned the repeated failures into checks. Until then, run it while you watch.
Four. You want different habits. /poteto-mode is one engineer's style. Lauren's. If yours is different, /automate-me reads your own recent chats and drafts a mode based on how you actually work.
Why it matters
A tool used everywhere gets resented and then dropped. Knowing when not to use it is part of using it well.
NextIn the same spirit, here are our own caveats.
Our honest caveat
What to say
Our run was not textbook, and you should know that.
We used the skills outside their normal setup. We never ran /setup-pstack, so the small settings file that tells pstack which models to use was missing.
pstack did not magically know that our host deployed by hand. Nothing does, until somebody asks.
And the checks only got good once we made "live", checked clause by clause on both addresses, part of the written definition of done.
What pstack gave us was not magic. It was a structure that pushed the right question to the surface. Eventually.
The figure-it-out skill has a line that sums this up: "Rigor is gates and artifacts, not 'try harder'." A gate is a check you must pass before you move on. An artifact is a thing you can open and inspect. Neither of them depends on anyone trying harder.
Why it matters
Being honest about the limits makes the rest believable. pstack is not a promise that agents will never miss. It is a way to make misses visible, and then to make them unrepeatable.
NextSo let us go back to where we started.
Back to the start
What to say
The pull request merged. The website did not change.
Before, "done" meant merged. Now "done" means a visitor on the live site sees it, in ten seconds, on both addresses. That is a meaning an agent can check by itself, with no one standing over it.
That is pstack in one idea. Lauren puts it in five words in her README: "if you want to go fast, go deep first."
Going deep means giving one agent a clear "done", a real check, and the lessons of past mistakes. Once you can trust one agent like that, you can run many of them. She calls that fearless parallelism. Remember the bottle from slide 5. Going deep is how you widen the neck.
Why it matters
This ties the whole story together. The slow step was checking. pstack makes checking something the agent can do. That is what lets you go fast.
NextLast, where all of this comes from.
Sources
What to say
Everything I said about what pstack is comes from Lauren's own sources. Her README for pstack version 0.15.15. The skill files for figure-it-out, reflect, and show-me-your-work. The principle files for Prove It Works and Encode Lessons in Structure. Her pstack guide. And her article, "How I Use Cursor".
The website story is not from Lauren. It comes from our own project notes from the malohacoast.com work.
If you want to go further, read the skills yourself. They are plain text. That is the best part.
Why it matters
Sources keep the line clear between what pstack claims and what happened to us. You can check both.