The PR merged.The website didn't change.
malohacoast.com rebuild · 5 Oct 2026
On 5 October an agent rebuilt my studio site. The pull request was reviewed and merged. Then I opened the live site, and it was the old page. Same as before.
What actually happened
An agent rebuilt the page and opened a PR. It was merged.
The live site still served the old consulting page.
The host, Cloudflare Pages, had no Git source connected. Merging deployed nothing.
Nobody lied. The agent checked the wrong thing.
The agent assumed merge to main means live. That was false here, because the Pages project was set up for manual uploads. The agent was not careless. It verified a proxy, the merge, instead of the real thing, the live site.
A first-principles guide
pstack, from the problem up
Lauren Tan's open-source skills for rigorous agent work, explained through one real project.
Use arrow keys to move. Press F for fullscreen.
This is a guide to pstack, a set of agent skills by Lauren Tan. No hype. We start with the problem, build up from fundamentals, then the mechanism, then where it does not fit. And we use one true story to hold it together.
The problem pstack solves
Agents are fast at producing work.They're unreliable at knowing when it's done.
"throughput without quality is not a goal i aspire to."Lauren Tan, pstack README
That is the core problem. Writing code is cheap now. Knowing it is correct, and actually done, is not. Lauren says it plainly in the README: throughput without quality is not the goal.
Why it compounds
"Verification is the slowest step in most agent work, because it's the step that usually waits on a human."pstack guide, "Verify the result and open a PR"
Add more agents without fixing this, and you just get more unchecked work to review.
And it gets worse as you scale. If every agent needs you to check its work, ten agents means ten times the checking. The guide makes the same point: skip verification and more agents only gets you more unchecked work.
First principles
What a good engineer brings that an agent doesn't
01
A way of working How this team debugs, designs, and ships.
02
A checkable "done" Something that can pass or fail, not a feeling.
03
Memory of misses Last month's mistake changes this month's habit.
A fresh agent chat starts with none of these. They have to be written down.
Strip it back. A strong engineer carries three things: a way of working, a clear idea of done, and the scar tissue from past mistakes. A fresh agent session has none of that. So if you want it, you write it down somewhere the agent will read.
So what is pstack?
"I've taken all the failure modes I've observed and turned them into skills."Lauren Tan, How I Use Cursor
Cursor plugin Plain markdown files MIT licensed /add-plugin pstack
pstack is exactly that, written down. Lauren took the failure modes she kept seeing in agents and turned each into a skill. It is a Cursor plugin, MIT licensed, and under the hood it is just markdown files.
Building block 1
A skill is a markdown file the agent reads
---
name: figure-it-out
description: "Design an auditable playbook when
no narrower one fits ..."
disable-model-invocation: true
---
A name, a "when to use" line, then instructions in prose. That flag means it runs only when you type it.
Our site: no skill told the agent how this host deploys.
The smallest unit is a skill: one SKILL.md file. A name, a description of when to use it, then plain instructions. Some, like figure it out, are marked so the agent never runs them on its own. You have to ask. And on our site, no skill anywhere told the agent how this particular host deploys.
Building block 2
Principles: 24 short rules with names
Prove It Works Subtract Before You Add Fix Root Causes Encode Lessons in Structure Never Block on the Human
The name is a steering handle. On our site, one line would have redirected the agent:
apply prove it works. open the live site.
The guide's own example: "apply prove it works. run the real import flow and show me the written records."
Next layer: principles. There are twenty four, each one rule with a memorable name. The trick is that the name becomes a steering handle. Instead of a paragraph of correction, you say apply prove it works, open the live site, and the agent already knows the full rule. That one line would have caught our miss.
Building block 3
Playbooks: 23 recipes for recurring tasks
bug fix feature refactoring perf prototype visual parity shipping autonomous run and more
bug fix: "reproduce a defect, root-cause it, and fix with runtime evidence."pstack README, playbook table
Our site: no stock playbook fit a redesign-and-deploy job.
Third layer: playbooks. A playbook is an ordered list of steps for a common kind of task. Bug fix, for example, means reproduce it, find the root cause, and fix it with evidence from the running app. Our job, a redesign plus a deploy to a hand-configured host, did not match any of them. Hold that thought.
The router
/poteto-mode ties it together
1 Read your request
›
2 Match a playbook
›
3 Copy its steps into a todo list
›
4 Call skills as steps fire
A skipped step stays in the list as skip: <reason>, so you can see what it chose not to do.
Our site: with a "verify live" step, skipping it would have shown up in plain sight.
You rarely call these pieces by hand. You type poteto mode and describe the goal. It picks the playbook, copies the steps into a visible todo list, and pulls in skills and principles as each step needs them. If it skips a step, it says so and why. If our list had a verify live step, skipping it would have been right there on screen.
When no playbook fits
/figure-it-out designs one first
A
Frame
Done as a falsifiable predicate
B
Design
Small units, riskiest first, checks before work
C
Loop
Hypothesis, smallest change, measure, keep or revert
D
Trail
Log every decision with evidence
E
Verify
Check the whole on the real product
"The deliverable before any code is the workflow itself." · figure-it-out SKILL.md
Our site job did not fit a stock playbook, so we used figure it out. Its first output is not code. It is the workflow: frame what done means, design the steps, run them as small experiments, keep a trail, and verify on the real product.
Back to our site · Phase A
Done, written so it can fail
On live malohacoast.com and www, within about 10 seconds:
A visitor can name the three products
Each product row has intentional media
The page no longer reads as a plain text list
Media stays small, and reduced motion gets a still
Hire link and contact email work
Zero client traces in frames, filenames, alt text, or copy
Here is the frame we wrote for the next round. Notice the first words: on the live site, on both hostnames. Not in the PR. Every clause can be checked, and every clause can fail. That is what makes it useful.
The principle doing the work
Prove it works on the real artifact
"It compiles" is not evidence. Neither is "it merged."
VERIFIED NOT VERIFIED INCONCLUSIVE
"Inconclusive is not a pass. Don't hide a negative." · figure-it-out SKILL.md
Our first "live verified" was one text check. It never tested the predicate clause by clause.
The principle behind this is prove it works. Check the real thing, not a stand-in. And a verdict has three values, not two. Inconclusive is not a pass. In our first run, even the fix was only checked by looking for product names on the page. Lucky, not proven.
Leave a trail
show-me-your-work: one row per decision
decision why evidence result
took screenshots of the old version before changing anything so we can compare old against new scripts/snapshot.sh, baseline/ saved 120 reference screenshots
Example row from the skill file. "The trail plus the diff is what lets the human come back and trust the work."
While it works, the agent keeps a decision log. One row per decision: what, why, a pointer to evidence, and the result. This example is from the skill itself. The point is that you can come back later and audit the run without rereading the whole chat.
Learn from the miss
/reflect turns a run into skill edits
3 REVIEWERS Judgment, Tooling, Divergent read the transcript
›
SYNTHESIZER Sorts into Accepted, Rejected, Backlog
›
YOU Approve which edits apply
"Skill changes affect every future agent in the org. Do not auto-apply." · reflect SKILL.md
After the miss we ran reflect. Three reviewers read the session in parallel, each with a different lens. A synthesizer sorts their proposals. And nothing changes until a human approves it, because a skill edit affects every future run.
What reflect found for us
Our playbook said "done is merged." When done is a live site, merged is a middle step.
Onboarding asked for repo and host. It never asked how the host deploys.
"Live verified" was a single lucky check, not the predicate.
Divergent reviewer findings, malohacoast.com session, 5 Oct 2026. Proposals, pending approval.
The most useful lens was the divergent one. It pointed at our own playbook: we had defined done as merged. It noticed the cheapest place to catch the problem was onboarding, by asking how the host deploys. And it called out that our check was luck.
Close the loop
Encode the lesson in structure, not more text
"Textual instructions are easy to miss."Encode Lessons in Structure, principle skill
For us: done now means live, verified clause by clause, on every hostname. It's in the predicate, so the next run can't skip it.
The last step is the one people skip. A lesson written as a reminder gets ignored. pstack says encode it: a check, a lint rule, a script. For us that meant baking live verification on both hostnames into the definition of done itself.
Trade-off
Rigor costs tokens
"pstack spends extra tokens on subagents and review panels. That's the price of the rigor."pstack guide, "Set up pstack"
Her advice: save /poteto-mode for work that needs rigor. A small, obvious edit doesn't.
Now the honest part. This is not free. Review panels and parallel reviewers mean more model calls. The guide says so directly, and its advice is to use poteto mode only where rigor is worth paying for.
When not to use it
A small, obvious edit.
No checkable finish line. "A duration is not a finish condition."
A loop you haven't earned trust in. It "only makes unchecked work faster."
You want different habits. It's one engineer's style; /automate-me drafts your own.
Quotes from the pstack guide, "Run work while you sleep" and "Recipes and pitfalls".
Skip it for trivial edits. Skip it when you cannot say what done looks like, because then there is nothing to verify against. Do not run it unattended in a loop before you trust its checks. And remember it encodes Lauren's opinions. If yours differ, automate me builds your own mode.
Our honest caveat
We ran it outside its usual setup. /setup-pstack was never run.
pstack didn't know our host deployed by hand. Nothing does until you ask.
The checks got good only once we wrote "live" into done.
"Rigor is gates and artifacts, not 'try harder'."
Our run was not textbook. We used the skills outside their normal Cursor setup. And pstack did not magically know how our host deployed. What it gave us was a structure that forced the right question, eventually. Lauren's line sums it up: rigor is gates and artifacts, not trying harder.
The PR merged. The website didn't change.
Now done means a visitor on the live site sees it, in ten seconds, on both hostnames.
"if you want to go fast, go deep first."
Lauren Tan, pstack README
So, back to where we started. The PR merged and nothing changed. Now done has a precise meaning that an agent can check by itself. That is pstack in one idea. If you want to go fast, go deep first.
Sources
pstack README, v0.15.15 · github.com/cursor/plugins/tree/main/pstack
figure-it-out, reflect, show-me-your-work SKILL.md · .../pstack/skills
Principles: prove-it-works, encode-lessons-in-structure · same folder
The pstack guide: Set up pstack, Verify the result and open a PR, Run work while you sleep, Steer with principle names, Recipes and pitfalls · .../pstack/docs/guide
Lauren Tan, "How I Use Cursor", X article · x.com/poteto/status/2058975157503570132
The malohacoast.com story comes from our own project notes (FIO2-FRAME.md, REFLECT-DIGEST.md, REFLECT-REVIEWERS.md), not from Lauren's sources. Sources read 6 Oct 2026.
Everything about what pstack is comes from Lauren's own repo, skill files, guide, and article. The site story is from our own notes.