
With AI coding agents like Claude Code, Codex, and Gemini CLI multiplying, you've probably had this thought at least once: "What if I just write a task in a note, and the agents build it, test it, and hand me the result?"
I actually built something like that. Markdown notes served as task requests, a file watcher read the notes and launched agents, and the results were written back as documents. Spoiler: the picture where you only write notes and everything else runs on its own didn't work as well as I hoped. Still, some of the design principles I settled on along the way carry over to other automation just fine. This post pulls out only those parts.
Humans step in only at the start and the end#
I split the whole flow into three stages.
- Plant the seed — A person writes an idea or requirement in a note. One markdown page is one task request
- Grow — Agents do the work, each in its own role
- Harvest — A person reviews the output that passed testing and marks it done on approval
The key is to pin human involvement to exactly two points: the start and the end. If people keep stepping in mid-process, the automation loses its point. And without an approval step at the end, unverified output flows straight through.
You can split the roles like this.
| Role | What it does |
|---|---|
| PM | Analyzes the note, structures it into a requirements doc, and breaks down the work |
| Dev | Writes code and unit tests from the requirements |
| QA | Runs the tests and leaves a result report |
| Release | Builds and carries it through to store submission |
It's worth adding a loop that sends failed QA results back to the dev stage. Passing on the first try is rarer than you'd think.
Narrow execution conditions with frontmatter#
If the file watcher launches an agent the moment it sees a note, you're in trouble. Half-written notes and scratch notes all get treated as tasks. So I made it run only when every frontmatter condition is met.
---
type: task # Is this a task document?
status: APPROVED # Has a human approved it?
dispatch: true # Is it OK to run now?
---Only when all three are satisfied does the note get handed to an agent. status: APPROVED acts as the human approval gate, and dispatch: true lets you express "approved, but don't run it yet." Separating approval from execution timing turns out to be more useful than you'd expect.
Don't run the same document twice#
A file watcher fires an event on every save. Tweak a note slightly and the same task can run again. That's why you need duplicate-run prevention: record execution history in a DB and skip tasks that already ran.
I kept tables like these in SQLite.
- tasks — Tasks created from notes and their current status
- agent_runs — Which agent ran when, and how it went
- events — A log of all activity: file changes, task creation, agent start/finish, QA passes
A table like events, which records every activity in chronological order, makes it much easier to trace back where things went wrong when something runs oddly.
Turn failures into documents too#
When an agent run fails, don't just leave a log and move on. I had it automatically create a QA document in a folder like docs/qa/. The failure becomes another task document.
The same way, you can add a branch that receives webhooks from an error monitoring service like Sentry and auto-creates QA or hotfix documents. Then a human-written request, a failure, and a production error all come in as documents in the same format. Whatever processes them doesn't need to care where the input came from.
Keep the core in a CLI, and the UI as a read-only viewer#
You'll eventually want a dashboard for a kanban board or run logs. Here's how I split it.
- The CLI is the core — File watching, agent execution, and DB writes all live in the CLI
- The dashboard is read-only — It reads the SQLite file the CLI created and only displays it
- Runs standalone — The CLI works completely on its own, without the dashboard
With this split, automation keeps running even if the dashboard breaks. UI code never touches the core logic. And since both share the same DB file, you don't need a separate API server.
I kept the commands simple, matching the flow.
tool seed <note-path> # Register a note as a task
tool grow <task-id> # Run an agent
tool status # Overall task status
tool harvest <task-id> # Approve the result
tool harvest --reject <id> # Reject and send back to the run stage
tool watch <vault-path> # Watch the vault and auto-registerWhy full automation didn't work#
The design principles held up, but the "just write a note and you're done" picture never materialized. There were three main reasons.
- Output quality — The context in a single note wasn't enough for agents to produce what I wanted. I kept having to fix the output myself, which broke the premise of "humans only step in at the end"
- Management overhead — Flipping frontmatter statuses and cleaning up auto-generated QA documents actually added work. A structure built to reduce work ended up creating new work
- Pipeline maintenance — Maintaining the automation tooling itself, like the file watcher, the DB, and the dashboard, cost more than I expected. At times I spent longer fixing the pipeline than building the product it was meant to serve
So these principles fit better when attached to small automation that assumes a human checks in along the way, rather than to "automate everything."
Summary#
Here's what to get right when automating agent work with markdown notes.
- Fix human involvement to two points: start and end
- Narrow execution with frontmatter and separate approval from execution timing
- Record execution history so the same task never runs twice
- Send failures and production errors back as documents in the same format
- Put logic in the CLI and keep the UI as a read-only viewer
The fully automatic pipeline didn't pan out the way I expected, but these principles apply just as well to small-scale automation. If you're automating anything around notes, start by checking your execution conditions and duplicate-run prevention.
Don't ruin the present with the ruined past.
— Ellen Gilchrist


