
Building a Personal Agent Task Board From Scratch
October 4, 202610 min read2,261 words
Text: Talia Bennett
Nail ownership, state tracking, and visibility before building the board itself.
Most personal agent task boards get abandoned within a week, and the reason has nothing to do with the model behind them. The board looked promising on day one: tasks moving, updates posting, something that felt like momentum. By day three, nobody's checking it anymore, because nobody can tell what's actually happening. The agent can execute tasks fine. What's missing is ownership, state tracking, and a way to see progress, and without those three things, execution has nowhere to anchor. Layering an agent onto a process that was never clearly defined just produces a faster version of that same undefined process: the board looks active, but the user has no way to tell what's done, what's stuck, or what's waiting on a decision.
The fix starts before any code gets written. Three questions need answers first: what the agent owns, how it keeps track of state across sessions, and how the user actually sees progress. The rest of this piece works through those three questions in order: ownership, state, and visibility, followed by the tooling and infrastructure that make all three hold up in daily use.
Deciding what tasks your agent owns versus what it assists with
The first decision, and the one that makes everything else work or not, is drawing a clear line between what the agent owns outright and what it merely assists with. This line has to get drawn before a single board column gets built, because every downstream design choice depends on it.
A task the agent owns is one it carries end to end: it pulls the context it needs, takes the action, updates its own state, and closes the loop without waiting on instruction at each step. A task the agent assists with looks different. It surfaces information or drafts something, then hands the decision back to a person. Mixing these two up is where boards start to fail quietly. An agent that half-owns a task ends up half-executing it, and the board lands in an ambiguous state nobody can read: is this done, or is it waiting on something?
The useful distinction here is between a workflow and a task in the harder sense. If the decision space is bounded, with a clear set of acceptable actions and a recognizable success state, that's a workflow, and an agent should execute workflows on its own. If the decision space is genuinely open-ended, where success can't be reliably evaluated in advance, that's a task the agent should reason through alongside a person, not run solo. Autonomy is the whole point of an agent. A system that asks for confirmation before every single action isn't an agent at all; it's just a UI with extra steps. But handing an agent unlimited autonomy over tasks it has no business owning doesn't fix that problem, it creates a new one, because small errors compound fast when nobody's checking the work.
Define three boundaries for every agent before writing any configuration: goal boundaries, which spell out what outcomes the agent is actually supposed to produce; tool access boundaries, which spell out what systems it can reach and what operations it's allowed to perform on them; and output boundaries, which spell out what it's permitted to return or act on. These three boundaries are the first real design artifact for the board, not an afterthought bolted onto a config file once the agent is already running. Get them wrong, and no one can ever say for certain who owns a given task, since the answer just gets guessed at differently every time the agent runs.
Mapping the agent loop so the task board reflects how the agent works
Once ownership is settled, the next question is how the board itself should be structured, and the honest answer is: not like a human project board. To Do, In Progress, Done is a metaphor built for people coordinating with other people. It doesn't represent any of the states an agent actually moves through, so a board borrowing that metaphor hides exactly the information the user needs.
An agent runs on a loop: goal, perception, reasoning, planning, action, observation, memory update, and back to reasoning again. Each pass through that loop is where the board's columns should come from, rather than from Trello convention. Take a single task moving through it: a task arrives and sits as queued, since the goal has been received but nothing has started yet. The agent decomposes that goal and picks its tools, which is planning. It starts making tool calls, which is acting, and if those calls are still in flight, the board should show that distinctly from a task that's finished its first call and stalled. If the agent hits something it can't resolve on its own, a missing piece of information or a decision outside its boundaries, the task moves to blocked, and this is the single most important state to design deliberately, because without it, the user can't tell the difference between an agent quietly working and one that's quietly stuck. Once the agent has an output, it checks that output against the original goal, which is a verifying state the board should expose rather than bury inside the execution step. Only after that check passes does the task move to complete.
Multi-step tasks need this same granularity applied to sub-tasks, not just the task as a whole. A task sitting at "in progress" across twelve tool calls looks, from the outside, identical to one that just finished its very first call, and a board that can't distinguish real progress from a stall stops being informative within a few uses, costing the user's trust.
Designing state persistence so the agent picks up where it left off across sessions
Ownership and loop states only matter if they survive past a single session. Without deliberate state persistence, every new session starts from zero: the agent loses whatever context it built up, the board loses continuity with what came before, and the user loses trust that the system is tracking anything real.
There are three tiers of memory worth knowing, each suited to a different scale of problem. In-context memory, which just means appending to the running message list, works fine within a single session, costs nothing extra, and disappears the moment that session ends. External file memory means writing a JSON or markdown file whenever the agent learns something it needs to hold onto, then reading that file back in at the start of the next session. It's cheap, simple, and handles cross-session persistence well right up until the volume of stored memory gets unwieldy. Vector database memory, using something like ChromaDB, Pinecone, or Weaviate, becomes appropriate once there are hundreds of facts that need semantic recall spread across many sessions.
For a personal task board, file-based memory is the right starting point rather than a placeholder for something more sophisticated. A state entry doesn't need to be complicated: a task ID, its current loop state, a timestamp, the decision or output the agent produced, and any open question still waiting on the user. That file stores task state, the agent's past decisions, and anything still pending, in a format the user can open and read directly, no database client required. Vector databases solve a real problem, but it's a problem that occurs at a scale most personal boards never reach, and reaching for one before file-based memory actually becomes unmanageable adds complexity the task doesn't need yet.
The dual-memory pattern that production agent systems rely on applies here too: active working memory for whatever session is currently running, and archived state for tasks that are paused or already finished. And the state file should do double duty as the board's actual data source. The interface the user looks at should read from the exact same file the agent writes to, so the board always reflects what the agent really did, rather than a separate record that can quietly drift out of sync. That schema, the fields each state entry carries, needs to get designed before any agent code gets written, because it's the contract between the agent and the board. If it changes partway through a build, both sides break at once.
Wiring real-world tools to the board so the agent can act, not just plan
A board can get ownership, loop states, and persistence exactly right and still be useless if the agent has nothing to act on. A model reasoning on its own can plan a response, but it can't open an issue or send an email by itself. For that, it needs tools wired to the real systems where the work actually lives, and the board's usefulness is bounded entirely by which tools are connected to it.
For most personal productivity use cases, four categories of tool cover nearly everything: email, with read, send, and label access; calendar, with read, create, and update access; a task or note store with read and write access; and at least one communication channel, something like Slack. Those four alone handle the bulk of daily coordination.
Wiring each of those up individually used to mean a custom integration per service. MCP, the Model Context Protocol, changes that by standardizing the connection between an agent and its tools: instead of building one integration per service, a single MCP endpoint reaches the full catalogue of connected tools at once. By 2026, that protocol has support across Claude Desktop, Claude Code, Cursor, and a growing list of other clients.
Two separate apps, one prompt, no manual routing between them.
Tool access should follow the same boundary logic laid out for task ownership. Read-only access is the right default for any system that doesn't strictly require writing to it, and write permissions get granted tool by tool, not handed to the agent wholesale. An agent holding broad credentials can be talked into using all of them, including ones it was never meant to touch for a given task. Every write action also needs to be idempotent: a tool call that runs twice by accident shouldn't produce a duplicate ticket, a duplicate refund, a duplicate email, or an unwanted delete. That's not optional for anything acting on a user's behalf. And the tool list is exactly where scope creep sneaks in. Wiring up more tools than the agent's defined ownership actually requires just hands it capabilities with no governed reason to use them. Wire only what the owned tasks genuinely need.
Surfacing progress in a way the user will check
A board the user has to go looking for gets checked once, out of curiosity, and then ignored. Progress needs to show up where the user already is, in a form that takes no interpretation to understand.
The real divide here is push versus pull. A pull-only board, one the user has to open a dashboard to see, is competing against everything else that's also asking for their attention that day, and it loses that fight quickly. A push model, where the agent sends a summary to a channel the user already checks, doesn't have that problem, because it doesn't ask for a visit, it just arrives there on its own. A morning digest delivered in Slack or email, listing what got done overnight, what's blocked, and what needs a decision today, replaces the habit of checking a dashboard with something the agent simply delivers.
The blocked state from the loop design earlier becomes the main trigger for these pushes. The moment a task moves into blocked, the agent should notify the user right away, through whatever channel they're already watching, and that notification needs to name the specific blocking condition and the exact input needed to unblock it. A generic "needs your attention" tells the user nothing and just adds another thing to decode.
A visual board, whether a structured task tracker or even a simple list, still has a place, but it belongs in a weekly review, not daily interaction. The agent should be handling day-to-day execution. The person's role is exception handling and decisions, not watching a status screen. And what gets reported matters as much as how often: a progress summary should describe outcomes, not activity. "Drafted and sent follow-up to three leads" tells the user something useful. "Called the email tool seven times" tells them nothing. The board's display layer has one job here, translating what the agent did into language the user can act on without decoding it first.
Running the board on infrastructure that does not require you to babysit it
A well-designed board built on fragile infrastructure inherits that fragility whether the design is good or not. Uptime, state recovery, and security are not things that get patched in later at the application layer. They have to be true of the infrastructure itself.
Moving from a prototype to something that runs daily exposes failure modes that don't appear in a demo: state lost on restart, an agent looping with no budget cap, tool credentials that quietly go stale. For a personal board, the "real user" hitting these problems is the person who built it, and these failures appear fast once the board is actually running day to day rather than being demoed once and set aside.
Four infrastructure requirements decide whether the board holds up under that kind of daily use. The agent needs to run continuously to own tasks that show up outside working hours. A board only active while a laptop happens to be open isn't an autonomous system, it's a script that runs when someone remembers to open their laptop.


