Documentation

How Polynomic works

Polynomic is a local-first, keyboard-first task manager for the agentic coding age. You manage Tasks; coding agents do the work behind them. A Task spans your repos, references the exact code it touches, runs an Agent Session to implement itself, and produces a reviewable diff before anything merges. This page explains the model and the loop end to end.

Core concepts

Twelve terms carry the whole model. Learn these and the rest of the app reads naturally.

Task
The unit of work, and the thing you actually manage. A complete development cycle — capture, plan, start, implement, review, merge — oriented around the work itself, not the agent that does it.
Ledger
A local, append-only SQLite record. Every status change, comment, edit, and code reference is an immutable Event; current state is a projection folded from the Ledger and rebuildable at any time.
Circuit
What a Task is worked by. Written as a tree of Steps you edit as an outline, or as a graph of Steps you wire on a canvas. It can be a single agent run, or retry, branch, and run steps in parallel — “implement, have two models review the diff, and fix what they find before they look again.” You pick one before you Start.
Step
One node in a Circuit — prepare a workspace, run a coding agent, review the diff, ask a question, merge, clean up. In a tree, a Step succeeds or fails with a reason, and the Control Step above it decides what that means. In a graph, the wire from the output it finished on decides what happens next.
Run
One execution of a Circuit for one Task. A Run is a fold of Ledger Events, not something held in memory — so it survives quitting the app, resumes where it left off, and renders on another machine that has never seen the Circuit.
Trigger
Something other than you pressing Start that starts a Circuit: a schedule, an Action finishing, Run now, or the chat. A Trigger makes a Task from a template and starts it, so what it starts is an ordinary Task with an ordinary Run.
Agent Session
How the work actually gets done: a real coding-agent process spawned in an embedded terminal, scoped to the Task over a loopback MCP server. It is one Step in a Circuit — the agent is an implementation detail behind the Task.
Harness
The coding-agent CLI a session runs — Claude Code or Codex. Not a swappable binary but a set of capabilities: Polynomic branches on what a harness can do, and says out loud when one can't (a Codex task won't quietly start without plan mode).
Initiative
A goal bigger than one Task, with Success Criteria that say when it is reached. Its Tasks need not share a parent; their work collects on one branch that lands once, when you approve its Goal Claim and merge it. Hand it to a Steward and a driver starts its Tasks while an agent, woken only when there is something to decide, plans and steers.
Workspace
When you Start a Task, Polynomic lays down one git worktree + branch per touched Repo side by side, so an agent gets full cross-repo context in a single isolated place.
Check
One named command a repo declares — build, test, lint, typecheck. Your own commands, run locally in the Task's Workspace by a Circuit, by you, or by the agent. Every run is recorded with its exit code and the commit it ran against; a red check ranks a review and tells a reviewer, and blocks nothing.
Code Reference
A pointer into source at file granularity, entered with @-syntax. References resolve live and are marked stale when their target moves, so a Task always knows the code it concerns.
Review
Every Task produces a reviewable per-repo branch diff. An independent AI Review session reads it and posts findings on the lines they concern — it never changes state or merges. You decide what lands.

Getting started

Polynomic is a desktop app for macOS, Windows and Linux. It drives a coding-agent CLI — Claude Code or Codex — for Agent Sessions, so you'll want at least one of them installed. Settings › Coding agent can install one and sign you in, and Polynomic disables Start with a banner if it can't find any.

  1. 1

    Install Polynomic

    macOS: a universal .dmg for Apple Silicon and Intel — drag Polynomic into Applications and open it. Windows: an installer. Linux: an AppImage. See the download page for first-launch notes on unsigned builds.

  2. 2

    Set up a coding agent

    Settings › Coding agent lists every harness, whether its CLI is on your PATH, and whether you're signed in. Install and Log in… run in your own terminal — an installer that asks for a password and a login that opens a browser both need one — and Re-check re-probes rather than repeating itself.

  3. 3

    Create a Project

    A Project points at one or more existing local git repos — Polynomic validates them and detects the default branch, but never clones. Give it a short key like POL; Tasks become POL-1, POL-2, and so on.

  4. 4

    Capture Tasks

    Press c to create a Task, ⇧C to make one from a Template, or open the command palette with ⌘K. Seed sample data any time with ⌘K → “Seed demo tasks” to explore the views.

  5. 5

    Start a Task

    Open a Task and press x to Start it. That runs the Task's Circuit — by default, the same thing Start always did: build a Workspace, spawn an Agent Session, and stream its terminal so you can watch, or console in.

  6. 6

    Review and merge

    When the session finishes you get a per-repo branch diff and an AI Review. Request changes, run an AI Fix, or merge — all from the keyboard. ⌘⇧R opens the queue of every diff waiting on you.

Tasks & states

A Task is an event ledger, not a row you mutate. Every change — status, comments, edits, code refs — is an immutable Event, and the state you see is a projection folded from those events. A Task moves through a small, fixed set of states:

Backlog Captured, but not ready to be worked yet.
Todo Scoped and ready to start.
Planning Started in plan mode: the agent is researching and proposing a plan, no code yet.
In Progress An Agent Session is actively working it in a Workspace.
In Review Work is done and a branch diff is ready to review.
Changes Requested You requested changes; the session resumes to address them.
Needs Attention Automation failed or got stuck, and needs you to sort it out.
Done Merged into the base branch. Verified against real commits.
Canceled Abandoned. Terminal, like Done — it resolves anything waiting on it.

Done and Canceled are terminal, and that makes them a guard as much as a label: moving a Task there while work is still in flight asks whether to stop it, and nothing automatic ever moves a Task back out of one.

View Tasks as a List or a Board (kanban, drag-to-change-state), and organize them with initiatives, labels, priority, and assignee — all filterable from the sidebar. The signature Code-Tree view rolls up Task counts per file and directory, so you can filter the list by the part of the system the work touches. Tick a set of rows with Space to change status, priority or labels on all of them at once.

Sub-tasks are Tasks

A Sub-task is a full Task with a parent pointer — its own state, its own Events, its own Agent Session — and it is Started on its own like anything else. The parent is organizational: it doesn't run its children and its state isn't a rollup of theirs. What it does do is wait for them, which is the one edge in the dependency graph you never have to draw.

Sub-tasks expand inline under their parent at any depth, appear as their own cards on the board, and carry the same right-click menu and Start controls as a top-level row. Re-file one under a different parent from the Parent field in its properties rail.

Dependencies

A Dependency is a directed edge between two Tasks: A depends on B means B's work has to land before A's can. Add one from a Task's Dependencies section; an edge that would close a cycle is refused. A parent also depends on each of its Sub-tasks, but that edge is derived from where the Task is filed rather than stored — views draw explicit edges solid and implicit ones dashed.

A Task with an unlanded dependency reads as Blocked, and one with none left as Ready. Both are computed on read — they are facts about other Tasks' states, so neither is a Task State and neither is stored. Terminal resolves a dependency, so a Canceled Task doesn't park a chain behind it forever.

  • Blocked informs, it never refuses. Start is never denied over a dependency. Every Start and Plan control raises a warning naming what hasn't landed, and starts the Task anyway on Enter — the decision is yours, only the noticing was missing.

  • The dependency graph. ⌘⇧G draws the order a set of Tasks has to happen in, laid out in columns of work order: left to right is the order to implement in, and the leftmost column is what can be started now. Unrelated chains are laid out and stacked separately rather than crammed into one graph, and a cycle is drawn in red with the loop named.

  • Agents read the graph both ways. A Task's opening prompt names what it waits on, so the session reads that work instead of reinventing it, and what waits on it, so it can say what you should tackle next. Ordering the work it just broke down is one of the few authoring powers a coding session has.

Initiatives & the Steward

An Initiative is a goal bigger than one Task. It has a Goal, written in prose, and Success Criteria that say how anyone can tell it got there. Tasks in an Initiative don't have to share a parent, so unrelated work can serve one goal. A Task's Sub-tasks come with it. Open them with ⌘⇧I or Open Initiatives in the command palette.

An Initiative has an end. It moves from Draft to Active, can be Paused, becomes Claimed when someone says the Goal is reached, and finishes as Reached or Abandoned. By default its Tasks merge into one Initiative Branch, and that branch lands on your base branch once, at the end, when you accept that the Goal is reached. Set it to land as it goes and each Task merges to the base on its own instead.

Criteria come in two kinds. An objective criterion is a command or a Check. Polynomic runs it against the Initiative's work when you press Run criteria. A judged criterion is a sentence, like “a first-time user can sync offline edits without reading docs”. An agent or a person judges it, with evidence, and the judgement is shown as a claim, not a fact.

Making one

Press New Initiative on the Initiatives page or in the command palette. It takes four steps: the Goal (a name, flagged if another Initiative already has it, and the prose), Success criteria, How it runs (where the work lands, and whether you work it by hand or a Steward does, with its budget, width and Check-in), and a Review that reads it all back. ⌘↵ goes on and Esc goes back.

Nothing you write in it is lost. Leave the wizard or quit the app and the draft waits under Drafts on the Initiatives page. Only Discard throws it away. If creating stops partway, the Initiative is kept as a Draft with what worked, and Finish creating it does the rest.

  • Shape with AI. On the Goal step it is already open: write a line about what you are after and press Draft it. An agent session on your harness reads the project's code (it changes nothing) and the board, and the Shape box shows what it reads and says as it works, with a Stop button. It drafts a Goal, Success Criteria and Tasks in order, and a name if you haven't given one. Keep the parts you want, or Try again with a note. Accepting fills in the wizard's own fields, which you can still edit, and nothing is saved until you press Create. The same box is on every Initiative's page. Shaping runs in the desktop app. A draft keeps shaping while you do something else, and its result waits in the draft.

  • Putting Tasks in. Add task, at the foot of an Initiative's Members, finds an open Task to bring in or files a new one under the title you typed. It stays open so you can add several. A Task already in another Initiative moves over. You can also set the Initiative field on any Task.

  • In the task list. An Initiative's Tasks are gathered under a header that names it, with its state and how many Tasks it has, at the place the first of them sorts to. Click the header to open the Initiative.

The page explains itself as you go. A Draft opens with numbered next steps: write a Goal, add criteria, then either ask for a plan and approve it, or put Tasks in and make it Active yourself. The next step is marked, with its control beside it.

Running the criteria

Run criteria runs every objective criterion against the Initiative Branch, or against the base if it lands as it goes. You can watch it on the page. Each row shows it is queued, then running with its elapsed time and last line of output (Output shows the rest), then whether it passed. Stop ends the running criterion and skips the rest. Results that finished before you pressed it are kept. Leaving the page doesn't stop a run. To run just one criterion, use the Run button on its row. The others keep their last results. A criterion you change starts untested.

The Initiative Branch

  • Refresh from base. Rebases the branch onto your base branch, so it reads as the base plus the Initiative's own work. Open Tasks cut from the old tip move onto the new one, keeping only their own commits. A Task with uncommitted work, or whose commits conflict, is left where it is and named, and it is rebased when it merges. Across several repos it is all or nothing: a conflict in one puts every repo back.

  • Open in terminal, Open in editor. On the branch panel and the Initiative page's sidebar. They open the branch's own worktree, so you can read or try the combined work. Open in editor opens it in the Editor with the Git panel. If this machine has no worktree for the branch yet, you're offered Make its worktree.

  • Merge into your base. The button names your base branch, and its question says how many commits come with it. It is refused while any Task in the Initiative is still open. Merging from the Goal Claim also makes the Initiative Reached.

By hand, or by a Steward

An Initiative works fine with nobody driving it: you start its Tasks yourself, like any other. Hand it to a Steward and it is worked toward its Goal for you. The Steward has two halves.

  • The driver, which is not AI. Plain code that runs on one of your desktops and costs nothing. It starts Tasks as they become Ready, one at a time by default, on the Unattended circuit unless the Task has its own. Each Task gets a coding agent and your checks, and its work merges into the Initiative Branch. The driver tracks spend against the budget. If a Task stops and can't finish, it waits for what is running to finish and then stops too. It looks again at least once a minute.

  • The Steward agent, which is woken. A short, fresh agent session, started only when there is something to decide. It reads the Goal, the criteria, where every Task stands and its own Log, decides, writes down what it saw and decided, and exits. It is never a conversation left running, so what it costs follows how many decisions there are, not how long the Initiative takes.

How a Stewarded Initiative runs

  1. 1. Write the Goal and criteria yourself, or use Shape with AI to draft them. Choose a Steward under How it runs in the wizard, or set Steward to Steward in the Initiative's properties. Nothing runs yet.
  2. 2. Press Create and ask for a plan at the end of the wizard, or Ask the Steward for a plan on the Initiative, on a desktop. Create as Draft skips this step for now. That desktop becomes the one that drives the Initiative. The Steward's first Wake reads the code and the board and proposes an Initiative Plan: the Tasks it would file, their order, and how each criterion will be tested. It files nothing.
  3. 3. Approve and start, or Send back… with what to change, and the next Wake proposes a new version. You can do this from a browser or your phone too.
  4. 4. Once you approve, the Initiative is Active. The Steward files the planned Tasks, and the driver starts each one when it is Ready and a slot is free. When a Task gets stuck, the Steward is woken to decide what to do: rewrite it, break it down, file what's missing, or ask you. When a Task's checks fail, the Steward reads what failed and can send the Task back to its own session to fix it; the checks run again and it merges once they pass, up to three rounds before it's left to you. By default each Task merges as soon as its checks pass. To review each one before it merges, set Review to As each lands in the Initiative's properties: a Task whose checks pass then waits under Waiting on you until you approve its diff, and requesting changes stops the Steward until you carry on.
  5. 5. When every Task has landed, the Steward judges the judged criteria and makes a Goal Claim, with evidence for each criterion. Objective criteria have to be passing as last run, so press Run criteria if they haven't been run.
  6. 6. Answering the claim takes two steps. Approve it to say the Goal is reached. That merges nothing, and you can do it from a browser. Then Merge into your base branch from the desktop. That lands the Initiative Branch and makes the Initiative Reached, once every Task in it is finished. If the Tasks landed as they went, there is nothing to merge, and approving makes it Reached. Reject the claim instead, say what's missing, and it goes back to Active and the Steward carries on toward that.

Every time AI runs

These six are the only places an Initiative calls a model. Everything else is ordinary code: starting Tasks, Check-in summaries, the budget, objective criteria, landing and abandoning.

  • Coding a Task. When the driver starts a Task, a coding agent works it on your configured harness. This is where nearly all of the cost is, and it counts against the Initiative's budget.

  • A Steward Wake. A short headless session on your harness, started when one of the events below needs a decision. It runs on the driving desktop, inside the Initiative Branch so it can read the actual work. Its cost is written in its Log entry and counts against the budget.

  • A Check-in, when you have Jev. With a TypeSafe key on the driving desktop, each Check-in asks Jev one question: does this need the Steward? It is a single typed judgement over where each Task stands, its progress report and how long since it moved. Jev is not shown the Success Criteria and never judges them.

  • Before the Steward files a Task, when you have Jev. Jev compares the new Task with up to 24 open Tasks outside the Initiative. If it judges an overlap 70% likely or more, nothing is filed. You get a Proposal naming the other Task instead, so the Steward doesn't step on other work.

  • Shape with AI. Only when you press Draft it, on an Initiative's page or in the New Initiative wizard. It runs a headless agent session on your harness, on the desktop. It reads your code without changing it, along with the board, and drafts a Goal, criteria and Tasks from your words. You can watch it and stop it. It doesn't count against the budget. Nothing is written until you accept, and then only the parts you kept, as you.

  • The chat panel. Only when you ask. The chat can create an Initiative, write its Goal and criteria, set how it lands, pause it and comment in its thread. It can't hand an Initiative to a Steward, approve a plan, answer a Proposal, approve a Goal Claim or abandon an Initiative. It tells you where to do those.

What wakes the Steward

The Steward is woken by something that needs deciding, and only that. Each of these wakes it once, and the Wake is told which happened:

  • you comment in the Initiative's thread, which is the way to steer it
  • you press Wake the Steward at the foot of the Initiative page, to have it look over where things stand now
  • you edit the Goal or the Success Criteria
  • you approve the plan, or send it back
  • you accept or reject one of its Proposals
  • you reject its Goal Claim
  • a Task stopped and the driver stopped starting new ones
  • every Task has landed
  • you asked for a plan and there isn't one yet
  • a criterion failed in a way that says its command is wrong, not the work: a program or dependency that isn't installed, a script that doesn't exist, a Check no repo declares. The Steward proposes the corrected criterion, usually with the missing install step in front
  • a Check-in that needs it (see below)

A Task finishing normally doesn't wake it; the driver just starts the next one. Nothing the Steward does itself wakes it either. It is never woken while the Initiative is Paused or Claimed, after it has ended, or once the budget is spent. Only one Wake runs at a time. While one runs, the Initiative page on the driving desktop shows it under The Steward is awake, live: what it says and each tool it calls, with a Stop button. When it ends, its session stays folded under its entry in the Log until the next Wake. Other devices say which desktop it is running on. Wake the Steward is disabled, with the reason, wherever no Wake could follow. After you press it, the page says where and when the Wake will happen. The events are read from the shared record, so a comment you leave from your phone wakes the Steward on the desktop the next time it looks, and a restart doesn't lose any.

Check-ins and Jev

A Task can run for hours without failing and still be going nowhere, so the driver also checks in on a schedule, whether or not anything has happened. A Check-in writes one line to the Steward's Log: how many Tasks have landed, which are working, stuck, ready or waiting, and what has been spent.

  • With Jev. The Auto schedule checks in every 20 minutes. Jev is asked whether the Check-in needs the Steward, and only a judgement of 50% or more wakes it. Below that, the Check-in is recorded as quiet and costs one Jev call. A run of quiet Check-ins means the Initiative is working.

  • Without Jev. The Auto schedule checks in hourly, and every Check-in wakes the Steward. The same happens if Jev can't be reached: when the gate can't answer, the Steward is woken.

  • Choosing a schedule. Pick a fixed cadence under Check-in instead of Auto. The shortest is five minutes. A desktop that was asleep checks in once when it wakes. It doesn't catch up on the ones it missed.

What the Steward may and may not do

The Steward can read the whole project. It can change only Tasks in its own Initiative, and only freely the Tasks it filed itself: it can edit, order, break down and cancel those. It can also choose which model each of those Tasks runs on. For a Task you filed, it can comment and add dependencies. Rewriting or canceling your Task, or bringing any existing Task into the Initiative, is a Proposal that you accept or reject on the Initiative page. So is any change to the Success Criteria: when one is broken, for example a command with a typo or a Check no repo declares, or doesn't make sense, the Steward proposes the fix and the criteria stay as they are until you accept it. A Proposal never holds up the other Tasks. Before you approve a plan, the Steward changes nothing on the board.

It never starts, stops or lands work, and it never edits code. Starting Tasks is the driver's job, and landing is yours. The Steward can say which Task the driver starts next, though: when it makes most sense to do one first, it puts that Task at the front, and the Members list marks it starts first. That only picks which ready Task takes a free slot, so a Task still waiting on a dependency keeps waiting. It can't cancel a Task that is being worked. It can have at most five open Tasks it filed at once, which you can change. If it changes the plan later, for example after you edit the Goal, the Initiative goes back to Draft until you approve the new version.

Budget and limits

  • Budget. An estimated cost covering every coding session on the Initiative's Tasks and every Steward Wake. At 80% you get a notification. At 100% nothing new starts, Tasks already running finish, and the Initiative waits until you raise the budget, pause it or abandon it. With no budget set, the page says so.

  • Width. How many Tasks run at once. The default is one, because Tasks merging into one branch in parallel tend to conflict.

  • When a Task stops. Once nothing else is running, the driver stops starting new Tasks and the page says which Task stopped. The Steward is woken to look at it. Fix or cancel that Task, then press Carry on. The page then says what happened: it's carrying on, it stopped again and why, or it will carry on at the driving desktop's next look.

One desktop drives

One desktop drives a Stewarded Initiative, normally the one where you asked for its plan. That desktop starts its Tasks, runs its Check-ins and Wakes, and estimates its spend, because it has the session transcripts. Every other device shows the Initiative and can make every decision on it. Runs on in the Initiative's properties names the driving desktop. If that desktop is asleep, the Initiative waits; press Take over on another desktop to drive it from there.

Pause stops the driver and the Steward, and the Tasks can still be worked by hand. Abandon asks once whether to land what's on the branch or discard it. A discarded branch is kept, so nothing is lost. Either way, open Tasks the Steward filed are canceled, and yours go back to the board as ordinary Tasks.

Circuits

Starting a Task used to mean one fixed sequence: cut a worktree, run an agent, review, merge. A Circuit makes that sequence yours. It is made of typed Steps, and you choose which one a Task runs before you Start it. A Circuit can be a single agent run — exactly what Start always did — or it can retry, branch, and run steps in parallel.

A Circuit comes in two shapes. A Tree Circuit reads top to bottom and is edited as an outline from the keyboard; it is what New circuit makes. A Graph Circuit is Steps wired together on a canvas. Both run, side by side, and the built-ins ship in both shapes.

The thing you can't get from a straight line is the interesting one: have two different models review the diff at the same time, and if either objects, fix what they found and have them look again — up to three rounds. As a tree, that is six lines.

Tree Circuits

Behaviour trees inspired the shape. Control Steps on the inside — In order, Together, Retry ×3 — decide when their children run and what their results mean. The ordinary Steps at the leaves do the work. This is the built-in Implement, evaluate, fix (tree):

Implement, evaluate, fix (tree) Built in
In order
  Prepare workspace
  Coding agent
  Retry ×4  Review until both pass
    Together  Review
      Review (correctness)            Opus
      Review (tests and edge cases)   Sonnet
    Before retrying
      Fix what review found           told {{failure}}: what each reviewer found
  In review
  If / Else  Approved?
    if    Approve the diff            Ask me: “Merge this work?”
    then  Merge · Done · Clean up workspace
    else  Changes requested · Parked (Stop)

There is no “Workspace failed → Needs Attention” wiring anywhere in it. A failure that reaches the top of a tree fails the Run, and the Task goes to Needs Attention with the reason of the Step that failed.

Control Steps

They are named as plain instructions rather than behaviour-tree jargon. The three that run children at once differ only in when they finish.

In order Run the steps one after another. Fails at the first that fails.
First that works Try the steps in turn until one succeeds.
Together Run every step at once and wait for all of them. Fails if any did, with every failing step's reason — so a fix pass hears from both reviewers.
Race Run every step at once; the first to finish decides, and the rest are stopped.
At least N of how many must succeed · 1 Run every step at once; succeed once N of them have, or fail the moment N can no longer be reached.
Retry ×N attempts in all · 3 Try the step until it succeeds, at most N times. An optional Before retrying runs between attempts; if it fails, the Retry fails at once.
Repeat ×N times · 2 Run the step N times.
Repeat until fails at most · 10 Run the step again and again until it fails.
Timeout seconds · 600 Fail the step if it takes longer than this. The step inside is stopped.
Budget dollars · $5 Fail the step once the agents under it have spent this much, estimated from their token usage. The step inside is stopped.
Invert Succeed when the step fails, and fail when it succeeds.
Always succeed Run the step and succeed whatever it did.
Always fail Run the step and fail whatever it did.
If / Else A condition, a step for when it holds, and optionally one for when it doesn't.
Use circuit Run another tree circuit here, as if it were written in place. Its tree is copied in when the run starts.

Success, failure, On failure and Stop

  • Every Step succeeds or fails, with a reason. A leaf's result is still recorded — “no changes”, “blocked” — but in a tree it is a failure's code, not a wire. Every leaf can say its own failure in its own words (“When it fails, say”); that sentence is what the Task shows if the failure reaches the top.

  • On failure cleans up, and does not rescue. Any Step, leaf or Control Step, can carry an On failure: steps that run when it fails, such as removing what it half-made. Once they have run the failure carries on upward unchanged; turning a failure into a success is First that works' job. A cleanup that fails adds its reason to the original.

  • Stopped by the tree counts as failing. Losing a Race, an At least N of decided without it, a Timeout or a Budget running out, or a Stop firing elsewhere ends a running Step as interrupted, and its On failure runs. Cancel or Stop Run is different: a person said stop, so no cleanup runs.

  • Stop ends the Run from anywhere. Success otherwise flows upward into the steps that follow. Stop is the early exit — “the diff was rejected, park it, don't mark it Done” — and ends the whole Run with the outcome it names, after anything it interrupted has cleaned up.

The steps that do the work

The same Steps sit at a tree's leaves and in a graph. Note what's on this list: creating the workspace and merging are ordinary steps, not machinery bolted on around the pipeline, and a circuit can ask an agent things that aren't code, make tasks and start them. Under each name are the results it can finish on; in a tree, the first is success and the rest are failures.

Prepare workspace ready · failed Cut a git worktree and branch for each of the project's repos.
Coding agent made changes · no changes · failed Run a real coding-agent session in a terminal you can watch and type into. Optionally in plan mode, on a named harness and model, or continuing the task's existing conversation.
Review with a model passed · failed review · failed A fresh, headless reviewer reads the diff and reports a verdict. Give two of them different briefs — or different harnesses — and run them at the same time.
Run checks passed · failed Run the project's own declared checks — build, test, lint, typecheck — in the task's workspace, and record each result. Passes only if everything it ran exited 0. A project that declares nothing passes straight through, so the step is safe in a shared circuit.
Run actions passed · failed · rejected Run the project's declared Actions — deploy, release, build. A gated one asks you first.
Work the sub-tasks all landed · stopped Start each sub-task in dependency order, unattended, and wait. Their work merges into this task.
Ask an agent answered · failed Ask a question, research (optionally on the web), classify into choices you list, or produce data. No code changes and no workspace needed. Later steps quote the answer, and a condition can branch on the choice.
Create a task created · failed Make a new task, plain or from one of your templates with its parameters filled in, optionally as a sub-task. An agent can fill in whatever you leave blank — your own text always wins.
Propose tasks proposed · failed An agent proposes a parent task with sub-tasks and the order between them. Nothing is filed until you keep it.
File the proposal filed · failed File what you kept of a proposal: the parent, its sub-tasks and their order.
Start a task started · failed Start another task on a circuit, exactly as pressing Start would — usually the one a step just made.
Condition yes · no Branch on what earlier steps reported — a reviewer's verdict, how many checks came back red, or how many times a step has been attempted.
Set status next Move the task to a lifecycle state: In Review, Changes Requested, Needs Attention, Done.
Ask me approved · rejected Pause until you approve or reject. Shows up on the task, and survives quitting the app. Pointed at a Propose tasks step, it reviews the proposal: keep, edit or discard each item.
Merge merged · blocked Fetch, rebase, and merge the task's branches into their base — locally, as always.
Clean up next Remove the worktrees and workspace, optionally deleting the branch.
Notify next Send a desktop notification. The seam for Slack and email later.
Stop — End the run here, as a success or a failure. A graph calls it End; in a tree it ends the whole run from wherever it sits.

Passing results down the tree

A prompt can quote what an earlier Step came to. Write a Step Reference in an agent's or reviewer's prompt, a question, a notification, a status reason, or a new task's title and description, and Polynomic fills it in just before the step runs:

{{node.<id>.reason}} Why that step failed.
{{node.<id>.port}} The result it finished on.
{{node.<id>.output.<key>}} Something it reported: a review's summary, an agent's answer, a new task's POL-42. A dotted key reads deeper — output.data.score.
{{failure}} Inside Before retrying and On failure: every step that failed to bring the run there, with its reason and summary. Under a Together, that is each reviewer that objected.

Values come from the step's latest attempt. A step that hasn't run reads as a short marker — (no result from “Review” yet) — never as nothing and never as an error. Typing {{ in the outliner offers the steps that are certain to have run first, and the run view shows what an agent was actually told under What it was told.

The designer

Press ⌘⇧C and then d — or Open the Circuit designer in the command palette — for a full-window editor: your circuits on the left, the circuit in the middle, the selected step's settings on the right. A Tree Circuit is an outliner you drive from the keyboard:

↵ Add a step — inside an open Control Step, else after this one
/ Change its kind; type to filter (“ret” finds Retry ×N)
⇥ / ⇧⇥ Indent into the step above / outdent
⌥↑ / ⌥↓ Move it among its siblings
W Wrap it in a Control Step
F / B Its On failure / a Retry's Before retrying
Space Edit its settings
I Insights: how each step has done

The editor checks your work as you go, using the same validator the runner uses — so if it says the circuit is fine, it will run. It flags what would actually strand a run: a Control Step with nothing in it, a Retry that tries zero times, a Repeat with no maximum, a Use circuit that leads back to itself, a reference to a step that can't have run yet. The toolbar says how many there are to fix.

Conditions are a form, not a code editor. You pick a step, a field it reported, and a comparison — or a limit on how many times it has been attempted. Nothing is evaluated as script, so a circuit is just data: it diffs, it exports, and it can't do anything you didn't write.

You can also just describe what you want. The Circuit Assistant (Design with AI) builds and changes a circuit from a sentence — “review the diff with two models, and give the fix pass three tries”. It edits the unsaved draft, tree or graph, rather than the stored definition, so Save stays a human action and a step you placed is never reverted underneath you. It has exactly two tools, touches no Task and no Ledger, and is grounded in no repo.

Simulate before a Task runs it

Simulate in the designer's toolbar plays the circuit, tree or graph, without running anything. The run is folded by the same engine a real one is, so what it shows is what would happen; only the work is imagined. Each step that would hand off to an agent, a reviewer, a merge or you waits for you to say how it went, and Step takes the likeliest outcome for you. A Timeout or Budget runs out only when you press Run out of time or Spend runs out.

Beside it, So far keeps what the run would have done — the Task's state, the workspace, the branch, the agents, the notifications — and a timeline you can click to rewind. ⌘Z undoes a choice, Cancel run cancels as a person would, and nothing is saved.

Watching a run

A Task running a Circuit shows its progress at the top of its detail view — a strip you can expand. A tree run is shown as the tree that was written, not as plumbing: it opens on the path to whatever is running or failed, steps that run at once sit side by side, and a Retry says which attempt it is on. Below, one reviewer objected on the first pass, the fix ran, and both are looking again. When a step is waiting on you, its question appears right above the steps:

Implement, evaluate, fix (tree) Running Stop
✓ Prepare workspace ready
✓ Coding agent made changes
◍ Retry ×4 · Review until both pass attempt 2 of 4
◍ Together · Review running
✓ Review (correctness) passed
◍ Review (tests and edge cases) running
↺ 1 earlier attempt failed review
✓ Before retrying · Fix what review found made changes
○ In review
○ If / Else · Approved?
  • Retry from here. A tree run that failed can be reopened at the step that failed instead of run again from the top. Everything that succeeded outside it stands: if the merge was blocked after two reviews passed, retrying the merge doesn't review again. The button sits on the failing row (or press r on it), in the Simulator, and in the Circuits home; a Timeout or Budget above it starts afresh. Run again is still there to start over.

  • Interrupted, and why. A step stopped by the tree reads interrupted, a step a Retry ran again folds its earlier attempts under it, and the failure's reason is shown on the row where it started — the same sentence the Task shows in Needs Attention.

  • Live budgets. A Budget shows what its subtree has spent against its limit — $1.20 of $5 — estimated from token usage on the machine driving the run, and read about every 20 seconds, so it can overrun by what is spent in between.

  • In the browser and on your phone. The web app and the Companion show the same tree, with the same words for each step.

Insights

A tree's steps keep their identity across edits, so Polynomic can say how each one has actually done. Press I in the outliner and every row gets its record — 92% ✓ · 14m · ×1.3: how often it worked, its median time, and how many rounds a Retry took. Select a row for the rest: what it came to across Runs, median and p90 time, how often its On failure ran, how often it was where the Run failed, its failures grouped by cause with the Tasks they happened in, and spend where it can be priced.

Each circuit also gets a headline — 11 Runs · 82% succeeded, and the step where it fails most. It is a reading of this device's Ledger: nothing is stored, nothing is sent anywhere, and a Run from an earlier version of the circuit counts only for the steps that still exist.

Graph Circuits

A Graph Circuit is the same Steps wired on a canvas: each finishes on a named output, and the wire from that output decides what happens next. Drag from a step's output dot to another step to wire them; select a wire and press ⌫ to remove it. Graphs keep working, and a new one comes from importing a file or saving a copy of a graph built-in.

Every Project still starts with the graph Standard, which is the old fixed pipeline drawn as a graph — plus the one branch that pipeline could not express: a merge blocked by a conflict offers an AI rebase and loops back to try again. If you never open the designer, nothing about Start changes — this is what runs:

Standard Built in
Start
  │
  ▼
Prepare workspace ──failed──────▶ Needs Attention
  │ ready
  ▼
Coding agent ──no changes│failed▶ Needs Attention
  │ made changes
  ▼
Set status: In Review
  │
  ▼
Ask me: “Merge this work?” ─rejected─▶ Changes Requested
  │ approved
  ▼
Merge ◀─────────────────────────────────┐
  │ merged   │ blocked                   │
  │          ▼                           │
  │      Any rebases left? ──no──▶ Merge blocked
  │          │ yes                       │
  │          ▼                           │
  │      Ask me: “Rebase with AI?” ─rejected─▶ Merge blocked
  │          │ approved                  │
  │          ▼                           │
  │      Rebase with AI ─────────────────┘
  ▼
Set status: Done ──▶ Clean up ──▶ End

What the graph buys you

  • Parallel steps. Two wires out of the same output run at the same time. That's the whole mechanism — no special syntax, no fan-out node. Two reviewers on one diff, on different models, concurrently.

  • Branches that rejoin. Each step finishes on a named output, and the wire from that output decides what happens next. Branches can merge back together and the step where they meet fires exactly once, whichever way the work went.

  • Bounded loops. Wire a step back to an earlier one and you have a loop. Guard it with a condition on attempts — “fewer than 3 tries” — and it stops on its own. Every circuit also has a hard step limit, so a mis-wired loop fails in seconds instead of burning tokens all night.

Implement, evaluate, fix, as a graph

The graph form of the tree above shows what wiring is for. Two reviewers read the same diff concurrently; a condition reads both verdicts; a failure loops back to a fix pass that resumes the original conversation rather than starting cold — bounded at three rounds.

Implement, evaluate, fix Built in
Coding agent
  │ made changes
  ├──────────────┬───────────────┐
  ▼              ▼               │  both run at once
Review (Opus)  Review (Sonnet)   │
  └──────┬───────┘               │
         ▼                       │
   Both passed? ──yes──▶ In Review ──▶ Ask me ──▶ Merge
         │ no                    │
         ▼                       │
   Attempts left? ──no──▶ Needs Attention
         │ yes                   │
         ▼                       │
   Fix what review found ────────┘  (resumes the same session)

A reviewer is a fresh headless session with a deliberately narrow tool surface: read the task, leave review comments, submit one verdict. It can't edit files, and it can't move the Task — it judges, and the graph decides what that means.

The built-ins

Built-ins are read-only — Save as a copy gives you an editable fork. The graph ones stay the default; each has a tree twin that does the same work, and one exists only as a tree.

Standard Cut a worktree, run a coding agent, review the diff, merge on approval — offering an AI rebase when the merge is blocked. The default.
Implement, evaluate, fix Two reviewers read the diff at the same time, and a fix pass that resumes the original conversation answers what they find — up to three rounds.
Work sub-tasks Work this task's sub-tasks in dependency order, unattended, merging each into this task's branch as it passes your checks. Nothing reaches the base branch until you approve this task's own merge at the end.
Unattended Cut a worktree, run a coding agent, and merge when the project's checks pass — no approval step. Meant for a sub-task inside a parent's Work sub-tasks run.
Weekly research Tree only. An agent researches a topic on the web and proposes a parent task with sub-tasks; you keep, edit or discard each one, and only what you keep is filed. Point a Trigger at it.

Choosing one

  • Per Task. A Circuit field sits in the task's properties rail, next to the model override and behaving the same way. Set it before you Start.

  • Per Project. Set a default and every Task in the Project uses it unless it says otherwise. Unset means the built-in Standard.

  • Global or project-scoped. A global circuit is offered in every Project; a project-scoped one only in the Project it belongs to. Built-ins are global, and read-only — “Save as a copy” gives you an editable fork.

  • Shared as a file. Export any circuit as JSON and import it somewhere else. Circuits don't sync between machines yet, so this is how you send one to a colleague.

Runs survive quitting the app

A run isn't kept in memory — it's recorded in the Ledger, one Event per step boundary, exactly like everything else about a Task. So closing Polynomic mid-run loses nothing. Reopen it and the run picks up at the step it was on: an agent that was mid-thought is offered for you to resume, an unanswered question is still waiting.

This is a real change from before, when a session that died with the app dropped its Task into Needs Attention for you to sort out by hand.

Each run also stores a snapshot of the circuit it started with, so editing a circuit never disturbs work already in flight — and a run still reads correctly on another machine that has never seen your definition. A run that is over isn't the end of the road either: Run again starts a fresh one from the first step.

Circuits home

⌘⇧C — or the sidebar's Circuits button — opens one full-window place that answers “what is my automation doing right now?” at a glance. It has three bands and a column, all walked with the keyboard: j k to move, ⇥ or 1–4 to jump between them, ↵ for the row's main action, and d for the designer.

Now Every active Run in the Project: Task, Circuit, where it is (Retry ×3 › Together › Review), how long, and spend. A Run waiting on a question leads the band and reads Waiting on you; ↵ opens the question in place, and y or n answers it.
Next Each Trigger's next fire: when, what fires it, and what it will make — with Run now (x), Skip next (s, which you can take back) and Pause (p).
Recently Finished Runs, newest day first, the ones that failed or got stuck first within a day. Each carries its reason, the step that failed, and Retry from here (r) on it.
By Circuit Each circuit that has run, with its Insights headline. ↵ opens it in the designer with Insights on.

Nothing on it is stored: every row is read from the Tasks and Triggers this machine holds. A Run another machine drives is shown with its owner named, and driving it is left to that machine. Empty bands are the good state and say so — Nothing running. Next: Weekly research, Mon 09:00. In the browser it opens with g c, and Triggers are read on the machine that fires them.

Triggers

A Trigger starts work without you pressing Start: every Monday at 09:00, when a deploy fails, or when you ask. A Run is always one Circuit working one Task, so a Trigger makes a Task from one of your templates and starts it on a Circuit. After that it is an ordinary Task — its timeline, Needs Attention, the Companion and sync all work without knowing a Trigger exists, and its first comment says which Trigger started it.

Open the Trigger editor from the command palette (Open Triggers). Each Trigger says when it fires and what it starts:

  • On a schedule. Written in plain language — “every Monday at 09:00”, “every weekday at 17:30 and 21:00”, “every 15 minutes”, “every month on the 1st at 08:00” — or as a five-field cron line, in the Trigger's own time zone. The editor says it back with the next fires. A fire missed while Polynomic was closed fires once on the next launch and says how late it was.

  • When an Action ends. When one of the project's Actions succeeds, fails, or either — a nightly deploy that failed can file and start its own investigation.

  • Only by hand. Nothing fires it on its own. Every Trigger, whatever its source, can also be run from the editor (⌘↵), from the command palette (Run trigger: …), or by asking the chat panel — which shows you what it will make, and fires only when you approve.

  • What it starts. A new task from a template, with its parameters filled in — including values only a Trigger has: {{date}}, {{time}}, {{weekday}}, {{week}}, {{month}} and {{trigger}}, plus {{action}}, {{action_status}}, {{exit_code}}, {{branch}}, {{commit}} and {{action_run}} when an Action fired it. Or a standing task, run again. Either starts on the Circuit you pick, or the task's own, or the project default.

Two guard rails. A Trigger doesn't fire again while the task it last started is still running, unless you let it. And Pause all holds back every schedule and Action on the machine — but not Run now or the chat, which are each a person deciding. Resuming doesn't catch up. Each Trigger keeps a History of its fires: who fired it, whether it started, was skipped or failed, and the Task it made. A task a Trigger made but couldn't start is left in Needs Attention with the reason, so it is there for you to find.

Weekly research that proposes work

Nothing an agent does on a schedule nobody is watching should file tasks on its own say-so. The Weekly research circuit has an agent research a topic on the web and propose one parent task with a sub-task per opportunity. You keep, edit or discard each item — on the desktop, in the browser, or on your phone — and only what you keep is filed, in its order. Next week's proposal is shown what is already filed, so it doesn't suggest it again.

  1. 1. In the template editor, choose Start from “Weekly research” and fill in its Topic.
  2. 2. Make a Trigger: every Monday at 09:00, a new task from that template, started on the Weekly research circuit.

Triggers live on one machine

Like Circuits and Actions, a Trigger is kept on the machine that made it, and only that machine fires it. In a cloud Project your phone still sees them: the Companion's Triggers screen shows each desktop's Triggers, when they fire next, and what they last made, as of when that desktop last synced. Ask Desk to run it (the button names the machine) is a request, not a run — the desktop fires it the next time it syncs, exactly like its own Run now. A request nobody answers within an hour expires, and the phone says nothing ran.

Task Templates

A Template is a reusable, parameterized blueprint for a Task: a name, a set of Parameters, and a partial description of the Task to make. Write {{name}} in any of its fields and filling that hole in is what turns one template into many Tasks. Press ⌘⇧T for the editor, ⇧C to make a Task from one.

  • Applying one writes ordinary Events. The Task that comes out is indistinguishable from one typed in by hand and holds no pointer back to the definition — so editing a template can't disturb a Task it made, and deleting one can't orphan it.

  • Unspecified is not empty. Every aspect is optional, and a field the template doesn't mention is the field ordinary creation would have given you — implemented literally by writing no Event for it.

  • Capture one from a Task you already have. The second task of a kind arrives when the first is on screen, so a Task can be turned into a Template — sub-tasks and all. The Task is untouched; no Parameter is invented, since which part varies is the one thing a finished Task can't tell you.

  • Global or project-scoped. Like Circuits, a Template is either offered in every Project or scoped to one, so it may name that Project's repos, labels and Circuits. No built-ins ship — a template encodes your workflow, not Polynomic's.

Code References

Reference code from a Task with @repo/path syntax — autocompleted over git ls-files. References resolve live and are marked stale when their target moves, so a Task always points at real code. This is what lets the Code-Tree view exist, and what gives an Agent Session the exact files the work concerns.

Code References on a Task

src/ledger/events.rs#append
src/ledger/store.rs

Today references are file-level; symbol-level granularity and @-mentions in comments are on the roadmap.

Agent Sessions

Starting a Task builds a Workspace — a git worktree and a polynomic/<id>-<slug> branch per touched repo — then spawns a real coding-agent process in an embedded terminal you can watch or type into. Sessions are resumable, and orphans are reconciled when you reopen a project. A worktree is bare unless the repo declares how to provision it — dependencies linked or copied across, a setup command run — which is what lets the agent run your build instead of only writing code.

An Agent Session is one Step in the Task's Circuit — which is why a Circuit can contain more than one of them, and why a fix step can resume the same conversation the first one started. You can also detach a running session into iTerm, Ghostty or Terminal.app and carry on there; it resumes with the same session and the same MCP connection.

Harnesses

The CLI a session actually runs is its Harness: Claude Code or Codex. It isn't only for Agent Sessions — the headless runs (AI Review, “generate sub-tasks”, “ask a question”, a Circuit's review step) and the conversations (a Task's Planning Session, the chat panel) resolve one too, so a Codex Task is reviewed, questioned and planned by Codex.

A harness is a set of capabilities rather than a swappable binary, and Polynomic branches on the capability, never on the name. A capability a harness lacks is visibly absent with a reason rather than silently inert: starting a Codex Task “in plan mode” is refused with that sentence, not quietly run without one. Harness resolves Step → Task → Project → default, and is recorded on the session — so one Run can implement with one vendor's agent and review with another's, which is the point. Two models from the same lab share blind spots; two from different labs share fewer.

An Agent Config picks which login a harness runs as — a named config home with its own credentials, settings, skills and plugins, chosen per Project. A work Project and a personal one can run the same agent as two different people. A Project stores the config's name, which is the part that travels; the path and the credential are facts about one machine.

What the agent can do to a Task

The agent talks to Polynomic over a loopback MCP server, scoped to its Task by a token. It can read and update the Task — but it only signals; Polynomic verifies against real commits before transitioning. A hollow “done” with no commits routes to Needs Attention instead of Done. A coding session's tool surface is deliberately narrow — it can read, annotate, break work down and order it, but it cannot curate how a Task is filed:

  • get_task
  • list_subtasks
  • add_comment
  • update_status
  • set_description
  • add_code_reference
  • copy_attachment
  • report_needs_attention
  • create_subtask
  • create_task
  • add_dependency
  • remove_dependency
  • request_start
  • run_check

Note what's missing: set_title, set_priority and set_parent say what work is called, how urgent it is and where it is filed — your ordering of your own board, and a Planning Session's job rather than a coding one's. request_start is a nomination, not a start: the suggestion is recorded on the Task, the dependency it implies is written into the graph, and you're asked about it once this Task's work has actually merged.

Setup & checks

A Workspace is a fresh git worktree, which means it starts with no node_modules, no .env, no warm build cache — so an agent working there can write code but can't run it, and a reviewer can only read. Two declarations per repo fix that: provisioning, which makes a new Workspace workable, and Checks, the named commands anything is allowed to run in it.

Both are declared, never guessed. Polynomic doesn't sniff your repo for a package manager or invent a test command: a symlinked node_modules is right for most JS projects and wrong for one with native addons, so it runs exactly what you wrote and nothing else. A repo with nothing declared behaves precisely as it did before this existed.

Declaring them

Open Project settings — the sliders button beside Projects in the sidebar — and each of the project's repos gets a card:

~/dev/polynomic

Base branch

main

Setup command

bun install

Link from primary checkout

node_modules
src-tauri/target

Copy from primary checkout

.env

Checks

test · bun test · 300s

typecheck · bun run typecheck · 9m

  • Link. Repo-relative paths symlinked from your primary checkout — node_modules, target, .venv. Free, instant, and shared: don't link anything two concurrent builds would corrupt.

  • Copy. Repo-relative paths copied instead, so each Workspace owns its own — .env, a warm cache. Costs disk, though on APFS the copy is a clone and costs nothing until it diverges. A path your primary checkout doesn't have is skipped and reported, not fatal.

  • Setup command. Run once in each fresh worktree, after the links and copies (which is what makes an npm install a no-op rather than a download). If it fails, the Prepare workspace step fails — a half-provisioned Workspace is the thing this exists to avoid, so the failure surfaces there instead of as a confused agent later.

  • Checks. A name, a command, and an optional timeout in seconds — default 9 minutes, after which the command's whole process tree is killed and the run is recorded as timed out. The name is what every other surface refers to, so keep it short: build, test, lint, typecheck.

Commands run through your shell with your login-shell environment, so a version manager on your PATH is there — a GUI-launched app otherwise has no PATH worth speaking of, and every check would fail with “command not found” for reasons that have nothing to do with the code.

Three ways to run a check

Once declared, the same command is reachable from all three places work happens — one implementation behind them, so they can't disagree about what “test” means:

  • In a Circuit. The Run checks step runs them and finishes on passed or failed. Its output — how many ran, passed and failed — is a number a Condition can read, so “if any check is red, loop back and fix it, at most twice” is a wire and a condition. Drop it after the coding agent, before review, or both.

  • From the Changes tab. A Run checks button in the diff toolbar, shown only when the project declares something. Results appear in a panel above the reviewer's verdict — a machine fact belongs above an opinion — with a failing check's output tail folded behind its row.

  • By the agent itself. run_check is in the coding session's MCP tools and the reviewer's, and both are told the declared checks exist. The agent verifying its own work with your authoritative command — rather than an approximation it invented — is most of the point of declaring one.

Every run is recorded

A run isn't a transient console line. It appends an Event to the Ledger — the command verbatim, the exit code or that it timed out, how long it took, the commit it ran against, whether uncommitted changes were present, and the tail of its output. So it is on the Task's timeline afterwards, it survives the branch being deleted, and it syncs to your other machine. The Task carries the latest result per repo and check; the history stays in the timeline.

Changes Run checks

Checks — 1 red

typecheck (polynomic) — passed in 21s against 4f1c9ab

test (polynomic) — failed (exit 1) in 48s against 4f1c9ab

lint (polynomic-www) — passed in 6s against 0b3ad12 — stale: the branch has moved since

A result speaks for one commit

Because each result records the commit it ran against, Polynomic can tell you when it no longer applies. A result the branch has moved past is marked stale and stops carrying weight anywhere: a stale green must not reassure you, and a stale red neither alarms you nor ranks your queue.

What red does do is order your reading and brief your reviewers. In the Review queue a diff whose own checks are red on the current tip ranks above one a reviewer merely flagged — a machine fact outranks an opinion. And an AI Review is now told what this Workspace can run and what already ran, so instead of the old “dependencies aren't installed, review statically”, its statement of what it could not check shrinks to what is genuinely unknowable.

And then it stops there. A red check blocks nothing — not the merge, not the Task's state. It ranks a row and tells a human, and the merge stays a human act, exactly like a reviewer's verdict. If you want a red check to actually stop work, that's a Circuit you draw — a Condition on the failed count, routing wherever you want it to go.

Tool permissions

An Agent Session can read, edit, and run commands on your behalf — so anything with side effects asks first. When the agent wants to use a tool that needs your sign-off, the request queues quietly rather than interrupting you. Press a from anywhere to open the Tool Approvals center and decide. The notification center flags any task that's waiting, so nothing blocks silently.

The approval center

The center is a triage surface, not a one-shot popup. A left rail lists every pending request — tool name, the Task it belongs to, and the command, path, or URL at a glance. The right pane shows the full tool input and the working directory it would run in, so you can see exactly what's about to happen before you decide.

Tool Approvals 2
  • Bash POL-42 npm install
  • Write POL-42 src/server/index.ts
Bash POL-42 Add the WebSocket transport
npm install

in ~/.polynomic/workspaces/POL-42/app

Deny d Allow always ⇧A Allow a

Three decisions

Every request resolves to one of three answers. The whole center is keyboard-driven: j / k move through the queue, then act on the selected request.

  • Allow a / ⏎ — Permit this one call. The agent unblocks immediately and keeps working.

  • Allow always ⇧A — Permit this call and remember the decision, so matching calls auto-allow for the rest of the session without prompting again. Bash grants are scoped to the program (allowing npm install allows all npm commands), and the file-editing tools share one grant.

  • Deny d — Reject the call. The agent gets the refusal and either tries another approach or routes the Task to Needs Attention.

Closing the center with Esc leaves requests queued — it doesn't answer them — so you can come back to them later. Summon it again any time with a, even when nothing is pending; it just shows an all-clear.

What never asks

To keep the center signal-rich, read-only and clearly safe tools skip it entirely: file reads, searches, web search, and a small allow-list of read-only git commands (status, diff, log…) run without a prompt. Anything you've already allow-listed in your Claude settings (~/.claude/settings.json or a repo's .claude/settings.json) is honored too. If you ignore a request long enough, it falls back to the agent's own permission prompt in the embedded terminal — so a request is never lost, just waiting on you.

Driving the approval center from outside the terminal is a harness capability, and Polynomic only claims it where it holds. Codex answers its own approvals in the session terminal instead — the picker says so when you choose it, rather than leaving you with a center that never opens.

Review & merge

Nothing reaches your main branch until you've seen it. Every Task produces a per-repo branch diff, read on its Changes tab. An independent AI Review — a fresh, headless session, never the one that wrote the code — reads it and reports two things: one anchored, rated finding per problem, and exactly one verdict for the diff as a whole. It can only annotate and judge; the profile withholds the tools that would let it change state or merge.

The Review queue

⌘⇧R opens the ranked list of Tasks whose changes are waiting on a human. The order is a ladder of what you can do next rather than of how alarming a Task looks: comments you wrote but never sent, then a diff whose own checks are red on the current tip — a machine fact outranks an opinion — then a diff a reviewer flagged, then one waiting with nothing said about it, then one a reviewer found nothing blocking in — and last, a diff still moving because an agent is writing to it, which sinks however long it has waited. Finishing one diff hands over the next (⇧J from anywhere in a Task), so the queue never has to be navigated back to. It decides nothing — the merge stays a human act on the Task's own Changes tab.

Working a diff

  • One finding, one place. A finding is anchored where it is about — a line, a range you dragged down the gutter, or a whole file — and carries an optional severity: blocker, should fix, or nit. n / ⇧N walk the unresolved ones, severity first and then in reading order; ⇧F lists them all on one screen, each quoting only the lines it concerns.

  • Resolve or dismiss. r marks the finding you are on as addressed, ⇧R dismisses it; each is a new Event rather than an edit, so a dismissal — exactly the decision someone may want to argue with later — is still there to argue with. Neither gates anything: unresolved findings are information, not a lock.

  • Viewed is yours alone. v folds a file away and moves to the next one still to read. That mark is a fact about a reader, not about the work, so it lives on your machine and is keyed to the file's current diff — a file the agent has since changed stops counting as viewed.

  • Read it against the intent. i puts the Task's description and Code References beside the code, so a diff is checked against what was asked rather than read cold. Files the Task's references name sort first; machine-written files arrive folded, with the count always stated.

Deciding

  • Request changes. Send the findings back and route the Task to Changes Requested.

  • Send to agent. Resume the original Agent Session with the unresolved findings, each with its severity spelled out — and it says what it is not sending, since an agent handed a finding you already dismissed goes and does the work anyway. Blockers only is offered beside it when there's a distinction to make.

  • Guarded merge. m fetches, rebases and merges --no-ff, locally. Conflicts route to Needs Attention; on success the worktrees are cleaned up. After the merge the Changes tab reads the diff back out of the merge commit, and says so — there is no branch left to diff.

  • Or open a PR. From an In Review Task, Create PR pushes the branch and opens one through the gh CLI. It is a parallel artifact: it doesn't replace the local diff or the local merge, and opening one doesn't change Task State.

  • Or wire it differently. Reviewing and merging are Steps in the Task's Circuit, so this order is a default rather than a rule. Add a second reviewer, gate the merge on both agreeing, or drop the approval step for work you're happy to land unattended.

Sync & sharing

A Project is local by default and stays that way unless you say otherwise. Sign in to an Account and a Project can sync its Ledger to a server, so a second machine — or a colleague you've added as a member — folds the same Events into the same state.

  • Conflict-free by construction. Every Event carries a Hybrid Logical Clock, and merging two machines' Ledgers is a union by Event id followed by a re-fold in HLC order — deterministic on every machine, with no central coordination and nothing to resolve by hand.

  • Still local-first. Sync is additive. The local SQLite Ledger stays the source of truth, the app works fully offline, and your code never leaves git — only Task data is on the wire.

  • The server names the author. On a shared Project the server stamps each Event's author from the pusher's token, because an author a replica asserts is a forgeable audit trail. An Event nobody can attribute renders as such, never as you.

  • One machine drives a Run. The Replica that started a Run is the one that can move it forward; the others render it. A Step started elsewhere would run with no Workspace, so Start, Stop, Merge and answering a pending question all refuse on another machine — naming the one that holds it.

Circuit and Template definitions live in the machine-local registry and don't sync — export one as JSON to share it. A Run does sync, and it carries a snapshot of its Circuit, so a colleague can read what a Task did even without the definition.

GitHub issues

Cloud Sync shares a Project with people who run Polynomic. A Tracker shares it with everyone else. Connect a Project to a repository under Project settings → Tracker and a Task can be Promoted to it: an issue is opened for it, and from then on the two are kept in step in both directions. Nothing is promoted for you — deciding that a particular piece of work is the whole team's business is the point of the feature.

It goes through the gh CLI you are already signed in to, so Polynomic never holds a GitHub credential. GitHub is the one Tracker that ships.

  • Promoting doesn't move the work. The Task stays in this Ledger, worked by this Project's agent with its repos, house rules, Circuit and Review queue. A tracked Task is marked wherever it is listed, because a colleague can now read it, comment on it and close it. Unlinking later leaves the issue exactly where it is — it is a statement about this board, not an instruction to anybody else's.

  • What travels, and when. Title, description, status, labels, assignee, comments — and priority and due date where the issue has fields for them. What you change here goes up as you make it; what a colleague changes over there arrives on a two-minute sweep, since a desktop app has no webhook to receive. Code References, sub-tasks, reviews and the agent's transcript stay here.

  • Status rides on labels, and you say which. GitHub issues are open or closed and Polynomic has nine states, so seven of them are carried by a label — a fact about your team's repo, so it is yours to map. Only the labels your mapping names are Polynomic's to add and remove; every other label on an issue is left to whoever put it there. A status you leave unmapped is never pushed at all, which is not the same as one mapped to no labels.

  • Who is who is asked, not guessed. A Task holds an account member, an issue holds a login, and no rule turns one into the other — so the pairing is yours to write, and an assignee Polynomic cannot name is left alone on the issue rather than unassigned. An assignee that already looks like a handle needs no entry.

  • Issues are offered, not adopted. An issue with no Task here shows as a dashed row at the top of Backlog with Take on and Dismiss, bounded by a filter that defaults to issues assigned to or mentioning you. Nothing is written to the Ledger until you take one on — binding a repo with nine hundred open issues should not deposit nine hundred rows on your board.

  • GitHub is another replica. A change made on the issue is not reconciled from a mirror table: it lands on the Task's own Ledger as an ordinary Event bylined GitHub, carrying GitHub's own timestamp, so conflicts settle by the fold that already exists. On a Project that syncs, one machine polling gh is enough for the whole team — a colleague who has never run it still sees “GitHub closed this”, attributed correctly.

Starting an agent on a tracked Task says so on the issue: the mapped status, the assignment if nobody over there has it, and one comment naming who started and which agent — once, never again. The agent's own running commentary stays on the Task unless you ask for it. All of it is best-effort: a GitHub that is unreachable never stops an agent starting. The web app can read a tracked Task but cannot promote or sync one, there being no gh in a browser to be signed in as.

Keyboard shortcuts

Polynomic is keyboard-first: capture, start, review, and merge without reaching for the mouse. Press ? in the app for this list any time — the help overlay and this page read the same source, so they can't disagree. On Windows and Linux, ⌘ is Ctrl.

Global
Open command palette ⌘K
Search tasks by any content /
Show keyboard help ?
Create a task c
New task from a template ⇧C
Go to the just-created task (while its toast is shown) g
Open the agent tool-approval center a
Back / forward through places you've been ⌘← / ⌘→
Toggle sidebar [
Toggle chat panel ]
Cycle appearance: System → Light → Dark ⌘⇧L
Global — full-window surfaces
Open Circuits: live Runs, what fires next, what finished (d for the designer) ⌘⇧C
Open the task template editor ⌘⇧T
Open the project's dependency graph ⌘⇧G
Open the Review queue ⌘⇧R
Open the project's Statistics page ⌘⇧S
Open the project's Initiatives — goals bigger than a task ⌘⇧I
Open Drafts — captures with no Project yet ⌘⇧D
List
Move cursor down j / ↓
Move cursor up k / ↑
Expand sub-tasks l / →
Collapse sub-tasks h / ←
Change status of task under cursor s
Change priority of task under cursor p
Start / resume task under cursor x
Open task under cursor Enter
Select / deselect task under cursor Space
Extend the selection down / up ⇧J / ⇧K
Select every visible task ⌘A
Clear the selection Esc
Board
Move cursor within a column j k / ↓ ↑
Move cursor between columns h l / ← →
Move task to prev / next column ⇧H / ⇧L
Change status of task under cursor s
Change priority of task under cursor p
Start / resume task under cursor x
Open task under cursor Enter
Select / deselect task under cursor Space
Select every visible task ⌘A
Clear the selection Esc
Detail
Edit title e
Change status s
Change priority p
Start task (Todo / Backlog) x
Schedule for planning — start in plan mode ⇧P
Switch to the Overview / Changes tab 1 / 2
Toggle between Overview and Changes d
Merge task (In Review / Changes Requested) m
Next / previous diff in the Review queue ⇧J / ⇧K
Copy task ID to clipboard ⌘.
Back to list Esc
Changes — inside the diff
Move between files in the diff j / k
Expand / collapse the file under the cursor l / h
Mark the file viewed, and move to the next unread one v
Comment on the whole file under the cursor f
Next / previous unresolved finding n / ⇧N
List the findings, without the code around them ⇧F
In the findings list: show the finding in the diff ↵
Resolve the finding you are on, and move to the next r
Dismiss the finding you are on, and move to the next ⇧R
Show the Task's intent beside the diff i
Circuits home
Move through the rows j / k
Jump between Now, Next, Recently and By Circuit ⇥ / 1–4
Answer the question, or open the row ↵
Approve / reject a question answered in place y / n
Open the row's task o
Retry from here, on a Run that failed r
On a Trigger: run now, skip next, pause x / s / p
Open the Circuit designer d
Tree Circuit outliner
Move between steps ↑ / ↓
Fold / unfold a step ← / →
Add a step after this one, or inside an open Control Step ↵
Indent / outdent ⇥ / ⇧⇥
Move the step up / down among its siblings ⌥↑ / ⌥↓
Change the step's kind /
Wrap the step in a Control Step W
Add or go to its On failure F
Add or go to a Retry's Before retrying B
Edit the step Space
Show or hide Insights I
Delete the step and everything under it ⌫
Undo / redo ⌘Z / ⇧⌘Z
Save ⌘S
Trigger editor
Move through the Triggers j / k
New Trigger n
Run the Trigger now ⌘↵
Save ⌘S
Close Esc
Review queue
Move between the diffs waiting on you j / k
Open the diff under the cursor Enter
Back to the task list Esc

Data & storage

Everything is local by default. Your tasks, events, and code references live in SQLite — one Ledger database per Project — kept entirely separate from your git repos. No cloud dependency, no seat fees. Your source stays in git; your task data stays separate; you own both.

~/.polynomic/
~/.polynomic/
├── registry.db                     project registry + your Circuits and Templates
├── projects/<project-id>/
│   ├── polynomic.db                 that project's event ledger + projections
│   ├── attachments/<task-id>/       files attached to a task
│   └── transcripts/                 archived session transcripts (gzipped)
└── workspaces/<task-id>/            one git worktree per touched repo

Circuit runs are Events in the Ledger like everything else, which is what makes a run resumable after a restart. Circuit and Template definitions live in registry.db and don't sync — export one as JSON to share it.

A session's transcript is not Ledger truth — the Ledger records that a session ran and links to it, never the turns themselves. Polynomic archives the harness's own transcript file into its data dir so it isn't lost to that tool's compaction or cleanup, and each session boundary stores a cursor into it, so the timeline can show exactly the turns that belong to one stretch of work.

Override the root with POLYNOMIC_DATA_DIR. Projections are pure caches: an append writes the Event and re-folds the projection in one transaction, and the projections can be dropped and rebuilt from the Ledger at any time.

FAQ

Which platforms are supported?
macOS as a universal build for Apple Silicon and Intel, Windows as an installer, and Linux as an x86_64 AppImage.
Do I need a coding-agent CLI?
Yes — an Agent Session spawns a real one. Claude Code and Codex are both supported; Settings › Coding agent can install one and sign you in, and Polynomic disables Start with a banner if it can't find any on your PATH.
Can I use Codex instead of Claude Code?
Yes. A harness is chosen per Step, Task or Project, and it applies to everything — a Codex Task is also reviewed, questioned and planned by Codex. Where Codex lacks something Claude Code has (plan mode, the in-app approval dialog), Polynomic says so at the point of choosing rather than quietly running without it.
Does Polynomic clone or push my repos?
It never clones — a Project points at existing local repos, and merges are local git operations. The one push is optional and explicit: Create PR on an In Review Task pushes the branch and opens a pull request through the gh CLI.
Can a Task span multiple repos?
Yes. A Project can include several Repos, and a single Task can touch more than one — Polynomic lays a worktree of each side by side in one Workspace.
What happens if an agent claims it's done but didn't commit?
Polynomic verifies against real commits before transitioning. A hollow “done” routes to Needs Attention rather than Done.
How do I approve the commands an agent wants to run?
Press a to open the Tool Approvals center and Allow, Allow always, or Deny each pending request. Read-only tools and anything you've allow-listed in your Claude settings run without asking. Codex answers its own approvals in the session terminal instead. See Tool permissions above.
Does Polynomic sandbox the agent?
No. Polynomic wraps nothing: a session runs as you, in its own git worktree, and every side effect goes through the approval flow. A sandboxing feature was built and then removed because it didn't work in practice — two optional runtimes that silently degraded to no sandbox at all when neither was available, so "sandboxing is on" told you very little. It is still wanted, and when it returns it will be one enforced boundary that refuses to start rather than quietly falling back. Codex brings its own sandbox, which is the whole sandbox story today.
Can an agent run my tests and my build?
Yes, once the repo declares them. In Project settings you give each repo a setup command and the paths to link or copy from your primary checkout, so a fresh workspace is workable, plus its own named checks (build, test, lint, typecheck). Then a circuit's Run checks step, the Changes tab's Run checks button, and the agent's own run_check tool all run the same command. Nothing is inferred — a repo with nothing declared gets the bare worktree it always did.
Do red checks block a merge?
No. A red check ranks the diff higher in the review queue, is told to reviewers, and sits on the task's timeline with the commit it ran against and the tail of its output — but it stops nothing, exactly like a reviewer's verdict. If you want red to halt the work, wire it: a Condition on the Run checks step's failed count routes wherever you want.
Is Polynomic a CI service?
No. Checks are your own commands run locally in the task's workspace, on your machine, with your login-shell environment — there's no runner, no image and no config file to write. Each run carries a timeout, and a check that hangs has its whole process tree killed and is recorded as timed out rather than wedging a circuit run.
Can I start a Task that's waiting on another one?
Yes. Blocked is information, not a lock — Start and Plan raise a warning naming what hasn't landed and then start the Task anyway on Enter. The decision is yours; only the noticing was missing.
Do I have to design a Circuit to use Polynomic?
No. Every Project comes with Standard, which does exactly what Start always did, and it's the default. Circuits are there when you want more than one agent pass — parallel review, a bounded fix loop — not as a setup step.
What happens if I quit the app while a Circuit is running?
The run resumes where it left off when you reopen. Run state lives in the Ledger, not in memory, so an agent that was mid-thought is offered for you to resume and an unanswered question is still waiting. Previously a session that died with the app dropped its Task into Needs Attention.
Can a Circuit loop forever?
No. In a tree, Retry ×N and Repeat ×N carry their own count, and the designer warns you about a Repeat until fails with no maximum. In a graph, guard a loop with a condition on attempts. Every circuit also has a hard step limit, and a run that trips it stops and routes the Task to Needs Attention with the reason.
Tree or graph — which should I write?
A tree, unless you already have a graph you like. New circuit makes a tree: it reads top to bottom, edits from the keyboard, and shows a running Run as the tree you wrote. Graphs keep running, and the built-ins ship in both shapes. A tree can't express arbitrary back-edges or a join across branches, which is why graphs stay.
Can a Circuit run without a Task?
No — a Run is one Circuit working one Task. That is why a Trigger makes a Task and starts it rather than running a Circuit on its own: what it did is then an ordinary Task you can read, retry and review.
Can I share a Circuit with my team?
Export it as JSON from the designer and import it on the other machine. Circuit definitions don't sync — a run does, and it carries a snapshot of its circuit, so a colleague can read what a Task did even without the definition.
Is my task data only on this machine?
Unless you say otherwise, yes. A Project is local-only until you sign in and make it a cloud Project; then its Ledger syncs and you can add members. The local SQLite Ledger stays the source of truth either way, and your code never leaves git.

Ready to manage the work, not the agents?

Local-first, keyboard-first, every change reviewable before it lands.