Paperclip: Organizing Multiple AI Agents into a Governable Autonomous Team
Paperclip: Organizing Multiple AI Agents into a Governable Autonomous Team
When several Claude Code, Codex, or other coding agents are running at the same time, the hard problem is no longer asking another model to write code. It is managing task context, permissions, cost, schedules, and accountability. Paperclip approaches this problem as an open source control plane for coordinating teams of AI agents.
The idea is simple: agents do the work, while Paperclip manages the organization. A user can define company or project goals, create roles and reporting lines, connect agents from different runtimes, and use one dashboard to manage tasks, budgets, approvals, and execution records. For teams moving AI from one shot conversations toward long running automation, this is an architecture worth examining.
Reading guide
1. Why multi agent systems need a control plane
2. Paperclip core models: goals, organization, tasks, and heartbeats
3. What Paperclip solves through cost control and governance
4. Local setup and a first experiment
5. Suitable use cases, limits, and adoption advice
First clarify this: Paperclip does not replace the agent
The Paperclip README describes the project clearly: it is a Node.js server with a React UI for coordinating a group of AI agents. It does not prescribe how to build an agent, and it is not a chatbot, prompt manager, or drag and drop workflow builder.
That distinction matters. Claude Code, Codex, Cursor, a Bash agent, or an HTTP bot can act as the execution layer. Paperclip handles the management problems above that layer. It provides a shared language across runtimes: who owns which responsibility, which goal produced a task, how much has been spent, and whether the next action needs human approval.
This separation also makes the architecture easier to evolve. When a team changes a model or an agent runtime, the organization, task records, and governance rules do not need to be rewritten together. When work grows from one operator to many agents, the team does not have to depend on scattered shell scripts, terminal tabs, and manual reminders.
Four core models
1. Goal: make the agent understand why
A task normally says what to do, but a multi agent system also needs to know why it is doing it. Paperclip connects companies, projects, goals, and issues so that a task retains its goal ancestry. The agent receives more than an isolated ticket; it can trace the task back to a higher level organizational objective.
This reduces a common failure mode: every agent completes a local task while the overall system moves in the wrong direction. When tasks are linked to parent goals, managers can check whether the work is still valuable, and agents can use a consistent priority order when tradeoffs appear.
2. Org chart: replace isolated sessions with roles and boundaries
Paperclip treats an agent as an organizational member with a role, title, reporting line, permissions, and budget. The README uses examples such as CEO, CTO, engineer, designer, and marketer. The point is not to imitate a human company; it is to make responsibility, delegation, and access control explicit data.
With an org chart, one agent can decompose a goal and delegate child tasks to other agents. Different agents can also be limited to a particular company, project, or tool set. For deployments that manage several products or customer environments, this is easier to isolate than putting every bot in one shared workspace.
3. Issues and workspaces: make work recoverable
Paperclip tasks store more than a title. The official README lists links to companies, projects, goals, parents, blocker dependencies, comments, documents, attachments, and work products. Tasks also use atomic checkout and execution locks so that two agents do not claim the same work.
The value is recoverability. A multi agent system should not assume every execution succeeds on the first attempt, and it should not bind context to one terminal window. Paperclip uses persistent issues, session state, structured logs, and workspace information to preserve execution state. After a restart, an agent can continue from the original task context instead of guessing what happened in the previous run.
4. Heartbeat: start work through events and schedules
Heartbeat is the mechanism that moves an agent from waiting for a command to working according to rules. An agent can wake on a schedule, inspect its work, or be triggered by events such as task assignment and mentions. The README also describes cron, webhook, and API triggers, together with concurrency and catch up policies.
This allows periodic work to be modeled as a routine rather than a reminder for a person. Daily support summaries, recurring reports, and project test checks can create tracked routine executions and related issues. Each run still enters the same budget, permission, and audit flow.
The governance layer: automation does not mean unrestricted execution
Once an agent can run for a long time, governance is part of the system boundary, not an optional extra. Paperclip exposes several practical controls.
Budget limits and hard stops
Paperclip can track tokens and costs by company, agent, project, goal, issue, provider, and model. It supports warning thresholds and hard stops. When an agent exceeds its budget, the system can pause it and cancel queued work. This is more useful than discovering a runaway loop only after the bill arrives.
During adoption, I would give each agent a small budget first and observe normal task costs. After heartbeat frequency, retry behavior, and tool calls are understood, the limit can be increased gradually. A budget is not only a cost saving mechanism; it defines the blast radius of an experiment.
Approval, pause, and audit log
The README lists board approval workflows, execution policies, review stages, decision tracking, and controls to pause, resume, or terminate an agent. High risk work can therefore follow a process in which the agent proposes, a human approves, and the system executes, rather than granting every permission at once.
Mutating actions, heartbeat state changes, cost events, approvals, comments, and work products can form durable activity records. When an outcome is unexpected, the team can investigate who made which decision and when instead of relying only on the model's final text.
Secrets and tool boundaries
Secrets and tool access are easy to overlook in a multi agent deployment. Paperclip describes instance secrets, company secrets, encrypted local storage, and scoped injection of secrets into a particular run. Its MCP tool gateway and plugin system provide a path for adding controlled capabilities.
In practice, I would never expose an entire environment file or a production token to every agent workspace. A safer pattern is to assign least privilege by role, limit each secret to a company, tool, and execution that actually needs it, and make every use traceable in the activity log.
Looking at the source: the shape of the system
Paperclip's control plane can be understood through several modules:
- Identity and Access manages board users, agent API keys, run JWTs, company membership, and invitations.
- Org Chart and Agents stores roles, titles, reporting lines, permissions, and budgets while connecting runtimes through adapters.
- Work and Task System handles issues, dependencies, comments, attachments, work products, and atomic checkout.
- Heartbeat Execution manages the wakeup queue, budget checks, workspace resolution, secret injection, skill loading, and adapter invocation.
- Governance and Approvals provides approvals, policies, decision records, hard stops, and audit logging.
- Plugins, MCP, and Observability make the system extensible and support opt in OpenTelemetry traces or Sentry error monitoring.
This architecture reveals an important design choice. The core is not one call to an LLM; it is placing an agent run inside a process that can be retried, approved, costed, and audited. That is also the key difference from a typical agent framework.
Local setup: start with a small and safe experiment
The official README provides two quick paths. For the installer, it recommends downloading the install script and its SHA 256 file, verifying the checksum, and then executing the script. It also notes that both files come from the same origin; for independent verification, use a release tag or a commit pinned copy from GitHub.
curl -fsSLO https://paperclip.ing/install.sh
curl -fsSLO https://paperclip.ing/install.sh.sha256
sha256sum -c install.sh.sha256
bash install.sh
A manual development setup is also available:
git clone https://github.com/paperclipai/paperclip.git
cd paperclip
pnpm install
pnpm dev
The requirements listed by the project are Node.js 24.11 or newer and pnpm 9.15 or newer. Development mode starts the API server at http://localhost:3100. The README says an embedded PostgreSQL database is created automatically, so an isolated first experiment does not require a separate database service.
I would keep the first test on the local loopback interface, use a low privilege agent, and set a small budget. Verify these four points before considering LAN, tailnet, or production deployment:
1. Does a task correctly carry context from a goal to an agent?
2. Do heartbeat, retry, and orphaned run recovery behave as expected?
3. Does execution actually stop when the budget limit is reached?
4. Are approvals, tool calls, costs, and work products recorded clearly?
Do not give an agent write access to a production repository merely because the interface looks like a task manager. An isolated workspace and test data are cheaper than cleaning up one mistaken delegation.
When Paperclip is worth using
I see three especially suitable situations.
First, a team already has several agent runtimes and is losing visibility into who is doing what. Paperclip can put agents from different providers or command line tools into one organization and task context.
Second, a team has many periodic or event driven jobs and wants agents to keep progressing instead of answering only when asked. Heartbeats, routines, and durable activity make this work closer to an operable service.
Third, a team needs a balance between autonomy and human control. Budgets, approvals, pause and terminate controls, secret scopes, and audit logs turn unrestricted automation into automation inside explicit boundaries.
On the other hand, Paperclip may be too heavy for one agent and a few manual tasks, or for a team that only needs a prompt chaining library. Its value comes from organization and governance. Without a multi agent coordination problem, the original agent tool is usually simpler.
Three adoption questions
Define what cannot be automated first
Not every action should be delegated. List the resources, data, and external side effects that require human approval, then encode those boundaries as policies before an incident forces the decision.
Treat cost as a product metric
Look beyond the total bill. Break costs down by goal, issue, provider, and model. If one task retries repeatedly or makes abnormal tool calls, cost data should help identify the cause rather than act as a report at the end of the month.
Make portability an early concern
Paperclip's company export and import, secret scrubbing, and collision handling suggest that the organization itself should be a portable asset. Keep secrets, environment differences, and agent adapters separate in templates so they can be copied to another project or isolated environment later.
Conclusion: the next bottleneck is operating AI agents
As agents become long running members of a team instead of one shot chat tools, the bottleneck moves from model capability to operations: queueing work, preserving context, limiting cost, constraining permissions, tracking decisions, and deciding when a human should intervene.
Paperclip provides an open source and self hosted control plane that puts goals, org charts, issues, heartbeats, budgets, and governance in one system. It does not claim to build the smartest agent. It tries to let different agents work inside clear organizational boundaries.
For an AI engineering team, the most useful idea may not be a particular UI feature but the separation itself. Keep agent runtimes separate from company scale collaboration, governance, and observability. Once we manage not just a few model calls but a group of digital workers that can keep acting, this control plane may become essential infrastructure.
References