Pi Agent Harness: Turning an AI Coding Agent into a Composable, Repeatable Engineering Toolchain
Pi Agent Harness: Turning an AI Coding Agent into a Composable, Repeatable Engineering Toolchain
AI coding tools are moving from “complete this snippet” toward “finish a verifiable engineering task.” As their feature sets grow, however, workflows can become tightly coupled to one product: the model, tools, context, sessions, permissions, and custom instructions are all packed into one interface. Change the model or deployment model and the team has to relearn the whole system. Pi Agent Harness takes a different approach. It keeps a small terminal agent core and lets users assemble their workflow with TypeScript extensions, skills, prompt templates, themes, and packages.
What makes Pi worth studying is not only that it is a coding-agent CLI. It separates the idea of an agent product into replaceable engineering layers. The same agent runtime can be used interactively or through print, JSON, RPC, and SDK modes. A team can validate a workflow in the terminal first and later embed the same capability in an internal service.
The short version: Pi addresses workflow lock-in, not a lack of model capability
Pi describes itself as a minimal terminal coding harness. Its default tools include read, write, edit, and bash, allowing a model to inspect a project, modify files, and run commands. At the same time, Pi deliberately does not treat sub-agents, plan mode, MCP, or complex permission prompts as mandatory core features. Those concerns can be handled through extensions, skills, external packages, or an engineering workflow designed by the team.
This choice has two direct consequences. First, the core is relatively easy to understand; users do not need to learn an enterprise control plane before running a first task. Second, a team does not have to accept the tool author's preferred collaboration abstraction. It can decide whether to add sub-agents, checklists, MCP, approval flows, or sandboxing. Pi fixes the smallest useful agent loop and hands the product shape back to the engineering team.
That does not make Pi suitable for everyone. If you want an out-of-the-box multi-agent workspace, a complete planning mode, detailed permission UI, or a built-in MCP ecosystem, Pi's restraint becomes additional work. For teams that want to put an agent inside an existing engineering process, however, the trade-off is useful.
Architecture: from a unified model interface to an embeddable agent runtime
Pi is more than one CLI. The official monorepo separates several related packages:
@earendil-works/pi-coding-agent: the interactive coding-agent CLI, including the terminal interface, command-line modes, sessions, and resource loading.@earendil-works/pi-agent-core: an agent runtime with tool calling, state management, and event streaming.@earendil-works/pi-ai: a unified multi-provider LLM API so application code does not hard-code every provider's request format.@earendil-works/pi-tui: terminal UI components and differential rendering.
This separation lets users choose the right abstraction. If the goal is terminal-based coding, install the coding agent. If the goal is an internal agent service, use agent-core. If the main requirement is switching among model providers, study the provider and model interfaces in pi-ai.
The basic agent-core model is state plus events. An Agent stores the system prompt, active model, tools, and messages. When prompt is called, it emits events such as agent start, turn start, message updates, tool execution, and agent end. The value is that a UI does not have to wait for one final answer. It can update progressively, record tool execution, create an audit trail, or stop the flow after a controlled event.
The official documentation also separates AgentMessage from the LLM messages understood by a provider. An application can keep UI-specific or domain-specific message types, then use transformContext and convertToLlm to prune, inject, and convert context before a model call. Context governance becomes a runtime hook instead of provider-specific glue code.
Tool execution is also explicit. The runtime supports parallel or sequential execution and exposes beforeToolCall, afterToolCall, and shouldStopAfterTurn hooks. A team can validate arguments before execution, attach audit information to a result, or decide to compact context after a completed turn instead of immediately making another model call.
Four execution modes: from a terminal to a service
Pi coding agent supports four main modes: interactive, print or JSON, RPC, and SDK embedding.
Interactive mode is for working with an agent in a terminal. You can ask Pi to inspect files, edit code, and run tests, while using commands such as /model, /session, /tree, and /compact to manage models and context.
Print or JSON mode is useful for scripts and CI. A task such as “check whether this change is missing tests” does not need a full TUI. The prompt, file references, and result can be connected to an existing pipeline.
RPC mode is useful when another process needs to control the agent. An external service can own the queue, permissions, observability, and user interface while Pi provides the agent process.
SDK mode embeds the runtime in an application. The official createAgentSession example shows how to create an in-memory session, send a prompt, and consume the result programmatically. This path makes it possible to validate prompts and tools with the coding agent before gradually turning them into a controlled product capability.
Getting started: run a small, verifiable task first
1. Install the coding agent
The official quick start uses a global npm installation and recommends --ignore-scripts so dependency lifecycle scripts are not executed during installation:
npm install -g --ignore-scripts @earendil-works/pi-coding-agent
The official installer is another option:
curl -fsSL https://pi.dev/install.sh | sh
Use the curl installer only after checking the source, execution environment, and supply-chain risk. Teams that need reproducible builds should prefer a pinned npm version and an internal package cache.
2. Authenticate with a model provider
Pi can use an API key from the environment or an existing subscription selected through /login. The example below shows the variable name without placing a credential in a script:
export ANTHROPIC_API_KEY="set this securely on your machine"
pi
For subscription-based login, enter Pi and run:
/login
Then select a provider and model. Pi maintains catalogs of tool-capable models. To refresh the catalog immediately, run:
pi update --models
3. Complete a small task with an observable result
Do not begin with a large refactor. Open a project with tests and ask Pi to perform a small, reversible task:
cd /path/to/project
pi "Find functions changed recently that lack tests. List the files and reasons. Do not modify files."
The first acceptance criteria should be what the agent read, what risks it identified, and whether it can explain the next step—not whether the prose sounds fluent. After the read-only behavior is acceptable, allow a tightly scoped edit:
pi -p "Add minimal tests for the functions that lack coverage. Run the relevant tests and list changed files and test results."
If you want to restrict capabilities, allow only read and search tools:
pi --tools read,grep,find,ls -p "Inspect this project's configuration and test coverage gaps"
You can also use --exclude-tools or --no-tools to make the boundary explicit. A command-line allowlist is easier to reproduce in scripts and CI than a prompt-only instruction such as “do not modify files.”
Compose the workflow with skills, extensions, prompt templates, and packages
Pi's composability comes mainly from four resource types.
Skills describe repeatable knowledge and procedures: how to run a database migration, check a framework's security settings, or generate tests according to team conventions. They can be placed in global or project directories and invoked with /skill:name.
Prompt templates turn frequent tasks into named entry points such as /review, /release-check, or /debug-api. This is easier to version-control than copying a long prompt into a chat window.
Extensions are TypeScript modules that can register custom tools, commands, keyboard shortcuts, event handlers, and UI. If you need to connect an internal deployment system, add an approval dialog, intercept tool calls, or register a custom provider, an extension is usually safer than modifying Pi's core.
Pi Packages are the distribution unit for sharing these resources. They can be installed from npm or git:
pi install npm:@foo/pi-tools
pi install npm:@foo/[email protected]
pi install git:github.com/user/repo@v1
pi list
pi config
Use -l for project-local installation under .pi; global resources live under ~/.pi/agent. A workflow can therefore be tested in one repository first and packaged for the wider team later.
The security boundary must not be overlooked. The official documentation states that Pi packages and extensions have full system access, and that skills may instruct the model to perform arbitrary actions. Review third-party source code, pin versions, restrict network and filesystem access, and keep sensitive credentials outside locations the agent can read. “Extensible” does not mean “trusted.”
Sessions and context: put repeatability first
Pi stores sessions as JSONL files with a tree structure. Each entry has an id and parentId, so /tree can move among branches in one session instead of creating unrelated copies of a conversation.
Useful commands include:
/session
/resume
/tree
/fork
/clone
/compact
/export
I recommend treating session management as part of the engineering process. Start each agent task with a stable name. At the end, record the goal, changed files, test results, and unfinished work. If the result is poor, return to the pre-edit branch rather than trying to reconstruct the old context from memory.
Context compaction should also be treated as an engineering operation. Pi supports manual and automatic compaction, but compaction is lossy; the full history remains in the JSONL file and can be revisited with /tree. For long-running agents, write a short structured state note before compaction with the current goal, verified facts, failed attempts, and next step.
Permissions and sandboxing: strong defaults require explicit boundaries
Pi runs with the permissions of the launching user by default. It does not provide built-in restrictions for filesystem, processes, network, or credentials. That is convenient for personal local development, but it is not an adequate boundary for CI, shared workstations, or sensitive repositories.
The official containerization guide describes three patterns. First, the Gondolin extension keeps Pi on the host but routes built-in tools and ! commands into a local Linux micro-VM. Second, Plain Docker places the entire Pi process in a local container. Third, OpenShell runs the entire process inside a policy-controlled sandbox with filesystem, process, network, credential, and inference controls.
A minimal Docker image can look like this:
FROM node:24-bookworm-slim
RUN apt-get update \\
&& apt-get install -y --no-install-recommends bash ca-certificates git ripgrep \\
&& rm -rf /var/lib/apt/lists/*
RUN npm install -g --ignore-scripts @earendil-works/pi-coding-agent
WORKDIR /workspace
ENTRYPOINT ["pi"]
Mount the project into /workspace and use a separate volume for agent settings and sessions. Do not mount the host's ~/.pi/agent unless the container is intentionally allowed to read the host's credentials and session history.
The core principle is to separate what the model believes it should do from what the process is technically allowed to do. System prompts, skills, and extensions govern the former. Containers, micro-VMs, filesystem permissions, network policy, and credential injection govern the latter.
Pi versus common agent abstractions
Pi, sub-agents, skills, agent teams, and workflows are not the same layer.
Pi is the harness for the execution loop and sessions. It manages models, messages, tools, events, and session history, but it does not decide the complete product workflow for you.
A skill is reusable knowledge or an operating procedure. It explains how to perform a kind of task, but it does not necessarily create another agent process.
A sub-agent is another agent instance, useful for isolated exploration or implementation work. Pi does not include sub-agents as a built-in feature; users can create them with tmux, extensions, or third-party packages.
An agent team is a higher-level collaboration model with roles, shared state, aggregation, and conflict handling. It requires more governance and observability than a single agent.
A workflow is the full coordination logic written as repeatable code: the input, the agents to start, the output validation, and the fallback path. Pi's value is that it supplies a small execution core that these higher-level abstractions can wrap.
Who should try Pi first?
I would start with teams that:
1. Already have mature CLIs, tests, and code review, and want to connect an agent to that engineering process.
2. Need to switch among model providers without coupling the model choice to the coding workflow.
3. Want skills, prompt templates, and extensions under version control as shared team packages.
4. Need to move the same agent gradually from an interactive terminal to CI, RPC, or an internal SDK-based tool.
5. Are willing to own permission, sandbox, supply-chain, and observability decisions.
If the first requirement is an immediately usable multi-agent workspace, or if the team cannot review third-party extensions and packages, Pi's minimal core may increase adoption cost. It is not a zero-decision product; it gives the decisions back to engineers.
My adoption advice: restrict the scope, then compose
Start with read-only tasks: summarize changes, search for vulnerability patterns, or identify test gaps. Use --tools read,grep,find,ls and store the output as a session or CI artifact.
Next, allow writes with a plan-first rule, a specified directory boundary, and a fixed test command. Use an extension or wrapper to record tool calls and changed files.
Only then add custom skills, packages, and model routing. Pin every package version, use the least-privilege execution environment available, and prepare a human handoff for failed tasks.
Finally, move stable workflows to RPC or the SDK. At that point it becomes worthwhile to invest in queues, retries, cost accounting, event storage, and multi-tenant isolation. Do not turn an interactive demo into an unattended deployment system on day one.
Conclusion: a small core can make an agent look more like an engineering system
Pi Agent Harness is not simply another chat interface that writes code. It demonstrates a useful layering model: pi-ai handles model providers, pi-agent-core handles state and events, the coding agent handles the CLI and sessions, extensions, skills, and packages handle workflow differences, and containers or micro-VMs provide the actual security boundary.
Once these responsibilities are separated, each layer can be tested, replaced, and reviewed. You can validate a small task in a terminal and later move the same runtime into RPC or an SDK. You can observe a model with read-only tools before granting write access. You can add provider, tool, and governance logic without forking the core.
If you are looking for the feature-richest AI coding platform, Pi may not be the shortest path. If you want an agent to become part of an existing engineering system rather than another isolated chat window, its minimalism is worth running and studying.
References