OpenCodeReview: Making AI Code Review More Controllable with a Deterministic × Agent Architecture
# OpenCodeReview: Making AI Code Review More Controllable with a Deterministic × Agent Architecture
AI code review is not only about whether a model can understand source code. A useful review must consistently inspect the intended files, anchor findings to the right lines, and balance noise against cost. Alibaba's open-source [OpenCodeReview](https://github.com/alibaba/open-code-review) takes a practical approach: deterministic engineering handles workflow constraints, while an agent handles judgment and context retrieval.
This makes OpenCodeReview more than a CLI that sends a diff to an LLM. It can run against a local workspace, connect to CI/CD and GitHub Actions, and integrate with coding agents. This article examines its architecture, operating modes, and adoption limits.
## Verified snapshot
At 2026-09-17 23:06 UTC, the GitHub API reported 34,626 stars and 2,461 forks for OpenCodeReview. The latest push was on 2026-09-17. The project is written in Go, uses the Apache-2.0 license, and publishes an npm package together with cross-platform release binaries. These numbers are a snapshot and change over time.
- GitHub:
- Releases:
- npm package:
## The boundary between engineering and the agent
A natural-language review skill often asks an agent to read a diff, find bugs, and report file names and line numbers. That is a fast prototype, but large changesets can expose incomplete coverage, position drift, and quality that changes with small prompt variations.
OpenCodeReview divides the workflow into two layers:
- **Deterministic engineering** selects files, filters exclusions, groups related files, and matches review rules.
- **The agent** makes dynamic decisions inside those prepared review units. It can read full files, search the repository, inspect other changed files, and produce findings.
The model is therefore not solely responsible for remembering the complete process. It focuses on semantic understanding, context retrieval, and judgment about whether a finding represents a real risk.
## Four engineering guardrails
### Precise file selection
The tool uses Git diff and repository operations to establish the review scope before the agent runs. This reduces the chance that a large changeset is silently sampled by the model and keeps exclusion logic testable.
### Group related files
Related files can be placed in one review unit, such as localized properties files. Each unit is handled by an isolated sub-agent. This divide-and-conquer strategy preserves useful context without forcing every file into one unbounded prompt.
### Match rules to file characteristics
Review rules are matched according to file characteristics and support multiple languages and configuration formats. Compared with putting every rule into natural-language guidance, template-based matching is easier to reproduce and target.
### Separate positioning and reflection
Finding a problem is different from placing a comment on the correct diff line. OpenCodeReview has independent comment-positioning and comment-reflection components, so correctness of the finding and location of the comment do not have to be delegated to one model response.
These mechanisms do not make AI review formal verification or guarantee zero false positives. They make unstable workflow steps more observable, testable, and repeatable.
## Three operating modes
### Review Git changes
Install the npm launcher:
```bash
npm install -g @alibaba-group/open-code-review
```
Configure an LLM provider and model:
```bash
ocr config provider
ocr config model
```
Then choose a review scope:
```bash
ocr review
ocr review --from main --to feature-branch
ocr review --commit abc123
```
Interrupted work can be listed and resumed:
```bash
ocr session list
ocr review --from main --to feature-branch --resume
```
### Scan files or directories
For a repository with no useful Git diff, use `ocr scan`:
```bash
ocr scan
ocr scan --path internal/agent
ocr scan --resume
```
`review` is change-centered; `scan` is content-centered. In a team workflow they should have different budgets and schedules: PR review should be low latency, while a whole-codebase scan is better suited to a scheduled or manual run.
### Delegate to an existing coding agent
Delegation mode lets a coding agent perform the review while OCR still handles file selection and rule resolution:
```bash
ocr delegate preview
ocr delegate rule src/main.go src/handler.go
```
This can give Claude Code, Codex, Cursor, Kimi Code, OpenCode, or another skill-compatible agent a shared review foundation instead of requiring every team to maintain a separate prompt.
## From CLI to CI/CD
The repository's `action.yml` provides a GitHub Action for PR review. Its inputs cover the LLM endpoint, model, authentication, output language, timeout, concurrency, custom rules, and JSON or stderr artifacts.
A sensible rollout has two stages:
1. **Observation:** emit JSON and preserve artifacts, then measure false positives, latency, and token use.
2. **Gating:** after the configuration is stable, publish selected findings as inline comments according to severity while retaining human review.
The tool should not become a merge blocker on day one. The README describes a precision-oriented trade-off with lower recall. That makes it a noise-reduction layer, not a replacement for tests, static analysis, or experienced reviewers.
## How to read the benchmark
The README describes AACR-Bench as containing 50 popular open-source repositories, 200 real pull requests, 10 programming languages, and 1,505 annotated issues cross-validated by more than 80 senior engineers. The project reports higher precision and F1, roughly one ninth of the token usage, and faster reviews than a general-purpose agent using the same underlying model. It also reports lower recall as a deliberate noise trade-off.
These are project-reported results, not an independent reproduction in this article. Teams should check whether the data resembles their own languages and repositories, whether model and prompt settings match, and whether precision and recall correlate with actual remediation cost. A small replay set of previously reviewed pull requests is a better adoption gate than a benchmark number alone.
## Implementation and extension points
The repository is split into several meaningful areas:
- `cmd/opencodereview`: Cobra CLI, provider configuration, review, scan, delegate, sessions, and output.
- `internal/diff`: Git diff parsing, hunks, and comment relocation.
- `internal/agent`: review-unit grouping, file selection, and agent execution.
- `internal/scan`: batches, budgets, providers, and resume behavior for full-file scans.
- `internal/llm`: OpenAI-compatible, Anthropic, and Bedrock clients with retry boundaries.
- `internal/mcp`: MCP client and external-tool integration.
- `internal/session`: review sessions, history, and result persistence.
- `internal/telemetry`: OpenTelemetry metrics and traces.
The `go.mod` file targets Go 1.25.5 and directly depends on Anthropic, OpenAI, MCP, Cobra, and OpenTelemetry SDKs. This separation gives engineering teams clear extension points for providers, rules, file allowlists, output, and telemetry without putting all behavior inside a prompt.
## Adoption limits and risks
### LLM configuration and cost remain real prerequisites
The default mode requires a provider and model, and review content is sent to the configured LLM endpoint. API keys should be kept in a secret manager rather than a repository or CI log. Teams must decide whether source code, configuration, and personal data may be sent to that provider. Delegation mode changes the execution path, but it does not remove the need to understand the agent's data flow and permissions.
### High precision is not complete coverage
A low-noise tool may miss real issues. Security-sensitive changes, database migrations, and authorization boundaries should also use compilers, tests, SAST, dependency scanning, and human review.
### Budget and latency need measurement
Grouping and concurrency can improve large-review stability, but each unit may still invoke several rounds of tool use. Track file count, tokens, wall-clock time, retries, and human-confirmed findings before tuning concurrency, timeout, and token budgets.
### The project is evolving quickly
The latest release observed in this research was v1.12.5, and the repository was still receiving recent updates. Pin the CLI version in production and make upgrades reversible.
## Conclusion
OpenCodeReview is notable not simply because it generates LLM review comments, but because it separates workflow correctness from semantic judgment. Engineering code owns file selection, grouping, rules, positioning, and sessions; the model focuses on context and dynamic decisions.
That is a useful signal for AI engineering: when an agent enters CI/CD, ask not only whether the model is capable, but also which steps can be constrained, tested, and replayed deterministically. OpenCodeReview is a strong Go case study for moving AI review beyond a one-off demo. Whether it should block pull requests still depends on a team's own replay set, cost data, and human validation.
## Sources
- [OpenCodeReview README](https://github.com/alibaba/open-code-review/blob/main/README.md)
- [GitHub repository metadata API](https://api.github.com/repos/alibaba/open-code-review)
- [Review CLI source](https://github.com/alibaba/open-code-review/blob/main/cmd/opencodereview/review_cmd.go)
- [Scan CLI source](https://github.com/alibaba/open-code-review/blob/main/cmd/opencodereview/scan_cmd.go)
- [Agent grouping source](https://github.com/alibaba/open-code-review/blob/main/internal/agent/grouping.go)
- [Git diff relocation source](https://github.com/alibaba/open-code-review/blob/main/internal/diff/relocation.go)
- [GitHub Action definition](https://github.com/alibaba/open-code-review/blob/main/action.yml)
- [Go module manifest](https://github.com/alibaba/open-code-review/blob/main/go.mod)