A CLI tool that leverages AI (Claude Code CLI or OpenAI API) to automatically review pull requests across GitHub and Azure DevOps, enforcing coding standards from configurable rule files.
- Multi-provider support via gitforge: GitHub and Azure DevOps
- Dual AI backend: Claude Code CLI (
claude --print) or OpenAI Chat Completions API - Rule-based reviews using Markdown files from guide (or any directory)
- YAML
frontmatterin rules for file-glob-based filtering (e.g.,paths: ["**/*.go"]) - Project-aware reviews: the reviewed repository's own
CLAUDE.mdis loaded automatically (on any provider) and forwarded to the AI as project-specific context, so the review honours the project's documented conventions — see Project Guidelines - Intent-aware reviews: the PR's title, branch names, description, and commit count are forwarded to the AI so it can judge whether the diff actually does what the author claims and flag undocumented scope creep — see Pull Request Context
- Inline and general PR comments posted back via gitforge
- Three modes: single PR review, batch review-all, and discover (list open PRs)
- Reviews each PR exactly once — subsequent pushes are no-ops. To request a re-review, post a PR comment that mentions
@code-guru(case-insensitive). On a re-review the bot acts as a reviewer who reads the existing conversation: it loads every prior bot inline thread plus every reply, classifies each asresolved/outstanding/outdated, posts one short reply nested inside each prior thread (via the provider'sReplyToThread, so the answer lands below the author's reply like a human reviewer rather than as a separate same-line comment), and auto-closes the threads it considers resolved (Azure DevOps thread statefixed). Net-new findings only land if the diff genuinely warrants one and it does NOT overlap a thread the bot already addressed — replacing the pre-existing failure mode where every re-review flooded the PR with reworded duplicates of every prior comment. The bot recognises its own prior comments by the built-incode-gurulogin shape, by self-detecting the account that posted its PR-wide review annotations on the PR, and by any identity listed inbot_identities(envCODE_GURU_BOT_IDENTITIES) — so re-reviews still read and resolve prior threads when the deployment posts under a service account
go install github.com/rios0rios0/codeguru/cmd/code-guru@latestOr build from source:
git clone https://github.com/rios0rios0/code-guru.git
cd code-guru
go build -o code-guru ./cmd/code-guru/Create a .code-guru.yaml file (searched in ., .config, configs, ~, ~/.config):
providers:
- type: 'github'
token: '${GITHUB_TOKEN}'
organizations:
- 'rios0rios0'
- type: 'azuredevops'
token: '${AZURE_DEVOPS_PAT}'
organizations:
- 'MyOrg'
ai:
backend: 'claude'
# When true, the bot also records a native pull request review (Approved /
# Changes Requested) on the platform's reviewer panel in addition to the
# text completion annotation. Defaults to true; set to false to opt out.
submit_native_review: true
# When false (the default), draft PRs are skipped entirely — set to true to
# opt back in.
review_drafts: false
# Times the AI backend is re-sampled per review when it returns a non-JSON or
# transient-error response before the review is marked failed. Defaults to 3;
# set to 1 to disable retries.
max_attempts: 3
# When true (the default), the reviewed repository's own root CLAUDE.md is
# loaded and forwarded to the AI as project-specific review context. Set to
# false to opt out.
project_guidelines: true
# When true (the default), the PR's description and commit count are fetched
# from the provider and forwarded to the AI as intent context (together with
# the title and branch names already in the prompt). Set to false to opt out.
pr_metadata: true
# Byte budgets for the two documents sent alongside the diff. The defaults
# (1 MiB / 64 KiB) assume a 1M-token context window — the guidelines budget
# is roughly a quarter of it. Lower max_guidelines_bytes on a small-window
# backend (e.g. a 128K-token model, or Anthropic with context_1m disabled).
max_guidelines_bytes: 1048576
max_pr_description_bytes: 65536
# When true (the default), a pull request too large for the model's context
# window is reviewed in batches instead of being skipped. Costs one model
# call per batch; max_review_batches caps how many a single review may use.
batch_large_reviews: true
max_review_batches: 20
claude:
binary_path: 'claude'
model: 'sonnet'
max_turns: 1
openai:
api_key: '${OPENAI_API_KEY}'
model: 'gpt-4o'
rules:
path: '${HOME}/Development/github.com/rios0rios0/guide/.ai/claude/rules'
categories: []
# Account identities the bot posts review comments under. Only needed when the
# deployment posts under a service account whose login does not start with
# `code-guru` (common on self-hosted Azure DevOps) — on a re-review the bot uses
# these to recognise its own prior threads and resolve them instead of
# re-posting. The bot also self-detects this from its own PR-wide review
# annotations, so this is optional. Override via CODE_GURU_BOT_IDENTITIES.
bot_identities:
- 'svc-codeguru@example.com'Tokens support three resolution strategies:
- Environment variable:
${GITHUB_TOKEN}expands from the environment - File path: if the resolved string is a file path, its contents are read
- Inline: literal token string
code-guru discover -c .code-guru.yamlcode-guru review-all -c .code-guru.yaml --dry-run
code-guru review-all -c .code-guru.yamlcode-guru review https://github.com/org/repo/pull/123
code-guru review https://dev.azure.com/org/project/_git/repo/pullrequest/456| Flag | Description |
|---|---|
-c, --config |
Path to config file (default: auto-discover) |
--backend |
AI backend: openai, claude, or anthropic |
--rules-path |
Path to rules directory |
--dry-run |
Run review without posting comments |
-v, --verbose |
Enable debug logging |
| Provider | Type Key | PR Comments | Inline Comments |
|---|---|---|---|
| GitHub | github |
Yes | Yes |
| Azure DevOps | azuredevops |
Yes | Yes |
| Backend | Key | How It Works |
|---|---|---|
| Anthropic | anthropic |
Calls the Anthropic Messages API directly via Go SDK |
| Claude Code | claude |
Invokes claude --print CLI as a subprocess |
| OpenAI | openai |
Calls the Chat Completions API with JSON response format |
Each review returns a verdict alongside comments:
| Verdict | Meaning |
|---|---|
approve |
No blocking issues, safe to merge |
request_changes |
Error-level issues that must be fixed |
comment |
Informational feedback only, not blocking |
The verdict is printed as VERDICT:<value> for machine parsing.
When trivial detection is enabled and CI has passed, PRs matching built-in adapters are handled without calling the LLM, saving tokens. In webhook mode, CI status is provided by the webhook event. In CLI mode, CI status detection is planned via gitforge's GetPullRequestCheckStatus() (not yet available).
There are two categories of trivial adapters:
These detect dependency update PRs and auto-approve them.
| Adapter | Matches When |
|---|---|
update-go |
Only go.mod, go.sum, CHANGELOG.md changed |
update-node |
Only package.json, lock files, CHANGELOG.md changed |
update-python |
Only pyproject.toml, requirements*.txt, CHANGELOG.md |
These detect version bump (release ceremony) PRs. If the repo contains an .autobump.yaml config file, the adapter validates that all version files declared in the config are present in the PR. Missing files result in a reject verdict.
| Adapter | Default Files | AutoBump Language Key |
|---|---|---|
bump-go |
CHANGELOG.md |
go |
bump-node |
package.json, CHANGELOG.md |
typescript |
bump-python |
*/__init__.py, CHANGELOG.md |
python |
| Adapter | Matches When |
|---|---|
docs-only |
Only *.md files changed (excluding a CHANGELOG.md-only change — see note) |
Note — a changelog-only change is a version bump. A PR that touches only
CHANGELOG.mdis the signature of a version bump / release ceremony, so it is matched exclusively by thebump-*adapters. Thedocs-onlyandupdate-*adapters decline it: the changelog may still accompany a documentation or dependency change, but it can never be the sole trigger. This means disabling thebump-*adapters reliably keeps version bumps out of trivial auto-merge, instead of having them silently fall through todocs-onlyorupdate-*.
Configure in .code-guru.yaml:
trivial:
enabled: true
adapters:
- 'update-go'
- 'bump-go'
- 'docs-only'
auto_merge: false # opt-in; true completes the PR after a trivial-approve verdict
merge_strategy: '' # 'merge' / 'squash' / 'rebase' — empty falls back to platform default
delete_source_branch: true # default on; deletes the source branch after an auto-merge (false keeps it)
auto_merge_allowed_authors: # restrict auto-merge to these PR authors; empty = any author
- 'autobump@example.com'
- 'autoupdate@example.com'Or via environment variables:
CODE_GURU_TRIVIAL_ADAPTERS=update-go,bump-go,docs-only
CODE_GURU_TRIVIAL_AUTO_MERGE=true
CODE_GURU_TRIVIAL_MERGE_STRATEGY=squash
CODE_GURU_TRIVIAL_DELETE_SOURCE_BRANCH=false # keep the source branch after an auto-merge (default deletes it)
CODE_GURU_TRIVIAL_AUTO_MERGE_AUTHORS=autobump@example.com,autoupdate@example.comauto_merge is intentionally off by default — it bypasses human review and merges cross-system, so the gate is "operator must explicitly opt in". A merge failure logs at warn and the trivial-approve verdict still stands; the PR author can complete the merge manually from the platform UI.
auto_merge_allowed_authors decides who is trusted to merge unattended, separately from triviality (which decides what is eligible). When non-empty, only PRs whose author matches an entry (case-insensitive) auto-merge — so a human's docs PR is approved but left for a human to merge, while a trusted automation account's PR (dependency bumps, version bumps, config refresh) merges on its own. Leaving it empty keeps the historical "any author" behaviour; that is not recommended together with policy bypass, because it force-merges every trivial PR — including a human's — past Required reviewers (the bot logs a warning in that case). A docs-only diff is not inherently safe: prose can carry a malicious install command, a poisoned package name, or a phishing link, and bypass means no human ever reviews it.
delete_source_branch is on by default: when a trivial PR auto-merges, its source branch is deleted once the merge completes — Azure DevOps removes it as part of completing the PR, GitHub deletes the head ref afterwards. It only has an effect when auto_merge actually fires (no merge, no branch to delete), and the deletion is best-effort — a failure is logged and never fails the merge, so the PR still shows as merged. Set it to false (or CODE_GURU_TRIVIAL_DELETE_SOURCE_BRANCH=false) to keep merged branches in place.
Code Guru can run as a long-lived HTTP server that receives webhook events from
GitHub Apps and Azure DevOps Service Hooks. Each event is enqueued onto a bounded
worker pool and the HTTP response returns immediately (202 Accepted), so the
review runs asynchronously and never blocks the sender.
code-guru serve --port 8080| Endpoint | Method | Auth | Notes |
|---|---|---|---|
/health |
GET | none | Liveness probe |
/webhooks/github |
POST | HMAC-SHA256 | Validates the X-Hub-Signature-256 header against server.webhook_secret. Acts on pull_request opened/synchronize/reopened. |
/webhooks/azuredevops |
POST | HTTP Basic | Username must be code-guru; password must equal server.webhook_secret. Acts on git.pullrequest.created/git.pullrequest.updated for active PRs. |
- GitHub -- the secret is the value configured on the GitHub App webhook
("Webhook secret" in the App settings). When
github_app.app_idandgithub_app.private_keyare configured the server signs an RS256 JWT and exchanges it for a per-installation access token, cached until 5 minutes before expiry. Withoutgithub_app.*the handler falls back to the configuredgithubPAT inproviders[]. - Azure DevOps -- ADO does not sign Service Hooks, so it uses HTTP Basic.
Configure the Service Hook with username
code-guruand password equal toserver.webhook_secret.
server:
port: 8080
webhook_secret: '${CODE_GURU_WEBHOOK_SECRET}'
workers: 8
queue_size: 100
shutdown_timeout: 30s
allowed_organizations:
- 'ExampleOrg'
allowed_projects:
- 'Platform'
github_app:
app_id: 123456
private_key: '${CODE_GURU_GITHUB_PRIVATE_KEY}'| Variable | Description | Default |
|---|---|---|
CODE_GURU_PORT |
HTTP port to listen on | 8080 |
CODE_GURU_WEBHOOK_SECRET |
Shared secret for HMAC (GitHub) and Basic Auth password (ADO) | |
CODE_GURU_SERVER_WORKERS |
Worker count draining the review queue | runtime.NumCPU() |
CODE_GURU_SERVER_QUEUE_SIZE |
Maximum buffered jobs before submitters get 503 Service Unavailable |
100 |
CODE_GURU_SERVER_SHUTDOWN_TIMEOUT |
Maximum drain time on SIGINT/SIGTERM |
30s |
CODE_GURU_SERVER_ALLOWED_ORGANIZATIONS |
Comma-separated allowlist of org/owner names (empty = allow all) | |
CODE_GURU_SERVER_ALLOWED_PROJECTS |
Comma-separated allowlist of ADO project names (empty = allow all) | |
CODE_GURU_GITHUB_APP_ID |
Numeric GitHub App ID | |
CODE_GURU_GITHUB_PRIVATE_KEY |
PEM-encoded RSA private key for the GitHub App |
docker build -t code-guru:latest .
docker run --rm -p 8080:8080 \
-e CODE_GURU_BACKEND=anthropic \
-e CODE_GURU_ANTHROPIC_API_KEY=$ANTHROPIC_API_KEY \
-e CODE_GURU_WEBHOOK_SECRET=$WEBHOOK_SECRET \
-e CODE_GURU_PROVIDER_TOKEN=$AZURE_DEVOPS_PAT \
-e CODE_GURU_SERVER_ALLOWED_ORGANIZATIONS=ExampleOrg \
code-guru:latestThe Dockerfile uses a multi-stage build (golang:1.26-alpine builder,
gcr.io/distroless/static-debian12:nonroot runtime) and runs as the
non-root user.
The image ships with a HEALTHCHECK directive that calls code-guru health against the local listener every 30 seconds. The health
subcommand can also be invoked directly for ad-hoc smoke tests:
code-guru health --url http://127.0.0.1:8080/health --timeout 4sExit codes: 0 on 200, 1 on any other status, network error, or
timeout.
When the serve controller starts inside a Kubernetes pod (detected
via the standard KUBERNETES_SERVICE_HOST env var that the kubelet
always injects), the dispatcher automatically swaps its default
per-pod in-memory webhook dedup for a cross-pod backend backed by
coordination.k8s.io/v1 Lease objects. This is required when the
deployment runs with replicas > 1 because Azure DevOps fires both
git.pullrequest.created and git.pullrequest.updated for every
new PR; without a shared lock the K8s Service round-robins one
delivery to each replica and the bot posts duplicate reviews. The
lease is named code-guru-{sanitised-key}-{hash} (the SHA-256 suffix
prevents collisions from the lossy character substitution) and is
created with leaseDurationSeconds: 900 (must exceed the bot's
maximum review wall-time so the takeover path never steals an
actively-held lease — the worst review observed was ≈8 minutes)
as freshness metadata —
Kubernetes does NOT auto-delete Lease objects when that duration
elapses, so the dedup contract relies on two explicit pieces of work:
- The owning pod
Deletes the lease after the worker finishes (success or failure), so a real follow-up push minutes later re-acquires immediately. - A subsequent webhook delivery whose
CreatereturnsAlreadyExistsruns a stale-lease takeover: itGets the holding lease, checks whetheracquireTime + leaseDurationSecondshas already passed, and if soDeletes the stale lease (with a UID precondition for race safety) and retriesCreate. This recovers from a pod crash mid-review — the maximum window during which a crashed lease blocks new work isleaseDurationSeconds.
The K8s API server's optimistic concurrency on Create is what
makes the dedup atomic across replicas: exactly one Create for a
given (namespace, name) succeeds and every concurrent Create
returns 409 AlreadyExists.
Required RBAC (apply once per namespace the bot runs in):
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
namespace: code-guru
name: code-guru-webhook-dedup
rules:
- apiGroups: ["coordination.k8s.io"]
resources: ["leases"]
verbs: ["get", "list", "create", "delete", "update", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
namespace: code-guru
name: code-guru-webhook-dedup
subjects:
- kind: ServiceAccount
name: code-guru
namespace: code-guru
roleRef:
kind: Role
name: code-guru-webhook-dedup
apiGroup: rbac.authorization.k8s.ioIf the RBAC is missing, the bot logs a Warn at startup and falls
back to the per-pod in-memory cache: still safe, but cross-pod
duplicates will not be suppressed. Outside a pod (local CLI runs,
unit tests) no Kubernetes API is contacted.
For CI/CD environments without a config file, all settings can be provided via CODE_GURU_* environment variables:
| Variable | Description | Default |
|---|---|---|
CODE_GURU_BACKEND |
AI backend | openai |
CODE_GURU_OPENAI_API_KEY |
OpenAI API key | |
CODE_GURU_ANTHROPIC_API_KEY |
Anthropic API key | |
CODE_GURU_RULES_PATH |
Path to rules directory | |
CODE_GURU_PROVIDER_TOKEN |
Git provider token | |
CODE_GURU_TRIVIAL_ADAPTERS |
Comma-separated adapter names | |
CODE_GURU_TRIVIAL_AUTO_MERGE |
Opt-in flag that completes the PR after a trivial-approve verdict | false |
CODE_GURU_TRIVIAL_MERGE_STRATEGY |
gitforge merge strategy (merge / squash / rebase); empty = default |
|
CODE_GURU_TRIVIAL_AUTO_MERGE_AUTHORS |
Comma-separated PR-author identities allowed to auto-merge; empty = any author (not recommended with bypass) | |
CODE_GURU_TRIVIAL_DELETE_SOURCE_BRANCH |
Deletes the source branch after a trivial auto-merge; set false to keep it |
true |
CODE_GURU_AI_SUBMIT_NATIVE_REVIEW |
Records a native review (Approved / Changes Requested) on the platform's reviewer panel; set to false to opt out |
true |
CODE_GURU_AI_REVIEW_DRAFTS |
When true, the bot reviews draft PRs as well — by default drafts are skipped |
false |
CODE_GURU_AI_MAX_ATTEMPTS |
Times the AI backend is re-sampled per review when it returns a non-JSON or transient-error response before the review is marked failed (1 disables retries) |
3 |
CODE_GURU_AI_PROJECT_GUIDELINES |
Loads the reviewed repository's own CLAUDE.md as project-specific review context; set to false to opt out |
true |
CODE_GURU_AI_PR_METADATA |
Fetches the PR's description and commit count as intent context for the AI; set to false to opt out |
true |
CODE_GURU_AI_BATCH_LARGE_REVIEWS |
Reviews a PR too large for the model's context window in batches instead of skipping it; set to false to get the "too large" notice and no review |
true |
CODE_GURU_AI_MAX_REVIEW_BATCHES |
Caps how many batches one batched review may consume; files left over are reported as unreviewed | 20 |
CODE_GURU_ANTHROPIC_CONTEXT_1M |
Requests the Anthropic 1M-token context window (context-1m-2025-08-07 beta) so larger PRs fit in one review pass; set to false for accounts/models that cannot use the beta |
true |
CODE_GURU_ANTHROPIC_REFUSAL_FALLBACK_MODEL |
Anthropic model to re-issue the review against when the primary model declines the content on content-safety grounds (stop_reason: refusal); empty disables the fallback |
(empty) |
CODE_GURU_BOT_IDENTITIES |
Comma-separated account identities the bot posts under (so re-reviews recognise its own prior threads); the built-in code-guru shape and self-detection apply when unset |
Rules are Markdown files loaded from the configured rules.path. Each file represents a rule category (e.g., security.md, golang.md, testing.md).
Rules can include YAML frontmatter with paths globs for file-specific filtering:
---
paths:
- "**/*.go"
---
# Go Conventions
Use `gofmt` for formatting...Universal categories (always included): architecture, ci-cd, code-style, design-patterns, documentation, git-flow, security, testing.
On top of the operator-configured rules, Code Guru reads the reviewed repository's own root CLAUDE.md — the file projects use to document conventions for AI tooling — and forwards it to the AI as project-specific review context. This works on every supported provider (GitHub and Azure DevOps) through the same file-access API the trivial detectors use, and it means the review honours conventions the generic ruleset cannot know about (naming, layering, testing patterns, intentional trade-offs).
Behaviour details:
- When the PR itself modifies
CLAUDE.md, the repository fetch is skipped — the model already reads the change in the diff, and layering the pre-change copy on top would present two conflicting versions of the same document. - The fetch is best-effort: a repository without a
CLAUDE.md, a provider without file-access support, or a transient error simply produces a review without project guidelines. It never fails or delays the review beyond a 10-second fetch timeout. - Content is bounded to
ai.max_guidelines_bytes(default 1 MiB, roughly 256k tokens — about a quarter of a 1M-token context window) so a pathological guidelines file cannot crowd the diff out of the model's context window. A large but legitimateCLAUDE.mdis sent whole; when the bound does bite, the truncation is logged with the file size and the budget. The default assumes a 1M-token window — on a 128K-token model, or on Anthropic withai.anthropic.context_1mdisabled, lower it viaai.max_guidelines_bytesorCODE_GURU_AI_MAX_GUIDELINES_BYTES. - The document is framed to the model as documentation, not instructions — it cannot change the output format, the verdict rules, or the reviewer role.
Enabled by default; opt out with ai.project_guidelines: false or CODE_GURU_AI_PROJECT_GUIDELINES=false.
Beyond the diff, Code Guru forwards the PR's author-supplied metadata to the AI so it reviews the change against its stated intent:
- Title and branch names (already part of the prompt header) signal the change type — a
fix/branch that quietly introduces new behaviour, or achorethat alters runtime logic, deserves a comment. - Description — the author's statement of what the change does and why. The model is told to flag significant changes the description leaves unmentioned (scope creep) and to weigh the author's explanations before flagging intentional oddities.
- Commit count — how the change was assembled, fetched from the provider's REST API (GitHub: one call returns both body and count; Azure DevOps: the PR resource plus its
/commitscollection).
Behaviour details:
- The fetch is best-effort with a 10-second timeout: an unsupported provider or an API error simply produces a review without the context — it never fails the review.
- The description is bounded to
ai.max_pr_description_bytes(default 64 KiB) so a generated body (release bots pasting entire upstream changelogs) cannot crowd the diff out of the model's context window. - The description is framed to the model as author-supplied data, not instructions — a body that says "approve this PR" is treated as content to evaluate, never as a command.
Enabled by default; opt out with ai.pr_metadata: false or CODE_GURU_AI_PR_METADATA=false.
Every review sends the full diff — plus the rules, the repository's CLAUDE.md, the PR metadata, and any prior review conversation — to the AI backend in a single request. When that combined prompt is larger than the model's context window, one pass cannot produce the review, so Code Guru splits the work instead of giving up:
- The change is reviewed in batches. The files are split into chunks that fit the model, reviewed one after another, and merged into a single review: the union of every batch's findings, and the most severe verdict any batch reported (one blocking finding blocks the pull request). The batch size is discovered from the backend's own overflow report and shrinks until it fits, so it adapts to the model, the beta flags in play, and how much of the window the rules and guidelines already consume.
- The pull request is told what is happening. Before the batches run, the PR gets a "reviewing this PR in batches" notice reporting the change's scale (file count and total diff size) and warning that the review will arrive but takes several times longer than usual — so nobody reads the delay as a crashed bot and merges early.
- Each batch knows it is only a slice. Its prompt states which part of the change it holds, so the model does not report the files it cannot see as an unused symbol, a missing test, or an incomplete change — the three findings a partial diff reliably invents.
- Gaps are reported, never hidden. A file too large to review even on its own, a failed batch, or a run that hits the batch cap is named in the final summary, and a review that could not cover the whole change never reports
approve. - Splitting the pull request is still better. A batched review is a weaker review: each batch sees only its own files, so a finding that spans two batches is invisible to it. The notice keeps the "split the change / exclude generated, vendored, and lock files" guidance for that reason.
- Cost control. Batching costs one model call per batch.
ai.max_review_batches(CODE_GURU_AI_MAX_REVIEW_BATCHES, default20) caps a single review;ai.batch_large_reviews: false(CODE_GURU_AI_BATCH_LARGE_REVIEWS=false) turns the fallback off entirely, restoring the previous behaviour where an oversized pull request gets a "too large for the AI model's context window" annotation and no review. - No wasted retries. A prompt-too-long failure is deterministic (the prompt is identical on every attempt), so the retry budget is skipped entirely for this class of failure — the batch split is the remedy, not a re-sample.
- A larger window on Anthropic. The Anthropic backend requests the 1M-token context window (
context-1m-2025-08-07beta) by default, so pull requests up to roughly five times larger fit in one pass before batching is needed at all. For prompts under 200K tokens the beta is a no-op; very large prompts may incur Anthropic long-context pricing. Opt out withai.anthropic.context_1m: falseorCODE_GURU_ANTHROPIC_CONTEXT_1M=falseon accounts or models that cannot use the beta.
The AI models Code Guru uses run content-safety classifiers that sometimes decline to review a change — most often security-related code (offensive-security tooling, exploit or credential-handling code, cryptography) and, for some models, certain biology content. On Anthropic this arrives as stop_reason: "refusal" (an HTTP 200 with no review), on OpenAI as a content_filter finish reason. Code Guru handles this explicitly rather than reporting a generic error:
- A clear, reassuring notice. The PR gets a "the AI model's content-safety system declined this pull request" annotation that names the cause, surfaces the provider's policy category when reported, and makes clear this is a limitation of the AI reviewer's safety filters — not a judgment that the change is malicious or a defect in the code. It points at the real remedies (request a human review; an operator can switch the model or backend) and does not suggest retrying, since the same content is declined the same way.
- No wasted retries. A refusal is deterministic on the same content, so the retry budget is skipped for this class.
- An optional fallback model. Because safety-classifier coverage varies by model, an operator can set
ai.anthropic.refusal_fallback_model(CODE_GURU_ANTHROPIC_REFUSAL_FALLBACK_MODEL) so the Anthropic backend re-issues the review once against a model that handles the content when the primary model refuses. It is off by default; a fallback that also refuses surfaces the original refusal.
Contributions are welcome. See CONTRIBUTING.md for guidelines.
See LICENSE for details.