BETAFor engineering teams merging Claude Code pull requests

Know which agent PRs need you. And which don't.

Clear the agent PR queue without reading every line.

Write your merge rules in one short YAML file. Verity checks every PR against them and against what the agent actually did, then marks which PRs need a reviewer and why.

verity · acme/mobile-api · open agent PRsas Verity marks them
NEEDS A REVIEWER · 2
#231Show order notes in the app● 2 worth a look✕ 1 rule failed! 2 rule warningsmobile-api-contract: the API contract changed
#229Retry failed webhooks with backoff● 1 worth a look"Retries are tested" has no evidence
NOTHING FLAGGED · 3
#228Paginate GET /orders● nothing flaggedrules passed · tests ran after the last change
#226Fix date parsing in the CSV import● nothing flaggedrules passed · every claim has evidence
#224Rename OrderState to OrderStatus● nothing flaggedrules passed · tests ran after the last change
PR #231 · opened by the agent

Show order notes in the app

✕mobile-api-contract OrderDto.cs · answered yesfails
!page-looked-at Detail.cshtml · unsurewarns
!migrations-changed 0042_order_notes.sql · changedwarns
✓tests-after-edit OrderDto.cs · a passing test after the last editpassed
@@ OrderDto.cs:12 @@
-    public string Status { get; init; }
+    public OrderStatus Status { get; init; }
PR description, written by the agent"The API contract is unchanged."not an input
✕ Verity · failingoverride: a reviewer who isn't the author
2 need a reviewer · 3 can merge on a quick lookevery PR, every push
WORKS WITHClaude Codeprivate GitHub reposbranch protectionyour AI code reviewer
01THE PROBLEM

Agents open PRs faster than your team can review them. So reviewers read every line and the queue grows, or skim the agent's own write-up and approve.

4.6×

longer before a reviewer picks up an AI-assisted pull request than a manual one.

LinearB 2026 benchmarks, 8.1M PRs ↗
61%

of 33,596 agent-written pull requests had no recorded human review.

AIDev dataset, ACM TOSEM ↗
36–76%

of agent failures were the agent reporting success on work that had failed.

Characterizing False Success in LLM Agents ↗
02WHAT CHANGES

Start the review where it's needed. The same queue, reviewed two ways.

WITHOUT VERITY
  • Five open agent PRs, and nothing in the queue says which one to read first.
  • #231's description says the API contract is unchanged. The reviewer believes it, or reads every file to check.
  • The Status type change ships, and the mobile apps break silently.
WITH VERITY
  • The queue shows two PRs that need a reviewer and three with nothing flagged.
  • #231's comment says mobile-api-contract failed on OrderDto.cs, where Status changed type.
  • The reviewer starts at that file. The other three get a quick look.
03A FLAGGED PR

When a PR needs you, it says why. PR #231 from the queue above. Pick a rule.

PR #231 · Show order notes in the app

select a rule, or a step on the session map
acme/mobile-api · 4 rules
FAILS THE CHECK · 1
WARNINGS · 2
PASSED · 1
promptreadeditrunclaim146811131517
RULE · FAILS THE CHECK · A QUESTION ABOUT THE DIFF
✕ mobile-api-contract
The mobile apps break silently on contract changes.
files: ["src/Api/Mobile/**"]
ask: "Does this change alter the API contract:
  fields, types, status codes or routes?"
about: diff
expect: no
severity: fail
MATCHED FILES
src/Api/Mobile/OrderDto.csanswered yes
THE PR DESCRIPTION · NOT AN INPUT
The description says the contract is unchanged. Verity never reads it. The question saw the diff, where Status changed type:
@@ OrderDto.cs:12 @@
-    public string Status { get; init; }
+    public OrderStatus Status { get; init; }
04GO DEEPER

Want the details?

05FAQ

What engineering leads ask first.

We already pay for an AI code reviewer. Why add this?

Your code reviewer reads the diff, and often the PR description the agent wrote. Verity checks your team's rules against the diff and the agent's session, and never reads the description. It doesn't judge code quality or correctness, so keep both.

What happens when a rule fires and it shouldn't have?

A reviewer who isn't the PR's author adds the label verity-override:<rule>, and the PR shows who did. If a rule fires too often, it's one entry to change in .verity/rules.yml.

Where do our session transcripts go?

Secrets are redacted on the developer's machine, then again on upload. Redacted transcripts are stored encrypted, separately for each org, and deleted after 30 days. Or keep them in your own bucket. A report opens only for people who can read the repository on GitHub.

Will it block our merges?

Only a fail rule that fires, and only if you make the check required. Starter rules are warnings, so nothing blocks until you choose.

What does it cost?

Free for design partners while we set pricing with them. It works with Claude Code on private GitHub repos today.

Spend review time on the PRs that need it.

Sign in with GitHub and add Verity to one repo. Starter rules are warnings, so nothing blocks until you choose.