- AI is now writing pull requests faster than humans can check them. The code usually works, so reviewers naturally stop scrutinising it as closely. That is exactly when the rare bad change can slip through.
- The answer isn't to "review harder." It's to redesign the pipeline itself: know with certainty which changes came from a machine, verify that claim cryptographically, and reserve meaningful human attention for the parts of the codebase where a mistake would actually matter.
AI coding tools have changed in the last two years. They used to suggest single lines of code while a developer typed. Today, tools like GitHub Copilot's coding agent or Claude Code work differently: you give them a task, and they write the code, open a pull request, run the tests, and even respond to review comments on their own. A pull request, or PR, is the standard way changes enter a codebase: someone proposes a change, someone else reviews and approves it, and only then does it get merged.
This shift is why AI code review has become one of the most urgent conversations in modern software development. The review process built for human developers is starting to break under the weight of AI-generated code, and here is why, along with how to fix it in practice.
The bottleneck: review queues are growing faster than reviewers
A developer might open two or three pull requests a day. An AI agent can open ten, and it never gets tired. So the volume of code changes waiting for human review grows several times over, while the number of human reviewers stays the same. Data from teams adopting AI coding agents backs this up: pull requests are 18% larger on average since AI adoption, and 27% of AI-generated pull requests combine multiple unrelated tasks into one PR — both of which make each individual PR review slower and more cognitively demanding.
Here is the uncomfortable part: AI-generated code is usually correct. That is exactly what makes it risky for human code review. When nine out of ten agent PRs are fine, reviewers quickly learn that approving without reading is almost always the right call. Psychologists call this automation bias. Within weeks, "review" becomes clicking Approve on autopilot. And the one PR in ten with a genuine security issue, an edge case bug, or a race condition sails through with an official human approval attached.
Large pull requests make this worse on their own. They take exponentially longer to review than small ones. They force constant context switching. They raise the cognitive load on whoever is on the hook for finding logic errors, input validation gaps, and architectural concerns buried in hundreds of changed lines. Delayed reviews require repeated reminders. That slows down the entire system, not just the one PR.
The fix: 3 steps to rebuild code review for AI agents
The fix is not telling people to "review more carefully". Attention does not scale. The fix is rebuilding the process around three questions:
- Who exactly wrote this change?
- Can we prove it?
- Does this particular change actually need human eyes?
Step 1: Give every agent its own identity
In most companies today, AI agents commit code either through a shared "bot" account or, worse, using the credentials of the developer who launched them. Both destroy the audit trail: you can no longer tell what the human wrote and what the machine wrote, and you cannot switch the agent off without also blocking the person.
Treat an agent like a new team member with a very limited badge: its own account, its own credentials, and a short list of permissions. On GitHub, the right mechanism is a dedicated GitHub App (not a personal access token), because Apps have scoped, short-lived tokens and show up in history as a distinct author like my-coding-agent[bot]. The permissions you grant should be minimal:
# GitHub App permissions for a coding agent (app manifest)
default_permissions:
contents: write # push to its own feature branches
pull_requests: write # open and update PRs
issues: read # read tasks it is assigned
# everything else stays "none" on purpose:
# no administration, no secrets, no workflows, no environments
Then make sure nobody, human or bot, can bypass review by pushing straight to the main branch. This is one Terraform resource:
resource "github_repository_ruleset" "main_protection" {
name = "protect-main"
repository = github_repository.app.name
target = "branch"
enforcement = "active"
conditions {
ref_name {
include = ["~DEFAULT_BRANCH"]
exclude = []
}
}
rules {
pull_request {
required_approving_review_count = 1
dismiss_stale_reviews_on_push = true
}
required_signatures = true # see step 2
}
}
The result: if the agent is ever compromised or simply goes off the rails, the worst it can do is open a bad PR that still has to pass the gates below. Its blast radius is small.
Step 2: Make authorship provable, not assumed
Git usernames are just text; anyone can put any name on a commit. So the next question is: how does the pipeline know a commit really came from the agent it claims to be?
The answer is signed commits. A signature is a cryptographic stamp proving which account produced a commit, like a sealed envelope instead of a plain postcard. The required_signatures rule above already rejects unsigned commits; what remains is making signing painless. Sigstore's gitsign does this without asking anyone to manage GPG keys, because it signs using the identity people and bots already log in with:
# One-time setup in the repo (humans and agents alike)
brew install gitsign
# or: go install github.com/sigstore/gitsign@latest
git config commit.gpgsign true
git config tag.gpgsign true
git config gpg.x509.program gitsign
git config gpg.format x509
From now on, git log --show-signature tells you, cryptographically, who authored what. This matters because it turns "was this written by an AI?" from a guess into a fact that software can check, and once your pipeline can check it, your pipeline can act on it.
Step 3: Send human attention where it actually matters
With identity and signatures in place, you can stop reviewing every change with the same intensity and start routing PRs by risk.
First, define which parts of the codebase are always high-risk, using a CODEOWNERS file. Whatever the automated checks say, changes to these paths require approval from the listed humans:
# .github/CODEOWNERS
/infra/ @platform-team
/charts/ @platform-team
/.github/ @platform-team # the pipeline itself
/src/auth/ @security-champions
/src/payments/ @payments-leads
Second, add a pipeline gate that combines steps 1 and 2: if a change was authored by the agent and touches a protected path, block auto-merge and demand a human, explicitly:
# .github/workflows/agent-gate.yml
on: pull_request
jobs:
route-by-risk:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with: { fetch-depth: 0 }
- name: Block agent changes to protected paths
run: |
AGENT="my-coding-agent[bot]"
PROTECTED='^(infra/|charts/|\.github/|src/(auth|payments)/)'
FILES=$(git log origin/${{ github.base_ref }}..HEAD \
--author="$AGENT" --name-only --pretty=format:)
if echo "$FILES" | sort -u | grep -qE "$PROTECTED"; then
echo "::error::Agent-authored change in a protected path. Human review is mandatory."
exit 1
fi
Third, let the boring changes flow. An agent PR that touches no protected paths, passes the full test suite, and carries valid signatures can merge automatically, with humans spot-checking a sample after the fact:
# Enable auto-merge on low-risk agent PRs (e.g. from a scheduled job)
gh pr merge "$PR_NUMBER" --auto --squash
The practical effect is a trade every engineering leader should like: reviewers go from skimming thirty PRs a day to properly reading the five that can actually hurt the business. The pipeline handles the rest, and it can be trusted to do so precisely because of the groundwork in steps one and two.
The takeaway
AI agents are already writing code in most organisations, whether the process acknowledges it or not. The pipelines those organisations run were designed on a quiet assumption: that every author is a human employee. That assumption is now false, and pretending otherwise means accumulating invisible risk with every merge.
The good news is that fixing it requires configuration, not new products: a scoped identity per agent, mandatory signed commits, and risk-based routing defined in a handful of files like the ones above. Three decisions, and your review process starts working with the robots instead of being quietly overwhelmed by them.
FAQs
No. AI code review works best as an advisory layer that summarises changes, flags likely issues, and runs mechanical checks while a human stays accountable for final approval. Treating AI review as the primary approval mechanism removes the accountability that a code review process is meant to provide.
Code produced by AI is generally syntactically correct and passes the tests, causing reviewers to be more likely to approve it without examining it carefully; this is referred to as automation bias. In contrast, code written by humans usually contains a greater number of small and obvious errors which attract the reviewer's attention, while code generated by AI is more prone to concealing subtle logical errors, security vulnerabilities, or the lack of handling of edge cases even when the code itself appears clean.
Yes, within limits. AI-assisted code review tools are good at mechanical checks, type checking, input-validation gaps, obvious race conditions, and summarising large diffs and can highlight issues a tired human reviewer might scroll past. They're less reliable on architectural concerns, business logic correctness, and false positives, which is why they work best paired with, not instead of, human reviewers.
Large pull requests take longer to review, increase context switching, and raise the odds that a reviewer skims instead of reading. When a coding agent bundles multiple tasks into one PR, it also becomes harder to isolate which part of the change caused a bug, slowing both PR review and rollback.
It will vary according to the tool and the size of the pull request, but when artificial intelligence is used to assist with the review it usually accelerates the more mechanical aspects of the process by summarising the PR descriptions, identifying possible problems, and cutting down the time needed to deal with simple changes so that human reviewers can direct their limited attention towards the more risky and architecturally important changes.
No, not for anything that will go into production. These tools can open, test, and even update pull requests by themselves. Because of this autonomy, it is even more important to have a human review process, limit agent permissions, and use signed commits once an agent can merge code without a person writing every line.
Besides summarising the code changes, a good PR description from a coding agent should include the original task or prompt, any assumptions made, and which parts of the codebase were changed. This gives human reviewers the information they need to quickly check the intent, instead of having to figure it out from the code alone.
Related Insights
Inconsistencies may occur.
The breadth of knowledge and understanding that ELEKS has within its walls allows us to leverage that expertise to make superior deliverables for our customers. When you work with ELEKS, you are working with the top 1% of the aptitude and engineering excellence of the whole country.
Right from the start, we really liked ELEKS’ commitment and engagement. They came to us with their best people to try to understand our context, our business idea, and developed the first prototype with us. They were very professional and very customer oriented. I think, without ELEKS it probably would not have been possible to have such a successful product in such a short period of time.
ELEKS has been involved in the development of a number of our consumer-facing websites and mobile applications that allow our customers to easily track their shipments, get the information they need as well as stay in touch with us. We’ve appreciated the level of ELEKS’ expertise, responsiveness and attention to details.