Reviewing AI-Generated Code: A Developer's Checklist
AI coding agents can produce a working feature in five minutes. That speed is the whole point. It's also the danger. When an agent generates 200 lines of code in thirty seconds, the temptation to glance at it, run the tests, and commit is strong. And the more you use agents, the stronger that temptation gets, because most of the time the code works. Until it doesn't.
The problem isn't that AI-generated code is bad. It's that it fails in ways you're not trained to catch. You've spent years developing instincts for human mistakes: off-by-one errors, copy-paste bugs, forgotten null checks. AI agents make a different class of mistakes, and your existing review instincts won't reliably catch them. You need a systematic process.
This guide is that process. A concrete checklist and a set of review patterns designed specifically for AI-generated code. Bookmark it. Reference it every time you review agent output. It takes five minutes and will save you from the kind of bugs that make it to production because the code "looked fine."
How AI Code Fails Differently Than Human Code
Before you can review effectively, you need to understand the failure modes. AI-generated code breaks in predictable patterns that are distinct from human mistakes.
Plausible but Wrong
This is the most dangerous pattern. The code looks correct. It compiles. It passes a quick visual scan. But there's a subtle logic error that only surfaces under specific conditions.
A human writing a sorting function might accidentally use > instead of >=. An agent is more likely to implement an algorithm that's conceptually wrong but syntactically perfect. For example, it might implement a rate limiter that resets the counter on every request instead of on a time window, and the code will read cleanly because the variable names, structure, and comments all describe a rate limiter.
Hallucinated APIs
Agents sometimes call methods that don't exist. Not random gibberish -- plausible method names that feel like they should be part of a library. response.getJSON() instead of response.json(). path.joinSafe() instead of path.join(). array.unique() instead of [...new Set(array)].
This also manifests as wrong function signatures: correct method name, wrong parameter order or wrong number of arguments. The code won't crash until that specific code path executes.
Over-Engineering
Agents love abstractions. Ask for a function that formats a date, and you might get a DateFormatterFactory with a strategy pattern, three interfaces, and a plugin system. This isn't just unnecessary complexity -- it's code you have to maintain, test, and reason about forever.
Watch for error handling that covers impossible cases, too. Try-catch blocks around operations that can't throw, fallback logic for scenarios that will never occur in your application, and defensive null checks on values that are guaranteed to exist by your schema.
Stale Patterns
Agents are trained on code from across time. They might use var instead of const, componentWillMount instead of useEffect, or the request npm package that was deprecated years ago. The code works, but it doesn't match modern best practices or your project's conventions.
Missing Edge Cases
The happy path works perfectly. The error paths are untested or broken. An agent building a file upload endpoint will handle the successful upload beautifully and completely ignore what happens when the disk is full, the file is zero bytes, or the user uploads a 10GB file.
Context Drift
The agent's change works in isolation but breaks assumptions elsewhere. It renames a function that's imported in twelve other files. It changes a return type that a downstream component depends on. It adds a required field to a database model without creating a migration. Each change is locally correct but globally broken.
The Review Checklist
Print this out. Tape it to your monitor. Run through it every time you review AI-generated code.
1. Does it actually solve the problem you described?
Re-read your original prompt. Then re-read the code. Agents sometimes solve an adjacent problem instead of the one you asked for. They might implement "search" when you asked for "filter," or add client-side validation when you needed server-side.
2. Did it modify only the files you expected?
Run git diff --stat before anything else. If you asked for a change to one file and the agent touched five, that's a red flag. Understand why every file was modified.
3. Are there any new dependencies?
Check package.json, requirements.txt, Cargo.toml, or whatever your project uses. New dependencies mean new supply chain risk, new bundle size, and new things to keep updated. If the agent added a library, ask yourself: is this necessary, or could it be done with what we already have?
4. Do the imports reference real packages and modules?
This catches hallucinated APIs. For every new import, verify the package exists and the imported symbol is a real export. A quick check:
# Node.js - verify a package exists
npm info <package-name>
# Python - verify a package exists
pip show <package-name>
For internal imports, verify the file path and exported function name actually exist.
5. Does the error handling cover realistic failure modes?
Look at every try-catch block and ask: what errors can actually happen here? Is the catch block doing something useful, or is it swallowing the error silently? Are there network calls, file operations, or database queries without any error handling at all?
6. Are there hardcoded values that should be configurable?
API URLs, timeout values, retry counts, file size limits, pagination sizes. Agents love to hardcode these because it makes the code work immediately. In production, you need them in environment variables or configuration files.
7. Does it follow the project's existing patterns and conventions?
If your project uses a service layer pattern, the new code should too. If you use camelCase for variables, the agent shouldn't introduce snake_case. If your error responses have a specific shape, new endpoints should match. This is one of the most common agent failures: the code works but doesn't belong in your codebase.
8. Are tests included? Do they test the right things?
"Tests pass" is not the same as "tests are good." Check what the tests actually verify. Are they only testing the happy path? Are they testing implementation details instead of behavior? Are they using mocks so aggressively that they're not testing anything real?
9. Is there dead code or unnecessary comments?
Agents frequently leave behind commented-out code, unused helper functions, or verbose comments that explain what the code does line-by-line. // increment counter by 1 above counter += 1 adds nothing. Strip it out.
10. Security: is user input validated and sanitized?
Check every place where user input enters the system. Is it validated? Is it sanitized before being used in SQL queries, HTML output, file paths, or shell commands? Agents sometimes build the functional logic perfectly and skip the security layer entirely.
Quick Review vs. Deep Review
Not every change needs the same level of scrutiny. Calibrate your review depth to the risk.
Quick Review (~2 minutes)
Use this for:
- Small changes (under 30 lines)
- Modifications to well-tested areas of the codebase
- Low-risk changes: UI tweaks, copy updates, logging additions
- Changes where the test suite provides strong coverage
Quick review process:
- Run
git diff --statto confirm scope - Scan the diff for obvious issues
- Run the test suite
- Check for unexpected file changes
- Commit if everything looks clean
Deep Review (~10-15 minutes)
Use this for:
- New features or new files
- Security-sensitive code (authentication, authorization, input validation)
- Database changes (schema, migrations, queries)
- API changes (new endpoints, changed contracts)
- Anything touching payments, user data, or permissions
- Changes over 100 lines
Deep review process:
- Read the original prompt to re-ground yourself on the intent
- Run
git diff --statto understand scope - Read every line of the diff
- Trace data flow from input to output
- Check edge cases: what happens with empty input, null values, maximum values?
- Verify error handling at every boundary (network, database, file system)
- Check that tests cover failure modes, not just success
- Run the test suite
- If there are database changes, review the migration carefully
- If there are API changes, verify backwards compatibility
The five minutes you spend on a deep review save hours of debugging in production.
Using Git Diff Effectively for AI Review
Git is your primary review tool. Learn to use it well.
See what changed and by how much
git diff --stat
This shows you every modified file and how many lines changed. It's the first thing to look at. If you expected changes to two files and the agent touched eight, stop and investigate before reading any code.
Read the actual changes
git diff
For staged changes, use git diff --cached. Read the diff line by line. Green lines (additions) are new code the agent wrote. Red lines (deletions) are code the agent removed. Pay special attention to deletions -- make sure the agent didn't remove something important.
Check for whitespace issues
git diff --check
Catches trailing whitespace, mixed tabs and spaces, and other formatting issues that agents sometimes introduce.
Look for unexpected file changes
Watch for these in git diff --stat:
- Lock files (
package-lock.json,yarn.lock) changing when you didn't add dependencies - Configuration files (
.eslintrc,tsconfig.json) being modified - Unrelated files appearing in the diff
.envfiles or other secrets
Selectively stage verified changes
git add -p
This walks you through each change interactively, letting you stage only the changes you've reviewed and approved. If the agent made five changes and you've verified four, stage those four and deal with the fifth separately. This is one of the most underused Git features for AI-assisted workflows.
Using a Second Agent as a Reviewer
One of the most effective review patterns is using a second AI agent to review the first agent's work. This sounds redundant, but it works for the same reason code review between humans works: the reviewer has fresh eyes and no attachment to the code.
The Pattern
- Agent A writes the code based on your requirements
- You open a separate session with Agent B
- You give Agent B the original requirements, the diff, and instructions to find problems
- Agent B reviews the code and reports issues
The Review Prompt
This prompt consistently produces useful reviews:
I asked an AI coding agent to make the following change:
[PASTE YOUR ORIGINAL PROMPT]
Here is the diff it produced:
[PASTE THE GIT DIFF]
Review this diff critically. Look for:
- Logic errors or subtle bugs
- Missing edge cases or error handling
- Security issues (unsanitized input, injection risks)
- Hallucinated APIs or incorrect function signatures
- Unnecessary complexity or over-engineering
- Violations of common best practices
- Anything that looks plausible but might be wrong
List every issue you find, with the specific line and a brief explanation
of what's wrong and how to fix it.
When This Works and When It's Overkill
Worth doing:
- Changes over 100 lines
- Complex business logic
- Security-sensitive code
- Unfamiliar territory (you're not an expert in this area)
- Code that's hard to test manually
Not worth the effort:
- Simple one-line fixes
- Changes you fully understand and can verify yourself
- Code with comprehensive test coverage that all passes
- Trivial additions like new log statements or copy changes
The second-agent review adds three to five minutes. For high-risk changes, that's a bargain.
Automated Safety Nets
Manual review catches what automated tools miss, but automated tools catch what manual review misses. Use both.
Type Checking
TypeScript in strict mode, mypy for Python, the Rust compiler -- these catch hallucinated APIs immediately. If the agent calls a method that doesn't exist or passes the wrong type, the type checker flags it before you even look at the code.
# TypeScript
npx tsc --noEmit
# Python
mypy src/
# Rust
cargo check
If your project doesn't use strict type checking, AI-generated code is a good reason to start. The number of bugs it catches from agent output alone justifies the setup cost.
Linting
Linters catch style violations, deprecated patterns, and suspicious code. They're particularly useful for catching the "stale patterns" failure mode.
# JavaScript/TypeScript
npx eslint .
# Python
ruff check .
# Go
golangci-lint run
Tests
Tests catch logic errors, but only if the tests are good. If the agent wrote the tests, review them with the same scrutiny you apply to the implementation. A test that always passes is worse than no test -- it gives you false confidence.
Run the full test suite, not just the new tests:
npm test
pytest
cargo test
go test ./...
Pre-Commit Hooks
Set up pre-commit hooks that run type checking, linting, and tests automatically. This creates a gate that AI-generated code must pass before it enters your repository.
# Example .pre-commit-config.yaml
repos:
- repo: local
hooks:
- id: typecheck
name: Type Check
entry: npx tsc --noEmit
language: system
pass_filenames: false
- id: lint
name: Lint
entry: npx eslint --fix
language: system
types: [javascript, typescript]
- id: test
name: Test
entry: npm test
language: system
pass_filenames: false
CI/CD as Backstop
Never skip the CI pipeline for AI-generated code. If anything, be more strict. Your CI should run the full test suite, type checking, linting, and security scanning on every push.
If an agent's code passes all automated checks and your manual review, you can be reasonably confident in it. If it fails any automated check, treat that as a signal that the code needs closer manual review too.
Red Flags to Watch For
These patterns in AI-generated code should trigger immediate scrutiny:
// TODO or // FIXME comments -- The agent is telling you it didn't finish the job. Don't commit code with TODO comments unless you're tracking them somewhere.
Try-catch blocks that swallow errors silently:
// Red flag
try {
await saveToDatabase(data);
} catch (e) {
// ignore
}
This hides failures. If the database save fails, you'll never know. At minimum, log the error. Better yet, decide whether the caller needs to know about the failure.
Commented-out code -- The agent left debugging artifacts or alternative implementations in comments. Remove them. That's what version control is for.
Unused imports or variables -- A sign the agent changed its approach mid-generation and didn't clean up. These aren't just messy -- they can mask actual missing dependencies.
Functions significantly longer than the project norm -- If your functions average 20 lines and the agent wrote a 150-line function, it probably needs to be broken up. Long functions from agents often contain duplicated logic or inline code that should be extracted.
Changes to security-sensitive files -- Any modifications to authentication, authorization, permissions, or .env files that you didn't explicitly request need careful review. The agent might have loosened a security constraint to make something else work.
Dependency changes you didn't ask for -- If package.json or requirements.txt changed and you didn't ask for new packages, find out why. The agent might have pulled in a dependency for something it could have done with existing tools.
Building Review Habits
Tools and checklists help, but habits are what keep your codebase safe over time.
Read every diff before committing. No exceptions.
This is the single most important habit. Even if you trust the agent. Even if it's a small change. Even if you're in a hurry. Read the diff. git diff takes ten seconds. The bugs it catches save hours.
Run the test suite after every agent change
Not just the new tests. The full suite. Agent changes can have unexpected ripple effects, and your existing tests are the safety net.
# Make this a reflex
git diff --stat && npm test
Keep a mental model of expected changes
Before you look at the diff, think about what should have changed. Which files? Roughly how many lines? What kind of changes? Then compare that expectation to reality. Discrepancies are where bugs hide.
Ask the agent to explain suspicious code
If something looks off but you can't pinpoint why, ask:
Explain why you used [specific approach] in [specific file] at [specific location].
What alternatives did you consider? What are the tradeoffs?
If the agent can't give a coherent justification, the code is probably wrong. Rewrite the original prompt with more constraints.
Maintain a list of things your agent gets wrong
Every agent has patterns it struggles with. Maybe your agent consistently gets date timezone handling wrong, or it always forgets to add database indexes, or it over-uses inheritance. Keep a running list and check those specific things during every review.
Over time, you can encode these checks into your prompts: "Remember to handle timezone conversion. Use UTC internally and convert at the display layer."
Treat review as non-negotiable
It's tempting to skip review when you're shipping fast and the agent's output has been reliable for the last ten changes. That eleventh change is where the bug hides. The cost of reviewing is constant and small. The cost of not reviewing is variable and occasionally catastrophic.
Agents UI's multi-session layout is designed for exactly this kind of workflow. Run your code-writing agent in one pane and a reviewing agent in another, each with separate context and history. The side-by-side view makes it easy to copy a diff from one session into a review prompt in the other without losing your place in either conversation.