Debugging with AI Coding Agents: A Practical Workflow
Most content about AI coding agents focuses on building: generating features, writing tests, scaffolding projects. But debugging is where developers actually spend a significant portion of their time, and it's where AI agents can be surprisingly effective — if you use them correctly.
The challenge is that debugging requires a different approach than code generation. You can't just tell an agent "fix the bug" and expect good results. Debugging is investigative work: reproducing the issue, forming hypotheses, isolating the root cause, and verifying the fix doesn't break anything else. AI agents excel at each of these steps individually. The trick is orchestrating them into a coherent workflow.
Why Agents Are Good at Debugging
AI coding agents have several advantages over manual debugging:
They read code faster than you. When a bug could be in any of 50 files, an agent can scan the entire codebase and identify suspicious patterns in seconds. You'd spend twenty minutes jumping between files.
They don't have assumptions. You wrote the code, so you think you know how it works. That's exactly why you miss bugs — your mental model doesn't match reality. An agent reads the code as-is, without preconceptions about what it "should" do.
They can hold more context. A race condition might involve interactions between a queue processor, a database transaction, and a WebSocket handler. Keeping all three in your head simultaneously is hard. An agent can read all three files and reason about their interactions.
They're systematic. When asked to investigate, agents methodically check error handling, edge cases, and boundary conditions. Humans jump to conclusions.
That said, agents also have weaknesses in debugging: they can't reproduce timing-dependent bugs, they sometimes "fix" symptoms instead of root causes, and they can be confidently wrong. The workflow below accounts for these limitations.
The Debugging Workflow
Step 1: Reproduce Before Investigating
The most common mistake when using agents for debugging is jumping straight to "find and fix the bug." Without a reproduction case, the agent is guessing — and agents are dangerously good at producing plausible-sounding fixes that don't actually address the problem.
Start with reproduction:
There's a bug where users intermittently get a 500 error when updating their profile.
Before trying to fix anything, help me reproduce this:
1. Read src/api/routes/profile.ts and src/services/user-service.ts
2. Look at the error logs I'll paste below
3. Write a test in src/api/__tests__/profile-bug.test.ts that reproduces the 500 error
Error from production logs:
TypeError: Cannot read properties of undefined (reading 'email')
at updateProfile (src/services/user-service.ts:47)
at handler (src/api/routes/profile.ts:23)
The agent now has a concrete task: write a failing test. This test becomes your source of truth for the rest of the debugging session. You'll know the bug is actually fixed when this test passes — not when the agent says it's fixed.
Step 2: Isolate the Root Cause
If the behavior worked in an earlier revision, use the reproduction to locate the change that broke it. This git bisect tutorial shows how to automate that search and give the agent a specific commit to investigate.
Once you have a reproducing test, ask the agent to investigate:
The test in src/api/__tests__/profile-bug.test.ts reproduces the 500 error.
Now investigate the root cause. Read through the call chain:
1. src/api/routes/profile.ts — the route handler
2. src/services/user-service.ts — the updateProfile function
3. src/db/queries/user.ts — the database query
The error is "Cannot read properties of undefined (reading 'email')" at
user-service.ts:47. Figure out under what conditions the value at line 47
would be undefined and explain the root cause before suggesting a fix.
Two critical elements in this prompt:
"Explain the root cause before suggesting a fix" — This forces the agent to reason about the bug rather than pattern-matching to a common fix. If the agent can't explain why the bug happens, its fix probably won't be correct.
You provide the call chain — Instead of letting the agent wander through the codebase, you point it at the relevant files. This speeds up the investigation and prevents the agent from getting distracted by unrelated code.
Step 3: Evaluate the Explanation
This is the step most people skip, and it's the most important. Before accepting any fix, read the agent's root cause analysis. Does it make sense? Does it match the error you're seeing?
Common red flags in agent explanations:
- "The issue might be..." — If the agent isn't sure, it's guessing. Push for specifics.
- "This could happen if the database returns null" — Could happen isn't the same as does happen. Ask for evidence.
- "I've added a null check to handle this case" — This is a symptom fix, not a root cause fix. Why is the value null in the first place?
If the explanation doesn't convince you, push back:
Your explanation suggests the database might return null, but the schema
has a NOT NULL constraint on the email column. Under what other conditions
could the value at line 47 be undefined? Check if there's a middleware or
transformation step between the database query and line 47 that could
strip the email field.
This back-and-forth is where agents shine. They'll re-read the code, check your constraint, and often find the actual root cause on the second pass. The first answer is frequently wrong for subtle bugs — the second or third is usually right.
Step 4: Fix with Constraints
Once you understand and agree with the root cause, ask for the fix with explicit constraints:
Good analysis — the issue is that the caching middleware returns a stale
object without the email field when the user updates their profile.
Fix this by invalidating the user cache in src/middleware/cache.ts when
the profile is updated. Specifically:
Constraints:
- Only modify src/middleware/cache.ts and src/services/user-service.ts
- Do not change the cache strategy for other endpoints
- The fix should be in the cache invalidation logic, not a null check
- Keep the existing cache TTL behavior for non-profile routes
After the fix, run the reproduction test to verify it passes.
Then run the full test suite to check for regressions.
Notice the constraint "the fix should be in the cache invalidation logic, not a null check." Without this, many agents will add a defensive null check — which silently hides the bug rather than fixing it. The null check means the 500 error goes away, but users still see stale data. The cache invalidation fix addresses the actual problem.
Step 5: Verify and Regression Test
After the agent applies the fix:
The reproduction test passes now. Good.
Add two more test cases:
1. Verify that updating a profile immediately reflects the new data
on the next GET request (no stale cache)
2. Verify that updating one user's profile doesn't invalidate
another user's cached profile
Run the full test suite after adding these tests.
These additional tests catch common mistakes in cache invalidation fixes: over-invalidation (clearing too much) and under-invalidation (not clearing enough). The agent writes them, runs them, and you have confidence the fix is correct.
Debugging Patterns for Common Bug Types
Race Conditions
Race conditions are notoriously hard to debug because they're timing-dependent. Agents can't reproduce them by running code, but they can analyze code for race condition patterns.
I suspect there's a race condition in src/workers/order-processor.ts.
Two orders for the same product sometimes both succeed even when only
one item is in stock.
Read the order processing code and the inventory management in
src/services/inventory.ts. Look for:
- Check-then-act patterns without locking
- Non-atomic read-modify-write operations
- Missing transaction boundaries
- Shared mutable state between async operations
Explain where the race condition could occur before suggesting a fix.
Agents are excellent at spotting check-then-act patterns:
// Classic race condition pattern the agent will identify
const stock = await getStock(productId); // Read
if (stock > 0) { // Check
await decrementStock(productId); // Act — another request could have decremented between check and act
await createOrder(userId, productId);
}
Memory Leaks
Memory leaks in long-running processes are hard to track manually. Give the agent specific symptoms:
Our Express server's memory usage grows from 200MB to 2GB over 24 hours.
The issue started after the commit that added WebSocket support.
Relevant files:
- src/ws/connection-manager.ts
- src/ws/handlers.ts
Look for:
- Event listeners that are added but never removed
- Maps or Sets that grow without cleanup
- Closures that capture large objects
- Missing cleanup in disconnection handlers
Check if the connectionManager Map in connection-manager.ts properly
removes entries when clients disconnect.
Performance Regressions
When something got slower, agents can compare before and after:
The /api/products endpoint went from 50ms to 800ms response time
after the last deployment. The only relevant changes were in
src/services/product-service.ts and src/db/queries/products.ts.
Read both files and identify potential performance issues:
- N+1 query patterns
- Missing database indexes (check the query against our schema in src/db/schema.sql)
- Unnecessary data fetching (loading full objects when only IDs are needed)
- Synchronous operations that could be parallel
If you find the likely cause, explain it and suggest a fix.
Do not modify any database schema — only application code changes.
Intermittent Failures
Bugs that only happen sometimes are the hardest to debug. Help the agent by providing patterns:
Our CI pipeline fails intermittently on the test "should process
webhook within timeout." It passes 90% of the time, fails 10%.
Test file: src/__tests__/webhook.integration.test.ts
Test implementation in: src/services/webhook-processor.ts
Common causes of intermittent test failures:
- Hardcoded timeouts that are too tight
- Tests depending on execution order
- Shared state between tests (database, global variables)
- Async operations not properly awaited
Read the test file and the webhook processor. Identify what makes
this test non-deterministic.
Multi-Agent Debugging
For complex bugs that span multiple systems, use multiple agents in parallel:
Agent 1 — Backend investigation:
Investigate why the payment webhook handler in src/api/webhooks/stripe.ts
sometimes fails to update the order status. Read the handler code,
the order service, and the recent error logs I'm pasting below.
Agent 2 — Frontend investigation:
The payment confirmation page sometimes shows "Payment Processing"
indefinitely. Read src/components/PaymentStatus.tsx and
src/hooks/usePaymentStatus.ts. Check if the polling logic has
edge cases where it could poll forever.
Agent 3 — Infrastructure check:
Read the webhook retry configuration in src/config/stripe.ts and
the queue configuration in src/config/bull.ts. Check if there's a
scenario where a webhook could be processed out of order or duplicated.
Run all three investigations simultaneously. Each agent explores a different layer of the system. Compare their findings to build a complete picture of the bug. This parallel approach works particularly well with tools like Agents UI that support multiple concurrent sessions.
The Debug Log Technique
When an agent can't figure out a bug from reading code alone, have it add strategic logging:
I can't reproduce the payment bug locally. Add debug logging to help
trace the issue in staging:
Add structured log statements to:
- src/api/webhooks/stripe.ts (log the incoming event type and order ID)
- src/services/order-service.ts (log before and after status updates)
- src/services/payment-service.ts (log the payment state transitions)
Use our existing logger (src/lib/logger.ts) with log level "debug".
Tag all logs with a correlation ID from the request headers.
Format: logger.debug('webhook:stripe', { correlationId, eventType, orderId, ... })
Do not change any business logic — only add logging.
This is a safe task for an agent: no business logic changes, clear scope, and the output (logs) gives you data for the next debugging step. After deploying and reproducing the issue, paste the relevant logs back to the agent:
Here are the debug logs from when the bug occurred in staging:
[logs...]
Based on these logs, what's happening? The order ID abc-123 received
two webhook events but only one status update was recorded.
Now the agent has concrete data to work with, not just code to speculate about.
Common Mistakes When Debugging with Agents
Accepting the First Fix
The agent's first suggestion is often a symptom fix, not a root cause fix. Adding a null check, wrapping something in a try-catch, or adding a retry mechanism are all band-aids. Always ask "why does this value end up in a bad state?" before accepting a fix that handles the bad state.
Not Providing Error Context
# Bad
Fix the bug in the login flow.
# Good
Users clicking "Login with Google" get redirected to /auth/callback
but receive "Invalid state parameter" error. This started after we
upgraded passport-google-oauth20 from 2.0.0 to 2.1.0.
Error: src/auth/google.ts:34 — InvalidStateError
Stack traces, error messages, version changes, and timestamps are not optional context — they're the debugging information the agent needs.
Letting the Agent Modify Too Many Files
A one-line bug fix that turns into a 15-file change is a red flag. If the agent is modifying files beyond the scope of the bug, it's probably refactoring rather than debugging. Constrain the scope:
Fix only the cache invalidation bug. Do not refactor, rename, or
reorganize any code. The diff should be as small as possible.
Skipping the Reproduction Step
If you can't reproduce it, you can't verify the fix. Period. Even if the agent's explanation sounds perfect, without a failing test that becomes a passing test, you're guessing. Take the five minutes to write the reproduction test. It will save you from deploying a "fix" that doesn't fix anything.
Building a Debugging Toolkit
Create a set of reusable debugging prompts for your project. Save these in your project's docs or in a shared directory:
Bug investigation template:
Bug: [DESCRIPTION]
Error: [ERROR_MESSAGE_OR_STACK_TRACE]
Frequency: [always / intermittent / once]
Started: [when it first appeared, relevant changes]
Relevant files:
- [FILE_1]
- [FILE_2]
Steps to reproduce:
1. [STEP_1]
2. [STEP_2]
Investigate the root cause. Read the relevant files, explain what's
going wrong and why, then suggest a minimal fix. Write a test that
reproduces the bug before fixing it.
This template forces you to gather the right information before involving the agent, which consistently leads to faster resolution.
Conclusion
Debugging with AI agents works well when you treat the agent as a tool for investigation and verification, not as a magic fix-it button. Reproduce first. Investigate the root cause. Constrain the fix. Verify thoroughly.
The workflow — reproduce, isolate, explain, fix, verify — mirrors how experienced developers debug manually. The agent just moves through each step faster, reads more code, and catches things you might miss. Your job shifts from doing the debugging to directing the investigation and evaluating the results.
Start with your next bug. Instead of diving into the code yourself, write a reproduction prompt, hand it to the agent, and compare how long it takes versus your usual approach. For most bugs, it's faster. For subtle ones, the agent's systematic analysis often finds root causes that would take hours of manual investigation.