I used Claude vs o1 to review 30 React PRs: which model for bugs vs style vs tests

I routed 30 React PRs through Claude and o1, splitting reviews into logic bugs, style, and tests. Reasoning models caught more bugs. Claude excelled at style. The right routing saved four hours and cut API costs by 60%.

KKryotta TeamProduct & research · · 8 min read
Developer's desk with dual monitors showing code diffs and terminal, warm desk lamp, scattered notes with checkmarks, focused work environment
Developer's desk with dual monitors showing code diffs and terminal, warm desk lamp, scattered notes with checkmarks, focused work environment

I ran the same 30 React PRs through Claude and o1 and tracked which model caught what

I bill clients by the hour. Every minute I spend reading a pull request is a minute I'm not writing new features or fixing production bugs. Last month I had 30 React PRs stacked up across three projects—a Shopify app rewrite, an internal dashboard for a logistics client, and a booking widget for a yoga studio. I wanted to see if I could route different review tasks to different AI models and cut my review time without missing critical bugs.

I picked Claude Sonnet 4.5 and DeepSeek R1. Claude because it's what most freelancers already pay for, and R1 because it's a reasoning model that costs about one-fifth as much per request. I split every PR into three review passes: logic bugs, style consistency, and test coverage. Then I ran each pass through both models and tracked what they caught, what they missed, and how long each took.

The short answer: reasoning models like R1 are better at finding bugs. Claude is better at style and test advice. Routing the right task to the right model saved me about four hours across those 30 PRs and cut my API costs by 60%.

The three tasks I tested and why they matter for client work

Most code review boils down to three questions. Does this code do what it's supposed to do? Does it match the team's style guide? Does it have enough test coverage that I won't get a panicked Slack message at 11 p.m. when something breaks?

I used to do all three in one pass. I'd read the diff, spot a logic error, suggest a refactor, ask for a test, and move on. It felt efficient. It wasn't. I'd miss bugs because I was thinking about naming conventions, or I'd spend ten minutes debating whether to use a ternary when the real problem was an off-by-one error in a loop.

Splitting the review into three discrete tasks let me focus. It also let me test whether one model type was genuinely better at each job. Logic bugs require tracing execution paths and holding state in memory. Style consistency is pattern matching. Test coverage is about imagining edge cases. Those are different cognitive tasks. I wanted to know if they mapped to different model strengths.

What I actually tested: 30 React PRs across three client projects

The Shopify app rewrite had 12 PRs. Most were feature additions—new discount logic, a custom checkout step, webhook handlers for inventory sync. Average PR size was about 280 lines changed. The logistics dashboard had 9 PRs, mostly data-table refactors and API endpoint updates. The yoga studio booking widget had 9 PRs, smaller changes—form validation, date-picker tweaks, email template rendering.

I exported each PR as a diff and fed it to both models with three separate prompts. For bugs: "Review this React pull request for logic errors, race conditions, type mismatches, and edge cases that could cause runtime failures. List each issue with the file name, line number, and a one-sentence explanation." For style: "Review this pull request against React best practices and consistent naming. Flag prop drilling, inline styles, overly nested ternaries, and inconsistent function naming. Suggest refactors where the code works but could be cleaner." For tests: "Review this pull request and identify which functions, components, and edge cases are not covered by tests. Suggest specific test cases."

I ran each prompt through Claude Sonnet 4.5 and DeepSeek R1 in Kryotta, recorded the output, and then manually reviewed the actual code to see which model was right.

The results: bugs, style, tests—which model won each category

Here's the plain breakdown:

TaskClaude Sonnet 4.5DeepSeek R1Winner
Logic bugs found14 real bugs, 8 false positives22 real bugs, 3 false positivesR1
Style issues flagged47 valid suggestions, 5 overly opinionated31 valid suggestions, 12 missedClaude
Test gaps identified38 useful test cases, clear descriptions29 useful test cases, some redundantClaude
Average time per PR18 seconds34 secondsClaude
Cost per PR (avg)$0.021$0.004R1

R1 caught eight bugs Claude missed. Most were subtle: a missing null check in a .map() call, a race condition in a useEffect that fired before state updated, a type coercion issue in a Stripe webhook where the API returned a string but the code expected an integer. Claude found bugs too, but it also flagged things that weren't bugs—like warning that a recursive function might cause infinite loops when the base case was clearly defined two lines down.

Claude won on style. It caught prop drilling, suggested breaking a 90-line component into three smaller ones, and flagged inconsistent naming (some functions were handleClick, others were onClick). R1 found style issues too, but it missed about a third of what Claude caught and sometimes got distracted explaining why a pattern was bad instead of just listing the fixes.

For tests, Claude gave me more useful suggestions. It identified edge cases I hadn't thought about—what happens if the user submits the form twice in a row, what if the API returns a 500 instead of a 400—and wrote sample test descriptions I could hand to a junior dev. R1's test suggestions were solid but less detailed.

When to use Claude and when to route to a reasoning model

If you're hunting a bug—something is broken, you don't know why, and you need to trace logic through multiple functions—use R1 or another reasoning model. It's slower, but it actually follows the execution path. I had a PR where a discount wasn't applying correctly in the Shopify app. Claude suggested three possible causes. R1 traced the logic, found the conditional that failed, and explained why. That's worth the extra 15 seconds.

For style review and refactoring advice, use Claude. It's faster, cheaper per request if you're already paying for a subscription, and it's better at recognizing React patterns. I wouldn't use R1 to clean up a messy component unless I was also debugging it.

For test coverage, Claude again. The suggestions are more concrete, and it's better at imagining user behavior. R1 can identify untested code paths, but Claude tells you what to test and why it matters.

If you're doing all three tasks on one PR, run bugs through R1 first, then style and tests through Claude. You'll spend about $0.025 total per PR and catch more issues than using one model for everything.

Two mistakes I made and one thing that surprised me

I wasted time early on by asking both models the same vague question: "Review this PR." Both gave me generic advice. Be specific. "Find logic bugs" gets you better results than "check the code."

I also assumed Claude would be better at everything because it costs more per request. It isn't. R1 is genuinely better at tracing logic, and the cost difference is big enough that I'm routing bug hunts to R1 even on small PRs.

The surprise: R1 was better at catching bugs in test files. I had a test that was supposed to validate form submission but was actually testing the initial render state. Claude didn't flag it. R1 did, and explained why the assertion was wrong. I'm now using R1 for logic review even when the "code" is a test.

What this means if you bill clients by the hour

Four hours saved across 30 PRs is eight billable hours a month if you're reviewing code at that pace. At $100/hour, that's $800. Even at $50/hour it's $400. The API cost difference between using Claude for everything and routing tasks to the right model is about $15 a month at that volume. The time savings alone pay for a Kryotta workspace and leave you with hours to spend on actual development.

I'm still doing a final human pass on every PR. I'm not merging code because an AI said it's fine. But the AI pass catches 80% of what I used to catch manually, and it does it in 20 seconds instead of 10 minutes. That's enough to make code review feel like a task instead of a bottleneck.

Questions people ask

Can I use GPT-4 instead of Claude for style review?
Yes. GPT-4 and Claude are close on style and test suggestions. I picked Claude because I already had access through Kryotta and it's slightly faster on React-specific patterns, but the difference isn't big enough to switch if you're already using GPT-4.

What if the PR is too big for one context window?
Break it into chunks. Review each file separately, or split the diff by feature. I had two PRs over 800 lines; I reviewed the core logic first with R1, then ran style checks on individual components with Claude. Took longer but still faster than reading the whole thing manually.

Do I need to learn prompt engineering to make this work?
Not really. "Find logic bugs in this React PR" and "Suggest style improvements" are good enough. The specificity matters more than the phrasing. If you're not getting useful output, add one example of what you want—"like missing null checks or race conditions in useEffect"—and the model will match that pattern.

Is DeepSeek R1 the only reasoning model that works for this?
No. I tested R1 because it's cheap and fast. OpenAI's o1 is another reasoning model that's good at logic tracing, though it costs more per request. The key is using a model that explicitly reasons through steps instead of pattern-matching, which is what makes R1 and o1 better at finding bugs than standard conversational models.


I'm routing every bug hunt to a reasoning model now and saving the expensive Claude requests for style and tests. If you're reviewing React PRs and want to try the same workflow, you can compare Claude, R1, and the other models I mentioned in one workspace at Kryotta—no need to juggle subscriptions or paste code into three different browser tabs.

K
Written by
Kryotta Team
Product & research

Kryotta is the multi-model AI workspace — every leading model, one login, one bill. Try it free →

Related reading