The Sunday night I had sixty invoices and two new models
I'm a sole trader. I run a small e-commerce consultancy—mostly Shopify setup, some email marketing, a bit of Google Ads management for independent retailers around Leeds and Belfast. About eight to ten clients at a time. They pay by bank transfer, I invoice through Xero, and every quarter I file a VAT return that I've usually left until the last week.
This January, I left it until the last weekend. I had sixty items to reconcile: client invoices, software subscriptions, a handful of travel receipts from a client meeting in Dublin, and three months of VAT that I'd been too busy to check. My Companies House deadline was Tuesday morning. I had two days, a Xero account that was showing six unmatched transactions, and a headache from staring at spreadsheets.
That Sunday night, OpenAI released o1 and Anthropic released Claude 3.7 Sonnet. Both companies said the same thing: better reasoning, better accuracy, built for complex tasks. I'd been using Claude Sonnet 4.5 for Xero work since November—pulling transaction notes, drafting invoice reminders, checking VAT calculations. It was fine. Sometimes it missed duplicate expenses. Sometimes it calculated the VAT backwards when a client paid in euros. But it was faster than doing it myself.
I decided to test both new models on the same sixty items. Not a rigorous lab experiment—I didn't have time for that—but a real reconciliation session with real consequences. I wanted to know if the extra cost was worth it, and whether "reasoning" meant anything when you're just trying to match a PayPal payment to an invoice before HMRC charges you a penalty.
What I actually tested
I exported my Xero transactions as a CSV. Sixty rows: forty-three invoices I'd sent to clients, twelve expenses (software, travel, one new laptop charger), and five VAT payments I'd made in October that Xero hadn't matched to the right quarter. I split the work into four tasks that take me the longest when I do them manually:
VAT calculation check. I gave each model ten invoices and asked it to verify the VAT. Three were standard-rated UK clients, two were zero-rated EU clients, three were UK clients who'd paid in euros, and two were Northern Ireland clients where the VAT rules changed mid-project. I wanted to see if the models could spot the one invoice where I'd charged 20% on a zero-rated service by mistake.
Duplicate detection. I'd accidentally logged the same £340 software subscription twice in November—once when the payment left my account, once when the receipt arrived by email a week later. I gave each model my full expense list and asked it to flag duplicates. No other context.
Missing receipt matching. Xero was showing six unmatched bank transactions. I gave each model the transaction descriptions (the ones your bank puts in the CSV, like "PAYPAL *ADOBECREATIV" or "TFL TRAVEL CHARGE") and my list of expenses, and asked it to suggest matches. This is the task I hate most. The descriptions are always truncated or misspelled, and I end up Googling half of them.
Invoice reminder drafting. I had four overdue invoices, all from the same client, a small homeware shop in Cardiff. I gave each model the invoice dates, amounts, and my previous email thread (three reminders already sent, all polite, all ignored), and asked it to draft a firmer follow-up that didn't sound like I was about to take them to small claims court.
I ran each task twice—once with OpenAI o1, once with Claude 3.7 Sonnet—and timed how long the model took to respond. Then I checked the answers against my own records and the Xero audit log.
VAT calculations: Claude was faster, o1 was pedantic
Claude 3.7 Sonnet found the mistake in eleven seconds. It flagged the zero-rated invoice, explained why it should have been zero-rated (services supplied to an EU business with a valid VAT number), and recalculated the correct total. It also noticed that one of the euro-denominated invoices used an exchange rate from the wrong week, which I hadn't asked it to check but was genuinely helpful.
OpenAI o1 took forty-one seconds. It found the same mistake, but it also wrote three paragraphs explaining the UK VAT rules for cross-border services, cited the HMRC VAT Notice 741A (which I didn't need), and suggested I check whether my EU clients were VAT-registered in their own countries (which I had, in 2023, and didn't need to do again). The answer was correct. The reasoning was exhaustive. The speed made it unusable for a Sunday-night reconciliation session.
o1 is designed to "think" before it answers. You see a message that says "Thinking…" and a progress bar. Sometimes it thinks for five seconds, sometimes for a minute. For a VAT calculation, that thinking didn't add value. Claude gave me the answer and the explanation in the time it took o1 to decide how much explanation I needed.
Duplicate detection: o1 missed one, Claude missed two
I gave each model my twelve expenses and asked it to flag duplicates. There were two: the £340 software subscription I'd logged twice, and a £28 train ticket to Dublin that I'd entered once as "Dublin Connolly" and once as "Irish Rail ticket 4 Nov".
OpenAI o1 caught the software subscription. It missed the train ticket because the descriptions were different and it didn't recognize "Connolly" as a station name. It did, however, flag a £15 Amazon order and a £15 Royal Mail postage charge as "possible duplicates" because they were the same amount and happened on the same day. They weren't duplicates. One was packing materials, one was postage. I had to check my bank statement to confirm.
Claude 3.7 Sonnet missed both duplicates. It told me there were no duplicates in the list. When I asked it to check again and look for similar descriptions or amounts, it found the software subscription but still missed the train ticket. It didn't hallucinate a false positive, which I appreciated, but it didn't do the job.
For this task, I ended up using o1 and ignoring its false positive. But I wouldn't trust either model to do duplicate detection without a manual check. If you've got a long expense list and you're worried about duplicates, you're better off sorting by amount in Xero and scanning it yourself. Takes three minutes.
Missing receipts: Claude guessed better
This was the task I cared about most. Xero had six unmatched transactions, and I didn't have time to dig through my email for receipts. The descriptions were:
- PAYPAL *ADOBECREATIV
- TFL TRAVEL CHARGE
- AMZN MKTP UK*2F8VQ3
- SHOPIFY *REAL GOOD D
- GOOGLE *WORKSPACE
- DD BRITISH GAS
I gave each model the descriptions and my list of logged expenses (which included Adobe, Amazon, Shopify, Google Workspace, and a gas bill, but not TfL) and asked it to match them.
Claude matched five out of six correctly. It paired "PAYPAL ADOBECREATIV" with my Adobe Creative Cloud subscription, "AMZN MKTP UK2F8VQ3" with an Amazon order for office supplies, "SHOPIFY *REAL GOOD D" with a Shopify subscription (I run a test store for client demos), "GOOGLE *WORKSPACE" with Google Workspace, and "DD BRITISH GAS" with my gas bill. It left "TFL TRAVEL CHARGE" unmatched and said it didn't see a corresponding transport expense in my list. That was correct—I'd forgotten to log a Tube fare from a client meeting in London.
OpenAI o1 matched four correctly. It got confused by "SHOPIFY *REAL GOOD D" and suggested it might be a client payment (because "Real Good" sounded like a shop name) rather than a Shopify subscription. It also flagged "TFL TRAVEL CHARGE" as a possible match for my Dublin train ticket, which made no sense—TfL is Transport for London, not Irish Rail. The reasoning was detailed, but the guesses were worse.
Claude won this round by a clear margin. For transaction matching, I'd use it over o1 every time.
Invoice reminders: o1 wrote a letter, Claude wrote an email
I asked each model to draft a follow-up for the overdue invoices. I gave them the context: four invoices, all between £850 and £1,200, all overdue by six to eight weeks, three polite reminders already sent.
Claude wrote a short, firm email. Two sentences acknowledging the previous reminders, one sentence with the total amount outstanding, one sentence saying I'd need to pause work on the current project until the invoices were settled, and a closing line offering to set up a payment plan if cash flow was an issue. It sounded like me. I sent it with one small edit (I changed "pause work" to "pause new work" because I didn't want to sound like I was downing tools mid-project).
OpenAI o1 wrote a three-paragraph letter. It opened with "I hope this message finds you well," which I never say. It included a sentence about "maintaining a positive working relationship," which felt passive-aggressive. It ended with "I trust we can resolve this matter amicably," which sounded like I'd hired a solicitor. The tone was polite, but it wasn't my tone. I didn't use it.
For drafting, I'll stick with Claude. o1's reasoning doesn't help when the task is about voice, not logic.
When the extra cost makes sense (and when it doesn't)
OpenAI o1 costs more than Claude 3.7 Sonnet. In Kryotta, o1 runs on a per-token model; Claude Sonnet is part of the standard workspace pricing. If you're doing one or two reconciliation tasks a month, the difference is negligible. If you're doing this every week, it adds up.
I'd use o1 for tasks where I need it to check its own work: VAT calculations on complex invoices, expense categorization when I'm not sure which category applies, or drafting a response to an HMRC query where I need to cite the right rule. The extra thinking time is worth it when the cost of getting it wrong is high.
I'd use Claude for everything else: transaction matching, duplicate checks, reminder emails, pulling data from Xero exports, and any task where speed matters more than exhaustive reasoning. It's faster, it's cheaper, and it gets the answer right often enough that I can check the edge cases myself.
I wouldn't use either model to file my VAT return or make a final accounting decision. They're tools for the repetitive part, not the judgment part. But for a sole trader who's reconciling sixty invoices on a Sunday night, Claude 3.7 Sonnet did the job I needed in the time I had. OpenAI o1 did a more thorough job, but I didn't need thorough—I needed done.
Questions people ask
Can I connect Xero directly to Claude or o1?
Not natively. You'll need to export your transactions as a CSV or copy-paste them into the chat. Kryotta doesn't have a Xero connector yet, but you can upload the CSV and work with it in the same workspace where you're running Claude or o1.
Which model is better for VAT if I sell into the EU?
OpenAI o1 is more thorough with cross-border VAT rules, but it's slower. If you're doing a quarterly reconciliation and you've got time, use o1. If you're checking one invoice before you send it, Claude is fast enough and usually correct.
Does Claude 3.7 Sonnet replace Sonnet 4.5 for accounting tasks?
No. 3.7 is a different model, not an upgrade. I found 4.5 better for transaction matching and duplicate detection. 3.7 is faster but less reliable on edge cases. Test both in your workspace and see which one fits your workflow.
Will either model catch every duplicate expense?
No. Both missed duplicates in my test. If you're worried about duplicates, sort your Xero expenses by amount and date, and scan them yourself. Takes less time than fixing the mistakes an AI makes.
I reconciled those sixty invoices by Monday afternoon. Filed the VAT return on Tuesday morning, ten minutes before the deadline. I used Claude for fifty-four of the tasks and o1 for six. Total cost in Kryotta: less than the £200 penalty I paid last quarter. If you're a sole trader staring at a Xero backlog, try Claude first. If it gets stuck, switch to o1. And if you're still doing this in a spreadsheet, stop—you're spending hours on work a model can do in seconds.
You can test both models in one workspace at Kryotta. No separate subscriptions, no API keys, just the task and the model that fits it.



