Self-improvement is the talk of the AI world right now. It means AI models taking on a growing share of the work of improving themselves, rather than depending solely on the ingenuity of human engineers. Media coverage mostly frames this as something frontier labs do: OpenAI and Anthropic improving their foundation models, GPT and Claude. But self-improvement is not limited to the neural network. It works just as well at the application layer, the layer where the model is actually put to work, whether that is customer support, coding agents, or an accounting system. This article covers how we build self-improvement into accounting, and why that positions us to become the best solution for accounting as AI evolves.
The ground the application layer stands on is rising at a remarkable rate. Stanford’s 2026 AI Index measured what a single year of model progress now looks like: on OSWorld, a benchmark of AI agents doing real work on a computer, accuracy rose from roughly 12% in 2024 to 66.3% in 2025, within six percentage points of human performance. On Humanity’s Last Exam, a benchmark built to be hard for AI and favourable to human experts, accuracy jumped from under 10% to 38.3% over the same year. And the curve has not slowed since the report went to print: as of July 2026, the top score on the OSWorld-Verified leaderboard stands at 85%, clearly above the 72% human baseline, and the leading score on Humanity’s Last Exam has climbed to 46.4%. A system designed to ride that curve improves even while you sleep; a system designed around the current models’ weaknesses becomes obsolete on the same schedule.
What exactly are we improving?
We see accounting as a simple input-output system: bank statements and documents in, bookkeeping and filings out. We see no benefit in complicating that picture, least of all when the goal is a recurring self-improvement loop. What matters is being crystal clear about the actual transformation you want to improve (all natural phenomena are transformations: inputs and outputs). That judgement is the one thing humans are still uniquely positioned for. It is the simple decision of what to do.
Once you have decided what to optimise, in our case turning raw customer financial data such as transactions, invoices, and documents into auditable ledgers and reports, two steps remain: formalising the evaluation and harnessing the data.
How does the self-improvement loop work?
Formalising the evaluation (the “eval”) is a simple process: run your system on a set of test inputs and compare the results to a corresponding set of known-good outputs. Then grade the result: how accurate was the transformation? Run this daily, or several times a day, and you build up a historical dataset of how your system performs and which changes affect the score, positively or negatively. An agent can then analyse the results, compare them with changes to the codebase, and suggest or even implement improvements to the system. This creates a simple but powerful self-reinforcing loop.
Why does data quality decide who wins?
Evals amount to nothing unless they are built on high-quality data. Data is king here. If we train our systems on low-quality, untrustworthy data, the result will be low-quality, untrustworthy systems. We currently train our systems with select design partners (never on our customers’ data without consent), and we envision a world where we conjoin this loop with audit firms, effectively creating a system that passes audits more efficiently than existing solutions.
What does this mean for your company?
The shift is no longer hypothetical for the accounting industry. In Blue J and CPA.com’s June 2026 survey of more than 1,000 tax professionals, weekly AI use had nearly doubled in a year:
| What tax professionals report | Share |
|---|---|
| Use AI for tax research at least weekly | 60%, up from 33% a year earlier |
| Agree AI-powered research saves time | 84% |
| Plan to adopt AI in the near future | 32% |
Source: Blue J and CPA.com, AI Tax Research Solution Outlook Report, June 2026.
Most of the industry is adopting AI as a feature. We are building the loop that compounds. If you’re a company that wants faster, cheaper, and more correct accounting, we’re the accounting firm to turn to. We’re an accounting firm on the outside, but an AI research company on the inside. As AI develops, we’re the only firm positioned to capitalise on it. By choosing us as your partner, you position yourself for the same growth. Join the waitlist.