In your country? See our local page →
Or choose your region

Original research · v1

How much routine finance work actually reaches completion?

Headline finding

In a synthetic 160-item corpus constructed by At Par to represent one trading month in an owner-led service business, 125 items (78.1%) reached a completed state and 35 (21.9%) stopped. Of those that completed, 113 were clean and 12 were recorded with an incomplete supporting document. The corpus is invented, not observed — At Par has no customers, and no client data exists to draw on.

The residue is not spread evenly. Stop rates run from 8.3% on payroll items to 50.0% on the journal-worthy events at period end, and 21 of the 35 stops (60.0%) could only be cleared by the business itself — a fact only it holds, a decision only it can take, or an authorisation only it can give. That is the finding worth arguing with: automation finishes most of the volume, and almost none of the part that carries authority.

The result

160 items, seven possible endings. Where they landed.

Each item in the corpus was given five plain properties — whether the supporting document is held, whether what it relates to can be determined from records already held, whether an existing rule settles how it is recorded, whether it needs a commercial judgement, and whether completing it would mean money leaving, a filing being submitted or a message going out in the client’s name. A rubric then assigned each item exactly one terminal state. The rubric was fixed before anything was scored.

Distribution of 160 synthetic finance work items across seven terminal states.
Terminal stateItemsShareWhat the state means
Stopped: consequential external act85.0%Completion would move money, submit a filing, or send a message in the client’s name. No additional information changes that; only the client’s authorisation does.
Stopped: needs a commercial decision31.9%The block is a judgement, not a fact — a concession, a write-off, a provision, a change of terms. Nobody outside the business is entitled to make it.
Stopped: missing fact106.3%What the item relates to cannot be determined from the item or from records already held. The fact exists, but only inside the business.
Stopped: missing evidence63.8%The supporting document is not held at all. It can often be obtained from a third party as well as from the client, which is why it ranks below a fact only the business holds.
Stopped: ambiguous treatment85.0%Facts and evidence are complete and two or more defensible treatments remain. A qualified accountant decides, and the decision becomes the rule for the next one.
Completed with an evidence gap127.5%Recorded correctly, and the supporting document is incomplete — cropped, illegible, or a confirmation standing in for an invoice. Counted as complete, and flagged.
Completed11370.6%Facts complete, evidence held, treatment settled by an existing rule, no commercial judgement, no external act. Recorded without a human decision.

Synthetic corpus, 160 items, authored by At Par · version 1, 29 July 2026. Not client data. Not a sample of any real population.

Read downward, the table splits into two unequal halves. 24 of the 35 stops are knowledge gaps — something is unknown, unevidenced or unsettled, and supplying it would finish the item. 11 are authority boundaries, where nothing is unknown at all: the work is finished up to the point where somebody has to decide or authorise, and no amount of further information moves it.

Excluding the 8 items that are outbound acts by construction, 125 of 152 routine items completed — 82.2%. The interesting number is not that figure. It is that the 35 stops cluster in a small number of places, and that 60.0% of them were never a provider’s to clear.
The distribution

The residue is concentrated, not scattered.

A single completion rate is the least useful thing this analysis produces, because it averages categories that behave nothing like each other. Sorted by how often they stop, the eight categories separate sharply.

Completion and stop rates by category of finance work item, synthetic corpus.
CategoryItemsCompletedStoppedMost common reason
Payroll items1211 (91.7%)1 (8.3%)missing fact
Bank transactions4843 (89.6%)5 (10.4%)missing fact
Inter-account transfers87 (87.5%)1 (12.5%)missing fact
Customer receipts2218 (81.8%)4 (18.2%)missing fact
Supplier invoices3024 (80.0%)6 (20.0%)missing fact · missing evidence · ambiguous treatment
Expense claims2016 (80.0%)4 (20.0%)missing evidence
Journal-worthy events126 (50.0%)6 (50.0%)ambiguous treatment
Outbound and filing acts80 (0.0%)8 (100.0%)consequential external act

Every figure above comes from a synthetic corpus authored by At Par. It is not client data, not a sample, and not a market statistic. Where two or more reasons stop the same number of items, all of them are shown rather than one — supplier invoices stop for three different reasons equally often, which is more informative than a single label would have been.

Across the routine categories the stop rate varies by a factor of about 6 — 8.3% for payroll items against 50.0% for journal-worthy events. The ordering is not random. The categories that complete most reliably are the ones where a rule already exists and the counterparty is known in advance: a salary at a contracted rate, a direct debit against a policy schedule, a receipt quoting an invoice number. The categories that stop most are the ones where the answer lives in a judgement rather than in a document.

One row deserves suspicion, and it is fair to say so before anybody else does. Outbound and filing acts stop at 100.0% because the rubric defines them that way: releasing money, submitting a filing and sending a message in a client’s name are treated as acts that require authorisation, so they can never be scored complete. That is a boundary written into the method, not a discovery made by it. It is included because leaving those items out of the corpus entirely would have understated how much of a finance month ends at an authority line rather than at a knowledge gap.

148 of 160 items (92.5%) were covered by an existing rule. That is the honest case for automation, and it is a strong one. The whole argument sits in the remaining 12 items, where an accounting decision has to be made for the first time — and where a system that cannot stop will make it silently.
The rubric

Seven endings, ranked by who can clear the block.

The rubric was written before scoring began, and its ordering is the one judgement in the method that changes the answer. The five stop states run from the block that no additional information can remove, down to the one a qualified accountant can settle without asking anybody. Each item is assigned the first state whose test passes, so an item carrying two blocking conditions is counted once, in the more binding of the two, and the stop counts never exceed the corpus.

  • Stopped: consequential external act. Completion would move money, submit a filing, or send a message in the client’s name. No additional information changes that; only the client’s authorisation does.
  • Stopped: needs a commercial decision. The block is a judgement, not a fact — a concession, a write-off, a provision, a change of terms. Nobody outside the business is entitled to make it.
  • Stopped: missing fact. What the item relates to cannot be determined from the item or from records already held. The fact exists, but only inside the business.
  • Stopped: missing evidence. The supporting document is not held at all. It can often be obtained from a third party as well as from the client, which is why it ranks below a fact only the business holds.
  • Stopped: ambiguous treatment. Facts and evidence are complete and two or more defensible treatments remain. A qualified accountant decides, and the decision becomes the rule for the next one.
  • Completed with an evidence gap. Recorded correctly, and the supporting document is incomplete — cropped, illegible, or a confirmation standing in for an invoice. Counted as complete, and flagged.
  • Completed. Facts complete, evidence held, treatment settled by an existing rule, no commercial judgement, no external act. Recorded without a human decision.

6 of the 160 items carry more than one blocking condition — an expense claim with no receipt and no project named, for instance. Under the precedence above each is reported once, under the more binding reason. Without a precedence rule those 6 items would be double-counted and the stop reasons would sum to more than the stops.

Grouping the stops by who is capable of clearing them produces the finding that matters commercially, and it is not flattering to anybody selling automation or to anybody selling a service.

The stopped items grouped by who is capable of clearing the block.
Who can clear itStopsShare of stopsWhy
Only the business2160.0%A fact only it holds, a commercial judgement, or an authorisation. No provider and no software can supply these.
The business or a third party617.1%A missing document. Often obtainable from the supplier, the bank or the platform without troubling the client at all.
A qualified accountant822.9%An unsettled treatment where the facts are complete. Decided once, then written down as the rule for the next one.

Every figure above comes from a synthetic corpus authored by At Par. It is not client data, not a sample, and not a market statistic.

60.0% of the stops in this corpus were not clearable by any provider or any system. They needed a fact, a decision or an authorisation that belongs to the business. A provider that promises to remove all of the residue is promising something the structure of the work does not allow. What can be removed is everything up to that line, and the reconstruction work of finding out which line it is.
Method

How the corpus was built — and how to check the figures.

The corpus contains 160 finance work items across eight categories, sized to represent one trading month in an owner-led service business of roughly fifteen to forty people: project-based work, a stable supplier list, staff expense claims, monthly payroll, a handful of period-end judgements, and the outbound acts a month generates.

The construction was deliberately ordinary. Each item was written as a one-line description of a real kind of finance work — a hosting invoice at the contracted rate, a round-sum transfer from a customer with several open invoices, a hotel folio photographed with the tax breakdown out of frame — and then given its five properties. Properties were assigned per item, by judgement, before any scoring took place. The terminal states were then derived from the properties by a script, not chosen by hand.

What the item mix is modelled on:

  • Category volumes reflecting a service business rather than a retailer or a manufacturer — 48 bank transactions, 30 supplier invoices, 22 customer receipts, 20 expense claims, 12 payroll items, 8 inter-account transfers, 12 period-end judgements and 8 outbound acts.
  • A stable supplier base. Most suppliers recur month to month, which is why most supplier items carry a settled treatment. A business changing suppliers constantly would score differently and the method would not notice.
  • Project allocation as a live requirement. Service businesses need costs against engagements, which is where several of the missing-fact stops come from. A business that does not allocate would complete more items and know less.
  • An unremarkable month. No year-end, no acquisition, no system migration, no seasonal spike, no dispute running through it.

What was measured, precisely: the count of items reaching each terminal state, overall and by category. That is the whole measurement.

What was deliberately not measured: the money value of any item, time taken, cost, error rates, the performance of any named software product or model, how long a stopped item stays stopped, and what happens after a stop is cleared. No economic claim of any kind is derived from this corpus — not hours saved, not headcount, not cash timing, not accuracy.

Reproducible: the corpus is scripts/authority/data/R01-corpus.json and the rubric and scorer are scripts/authority/data/R01-score.mjs in the site’s own repository. Running the scorer prints every figure on this page. This page renders whatever the scorer returns, so a changed item changes the published numbers — there is no hand-typed statistic anywhere on it.

Limitations

What this benchmark cannot tell you.

Written against At Par’s own interest, because a benchmark that only flatters its publisher is marketing with a table in it.

  • The corpus is invented. At Par wrote all 160 items and assigned every property. The scoring is deterministic and the inputs are judgement. A different author, with the same rubric, would produce a different headline.
  • There is no statistical validity here at all. This is not a sample of anything. There is no population, no sampling frame, no confidence interval and no margin of error. A figure from this page cannot honestly be quoted as a proportion of real businesses, real transactions or real accounting systems.
  • The composition drives the answer. Weight the corpus toward high-volume bank activity and the completion rate rises; weight it toward work in progress and unbilled delivery and it falls. The mix is a claim about what a service-business month looks like, and that claim is contestable.
  • No product was evaluated. Nothing here scores QuickBooks, Xero, Odoo, any bank feed, or any model. Benchmarking a named product honestly would require primary-source capability research and pinned version numbers, and that work was not done.
  • Counting items ignores consequence. A subscription renewal and a period-end revaluation each count as one. The stops are almost certainly weighted toward the items that matter most, which means the 21.9% figure understates the importance of the residue — but this analysis cannot show that, because it measures no value.
  • One month, one business shape. No seasonality, no year-end, no group structure, no jurisdiction-specific obligation, no audit cycle, no migration between systems.
  • The precedence is a choice. Reordering the five stop states would redistribute the reasons — 6 items would move — without changing the completed-versus-stopped split. Anybody who disagrees with the ordering can reorder it in the scorer and see exactly what moves.
  • The authority boundary is At Par’s own. The rubric treats releasing money, submitting a filing and sending a message in a client’s name as acts that stop for authorisation, because that is how At Par operates. A provider that executes those acts would score them complete, and the two numbers would not be comparable.
The strongest objection to this benchmark is that the properties and the desired conclusion were produced by the same author. That objection is not resolved by publishing the method, and this page does not pretend otherwise. What publishing the corpus does allow is disagreement at the level of individual items rather than at the level of a claim — and version 2 will be worth more than version 1 mainly because of the items people argue about.
Where we fit

Where At Par fits — and the limits.

At Par is an outsourced accounting and finance-operations service for owner-led service businesses, and it is built around the residue this analysis describes. Machine work handles the volume that a rule already covers. What a rule does not cover is held rather than posted, with a plain-language reason, a named owner and a next step — and a qualified accountant (ACCA) is accountable for the treatments applied.

That is the whole shape of the service. It sits across bookkeeping and month-end close, receivables and invoice follow-up, payroll and management reporting. What a provider should own to completion is set out separately in what an outsourced accounting provider should actually own, and how records and evidence are protected is covered under security.

At Par prepares payments and never moves, releases or executes client money. It prepares filings and does not submit them on a client’s behalf. Anything leaving the business in the client’s name stays subject to the client’s authorisation, and commercial decisions — write-offs, concessions, payment terms — stay with the owner. Those are the same boundaries the rubric above scores as stops, which is why 11 of the 35 stops in this corpus can never be removed by better software. At Par does not provide audit or assurance, and is not a collection agency.

For the operating standard behind these numbers rather than the measurement, see what should happen when bookkeeping automation is not sure and what should never happen automatically. For one stopped item followed all the way to its resolution, see one unmatched bank credit. The rest of the library is at guides.

Questions

Asked by people who intend to argue with the numbers.

How much routine finance work can bookkeeping automation complete? +

In this analysis, 125 of 160 items — 78.1% — reached a completed state, and 35 stopped. Excluding the outbound acts that stop for authorisation by definition, 125 of 152 routine items completed. Those figures come from a corpus At Par invented and scored against its own published rubric, so they describe what a defined method does to defined items. They are not a measurement of real businesses, real accounting systems, or any commercial product, and should not be quoted as one.

What is the Finance Work Completion Benchmark? +

It is a controlled analysis published by At Par in July 2026. A synthetic corpus of 160 finance work items — bank transactions, supplier invoices, customer receipts, expense claims, payroll items, inter-account transfers, period-end judgements and outbound acts — was given five explicit properties per item, then scored against a seven-state rubric fixed in advance. The output is the count of items reaching each terminal state, overall and by category. The corpus, the rubric and the scorer are all published so the figures can be reproduced or disputed.

Is this benchmark based on real client data? +

No. Every item was written by At Par, and every property was assigned by judgement before scoring. At Par has no customers, so no client records exist to draw on and none were used. The corpus is labelled synthetic in the file, on the page, and beside each table. The value of the exercise is that the rubric and the inputs are both open to inspection, not that the inputs were observed in the wild.

Why does finance work stop even when nothing is unknown? +

Because some items end at an authority boundary rather than a knowledge gap. In this corpus 11 of 35 stops were of that kind: releasing a payment run, submitting a prepared filing, sending a reminder in the client’s name, or deciding whether to provide against an old debt. Nothing about those is uncertain. They stop because the act belongs to the business owner, and no additional information transfers that authority to a provider or a system.

Which kinds of finance work stop most often? +

In this corpus, journal-worthy events at period end stopped 50.0% of the time and payroll items 8.3% — a spread of roughly 6 times across routine categories. The pattern is consistent: items governed by an existing rule against a known counterparty complete reliably, and items requiring a first-time accounting decision or a fact held only inside the business do not. Supplier invoices and expense claims sat in between, stopping around a fifth of the time, mostly for missing evidence or an unstated project.

Can a figure from this benchmark be quoted as an industry statistic? +

No, and doing so would misrepresent it. There is no sampling frame, no population, no confidence interval and no external data. The correct way to cite it is as a structured analysis of a corpus At Par constructed — for example, that in a 160-item synthetic corpus scored against a published rubric, roughly four items in five completed. Any phrasing beginning "X per cent of businesses" or "X per cent of transactions" would be false.

How were the seven terminal states decided? +

They were written before scoring and ordered by who is capable of clearing the block, from the most binding to the least: a consequential external act, then a commercial decision, then a missing fact, then missing evidence, then an ambiguous treatment, then completion with an evidence gap, then clean completion. Each item is assigned the first state whose test passes, so an item with two blocking conditions is counted once under the more binding one. Ordering is the single judgement in the method that changes the reported reasons.

Does this benchmark compare accounting software products? +

No. No product, platform or model was tested, named or scored. Measuring what a category of automation completes is a question about the structure of finance work, and it can be answered from a defined corpus. Scoring a named product would require primary-source capability research against pinned versions, repeated as those versions change, and that work was not done here. Any comparison drawn between these figures and a specific vendor would be invented rather than measured.

How can the numbers on this page be checked? +

The corpus and the scorer are files in the site repository: the item list with its properties, and the rubric that turns properties into terminal states. Running the scorer prints every figure that appears here, and the page is rendered from the scorer output rather than from typed-in text, so the two cannot diverge. Disagreement is best expressed at the item level — change a property, rerun it, and see which figures move and by how much.

Run it on a real month

Score one of your own months. We’ll show you where it stops.

Send a month of real activity and we will tell you which items an existing rule already covers, which stop for a missing fact or a missing document, and which are simply not ours to clear. The last group is usually smaller than people expect and matters more.

A few seats this cohort No lock-in · billed monthly Clean exit — records & evidence, always yours A person, not a bot

Prefer to reach us directly? Tell us a little and we'll come back with a time.

Score a month

We use your details only to prepare for and hold this call. No spam, ever.

Reviewed by an ACCA on the At Par team · Last updated 29 July 2026 · Version 1, 29 July 2026 — a 160-item synthetic corpus authored and scored by At Par, not client data and not a market study. See what we actually do.

Score a month The finding