Blog

Close management is becoming a harness for AI agents

Close management is becoming a harness for AI agents

Written by

Artifact Team

Artifact Team

Artifact Team

Close management software brought order to the close. It helped firms standardize working papers, sequence procedures, assign preparers, and dynamically track a checklist’s progress. But it never did the work itself.

Agents can now do the work. Not all of it, and not unsupervised, but enough that the question facing a firm has changed. Software that used to coordinate humans to run the close now has to instruct agents, check what they produce, and know when to stop them.

Every firm will run agents, and nearly all of them on the same few models. The difference will come from what sits around those models.


What a harness is

A harness is everything built around an AI model to make its work dependable. A model inside a harness, doing work, is what people mean by “an agent”. The model thinks. The harness sets the terms it works under: what job it takes, what information it sees, what tools it can use, what it may change, what it must leave alone, how its work gets checked, what it does when it is unsure, and what happens when something breaks.

A horse's harness is a useful analogy. It does not make the horse stronger or faster. It connects the horse's power to the load and puts the reins in the driver's hands. The horse is unchanged, but the harness makes it more useful for a specific purpose.

The same holds for AI models. ClawsBench, published in April, ran 6 models against the same tasks with and without a harness around them. Bare, every model scored between 0% and 8%. With a harness, the same models scored between 39% and 63%. That is a lift of 39 to 63 points from the harness alone. Swapping the model, with the harness held constant, moves the result by just 10 points. The choice of harness matters more than the choice of model.

Consistency is the hard problem which harnesses solve, and consistency is the problem at the heart of accounting AI. APEX-Accounting, published in July, is a 160-task test in which every task was written, solved and given a grading rubric by qualified accountants and bookkeepers. Across the whole field, no model got more than 21.5% of tasks right even once in 8 attempts, and no model got more than 2.6% right on all 8. What that suggests: Bare foundational models on their own are a long way from output a firm can rely on every month.

The gap between getting it right once and getting it right every time is familiar territory for accountants. A junior who reconciles an account correctly once is not yet someone you let sign off alone. And the structure of close management partly exists to mitigate that gap, to ensure that a firm gets a dependable result from people who are not always right.

A close checklist breaks a large job into defined pieces. A workpaper defines what the finished output should look like and leaves an evidence trail back to the general ledger. Dependencies stop step four running before step three. Materiality decides what deserves attention. A tie-out either passes or it does not. A manual review at the end is the ultimate sign-off gate.

Firms filled that structure with people, because people were the only option. The checklist described the terms of the work, and humans went and did it somewhere else: in a workpaper, in the ledger, in their heads.

In that sense, close management has always been a harness. It has just been a harness built for humans. Now that agents can do part of the work, close management software has to become a harness for agents too.


Artifact’s close management tool is a harness

Most close management software does four things:

  • Standardization: structured formatting for each working paper

  • Sequencing: dynamic tracking and dependency-based ordering of tasks

  • Ownership: defining preparers and reviewers

  • Reporting: oversight into the status of each item in a close

All four of these things concentrate on coordinating the work. None of them touch the work itself.

A modern harness has to do more than coordinate the work, because an agent alone brings none of what a firm’s accountant brings to a checklist item. An accountant arrives with training, firm context, client history, and the instinct to stop when something looks wrong. An agent without a harness arrives with nothing other than a language model.

Our approach at Artifact is to build the rest: a harness that gives each close agent what it needs to do the work dependably every time – Context, Structure, Evidence, Orchestration.

Context: what the agent knows before it starts

A foundational model knows accounting. It does not know that one client books its intercompany eliminations in a set order, or that the firm settled on a treatment for a category of cost after two partners talked it through in 2019. That institutional knowledge is the firm's real asset, and it usually lives in people's heads.

In Artifact, that knowledge lives in two places. Client Intelligence holds what is true of one client: the profile, what a normal month looks like, recurring patterns, the things to watch. Firm Intelligence holds the firm's own conventions and standard treatments, and applies to every client the firm runs. Both are seeded from the client's own ledger, taught by the documents a firm uploads, and extended after every close. Close agents read both before they touch anything.

Client intelligence: the profile, what a normal month looks like, and the recurring patterns, held per client and read before any work starts.


The firm level is the half that changes a practice. A convention written down once reaches every close the firm runs. A new preparer inherits it on day one. A partner stops answering the same question every quarter.

It also compounds automatically. After each close, a pass reads what the reviewers changed: which cells they edited, which journals they rewrote or dismissed, and the reason they gave. Those reasons are captured as structured categories, which is what makes them usable next period, not just auditable. 

This is where working with agents resembles working with humans. Marking a figure wrong corrects one figure. Explaining why it was wrong, that the accrual should have waited for the invoice, or that this client books that cost elsewhere, increases the likelihood that the agent/’s output will be dependable next month.


Structure: what the job actually is

A checklist item written for a person can afford to be vague. "Reconcile the operating account" is plenty for a preparer who has done it forty times. An agent needs the job specified.

Every item an Artifact agent runs carries standing instructions in three fixed parts: what to collect, how to work it, and what to produce. Those instructions live in a template read out of the workbook the firm already closes with, so each item arrives tagged with the cell it came from and bound to the workpaper tab it belongs in. What a firm signs off at setup is its own close, line by line, against the sheet it was read off.

A template's dependency map, with each item traceable to the workbook cell it was read from and bound to the working paper it produces.


The property that matters most: an Artifact close template is a firm asset, not a client one. Build it once and run it against every client it fits, supplying the client and the period each time. Clone it and specialize it for construction clients, for nonprofits running fund accounting, for SaaS companies with deferred revenue. Close logic that genuinely differs by industry gets written once per industry.

This is the part with no equivalent in prompting a foundational model. A good prompt is a good result once. An Artifact close template is the same result every month, for every client it fits, improved by everyone who touches it.


Constraints: what the agent may not do

Every vendor in this market now promises human oversight, in almost identical words. The real question is whether the oversight is a promise or a rule the software cannot break.

In Artifact’s close management harness, each agent holds a short list of the things it may do, and anything off that list it cannot do at all. These are rules which the software cannot break. Anything that only reads data proceeds on its own. Anything that would change data stops and asks, showing the specific change it wants to make.

Four rules follow:

  • An agent cannot mark its own work complete. A finished run leaves the item awaiting review, and the progress figure a partner reads counts approvals, so a close can never look further along than what somebody has read.

  • An agent will not overwrite a person. One writer holds a workbook at a time. If an accountant has it open the agent waits, and the hold lapses when they walk away, so a forgotten tab cannot stall a close.

  • An agent that is unsure stops. Two overlapping depreciation entries that would double-count produce a question, not a decision. The item waits as long as it takes, and the close reports that it is waiting on a person.

  • An agent cannot post to the ledger. When you approve a journal, it posts under your login. That permission belongs to a person, and the software has no way around it.

Input required: the agent found two overlapping depreciation entries that would double-count $224.92, and asked which to keep.


The rules bind the reviewer too. Discarding a proposed journal requires a reason, and editing one revalidates it, so a journal cannot leave review with debits and credits out of balance. A sign-off can be withdrawn while the close is open, and once the close is done the software refuses, so the books do not get quietly reopened later.


Evidence: how anyone knows the number is right

A believable but inaccurate workpaper, produced faster than a person could produce it, is a new kind of risk. A number can be arithmetically perfect and still impossible to review.

In AccountingBench (2025), a research team had AI agents close a real company's books month after month for a year. Accuracy held above 95% for the first few months, then slipped. When the figures would not tie to the bank balance, the agents started pulling in unrelated transactions to make up the difference, which the researchers had explicitly instructed them not to do. What the agents were measured on was whether the reconciliation passed, and a forced tie passed that test.

In Artifact, the evidence an agent relies on to generate an output gets shown to the user first. Every material figure carries its own support: where the number came from, the calculation as a formula and in plain English, the reasoning, whether it tied, and how material it is. Click the cell and you can read it. Findings are ranked by materiality, because a close that flags everything equally is a close nobody reviews properly. If an accountant edits a figure that already has support recorded against it, the entry is marked edited since audit, because the recorded support no longer describes what is in the cell.

The support behind one material figure: the calculation, the reasoning, the source, how material it is, and whether it tied.


The harness also knows when support is missing. A workpaper produced with nothing recorded behind its figures is marked as exactly that, so an unsupported workpaper arrives looking unsupported. That is what makes review scale: a reviewer spends attention on the items that need it, not on the ones that arrive explained.


Orchestration: many agents, many people, one close

Starting a close releases the whole checklist at once. Items with dependencies wait, and they show as waiting, so an item stuck behind the fixed-asset roll-forward is never dressed up as in progress. Several people can work the same close at once on different items. Preparation happens in parallel. Sign-off does not, and should not.

The board is where those rules become visible, because it only offers a move the software will honor. Drag an item to “approved” and it is signed off. Drag one back and the agent picks it up again.

One close, many items in flight, each sitting in a state a person can act on.


The close is also steerable in plain language. You change a template by describing what you want changed, and you correct a journal by saying what was wrong with it. Redirect a running item and you are talking to the agent doing the work, mid-job, so it carries on from what it has instead of starting over. Natural language is what keeps the harness adjustable by the people who own the process.


What a harness gives a firm that a prompt cannot

A model with nothing built around it can still do impressive work on one item, once, for someone who knows how to ask. That is riding a horse bareback. You may well arrive at your destination. You would not put the firm's month-end on it.

Three things follow from having a customisable harness that is built for close management:

  • Institutional knowledge stops walking out of the door. The treatment a senior partner carries in their head becomes firm intelligence that every agent reads. The judgment a reviewer applies becomes an instruction inside a template. Expertise ends up recorded in the system that does the work, and it survives the person who had it.

  • The unit that improves is the firm, not the engagement. An instruction sharpened on one client's close improves the next close for every client on that template. Thirty clients on one template means one improvement landing thirty times. That is horizontal scale, and it is the thing a prompt cannot give you, because a prompt improves one conversation.

  • Preparation scales without hiring. Preparation runs in parallel across a book of clients overnight. The constraint moves to review, which is the work you actually want your seniors doing.


The close itself has to be redesigned

This is the part most firms have not started, and it matters more than which software they buy.

A close checklist is a design document, and every item in it was written for a person. The granularity, the sequencing, the wording, the amount left unsaid. "Investigate significant variances" assumes a reader who knows what significant means for this client. Items are batched the way they are because that is how work gets handed between people.

None of those assumptions hold for an agent. An agent does not need items grouped to fit somebody's schedule. It does not tire, so sequence can follow real dependency instead of convenience. It cannot infer what a partner would have inferred, so anything implicit has to be made explicit. And it can run thirty items at once, which makes serial ordering a product of staffing instead of a control.

So the useful question for a firm is not which tool to buy. It is which items on your checklist exist because the work requires them, and which exist because a person was doing the work.


The question that is left

A year ago the question was whether an agent could prepare a working paper at all. The benchmarks above show why that was the wrong thing to get stuck on. Capability arrived before reliability, and reliability is a property of the harness.

The questions left are harder and more interesting. How do you coordinate a fleet of agents across a whole book of clients? How much of your firm's judgment can you actually write down? Which of your close procedures were built around human limits, and what do they look like once those limits are gone?

Close management used to mean arranging people around the work. It is becoming the practice of designing the conditions the work runs in. The firms that move first will be the ones that treated their own procedures as the thing to redesign.

Ready to Delegate

the Busy work?

Let Arti do the work so your team can lead.

Ready to Delegate

the Busy work?

Let Arti do the work so your team can lead.

Ready to Delegate

the Busy work?

Let Arti do the work so your team can lead.