An assistant that answers questions about your own figures is easy to demonstrate and hard to trust. The demo goes well, because the questions were picked and someone was watching. The risk sits in the answer that looks right: an amount that is plausible, neatly explained, with a reference to a ledger category, and nowhere to be found in your administration.
That is not a model defect that a better model removes. A language model predicts text, and a number is text. So you have to measure it, and measuring means something other than "we tried it for a week and it seemed fine".
I built a test harness for this. Not as a demo but as evidence: thirty questions, seven checks per answer, one log line per answer, and a test suite that at the time of writing has 190 tests green. What follows is the method out of it. The approach is not tied to one kind of administration.
Do not let the model do arithmetic
The most important decision sits before the model. Every figure comes from one layer that queries the source, and the model gets them handed to it. It may explain, summarise and refer on. Adding up is not its job.
That sounds like a restriction and it is the only reason the checking afterwards can work. If the model does its own arithmetic, you cannot tell a wrong answer apart from a fault in your integration. By making that cut you know, for every number in the answer, where it is supposed to come from.
One practical consequence: pass fewer figures, not more. The more numbers in the context, the more combinations of sum, difference and percentage are valid, and the bigger the chance that an invented number lands on one of them by accident.
Build a question set whose answers you already know
This is the work nobody wants to do and it decides everything. Thirty questions a colleague would recognise, each with the facts that have to appear in the answer.
One rule makes the difference: you do not type those expected facts, you compute them from the source. A typed benchmark goes quietly stale after the first data refresh, and then you are measuring your own typo instead of the assistant.
Include questions whose correct answer is "I don't know". A forecast, a benchmark, a financial year that is not there. An assistant that never refuses is not a careful assistant, it is an assistant that guesses at the questions you have not asked yet.
Seven checks per answer
One check for invented numbers is not enough, because "wrong" has several shapes here. My harness runs seven per answer. The form carries over, the content depends on your own domain.
Every number in the answer appears in the retrieved data, or follows from it with one addition, subtraction or percentage. The margin is half a cent, not a percentage. A question outside the data gets no estimate, and the other way round, a question that is in the data may not be refused. A figure comes with a source reference, and that reference has to really exist: a model knows category numbers from other charts of accounts and cites them just as fluently. A period that has not been closed never leaves the assistant as a fact. If you work for several clients in one system, no amount, invoice number or name belonging to one of them turns up in an answer about another. And an instruction somebody typed into an invoice description is read as text, not carried out.
That last one is not theoretical. A free-text field in your own administration is an input channel you built yourself.
Plant the failures yourself
A green suite says nothing until you know what turns red. Of the thirty recorded answers in my set, nine are deliberately written wrong, each on one failure path. Those nine produce nine findings and the twenty-one good ones produce none. That second half matters just as much: a check that also fires on correct answers gets switched off within a week.
One step further: remove the most important check in a throwaway copy and run the same tests. Seven of mine fall over. Without that proof, all a green suite tells you is that it is green, not that it enforces anything.
One log line per answer
Per answer, one line with the question, the retrieved source values, the answer, the outcome of each check, and which vendor, which model and which region produced it.
Those last three fields are why the log still has value later. A vendor swaps a model out without you changing anything. Without that field in your log, "it was better last month" is a feeling. With it, it is a line you can look up.
What a harness like this does not catch
This belongs in the story, otherwise you are selling certainty you do not have.
It does not measure whether an answer is complete. If a question has two causes in the data and the answer names one, that is not untrue and it passes clean. It does not measure whether an answer is understandable. It does not measure whether your source data is correct: if the integration delivers a wrong amount, the assistant explains that wrong amount correctly and traceably. And I recognise a refusal by a list of phrasings, so a model that refuses in a new way gets counted as "did answer". That is the first list that has to grow in real use.
Grounding protects you against invention. Not against an integration fault, and not against the wrong question.
Where to start
Pull twenty questions out of last week's work whose answers you already know, write down per question which figures belong in the answer, and record that list before you switch an assistant on. That is an afternoon of work and it is the only benchmark that survives a later change of model or vendor.
If you want to know how this plays out in your own administration, see AI automation. On where the data sits during a project like this, read keeping data under your own control.