Every month, before a model may keep serving Toby customers, we put it through the same test: a fixed set of accounting questions, asked three times over, scored on whether it sends each one to the right tool and whether its explanation says only what the source says. We run it because of how a model writes. It sounds just as sure when it is wrong as when it is right, and nothing in its reply tells the two apart.

Ask your AI assistant what a threshold is and the answer comes back in complete sentences, in a settled tone. Ask again and you may get the same figure in different words, or now and then a different figure in the same tone. Both follow from how the answer is made.

One piece at a time

A language model writes an answer one token, roughly three quarters of a word, at a time. At each step it scores every possible next token by how likely it is to follow the text so far, picks one, adds it, and repeats. The likelihoods come from training on a very large amount of text. The model never looks anything up. When it writes a dollar figure, that is because those characters were the likely continuation, not because it consulted a register.

Why the same question gets different answers

Most assistants do not always choose the most likely token. They sample: usually the likeliest, sometimes the second or third, so the writing doesn't read like a form letter. A setting called temperature controls how often. On an email draft that variation is a feature. On a figure it means the answer can change between two runs of the same question, and neither run tells you which one to trust.

Why it sounds sure

Three things push a model towards a confident tone, and none of them is evidence.

First, the text it learned from. Reference material, guidance notes and forum answers are written by people stating things plainly. A model that predicts that kind of text reproduces the plainness along with the content.

Second, the finishing training. After the first round, assistants are adjusted towards the answers people rated highly, a step called reinforcement learning from human feedback. People tend to prefer a direct answer to a hedge, so the model learns to give one.

Third, the way models are tested. Researchers at OpenAI argued in 2025 that most benchmarks give no credit for "I don't know", so a model that guesses scores better than one that abstains. Training towards those scores teaches the model to guess.

The result is that the confidence in the wording has almost no connection to whether the content is correct. A model can carry some internal signal of its own uncertainty, and research has shown that signal can be measured. It rarely reaches the tone of the reply.

What this means in practice

For drafting, summarising and explaining, none of this matters much. You read the result and correct it. For a rate, a threshold or a due date, it matters a great deal, because there is no difference on the page between a figure the model got right and one it did not. A confident tone is not a reason to rely on an answer, and a hesitant one is not a reason to doubt it.

The useful checks are the ones from our earlier piece on reading a compliance answer: is there a working, does the citation point to a paragraph that says what the answer says, is the date the period you asked about, and who signed off the inputs.

How we handle it at Toby

At Toby we do not let the model produce figures. Rates, thresholds and dates come from tables a named person has signed off against the published source, and the calculation is ordinary code that gives the same answer every time. The model writes the explanation around that answer and chooses which tool a question goes to.

Choosing the tool is still a model decision, and that is what the monthly test measures. In the 30 September run, one model sent "director drew on the loan account, is it a deemed dividend" to our Division 7A loan-repayment tool instead of to research. It did so in the same even tone it used for the questions it routed correctly, and nothing in the reply said it was unsure. The tool's description did not say what the tool does not cover. We rewrote the description and re-ran the full set, and the model kept its place only once it routed every question correctly.

Next

Where your data goes: what leaves your accounting file when you ask your AI assistant a question, and the residency questions worth putting to a vendor.

Sources

  1. Kalai, Nachum, Vempala and Zhang, Why Language Models Hallucinate (OpenAI, 2025).
  2. Kadavath et al., Language Models (Mostly) Know What They Know (2022).
  3. Ouyang et al., Training language models to follow instructions with human feedback (2022).