Training means using examples to develop or adapt a model so the change can affect future responses. Processing a request to produce one answer is a separate step. Product terms may also describe retaining prompts for security, support or analytics. Those questions matter, but keeping data is not automatically the same as using it to train a model. Ask which process the provider means by improve the service, which data enters it, and which model is affected.
If you first need to know what can leave a practice file, read our piece on where your data goes. This article looks at what may happen after information reaches a provider.
Start with the exact plan
Write down the product name, account type, subscription plan and features you intend to use. A public chatbot, a business workspace, an API and a feature inside accounting software may have different terms or controls. Do not infer one policy from another. Check whether the setting covers prompts, uploads, outputs, ratings, feedback and data accessed through connected tools.
The OAIC's guidance on commercially available AI products recommends due diligence on whether a product suits its intended use, whether it has been tested for that use, how human oversight will work, the privacy and security risks, and who can access personal information entered into or generated by the product. Apply those questions to the exact service and plan your team will use.
Check whether a training control is enabled by default, who can change it, and whether it applies across every user and feature. Save the terms and settings you reviewed, with the date. A staff member using a personal account may not be covered by the workspace terms the practice checked.
Separate training from retention
Ask whether prompts, uploaded files, answers, feedback or connected data can be used to train or fine-tune a model. Find out whether that would improve a shared model, a model for one customer, or another product feature. Ask which controls apply to each type of information and whether the control covers future use only.
Then ask how long each type of information is kept, including service logs, feedback and copies held for safety or support. Ask who can access it, what deletion means, and whether a record remains after a chat or account is removed. A no-training setting does not answer those retention questions. A short retention period does not, by itself, explain whether information is used before deletion.
Read the training terms carefully
The OAIC's guidance on developing and training generative AI models explains that privacy obligations apply when personal information is used for training. For a developer using information already held for a different training purpose, the guidance says to consider whether that use would be reasonably expected and related to the original purpose. Where that cannot be established, it discusses consent, a meaningful, informed opt-out, or both.
That guidance is about privacy obligations for model developers, not a universal description of every vendor contract. It is still a useful prompt for a practice: ask the vendor what happens if client information enters a training process, and whether an opt-out prevents future use, covers information already submitted, or does both.
Ask for written answers
Send the vendor a short list tied to your exact plan and features:
- Can prompts, uploads, answers, feedback or connected data be used to train or fine-tune any model?
- Which setting or contract term controls each data type, and is it on or off by default?
- Does the control apply to every user, product feature and model in our workspace?
- How long are inputs, outputs and logs retained, and who can access them?
- How do we opt out, what happens to data already submitted, and how can we verify the change?
Ask for the answer in writing, with the product name, plan, applicable terms and date. Keep the response with your vendor review. If the answer relies on a setting, ask for the exact setting name and where an administrator can confirm it. Recheck when the plan, product or terms change.
If the answer is unclear
Treat an unanswered question as unresolved, not as confirmation that data is excluded. Use the least information needed, remove client identifiers where possible, and test workflows with synthetic records. The OAIC recommends that organisations do not enter personal information, particularly sensitive information, into publicly available generative AI tools as a matter of best practice.
How Toby states its position
Our current security page says nothing sent to TobyAI is used to train a model and that tax file numbers, bank account numbers, salary figures, employee names and raw organisation identifiers never reach any AI provider. Our FAQ also says nothing sent, uploaded or connected trains a model. Those are statements about Toby's current product, not a shortcut for checking another service.
Related reading
For the earlier step in that check, read where your data goes when you use AI in a practice.
Sources
- Office of the Australian Information Commissioner, Guidance on privacy and developing and training generative AI models. Supports the points on training personal information, secondary-purpose expectations, consent and opt-out.
- Office of the Australian Information Commissioner, Guidance on privacy and the use of commercially available AI products. Supports the due-diligence checks on product suitability, oversight, privacy and security risks, and access to input and output data.