OAIC guidance: developing and training generative AI models
The privacy regulator's October 2024 guidance for anyone who builds, trains, adapts or fine-tunes a generative AI model with personal information. Public data is still personal information; reuse of customer data needs a reasonable expectation or consent.
- Status
- In force · Australia (Commonwealth) · OAIC
- Applies to
- Entities that develop, train, adapt or fine-tune generative AI models using personal information, including fine-tuning a vendor's model on your own data
- Primary source
- Official document →
- Last reviewed
- 28 September 2026
What it is
The companion to the OAIC’s guidance on using AI products, published the same day, 21 October 2024, and aimed at developers. The OAIC defines a developer broadly: “any organisation who designs, builds, trains, adapts or combines AI models”, which includes a business that fine-tunes a commercial model on its own customer data. Like its companion it creates no new law; it applies the Australian Privacy Principles to the act of building and training, and it is noticeably stricter in tone, telling developers to “take a cautious approach” and, where in doubt about whether the Act applies, to “err on the side of caution and assume it applies”.
What it requires
Four positions carry the weight. Publicly available data is still personal information: “just because data is publicly available or otherwise accessible does not mean it can legally be used to train or fine-tune generative AI models”, and scraping is a covert collection that must still be lawful and fair. Sensitive information needs consent: developers “must not collect sensitive information without consent unless an exception applies”, which catches photographs and recordings of people, and sensitive data collected without consent “will generally need to be destroyed or deleted from a dataset”. Existing customer data cannot simply be repurposed: training is usually a secondary purpose, “in many cases it will be difficult to establish that such a secondary use was within reasonable expectations”, so developers must seek consent and offer a meaningful opt-out. Third-party datasets need assurances about how they were collected. Beyond those, accuracy under APP 10 means high-quality data, testing for bias and clear statements of a model’s limits, and privacy by design under APP 1 means a privacy impact assessment before the work starts.
Does this reach your business?
If you build models, yes. If you fine-tune a vendor’s model on your own data, yes, and this is the case that surprises businesses: the moment customer records, support transcripts or employee data go into a training or fine-tuning run, you are a developer in the OAIC’s sense and the reasonable-expectations test applies. If you only use AI tools and never train or adapt one, the companion guidance on commercially available products is the one that reaches you. Businesses below the Privacy Act threshold are not bound, but a vendor or customer will usually require the same standard by contract.
What we recommend
Our advice is to treat “we will fine-tune on our data” as a privacy decision before it is a technical one. Before any training run: list the personal information in the dataset, remove or gain consent for anything sensitive, ask whether the people concerned would expect their data to train a model (for most customer data the honest answer is no), and if not, get consent and build an opt-out. Record where every third-party dataset came from and what assurances came with it. Run a privacy impact assessment and keep it; it is the document the OAIC will ask for first. If the training would have to happen without those steps, do not train.
Questions people ask
Yes. The OAIC's definition of a developer includes any organisation that adapts or fine-tunes a model, so fine-tuning on your customer data brings the guidance into play.
Not automatically. The OAIC says public availability does not make collection lawful; scraping is a covert collection that must be lawful and fair, and sensitive information such as photographs generally needs consent.
Only if the people would reasonably expect it, which the OAIC says will often be hard to establish. Otherwise seek consent and offer a meaningful opt-out.
The OAIC's position is that it will generally need to be destroyed or deleted from the dataset.