What is de-identification under the Privacy Act?
De-identification is the process of treating information so that an individual is no longer reasonably identifiable. The starting point is the definition of personal information in the Privacy Act 1988, which covers information about an identified individual, or an individual who is reasonably identifiable. De-identified information is information which has undergone a process of de-identification and no longer falls within that definition.
The threshold is expressed as risk, not certainty. Information will be de-identified where the risk of an individual being re-identified is very low in the relevant release context, also described as the data access environment. Put another way, there must be no reasonable likelihood of re-identification occurring. The Privacy Act does not require the risk to be removed entirely. It requires the risk to be mitigated until it is very low.
Two features of that test are easy to miss. First, it is contextual. The same dataset can be personal information in one environment and de-identified in another, which makes de-identification a property of the data and its setting together, not of the data alone. Second, the OAIC is explicit that this is a risk management exercise, not an exact science, and that a bright line between personal and de-identified information cannot always be drawn.
A de-identification process generally involves two steps. The first is removing direct identifiers such as name and address. The second is either removing or altering other information that could allow an individual to be identified, such as a rare characteristic or a combination of quasi-identifiers, or putting controls and safeguards in place in the data access environment, or both. Techniques the OAIC describes include sampling, removing quasi-identifiers such as significant dates, profession and income, and rounding values into ranges rather than exact figures.
Terminology is not settled. The OAIC notes that anonymisation and confidentialisation are used in Australia to describe similar processes, and advises confirming that all parties understand the terms the same way. What matters legally is whether an individual remains reasonably identifiable, not which label is applied.
Who does de-identification apply to?
De-identification is relevant to every APP entity, and it operates in two distinct modes.
As an obligation, APP 11.2 requires an entity that no longer needs personal information for any purpose for which it may be used or disclosed to take reasonable steps to destroy or de-identify it. Two conditions switch that duty off. It does not bite where the information is contained in a Commonwealth record, which matters for agencies and for Commonwealth contracted service providers, and it does not bite where the entity is required by or under an Australian law, or by a court or tribunal order, to retain the information. Retention schedules are therefore the first thing to check, not the last. Other principles also turn on de-identification, including APP 4.3 on unsolicited personal information and APP 6.4 on disclosure.
As an enabler, de-identification can permit sharing or secondary use that the Australian Privacy Principles would otherwise prevent. Because properly de-identified information is not personal information, the APP 6 limits on use and disclosure fall away. That is why de-identification carries so much weight in workers compensation, work health and safety and human resources work, where the underlying records are sensitive and the analytical appetite is high.
The practical caution is that the second mode depends entirely on getting the first right. An entity that treats direct-identifier removal as sufficient, then relies on the result to justify a secondary use, has not moved the data outside the Privacy Act at all.
Where does de-identification fit with AI?
De-identification is the control most often reached for when someone wants to put regulated data into an AI tool. It is a legitimate control, but the OAIC has been direct about its limits in this setting.
Information input into an AI system may be at risk of re-identification even when de-identified or anonymised. Once personal information has been entered into an AI system, particularly a generative product, it will be very difficult to track or control how it is used, and potentially impossible to remove it from the system. Large models are also good at exactly the inference that de-identification is meant to defeat, which means a release context that was low risk in a spreadsheet may not be low risk in a prompt.
Two OAIC positions follow. As a matter of best practice, organisations should not enter personal information, and particularly sensitive information, into publicly available AI chatbots and other publicly available generative AI tools. Where AI is used with personal information at all, only the minimum amount sufficient for the purpose should be used, and privacy-preserving techniques should be considered to reduce what goes into a prompt without compromising the accuracy of the output.
De-identification also appears as a remedy. Where an AI system generates or collects personal information the organisation is not permitted to collect under APP 3, including incidental capture in an AI meeting transcript, that information needs to be destroyed or de-identified.
What should practitioners do about de-identification?
Define the release context before choosing a technique. The question is never simply whether a dataset is de-identified, but whether it is de-identified for the environment it is about to enter. A controlled internal analytics environment and a third-party AI service are different environments and may need different treatment of the same records.
Use placeholders rather than raw records in prompts. For claims and case work, structured tokens such as claimant name, claim number and date of birth placeholders keep the reasoning intact while keeping identifiers out of the tool entirely. That is a stronger position than de-identifying a document and hoping the residual detail does not re-identify anyone.
Assess re-identification risk on the full picture, including attribute disclosure, spontaneous recognition by someone who knows the individual, and the gravity of harm if re-identification occurred. Small cohorts, rare injuries and unusual role titles defeat naive de-identification quickly.
Finally, record the reasoning. If de-identification is the basis on which data was shared, used for a secondary purpose or fed to a model, the assessment that supported it should be documented and revisited when the environment changes. Related material sits under the privacy topic hub and across the GRC section.
Bottom line
De-identification is a risk judgement about data and its environment together, not a checklist applied to a spreadsheet. The Privacy Act asks whether an individual remains reasonably identifiable, and the OAIC's threshold is that re-identification must be no more than a very low risk in the release context. Stripping names does not reach that threshold on its own. AI makes the context harder rather than easier, because a model can infer what a redaction removed, and because data entered into a generative system is difficult to control and potentially impossible to retrieve. Placeholders in prompts, minimum necessary information and a documented assessment are what hold up later.
TheAICommand. Intelligence, At Your Command.*
TheAICommand. Intelligence, At Your Command.
