Every dataset that has ever been called «anonymous» carries the same assumption: strip out the names and the people disappear. The evidence says otherwise. Data is a fingerprint, and reidentification — putting the names back — has become routine enough that the GDPR treats it as a live risk rather than a theoretical one.
In this article we will discuss...
The Australian medical records
A few years ago the Australian government published an «anonymised» dataset covering medical billing records, including every prescription and surgical procedure, for 2.9 million people. Names and other identifying features had been removed to protect privacy.
A research team at the University of Melbourne found it straightforward to identify individuals and read their medical histories — without any consent, simply by comparing and cross-referencing the dataset against other publicly available information. The government withdrew the dataset after it had already been downloaded 1,500 times. By then it was too late.
The browsing data that was not anonymous either
German researchers demonstrated the same point from a different angle. Working with supposedly anonymous browsing records — no names, no account identifiers, only the pattern of pages visited — they were able to attach real identities to individual users.
The method was mundane. A public social media profile, a page visited while logged in, a small set of sites that only one person in a dataset would plausibly visit in that order: the combination is unique long before it looks unique. Among the people they identified were a judge and a politician, along with their browsing histories.
Why removing the name is not anonymisation
The two cases share a structure. In both, someone removed direct identifiers and concluded the job was done. In both, the remaining data was still a unique pattern — a fingerprint — and a fingerprint only needs one matching reference to become a name.
The academic literature has been consistent about this for over a decade. A handful of data points about location, timing or consumption is usually enough to single out one person in a population of millions. The more granular the data, the fewer points you need.
This is why the distinction between anonymisation and pseudonymisation matters so much in practice, and why so many organisations get it wrong.
- Pseudonymised data is still personal data. The identifiers have been replaced, but reidentification remains possible with additional information. The GDPR applies in full.
- Anonymised data falls outside the GDPR — but only if reidentification is genuinely impossible, taking account of all the means reasonably likely to be used, by the controller or by anyone else, and of technology as it develops.
That second definition is far more demanding than the way the word «anonymised» is used in most internal conversations. Replacing a customer number with a hash is not anonymisation. Aggregating into groups of five is not necessarily anonymisation either, if the groups are small enough to single someone out.
The European position on AI models
The same reasoning has now reached artificial intelligence. In its Opinion 28/2024, the European Data Protection Board held that a model trained on personal data cannot be presumed anonymous. It has to be assessed case by case, and the controller must be able to show that both direct extraction of personal data from the model and retrieval through queries are insignificant possibilities.
It is the Australian dataset argument transposed: the fact that you cannot see the personal data does not mean it is not there.
What this means for a company
Most organisations hold at least one dataset they describe internally as anonymous and treat accordingly — sharing it with partners, using it for purposes never disclosed, keeping it indefinitely, leaving it out of the record of processing activities. If that dataset is in fact only pseudonymised, every one of those decisions is a processing operation without a legal basis.
Four questions usually settle it:
- Could any single row be traced back to one person by combining it with data that is publicly available or that a partner already holds?
- Does anyone still hold the key — the mapping table, the salt, the original export? If so, the data is pseudonymised, not anonymous.
- How granular is it? Precise timestamps, exact locations and complete event histories are the fields that make records unique.
- Is the assessment written down? Under the accountability principle, the burden of proving anonymisation sits with you.
What to do about it
Where data genuinely needs to be anonymous — for publication, for sharing, for use beyond the original purpose — anonymisation has to be designed, not asserted. That means generalising granular fields, suppressing rare values, adding noise where the analysis tolerates it, and testing the result against the reidentification attacks that are actually plausible for that dataset.
And where it does not need to be anonymous, the honest answer is usually simpler: keep treating it as personal data, document the legal basis and set a retention period. Most «anonymisation» projects exist to avoid a conversation that would have taken an afternoon.
Checking whether the datasets a company calls anonymous really are is one of the things we verify in a data protection audit. It is also, more often than not, where the surprises are.





Leave a Reply
Want to join the discussion?Feel free to contribute!