PDPA: Does Removing a Name Make Data Anonymous?

ai graphic imaging

PDPA: Does Removing a Name Make Data Anonymous?

Organizations increasingly remove names and other obvious identifiers from datasets before using the data for analytics, research, artificial intelligence development, or sharing it with third parties. A common assumption is that once names and identification numbers have been removed, the information is no longer personal data and therefore falls outside the Personal Data Protection Act B.E. 2562 (2019) (PDPA).

That assumption can be dangerous. Recent guidance from the Personal Data Protection Committee (PDPC) provides an important reminder: removing a person’s name does not, by itself, make data anonymous.

When does data become anonymous?

The key question is not simply whether direct identifiers have been deleted, but whether an individual can still be identified from the remaining information.

The PDPC’s approach indicates that data may be treated as anonymized where the data subject cannot be identified without additional information and appropriate technical and organizational measures have been implemented to ensure that identification is not reasonably possible in practice.

This distinction is particularly important where a dataset contains multiple indirect identifiers. Removing a person’s name while retaining information such as age, date of birth, location, gender, occupation, accident location, medical information, or other characteristics may still allow that person to be identified when those data points are considered together or combined with information from other sources.

Accordingly, de-identification is a question of substance, not merely the removal of specified fields.

Pseudonymization is not anonymization:

Organizations should also distinguish anonymization from pseudonymization.

If a person’s name is replaced with a code but the organization retains a separate table linking that code to the individual’s identity, the information remains capable of being attributed to that person. The dataset is therefore pseudonymized rather than truly anonymized.

Pseudonymization can be an important security and privacy measure, but it does not automatically take the data outside the PDPA. By contrast, properly anonymized information that can no longer reasonably be linked to an identifiable individual is no longer personal data for PDPA purposes.

This distinction has significant practical consequences. An organization cannot simply label a dataset “anonymous” or remove names and assume that the PDPA no longer applies.

Research provides a useful illustration:

The issue arose in the context of research into the causes of motorcycle accidents. Such research may involve ordinary personal data as well as sensitive personal data, particularly health information concerning injured persons.

The PDPA expressly recognizes research and statistical purposes as circumstances in which personal data may be processed without relying exclusively on consent. Section 24(1) provides a basis relating to research or statistical purposes, subject to appropriate safeguards protecting the rights and freedoms of data subjects. For sensitive personal data, Section 26(5)(d) similarly permits processing where necessary for scientific, historical or statistical research, or other public-interest purposes, subject to necessity and appropriate safeguards.

The PDPC has also prescribed specific safeguards for processing personal data for research and statistical purposes.

The practical significance is that organizations should not assume that anonymization is the only way to conduct research lawfully. Personal data may remain subject to the PDPA and nevertheless be processed for legitimate research purposes where the applicable statutory requirements and safeguards are satisfied.

What should organizations do in practice?

For organizations seeking to take datasets outside the scope of the PDPA, anonymization should be treated as a risk-based technical and governance process, rather than a simple data-cleaning exercise.

Direct identifiers should be removed, but organizations should also assess combinations of indirect identifiers and consider whether information could be matched against other reasonably available datasets. Where coded identifiers have been used, organizations should consider whether any linkage mechanism remains available. If a mapping table between codes and identities continues to exist, the resulting dataset is likely to remain pseudonymized rather than anonymous.

Depending on the nature of the dataset, additional techniques may be necessary, including aggregation, generalization, suppression, masking, reducing geographic or temporal precision, and other techniques designed to reduce re-identification risk.

Organizations should also document the anonymization process, the methodology used, the potential sources of re-identification, and the conclusion reached regarding residual risk. Technical measures should be accompanied by organizational controls restricting access and preventing attempts to re-identify individuals.

Why this matters for AI, analytics and data sharing:

The distinction has implications well beyond academic research.

Businesses increasingly want to use existing customer, employee, patient, transaction, location, behavioral, and operational datasets for AI development and training, statistical analysis, product improvement, or collaboration with external service providers and research institutions.

Where the information remains identifiable, the organization must continue to consider the PDPA requirements applicable to its collection, use, disclosure, retention, security, and other processing activities.

Where information has been effectively and irreversibly anonymized so that individuals are no longer reasonably identifiable in practice, however, the resulting dataset may fall outside the scope of the PDPA.

This makes anonymization potentially valuable for data-driven businesses, but it also means that organizations should be cautious about treating anonymization as a shortcut around data protection obligations. A dataset that can realistically be reconstructed, linked, or matched back to individuals remains exposed to PDPA risk regardless of what the organization calls it.

Key Takeaways:

  • Deleting names does not automatically anonymize personal data.
  • Organizations must consider whether individuals remain identifiable from other information or combinations of information in the dataset.
  • Pseudonymized data remains personal data where a person can be re-identified using additional information.
  • Proper anonymization requires technical and organizational measures designed to make re-identification impracticable.
  • Research and statistical processing may have specific legal bases under Sections 24(1) and 26(5)(d), subject to appropriate safeguards.
  • Organizations using data for AI, analytics, research or external data sharing should conduct and document a re-identification risk assessment before treating a dataset as outside the PDPA.
  • Anonymization should be viewed as an ongoing risk-management and governance exercise, not simply the deletion of names or identification numbers.

Author: Panisa Suwanmatajarn, Managing Partner.

Other Articles

Posted in