Finetuning purpose of data protection law in Korea: data subjects’ rights on pseudonymized data – Keynote at APSN

by | Sep 28, 2026 | Free Speech, Open Blog, Privacy | 0 comments

K.S. Park's Keynote speech at 10th Asian Privacy Scholars Network held at Chinese University of Hong Kong, September 25, 2026


AI training requires collection and use of enormous amount of personal data already collected for other purposes or newly collected for the training purposes.  As AI training will surely count as processing, lawful bases have been developed, for instance, by CNIL to justify that. 

AI training is a learning.  A data controller can simply store 10,000 data sets or hold them in the form of vectors in the learning of AI.  As to the data collected and used independently justified by a lawful ground, whether data controllers process with or without AI in principle should be within the reasonable ambit of the original permitted purpose.  Whether data protection law’s purpose is privacy or informational self-determination may affect how we proceed on this issue.

In the meantime, Korean data protection law has created pseudonymized data as a saving grace that will balance data protection and data innovation.  Privacy concerns on the re-identification risk of pseudonymized data led the legislators to absolutely ban reidentification at all circumstances, making pseudonymized data legally indistinguishable from anonymized data. However, data subject rights such as access, erasure, correction, etc., cannot be protected if pseudonymized data are not reidentified.  As a result of the absolute ban on reidentification, data subject rights were all paralyzed as to pseudonymized data.  In reality, data privacy advocates’ campaign to stop telecom operators from selling the customers’ personal data to insurance companies was stopped in court as the telecom operators complained that they could not extract the advocates’ personal data from the database as reidentification is prohibited. 

The campaign was modified to request to cease such processing for future (i.e., with respect to the advocates’ personal data not yet pseudonymized):  The cease-processing right can be exercised without reidentifying pseudonymized data.  However, the Supreme Court intervened and ruled that pseudonymization, being a privacy-enhancing processing, is not a processing that is subject to the cease-processing right, and in the end that the advocates could not cease pseudonymization of their personal data and the subsequent use of the telecom data by insurance companies.  Once pseudonymized, the data can be used for ARS purposes and the advocates cannot stop such further processing, generating some anger among the privacy activists. 

__________________________________

A question to be generalized for all jurisdictions is whether reidentification of pseudonymization should be banned.  Pseudonymization is a security measure and therefore it is basic that pseudonymization should be deemed within the reasonable scope of original collection.  Then reversal of pseudonymization should be equally freely allowed or should be controlled like a ratchet?  What does “unauthorized reversal of pseudonymization” mean in GDPR?  If pseudonymization is a one-way street, there does not seem to be any problem with the Korean situation.  However, if the goal of data protection law is informational self-determination as opposed to privacy, sacrificing data subject rights so quickly seems too hollow.

A second question, modified to accommodate informational self-determination,  is whether reidentification of pseudonymization should be allowed at least for the purpose of affording data subjects’ rights such as access, cease-processing, etc.    

GDPR Article 11

1. If the purposes for which a controller processes personal data do not or do no longer require the identification of a data subject by the controller, the controller shall not be obliged to maintain, acquire or process additional information in order to identify the data subject for the sole purpose of complying with this Regulation.

2. Where, in cases referred to in paragraph 1 of this Article, the controller is able to demonstrate that it is not in a position to identify the data subject, the controller shall inform the data subject accordingly, if possible. In such cases, Articles 15 to 20 shall not apply except where the data subject, for the purpose of exercising his or her rights under those articles, provides additional information enabling his or her identification.

Under the above provision, data subject rights are deemed derogated when they cannot identify the data.  However, pseudonymized data are different:  as long as data subjects submit their identification data, the rights are revived.  Compared to GDPR, Korean law went too far in the direction of privacy (as opposed to informational self-determination) in derogating data subjects’ rights by banning reidentification even when data subjects present their identifying data.  Indeed, Korean law Article 28-7 legally derogates data subjects’ rights for pseudonymized data on top of Article 28-5 making such affordance factually difficult. 

As a result, pseudonymization becomes a dangerous processing for Korean data subjects because it obliterates data subjects’ rights. This means that data subjects will have interest in ceasing pseudonymization to protect their rights. 

Now, the civil society demanded that the Article 28-7 derogation of data subject rights (Article 28-5 reidentification ban) apply only to pseudonymization done for the purpose of ARS purposes, arguing that such absolute derogation should be reserved only in exchange for public interest imbued in ARS processing.  However, this creates other anomalies: pseudonymization can be done for other purposes than ARS processing such as safety and security.  For those non-ARS-bound pseudonymizations, data controllers are under obligations to afford all data subject rights.  It means that, for the socially beneficial processings (ARS), the dial is turned more toward privacy while data subjects must sacrifice their informational self-control.     

The fact that pseudonymization may be disfavored by data subjects leads to our third question:  Should data subjects have the right to cease processing even if the processing is pseudonymization and therefore privacy-enhancing?  It is basic that pseudonymization should be deemed within the reasonable scope of original collection, so initial processing is already lawful.  But cease-processing right probably means that data subjects are entitled to cease processing even if the initial processing was lawful without their consent. If so, what if the contemplated processing is privacy-enhancing?  What is the nature of cease-processing right?

This again deeply interrogates the purpose of data protection law.  If the Holy Grail of data protection law is privacy, we should probably allow the lawful processing to move forward if it is privacy-enhancing.  If data protection law has its own purpose such as informational self-determination, the privacy-enhancing processing should be ceased even if it was initially lawful. 

In following such interrogation, we should note that GDPR allows the cease-processing right only when data was to be initially processed for (1)  public interest, (2) data controllers’ legitimate interest, and (3) ARS purposes. Korean law allows the cease-processing right only when data was to be initially processed for (1) data controller’s legitimate interest and (2) ARS purposes and (3) data subjects’ consent. (Under GDPR, one can obtain the effect of “cease-processing” by exercising the erasure right which can be exercised against the processing done under data subject’s consent. Under Korean law, erasure right can be exercised against all processings except the one done under other mandatory laws and regulations while GDPR allows erasure to be exercised against the processing “no longer necessary”, consented to, or other limited processings.)  Cease-processing right does not seem to entitle data subjects to absolute control on processing (i.e. informational self-determination) as there is already a cross-border difference on the scope of cease-processing right.

A fourth question generalized for all jurisdictions is the nature of ARS processing. GDPR does not make pseudonymization an absolute condition of ARS purposes.  Korean law did,  ostensibly for the purpose of protecting privacy of data subjects whose data are used for ARS purposes.  However, there are times pseudonymization of data make ARS especially scientific research difficult such as longitudinal studies of long duration.  In Korea, especially where reidentification is absolutely banned, this problem becomes even severe.  It is as late as 2026 that the government recognized the problem and began a legislative amendment but the civil society already entrenched in the thought of pseudonymization only as a prerequisite to non-consensual ARS processing have registered opposition.  Within EU, CNIL and other data protection authorities have allowed non-pseudonymized AI training even for non-ARS purposes under data controllers’ legitimate interest as a lawful basis.  The current civil society’s position becomes paradoxical as ARS processing becomes more difficult (as it must be done through pseudonymization) than non-ARS processing which can theoretically be done without pseudonymization under the CNIL-sanctioned ‘legitimate interest’.   This may have all begun from the time when ARS processing (generally a socially desirable thing compared to non-ARS processing) was conditioned upon pseudonymization.  If we assume that ARS processing originates from collective informational self-determination, we are losing more of that in favor of privacy.  

In both situations (1) absolute banning of reidentifcation of pseudonymized data and (2) making pseudonymization an absolute condition of ARS processing, the policymakers’ intent was to provide certainty by erring in the direction of privacy enhancement. However, such little tweak caused a set of ripple effects: (1) makes pseudonymization dangerous for data subject rights (2) allows data controllers to reidentify non-ARS-related pseudonymization freely; and (3) makes ARS processing more difficult as it requires all AI training to be done on a pseudonymized basis.  

Data protection law is unique in a sense that most remedies or laws expanding human autonomy require generation of personal data but data protection law restricts such data generation.  For instance, to sue for racial discrimination, the racial profiling of the plaintiff and others similarly situated are central to such action.  Data protection law restricts collection and production of the very data needed for justice.  The Korean fiasco on pseudonymized data is a special case of a conflict between remedies vs privacy and brings us back to a debate on the purpose of data protection law – is it privacy or is it informational self-determination?

The debate carries even higher a stake in the age of artificial intelligence as AI becomes powerful and its fair distribution among people seems to be one of the few safeguards from the existentialistic fear of civilizational destruction (a la Zuckerberg’s 2026 Declaration).  Data portability is designed to reduce the network effects holding together data oligopolies by strengthening data subjects’ informational self-determination but has been also criticized for encouraging undesired data transactions among individuals whereby powerless people sell their data out for pennies to the oligopolies and end up with a privacy hell.  Yes, complete open source AI movement is being discussed but does it open up another hell gate? We didn’t succeed in non-proliferation of nuclear weapons by open sourcing nuclear technology while we did succeed in near non-use of nuclear weapons because several owners are checking and balancing one another.  How about AI, which is being sized up to be as dangerous as nuclear weapons?  If we go for non-(decimating)-use of AI by checking and balancing, at what level of open sourcing it do we stop? What is the role of data protection law in that project?

0 Comments

Submit a Comment

Your email address will not be published. Required fields are marked *

Recents