Blog post
|
22.09.2026
Currently, the „Marketplace of AI Opportunities“, set up by the Federal Ministry for Digital Affairs and State Modernisation (BMDS), already lists 399 AI systems that could be used in public administration. According to a recent study, around 90 per cent of lawyers in Germany use at least one AI tool in their day-to-day work. In this context, similar questions arise for public authorities and the legal profession: what can be done with personal data when today’s language models (LLMs) come into play?
From an administrative perspective, this raises, amongst other things, the question of whether and to what extent employees’ data may be used to train the organisation’s own AI systems. From the legal profession’s perspective, it is necessary to consider whether anonymising documents – by redacting client-related identifiers such as names and addresses – is sufficient to ensure that, as holders of professional secrecy, they do not breach their duty of confidentiality.
When the authorities develop AI
Even from the perspective of data sovereignty, developing in-house AI systems within the public sector can make sense. Potential applications include those for human resources management itself, for example for allowances and business travel, performance appraisals or disciplinary matters. However, anyone wishing to train such systems themselves will need data. In human resources management, this primarily consists of employee data.
The European Regulation on Artificial Intelligence (AI Regulation) generally classifies human resources management systems as so-called high-risk AI systems in accordance with Article 6(2) in conjunction with Annex III, No. 4(b) of the AI Regulation, although such a classification must be assessed on a case-by-case basis. From 2 August 2027, these systems will be subject to specific documentation, quality management and compliance obligations, which also apply to a public authority where it acts as a provider of an AI system. In addition, there are the requirements of Article 10 of the AI Regulation regarding the quality of the training data used: this must be as representative, error-free and complete as possible.
The real hurdle lies in data protection law: Valid consent from employees to data processing under Article 7 of the General Data Protection Regulation (GDPR) will regularly fail due to the lack of voluntariness resulting from the hierarchical nature of an employment relationship (Article 7(4) GDPR). The requirement for informed consent is also likely to be difficult to meet given the lack of transparency in training systems. Similarly, withdrawing consent under Article 7(3) of the GDPR will be virtually impossible to implement, as data that has been trained into an AI model can hardly ever be removed from it.
Furthermore, training using employee data will not be covered by Section 26(1), first sentence, of the Federal Data Protection Act (BDSG). According to this provision, data processing for the purposes of the employment relationship may, in principle, be permissible. The AI—
However, training does not serve such a purpose. The general provision in Section 3 of the BDSG and the corresponding provisions in the state data protection laws, in turn, only permit processing of a limited scope, which is not the case when training an AI system.
Nor may an external provider use the employee data entrusted to it as a data processor to improve its AI system. The interests of the employees take precedence, as they cannot be expected to anticipate such use – in the case of civil servants, for example, due to the confidentiality of their personnel files.
The only viable option therefore remains training using anonymised employee data: the anonymisation itself can be justified as a minor interference under Article 6(1), first subparagraph, point (e) of the GDPR in conjunction with Section 3 of the Federal Data Protection Act (BDSG) or the relevant state general clause. For identifiable personal data, however, the legislator would have to establish a legal basis that not only assigns the task of developing AI systems to the public administration but also grants it the authority to access its employees’ data for this purpose.
For public authorities planning an AI project in human resources management, this has two implications: they must incorporate the anonymisation of training data into their project planning from the outset, as a legal prerequisite for the entire project. If a system from an external provider is used, the data processing agreement should be reviewed to ensure that it excludes the use of the data provided for training purposes.
When anonymisation is no longer enough
But when can data actually be regarded as anonymised? To this day, it is frequently assumed – including within the legal profession – that redacting direct identifiers such as name, address or job title is sufficient to anonymise the documents, thereby making it unnecessary to enter into a separate confidentiality agreement with the provider.
Such redactions stem from a traditional concept of confidentiality that can be described as „practical obscurity“. It is assumed that the effort involved in de-anonymising a redacted document is not economically proportionate to the benefit if the direct identifiers have been removed. Although the document remains theoretically reconstructible, in practical terms it is obscure.
This approach will no longer work in 2026. Recent studies on the use of LLMs show that commercially available language models are already achieving significant success rates in de-anonymisation: In a study conducted by ETH Zurich and the developers behind Anthropic’s Claude language model, for example, a language model correctly matched a good third of around 89,000 LinkedIn profiles to anonymous user posts on another internet forum. [Link to the study].
A traditional – and manual – keyword search achieved a hit rate of just 4.2 per cent under the same conditions.
Anyone who uploads documents that are only supposedly anonymised to a cloud-based language model, but which the model can identify – despite the redacted names – on the basis of the remaining quasi-identifiers – that is, factual details that appear insignificant at first glance, such as a reference to an industry or stylistic features – risks breaching their duty of confidentiality. For lawyers, these obligations are set out in Section 43a(2) of the Federal Lawyers’ Act (BRAO). In addition, there is also a risk of criminal liability for the breach of private secrets under Section 203(1)(3) of the Criminal Code (StGB).
Manual attempts at anonymisation are no longer sufficient
In an era where large language models are widely available, the protection of confidential information must therefore shift from manual anonymisation attempts to legal safeguards:
· The use of cloud-based language models by solicitors and other professionals bound by professional secrecy remains permissible only on the basis of an agreement with the provider, for example in accordance with Section 43e of the Federal Lawyers’ Act (BRAO), supplemented by the provider’s obligation under Section 203(4) of the Criminal Code (StGB), which covers confidentiality, a notice of criminal liability, the requirement for written form and the extension of these obligations to subcontractors.
· Where such an agreement cannot be reached, the client must give their consent on a case-by-case basis. A mere unilateral assurance by the AI provider in its general terms and conditions or privacy policy is not, however, sufficient.
Solicitors and other professionals bound by professional secrecy must review their existing contracts in this regard. Standard consumer subscriptions offered by major providers generally do not meet the requirements of Section 43e of the German Solicitors’ Act (BRAO) (or Section 203(4) of the German Criminal Code (StGB)).
My recommendation
- When using AI, public authorities and those bound by professional confidentiality should always determine the requirements for adequate anonymisation in light of the current state of the art.
- As long as there is no legal basis that both assigns the task of AI development to the administration and expressly grants it the authority to use its employees’ data, training AI using its own data sets entails significant legal risks.
The mind behind the article.
These topics were also discussed at this year’s Autumn Academy organised by the German Foundation for Law and Informatics (DSRI): Rechtsanwalt Jessen-Lieberum spoke on the „AI in Professions and Sectors“ panel about the end of practical obscurity and the implications for the legal profession; Rechtsanwalt Frenken explained, on the „Administration“ panel, the legal requirements for training AI systems using data held by public authorities.
They have written papers on this subject, which have been published in the proceedings of the Autumn Academy: Bernzen/Heinze/Steinrötter (eds.), Digital Economy 2030 – IT Law at a Turning Point?, Proceedings of the DSRI Autumn Academy 2026, C.H. Beck 2026.