Back to blog

AI in the grip of GDPR — and why compliance is decided by where you deploy

In 2026 the question is no longer whether to use AI, but where the data ends up when you do. We walk through what is going live right now under GDPR and the EU AI Act, where the real pressure points are, and why a surprisingly large part of compliance is not a legal question but an architectural one.

Today every company is one „copy–paste” gesture away from using artificial intelligence. A contract, an incoming invoice or a customer letter can be pasted into a chatbot, and seconds later the summary or the extracted data is ready. It is precisely this ease that hides the real risk: the same gesture is also a data-protection decision — because in that second the personal data in the document leaves the company and lands on a third party's server.

In 2026, therefore, the question is no longer whether to use AI, but where the data ends up when you do. And although the topic arrives dressed in legal robes — GDPR, the EU AI Act, regulatory fines — a surprisingly large part of compliance is not a legal question but an architectural one. It is decided by where the system runs.

Let's go through what applies to us, what is going live right now, and where, in practice, it is decided whether an AI rollout will be compliant.

Two rules that apply at the same time

The most common misconception is that the AI Act „replaces” or „overrides” the GDPR. It does not: the two apply in parallel. If the AI touches personal data — and a document-processing system almost certainly does — then both frameworks must be satisfied at once.

The two ask different questions. The GDPR asks: are you processing the personal data lawfully and fairly? — do you have a legal basis, did you inform the data subject, can you fulfil their rights. The AI Act asks: is the system itself safe, transparent and risk-proportionate? — can it be assigned to a risk category, is there human oversight, do users know they are dealing with AI.

It is worth clarifying the roles early too, because this determines who bears what. In its own rollout, most companies are the deployer under the AI Act and the controller under the GDPR — meaning they are responsible for what they use the system for and on what basis. Whoever supplies the system is the provider and the processor — they are responsible for making sure the system itself is compliant and secure. A good solution stands or falls on a healthy division of labour between these two roles.

And the stakes are not symbolic. The GDPR's fine ceiling reaches 4% of global annual turnover or EUR 20 million; for prohibited practices the AI Act goes even higher, up to 7% or EUR 35 million, and for high-risk systems up to 3% or EUR 15 million. The two sanctions can stand side by side for the same system. Hungary's data-protection authority (NAIH) is also actively watching the area — it published a separate sector review of AI use in the Hungarian banking sector — so the topic is on the agenda at home too, not merely Brussels theory.

What is going live right now? The 2026 schedule

Here is the article's „news value”, because the schedule changed recently. At the end of 2025 the European Commission came out with the Digital Omnibus package, which simplifies some AI Act rules and — most importantly — postpones the hardest obligations. The Council gave final approval to the AI part of the package on 29 June 2026 (Parliament voted on 16 June); publication in the Official Journal is expected in July.

The key message that many misread: the high-risk obligations were pushed back, but transparency was not.

The updated EU AI Act timeline between 2025 and 2028 after the Digital Omnibus: a timeline from prohibited practices to the high-risk system deadlines, highlighting the transparency obligation on 2 August 2026.

Figure 1 — The updated EU AI Act deadlines after the Digital Omnibus.

Concretely:

  • 2 August 2026 — Article 50 of the AI Act, the transparency obligation, goes live. If your system interacts with people (e.g. a chatbot) or generates AI content, you must signal that it is AI. The Omnibus did not push this deadline back. From this point the supervisory authorities' full power to impose sanctions also applies.
  • 2 December 2026 — for systems already on the market, the grace period for machine-readable marking of synthetic content (watermarking) expires, and the new prohibitions take effect (e.g. systems generating non-consensual intimate content, or abusive content depicting children).
  • 2 December 2027 — from this date, stand-alone high-risk (Annex III) systems must comply — for example HR screening, credit scoring, biometric identification, educational assessment. This deadline was previously 2 August 2026; the Omnibus moved it here.
  • 2 August 2028 — the deadline for high-risk systems embedded in products (Annex I) (previously August 2027).

An important, sober conclusion: the postponement is not a rest break. High-risk compliance — a risk-management system, technical documentation, conformity assessment — is months, not weeks, of work. The time gained is best spent on preparation, not procrastination.

There was an important turn on the GDPR side too. The package's original draft would have narrowed the concept of personal data (certain pseudonymised data could have fallen outside the regulation's scope), but this proposal was ultimately removed. So the concept of personal data stayed broad: since even seemingly „anonymous” data can often be re-identified or inferred, most data in documents still falls under the GDPR. The European Data Protection Board (EDPB) had, in any case, cautioned earlier: an AI model trained on personal data cannot automatically be considered anonymous.

Where does it hurt in practice? Five pressure points

When we let AI loose on our documents, five concrete questions arise — these are the actual „pain points”:

1. Legal basis. On what basis are you processing the personal data in the document? Performance of a contract, a legal obligation, perhaps legitimate interest? If you rely on legitimate interest, the EDPB requires a three-step balancing test: is there a real, lawful interest; is the processing necessary for it; and are the data subject's rights not overriding. This is not a formality — it must be documented.

2. Transparency and information. Does the data subject know that AI is involved in the process, and what happens to their data? Lack of information is one of the most common regulatory objections — and from August 2026 the AI Act's separate transparency layer is added on top.

3. Data-subject rights. Can you disclose, rectify or erase the personal data if the data subject asks? This looks easy while the data sits in a database — but it is more complicated if it has „baked into” a model. That is exactly why it matters where and how your system stores the data.

4. Automated decision-making. Article 22 of the GDPR limits legally significant decisions about people being made solely in an automated way, without human oversight. The correct pattern is human-in-the-loop: the AI prepares and suggests, a human makes the decision.

5. Transfer to a third country. This is the biggest and most often underestimated risk. If the document goes to a foreign cloud AI, the data leaves the EU — and we step into the legal quagmire well known since the Schrems II ruling.

Above all this hover a few general obligations: for large-scale or systematic processing a data protection impact assessment (DPIA) is often mandatory, a data processing agreement (DPA) is required with every processor, and appropriate data security is always needed.

Of the five points, the sharpest — transfer — depends entirely on where the data goes. And this is where the article's real thesis comes in.

The decisive question: where does the data go?

With every cloud-based AI service, when you send a prompt (and a document inside it), the data leaves your network and is processed on a third party's infrastructure. For most everyday tasks this is acceptable. For regulated, personal or commercially sensitive data, it is not.

This is also where the most common myth fails: „but I use an EU data centre, so I'm fine.” Unfortunately, not necessarily. The US CLOUD Act extends to a US provider's data stored anywhere — including its servers operating in the EU. So an EU region of a US cloud solves the data's geographic location, but does not solve sovereignty: the data can still fall under a foreign jurisdiction.

The same document in two architectures: with cloud AI the personal data crosses the company and EU border to an external provider under a foreign jurisdiction; with on-premise the data stays within the company's own infrastructure throughout.

Figure 2 — The same document, two architectures. With cloud AI the data crosses the border; with on-premise it stays within the company throughout.

The clean architectural answer to this is on-premise (locally running) deployment: the model runs on the organisation's own infrastructure, and the data never leaves the network. With this:

  • the risk of transfer to a third country disappears (there is nothing to carry across the border);
  • this is the practical implementation of GDPR Article 25 — data protection by design and by default; protection is not bolted on afterwards but is part of the architecture;
  • if you use an open-weight model that is not retrained on your data, then the data does not „bake into” the model, and your documentation burden is far cleaner too.

Here comes the expert honesty, though, that separates real advice from marketing: on-premise is necessary but not sufficient. Just because the system runs locally, you still need a legal basis (Article 6), an impact assessment where required (Article 35), the technical feasibility of data-subject rights, plus proper access management, logging and security (Article 32). On-premise is the strongest starting point, not the finish line.

How far this is from a narrow technical topic is well illustrated by the fact that sovereign AI has become a boardroom-level question. Gartner forecasts that by 2030 the majority of European enterprises will „repatriate” their sensitive AI workloads — while today that share is still below 5%.

How we approach it

We build DocAI around exactly this principle: Hungarian-language, locally running document intelligence in which the data stays at the company. In practice this means processing happens on the client's own (or a dedicated on-premise infrastructure we operate) — the documents do not go to an external cloud AI, there is no OpenAI, Anthropic or Google API in the picture.

We run the model self-hosted, on open-weight foundations, and we do not train on the client's data — so the data does not bake into the model. Hungarian-language personal-data recognition (NER/PII) is built into the process, so personal data is identifiable and manageable within the document, which supports data minimisation and purpose limitation. Where possible, we use deterministic processing paths to reduce the „black box”. And the system helps the human decide, it does not replace them — in line with the logic of Article 22.

And because technology alone is not enough, the legal frame is in place too: a data processing agreement (DPA) under GDPR Article 28, with references built on the EU AI Act and an orderly liability framework.

One thing is important to state clearly: this is not „automatic GDPR compliance”. No such thing exists. This is the best possible starting point, onto which the company builds its own legal basis, impact assessment and processes — and in that we are partners.

A practical checklist for the company

If you are thinking about an AI rollout, these nine points are a good starting point:

  • AI inventory — in which process do you use or plan AI, and on what data?
  • Risk classification — is the use prohibited, high or limited risk? (Annex III?)
  • Legal basis recorded for every AI-based processing activity (Article 6)
  • Information — does the data subject know AI is involved? (Articles 13–14 + AI Act Article 50)
  • Data transfer — does the data leave the EU or the organisation? If so, is there a guarantee for it? (Schrems II)
  • DPIA where processing is large-scale or systematic (Article 35)
  • DPA with every processor (Article 28)
  • Human oversight for decisions with significant impact (Article 22)
  • Logging and storage limitation — do not keep personal data longer than necessary

On-premise deployment makes several points more favourable „by default” — above all the transfer and the data processing agreement to be signed with a cloud AI —, but the legal basis, the impact assessment and data-subject rights remain your responsibility.

Summary

Regulation is not the enemy of AI — much more so of bad architecture. Whoever asks the right questions — where does the data go, who makes the decision, what is the legal basis — will find that compliance is not a brake but a competitive advantage: a trust the client can verify too. There is time to prepare until the end-of-2027 deadline for high-risk obligations, but preparation is best started now. And the good news is that the hardest-looking question — protecting personal data — largely comes down to a single, well-made decision: where the system runs.

This article is for information only and does not constitute legal advice. In a specific situation it is worth involving a data-protection expert or lawyer.