← Blog

Dutch Spy Agencies Accused of Training AI on Leak Data

Share on X

July 1, 2026: Dutch digital-rights group Bits of Freedom accused the Netherlands' intelligence services of training in-house artificial intelligence on massive bulk datasets—including information they "appear to be purchasing" from data leaks—while the independent watchdog CTIVD separately concluded that AIVD and MIVD do not always comply with the law when processing those collections, per NL Times.

What the CTIVD reported

The Commission for the Oversight of Intelligence and Security Services (CTIVD) reviewed how the AIVD (civilian intelligence) and MIVD (military intelligence) handle so-called bulk datasets—collections that can contain millions of entries with names, phone numbers, location information, and other sensitive fields.

These datasets may come from other public authorities, private data brokers, or, as CTIVD notes in general terms, material illegally obtained by hackers and sold or shared online. Intelligence agencies use such data to investigate terrorism and espionage, but Dutch law requires strict retention limits, narrow access, and deletion of irrelevant material.

CTIVD found:

  • More agency departments are working with bulk collections without centralized compliance oversight
  • Some employees accessed sensitive data without required authorization
  • Processing constitutes a privacy invasion for people in the datasets who "have nothing to do with espionage or terrorism," chair Hugo Hillenaar said

Defense Minister Dilan Yeşilgöz and Interior Minister Pieter Heerma told NL Times they are broadly following CTIVD recommendations, with improvements already underway, while defending bulk processing as "essential" to national security work.

Bits of Freedom: AI plus leak-market data

Bits of Freedom director Evelyn Austin argued the services have a history of over-collecting citizens' data and were forced to delete excessive holdings in prior years—yet "apparently have not learned from it."

The group's July 2026 allegation goes further: intelligence agencies may be training proprietary AI systems on citizens' data, and "they even appear to be purchasing data that comes from data leaks." If substantiated, that would mean breach victims whose information circulated on criminal forums could see their data repurposed for state surveillance models—a novel privacy harm distinct from ordinary leak resale to fraudsters.

At catalog time, BreachHistory has not indexed a named corporate victim breach row for this story; it is an oversight-and-advocacy dispute about how the Dutch state uses third-party and potentially leak-derived bulk corpora.

Why this matters for breach victims globally

When a company loses your Social Security number or address history to a forum dump, conventional advice focuses on credit freezes and phishing awareness. The Dutch debate adds a policy layer: leaked datasets may re-enter the economy not only through criminals but through lawful-ish bulk purchases by governments building AI analytics pipelines—often with limited public visibility.

Austin warned that expanding spy powers while oversight finds ongoing compliance failures is "particularly concerning" because intelligence services "are designed to operate in secret."

What defenders and policymakers should watch

  1. Provenance requirements for any bulk dataset used in ML training—ban or restrict leak-sourced purchases
  2. Centralized access governance when multiple departments ingest the same bulk collection
  3. Deletion audits to ensure irrelevant citizen rows are purged, not retained for model fine-tuning
  4. Transparency reports on AI systems trained on personal data, analogous to breach-notification ethics

Related reading

For context on how massive broker breaches feed downstream abuse, see BreachHistory's canonical record for the National Public Data ~2.8B-row exposure—the kind of corpus intelligence and fraud ecosystems alike find valuable.

Primary source: NL Times — Dutch intelligence agencies accused of privacy breaches (July 1, 2026).