Home Research Russian Big Tech Wants Users’ Data. That’s Also a Censorship Move

Russian Big Tech Wants Users’ Data. That’s Also a Censorship Move

The Association of Big Data asked the Russian government to let it train AI on anonymized data about Russian citizens.

Sber, Yandex, VK, Rostelecom, Avito, and HH (hh.ru, a Russian job-search platform) asked Mintsifry (Russia’s Ministry of Digital Development) to let them train AI on anonymized data about Russian citizens without their consent. We look at what the law already allows, what business says is missing, and why the same companies approved standards that same week that make this request unnecessary. This piece continues our coverage of AI regulation. Our previous piece on the topic focused mainly on who would get “national” model status. This one is about what these models will actually learn from.

Summary

  • What happened?
    On August 3, 2026, the Association of Big Data (an industry group representing major Russian tech companies) sent Mintsifry a request to allow processing anonymized personal data without citizens’ consent for developing and training AI.
  • Who’s asking?
    The association includes Avito, Sber, Rostelecom, Yandex, HH, and VK.
  • What’s already allowed?
    Since September 1, 2025, processing anonymized data without consent has been legal. A state system is already running, and operators must hand over such data to it for free whenever Mintsifry demands it.
  • What else do they want?
    To take data processed with privacy-enhancing technologies out of the personal-data category entirely. Plus, to introduce a state-accredited “trusted intermediary” and split data into paid and free categories.
  • What’s the real twist?
    On the exact same day, August 6, the same association announced that state standards for synthetic data had been approved. According to Sber’s estimate, models trained on synthetic data lose 2-3% accuracy compared to models trained on real data.
  • Where’s the document?
    There isn’t one yet. The request itself hasn’t been published. We know its contents only from Kommersant’s reporting.

What’s happening

On August 6, the newspaper Kommersant reported on the Association of Big Data’s request to Mintsifry. The letter is dated August 3. The newspaper reviewed it, but the document never became publicly available. The ministry didn’t respond to the newspaper’s inquiry.

The association is asking for permission to process data using privacy-enhancing technologies without citizens’ consent, provided three conditions are met:

  • the risk of identifying a person is zero or a small fraction of a percent,
  • the data goes only toward developing and training AI, and
  • Russian companies handle the processing inside the country.

Privacy-enhancing technologies (PETs) are a set of mathematical and software methods that let you process data without exposing the original information about people. The association’s argument boils down to the fact that the law does not carve out a separate category for these technologies. Any algorithmic transformation of data counts as anonymization, so mature protection methods end up legally conflated with weak ones.

The association also proposes an organizational layer: a “trusted intermediary” accredited by government agencies that would run the information system for storing anonymized data and controlling access to it. It proposes splitting the data into two categories, paid and free.

The association hasn’t specified exactly who would get access to these datasets.

The existing legislation

The phrase “asking to simplify” creates the impression that anonymized data is currently banned. It isn’t. Federal Law No. 233-FZ of August 8, 2024, added Article 13.1 to the personal data law, introducing the concepts of “anonymized personal data” and “the composition of anonymized data.” Since then, processing such data without a person’s consent has been allowed, including for training AI. The only remaining ban covers using it for advertising, promoting products, and political campaigning.

The anonymization methods originate from Roskomnadzor Order No. 140 of June 19, 2025 (Roskomnadzor is Russia’s communications regulator). It specifies identifier substitution, changing the data’s composition, decomposition, and shuffling. Removing a name and phone number isn’t enough. The operator must make it impossible to reconstruct someone’s identity even from indirect clues.

At the same time, a state system has appeared where these datasets flow. It runs on the Unified Platform of the National Data Management System, and people in the industry call it the “state data lake.” Mintsifry can order any personal-data operator to anonymize information and hand it over there, and for free.

Business wants data that has passed through privacy-protection technologies to stop counting as personal data altogether, along with all the operator obligations and individual rights that come with that status. Operators already hand anonymized datasets to the state for free on demand. Among themselves, they propose trading it, splitting the data into paid and free categories.

It is censorship

Censorship has long ceased to be just about outright bans. In other pieces, we’ve looked at how it works through registries of what’s permitted, through fees for the right to make contact, through requirements to identify yourself. AI law gives us another version of censorship (just a less visible one).

A model is a filter

A neural network answers based on what was in its training set. Whatever wasn’t in the data simply doesn’t exist for the model. When someone asks an assistant about events, laws, or organizations, they’re not getting an internet search, they’re getting a retelling of the training set. Whoever decided what went into that set also decided the boundaries of the answer. This, in fact, became one of the basic problems with neural networks built in China, which reluctantly answer or don’t answer at all questions about various unpleasant events in China’s history.

The training set’s contents stay hidden

The registry of blocked websites at least exists as a list. You can download it, check it, and challenge a specific entry. You can even bypass the blocks. Here, there’s no list at all. What ends up in the state system of anonymized data, and what comes out of it, isn’t public knowledge. Checking what’s missing from the training set is impossible in principle, because an absence leaves no trace.

Access gets distributed administratively

The proposal for a state-accredited “trusted intermediary” repeats a construction we’ve seen in the registry of socially significant services and in the registry of trusted virtual PBXs. A list of who’s allowed appears. First, they rationed access to the internet and the right to make phone calls. Now they’re rationing the raw material used to build machines that answer questions.

A loop closed from both sides

We’ve already covered the law that checks models for compliance with “traditional values,” handing the review to state-accredited organizations. That’s output control: what a model is allowed to say. The current request fills in the input side: what Russian AI is even allowed to learn from. Between these two points, there’s no room left for content to be shaped by anything other than an administrative decision.

In this setup, a person’s consent was the last point where they had even a formal say in the decision. And it looks like that point is about to disappear too.

A curious coincidence

While Kommersant was preparing its story, the Association of Big Data announced that Russia’s first series of national standards for data synthesis had been approved. The documents take effect on September 1, 2026. The news reached Kommersant at the same time as big tech’s request to Mintsifry.

Synthetic data gets generated to reproduce the statistical properties of real data without belonging to any actual living person. In the same statement, the association’s executive director, Alexei Neiman, says this approach lets companies develop services “without excessive reliance on real personal data”.

Here’s how it all looks in the end. Publicly, the association celebrates the synthetic data standards. Privately, it wants to train AI on real people.

Estimate from Kirill Menshov, senior vice president of Sber, from the Association of Big Data’s statement dated August 6, 2026

Who’s saying what

The ministry is staying silent for now. Industry experts believe Russian big tech faces a data shortage, which is why the association is writing letters to the government. For instance, Zarya Ventures partner Alexander Ponomarev notes that Russia’s data market remains underdeveloped, with most datasets locked inside major tech ecosystems. Just AI representative Svetlana Zakharova adds that companies have exhausted their own sources and are looking for new ones.

We’ve seen this somewhere before

In our piece about call blocking, we explained how the Association of Big Data proposed creating a registry of trusted virtual PBXs. The owner adds their addresses after authenticating through Gosuslugi (Russia’s state e-services portal), and carriers accept voice traffic only from addresses on that registry. And again, it’s the same pattern: business creates an intermediary under government oversight and keeps outsiders away from the data.

We also previously covered the criteria for “national” model status, which, according to experts, only large ecosystem players can meet. The same companies are now asking for access to citizens’ data.

International practice

The question of whether you can train a neural network on people’s data without asking them has been one of the biggest debates of the last five years. For example, in December 2024, the European Data Protection Board, responding to a request from the Irish regulator, stated: legitimate interest can serve as a legal basis for developing and operating AI models.

But usually the burden sits on the company’s side. It must carry out and document a legitimate-interest assessment, and the person keeps the right to object to the processing. Then a regulator comes along and checks whether the assessment was done honestly. Analysts at the International Association of Privacy Professionals calculated that the phrase “case by case” appears sixteen times in the document, and the words “may” and “might” appear 161 times. The regulator is being extremely cautious, treating plain water like hot soup.

The US has no federal law on personal data. Industry rules, individual state laws, and Federal Trade Commission practice fill the gap instead. No single answer has emerged there either. It’s not that nobody cares (each state just handles it however it can).

Russian big tech wants to avoid scrutiny every time it touches people’s data.

Implications for the public

If this initiative gets adopted, it would have several practical consequences.

  • We’ll lose any control over AI models in Russia. Right now, a person at least formally decides what happens to their data before it gets anonymized. The proposal removes that checkpoint for AI training purposes.
  • The category of “no longer personal data” keeps expanding in Russia. Along with that status go the operator’s obligations: to notify people, to provide access, and to delete data on request.

There’s a separate angle for NGOs and independent projects. Your data lives in these ecosystems too: email, classified ads, banking, search, resumes. If the access rules loosen, that data will end up in training sets without anyone talking to you about it separately. You can’t assess this in advance, because the contents of these datasets never get published.

There’s a second angle we already touched on above: Russian big tech’s AI models will be able to give out only information that’s convenient for the authorities, expanding censorship.

Recommendations

  • If you work with people’s data or train models on it, some of these obligations already landed on you a year ago, regardless of what happens to this request.
  • Check whether you’re ready for a demand from the ministry. Since September 2025, Mintsifry has had the authority to order any personal-data operator to anonymize its information and hand it over to the state system for free. It’s worth figuring out in advance which of your datasets this covers and exactly what would leave your hands.
  • Don’t confuse removing a name with anonymization. Roskomnadzor Order No. 140 specifies concrete methods: identifier substitution, changing the data’s composition, decomposition, and shuffling. What gets checked is the outcome: a person shouldn’t be identifiable even from indirect clues. A birth date plus a city plus a profession can identify someone just as well as a passport.
  • Track the provenance of your datasets. For every dataset, record its source, the legal basis for processing it, and any usage restrictions. As long as regulation keeps changing every few months, this is the only way to assess your risk quickly without digging through the archive by hand.
  • Keep raw data and derived data architecturally separate. Embeddings, aggregates, and model weights should live apart from raw records. If everything sits mixed together, you won’t be able to prove anonymization.
  • Test your model for memorization. Large models can reproduce chunks of their training data verbatim. If someone can extract a fragment of the source data from your model, anonymization collapses along with your legal basis for using the data. Testing for training-data extraction belongs in your regular test suite, not on a wish list.
  • Consider synthetic data as a practical tool. Starting September 1, 2026, national standards PNST (Russia’s preliminary national standard) 1064-2026, 1065-2026, and 1066-2026 will take effect, and you will be able to cite them in your documentation. Market participants estimate the accuracy loss at 2-3%, and synthetic datasets raise noticeably fewer legal questions.
  • Remember the purpose limitation. Anonymized data is allowed for analytics and model training but banned for advertising, product promotion, and political campaigning. If your product is about marketing, that path is closed no matter how good your anonymization is.
  • Collect less. This advice sounds dull, but it’s the only one that holds up no matter how regulation evolves. Nobody can demand data you never collected, so it can’t leak and can’t end up in someone else’s training set. For projects working with vulnerable people, this is the single most important architectural decision, and you make it at the very start, not after the first request comes in.

Don’t miss the next Riposte!

We don’t spam! Read more in our privacy policy