AI Briefing
KO

Training Data for Gemini 3 and Gemma 4 Disclosed

·2026.08.28 18:30

Key point

In compliance with EU AI Act obligations, the scale of training data for Gemini 3 and Gemma 4 has been disclosed, but specific dataset names and sources are not included.

1 / 4

Details

Google has published the Public Summary of Training Content documents for the Gemini 3 Pro and Gemma 4 models in accordance with Article 53 of the EU AI Act. These are regulatory submission documents prepared according to the template distributed by the EU AI Office, aiming to ensure copyright holders' rights and transparency.

Disclosed Information and Limitations

The two documents clearly reveal the volume of training data. Text is listed as over 10 trillion tokens, images as over 1 billion, and audio and video as over 1 million hours each. However, not a single specific dataset name or unique identifier is included, such as the list of large-scale public datasets or the top 10% of crawled domains required by the EU template. This is interpreted as a result of Google reducing specificity to protect trade secrets, alongside the structural limitation that the template does not require data mixture ratios.

Regulatory Scope and Timeline

Although Gemma 4 is an open-source model under the Apache License 2.0, the obligation to disclose training content under the AI Act does not apply the open-source exception, requiring compliance identical to closed-source models. The regulation applies from August 2, 2025, and the AI Office's enforcement powers commence on August 2, 2026. Violations may result in fines of the higher of 3% of global turnover or €15 million.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.