Generative AI models require enormous amounts of data for training in order to develop their creative capabilities. The manner in which developers obtain this data is not always legitimate, and we have already seen many face legal challenges because of this, as is the case with Meta or OpenAI.
The latest controversy in this regard centers on Google. It turns out that the technology giant is utilizing its vast YouTube library, which contains 20 billion videos, to train Gemini, VEO3, and other AI models.
This fact was revealed by CNBC after obtaining confirmation from Google, which states that it only uses a subset of these videos and does so legally, in accordance with agreements with creators and media outlets.
“We have always used YouTube content to improve our products, and this has not changed with the advent of AI. We also recognize the need for safety measures, which is why we have invested in robust safeguards to enable creators to protect their image and likeness in the age of AI, something we are committed to continuing,” a YouTube spokesperson explained.
Nevertheless, the news has caused skepticism and concern about whether creators, media outlets, and companies are truly aware of these agreements and how Google uses their content. Likewise, there are no available settings that allow users to prevent the company from utilizing published videos on its platform to train its AI models.
Within the YouTube Terms of Service, the platform makes it clear that “by uploading Content to the Service, you grant YouTube a worldwide, non-exclusive, free of charge and royalty-free, transferable license with the right to sublicense to use such Content (including to reproduce, distribute, modify, transform, display, communicate to the public, and perform it) for the purpose of operating, promoting, and improving the Service.”
The provision regarding “improving the service” could be where the loophole lies if Google considers that training its AI models falls under this category. Furthermore, in Google’s Privacy Policy, it is indicated that the legal bases for processing your information may include using user information “publicly available online or from other public sources to train new machine learning models and to develop foundational technologies underlying a range of Google products, such as Google Translate, Gemini applications, and Cloud’s AI capabilities.”
It should be noted that Google’s Privacy Policy “applies to all services offered by Google LLC and its affiliates, including YouTube, Android, and services provided on third-party websites, such as advertising services.”
This reality reopens the debate about the use of content generated by individuals to refine models that could become direct competitors to these very creators. Moreover, the lack of control over how companies handle intellectual property emphasizes the unfairness inherent in this system.
Luke Arrigoni, CEO of Loti, a company focused on protecting the digital identity of creators, explained to CNBC: “It is plausible that they are taking data from many creators who have invested significant time, energy, and thought into these videos. It is helping the Veo 3 model to create a synthetic version—a poor imitation—of these creators. That is not necessarily fair to them.”
Photo: Canva
Your email address will not be published. Required fields are marked *
Δ