Microsoft Explores Crediting AI Data Contributors

Microsoft is embarking on a new research project aimed at understanding how specific training examples influence the outputs of generative AI models, such as text, images, and other media. This initiative was highlighted in a job listing from December that recently resurfaced on LinkedIn, seeking a research intern to join the effort.
The project's goal is to develop a method to train models so that the impact of particular data, like photos and books, on their outputs can be "efficiently and usefully estimated." The job listing points out that current neural network architectures lack transparency in tracing the origins of their outputs, and there are compelling reasons to address this issue. One reason mentioned is the potential for providing incentives, recognition, and even compensation to individuals who contribute valuable data to future AI models.
The backdrop to this research is the ongoing legal battles involving AI companies, including Microsoft, over intellectual property rights. AI models are often trained on vast datasets scraped from public websites, which can include copyrighted material. While AI companies often claim protection under fair use doctrine, creators across various fields—artists, programmers, authors—dispute this stance.
Microsoft is currently facing legal challenges, including a lawsuit from The New York Times, which alleges that Microsoft and OpenAI infringed on its copyright by using its articles to train their models. Additionally, several software developers have sued Microsoft over its GitHub Copilot AI coding assistant, claiming it was trained on their copyrighted code.
The research project, referred to as "training-time provenance," involves Jaron Lanier, a notable technologist at Microsoft Research. Lanier has previously written about "data dignity," advocating for a system that connects digital content with its creators and potentially compensates them for their contributions to AI outputs.
While Microsoft's project is still in its early stages, other companies like Bria, Adobe, and Shutterstock are already experimenting with compensating data owners based on their contributions to AI models. However, large AI labs have generally not established individual contributor payout programs, opting instead for licensing agreements or opt-out mechanisms for copyright holders, which can be cumbersome and limited in scope.
Microsoft's initiative might remain a proof of concept, similar to OpenAI's yet-to-be-released tool for creators to control how their works are used in training data. There's also speculation that Microsoft might be attempting to "ethics wash" its AI practices or preempt regulatory and legal challenges.
This move by Microsoft is particularly noteworthy given the recent calls from other AI labs, like Google and OpenAI, for the U.S. government to relax copyright protections for AI development. Microsoft has not yet responded to requests for comment on this project.
Related article
Musk Considered Leaving OpenAI to His Kids as Altman Testifies
This morning, OpenAI CEO Sam Altman took the stand to address former co-founder Elon Musk’s lawsuit challenging the company’s corporate structure.When asked about Musk’s claim that other founders “stole a charity” by launching a for-profit subsidiary
Sam Altman Sparks Debate Over AI's Deceleration
Listen onApple PodcastsListen onSpotifyOpenAI CEO Sam Altman recently suggested that it may be time to “pace the rate of AI development” to allow society to “harden around some of these new capability levels.”On the latest episode of TechCrunch’s Equ
Anthropic Opens Doors to EU Cybersecurity Agency as Mythos5 Model Faces Compliance Exam
Artificial intelligence compliance regulations are advancing significantly. Leading AI firm Anthropic has officially granted the European Union's cybersecurity authority access to its Mythos AI model, a pivotal move for this advanced large language m
Related Special Topic Recommendations
Comments (37)
0/500
Interesante proyecto, pero lo que realmente necesito es que alguien en la IA me explique por qué mi asistente virtual aún no puede organizarme el escritorio 😅. ¿Esto de la atribución podría cambiar cómo las empresas comparten datos? Me preocupa un poco la transparencia de todo el proceso. Ojalá no sea solo un gesto de relaciones públicas.
Lo de Microsoft parece interesante pero, ¿y la privacidad de los datos? 🤔 A veces siento que con toda esta exploración de IA, estamos perdiendo el control de lo que se usa para entrenar los modelos. ¿Habrá una compensación justa para quienes contribuyeron? No quiero que esto se convierta en otro caso de 'big tech' aprovechándose de 'datos gratis'...
이런 연구는 AI 데이터 기여자들에게 공정한 보상을 제공하는 데 중요한 단계가 될 수 있겠네요. 😊 근데 MS가 과연 저작권 문제를 해결할 수 있을지 의문이 드네요. 데이터 소싱 방식이 좀 더 투명해져야 할 시점인 것 같아요!
This is super intriguing! Microsoft's diving into how AI training data shapes outputs—mind-blowing stuff. Wonder how they'll credit contributors fairly? 🤔
This Microsoft AI project sounds intriguing! Crediting data contributors could reshape how we value creative input in AI. Curious to see if it'll spark ethical debates or just be a tech flex. 🤔

Musk Considered Leaving OpenAI to His Kids as Altman Testifies
This morning, OpenAI CEO Sam Altman took the stand to address former co-founder Elon Musk’s lawsuit challenging the company’s corporate structure.When asked about Musk’s claim that other founders “stole a charity” by launching a for-profit subsidiary
Sam Altman Sparks Debate Over AI's Deceleration
Listen onApple PodcastsListen onSpotifyOpenAI CEO Sam Altman recently suggested that it may be time to “pace the rate of AI development” to allow society to “harden around some of these new capability levels.”On the latest episode of TechCrunch’s Equ
Anthropic Opens Doors to EU Cybersecurity Agency as Mythos5 Model Faces Compliance Exam
Artificial intelligence compliance regulations are advancing significantly. Leading AI firm Anthropic has officially granted the European Union's cybersecurity authority access to its Mythos AI model, a pivotal move for this advanced large language m
Interesante proyecto, pero lo que realmente necesito es que alguien en la IA me explique por qué mi asistente virtual aún no puede organizarme el escritorio 😅. ¿Esto de la atribución podría cambiar cómo las empresas comparten datos? Me preocupa un poco la transparencia de todo el proceso. Ojalá no sea solo un gesto de relaciones públicas.
Lo de Microsoft parece interesante pero, ¿y la privacidad de los datos? 🤔 A veces siento que con toda esta exploración de IA, estamos perdiendo el control de lo que se usa para entrenar los modelos. ¿Habrá una compensación justa para quienes contribuyeron? No quiero que esto se convierta en otro caso de 'big tech' aprovechándose de 'datos gratis'...
이런 연구는 AI 데이터 기여자들에게 공정한 보상을 제공하는 데 중요한 단계가 될 수 있겠네요. 😊 근데 MS가 과연 저작권 문제를 해결할 수 있을지 의문이 드네요. 데이터 소싱 방식이 좀 더 투명해져야 할 시점인 것 같아요!
This is super intriguing! Microsoft's diving into how AI training data shapes outputs—mind-blowing stuff. Wonder how they'll credit contributors fairly? 🤔
This Microsoft AI project sounds intriguing! Crediting data contributors could reshape how we value creative input in AI. Curious to see if it'll spark ethical debates or just be a tech flex. 🤔





Home






