Amazon has reportedly bought large quantities of printed books, including rare and foreign-language titles, for scanning into AI training data. A 404 Media investigation tracked a rare book with a tracking device to Amazon’s VGT3 warehouse in Las Vegas, Nevada. Employees at the site reportedly said shipments of books are received there, their bindings are cut off to speed scanning, and the physical copies are destroyed. A bookseller involved in the tracked shipment said rare books can carry historical, intellectual, sentimental, and monetary value beyond their text. The Internet Archive, founded by Brewster Kahle in 1996, digitizes books while retaining their physical bindings. Its scanning centers use operators known as scribe operators to capture books one page at a time. The organization says it does not destroy books by cutting off their bindings and instead adapts machines and software for nondestructive digitization. The Internet Archive contrasts its process with Google Books, launched in 2004, which it says used industrial-scale destructive scanning to increase processing speed.
observationalepidemiology.blogspot.com
2 min
8/21/2026
The GitHub repository for Alignment Whack-a-Mole provides the data preprocessing pipeline, finetuning scripts, memorization evaluation code, and analysis scripts related to the study of large language models' verbatim recall of copyrighted books. Example files include excerpts and generations from "The Road" by Cormac McCarthy, but full book content is not included.
github.com
4 min
4/30/2026
Universal Music Group is engaged in a legal dispute with AI music platforms Suno and Udio, claiming they used copyrighted music for training without authorization. Major record labels argue that these AI systems were developed using significant amounts of copyrighted material.
pocketables.com
8 min
4/10/2026
NEW: Musician Murphy Campbell says she isn’t making money on YouTube because an AI company is cloning her music and filing copyright claims against her own videos “An entity called Timeless Sounds IR uploaded AI-generated versions of my songs to all major streaming platforms... They used a distributor, which I just discovered, and that distributor’s name is Vydia. They used Vydia to upload all the...
twitter.com
1 min
4/5/2026
An AI company is generating songs that replicate the voice and style of a musical artist, leading to the artist filing a copyright claim against the AI for using her voice on her original songs. As a result, the original artist is no longer earning revenue from her music.
twitter.com
1 min
4/5/2026
Top AI models can generate near-verbatim copies of bestselling novels, indicating that they memorize more training data than previously understood. This memorization capability raises legal concerns regarding copyright and the implications for AI developers.
arstechnica.com
1 min
2/23/2026
Amazon has reportedly bought large quantities of printed books, including rare and foreign-language titles, for scanning into AI training data. A 404 Media investigation tracked a rare book with a tracking device to Amazon’s VGT3 warehouse in Las Vegas, Nevada. Employees at the site reportedly said shipments of books are received there, their bindings are cut off to speed scanning, and the physical copies are destroyed. A bookseller involved in the tracked shipment said rare books can carry historical, intellectual, sentimental, and monetary value beyond their text. The Internet Archive, founded by Brewster Kahle in 1996, digitizes books while retaining their physical bindings. Its scanning centers use operators known as scribe operators to capture books one page at a time. The organization says it does not destroy books by cutting off their bindings and instead adapts machines and software for nondestructive digitization. The Internet Archive contrasts its process with Google Books, launched in 2004, which it says used industrial-scale destructive scanning to increase processing speed.
observationalepidemiology.blogspot.com
2 min
8/21/2026
The GitHub repository for Alignment Whack-a-Mole provides the data preprocessing pipeline, finetuning scripts, memorization evaluation code, and analysis scripts related to the study of large language models' verbatim recall of copyrighted books. Example files include excerpts and generations from "The Road" by Cormac McCarthy, but full book content is not included.
github.com
4 min
4/30/2026
NEW: Musician Murphy Campbell says she isn’t making money on YouTube because an AI company is cloning her music and filing copyright claims against her own videos “An entity called Timeless Sounds IR uploaded AI-generated versions of my songs to all major streaming platforms... They used a distributor, which I just discovered, and that distributor’s name is Vydia. They used Vydia to upload all the...
twitter.com
1 min
4/5/2026
Top AI models can generate near-verbatim copies of bestselling novels, indicating that they memorize more training data than previously understood. This memorization capability raises legal concerns regarding copyright and the implications for AI developers.
arstechnica.com
1 min
2/23/2026
"Better to beg forgiveness" emphasizes caution in interactions with unapproachable entities. The content includes various topics such as digital economy legislation, copyright issues, and profiles of notable figures.
pluralistic.net
18 min
7/31/2026
Universal Music Group is engaged in a legal dispute with AI music platforms Suno and Udio, claiming they used copyrighted music for training without authorization. Major record labels argue that these AI systems were developed using significant amounts of copyrighted material.
pocketables.com
8 min
4/10/2026
An AI company is generating songs that replicate the voice and style of a musical artist, leading to the artist filing a copyright claim against the AI for using her voice on her original songs. As a result, the original artist is no longer earning revenue from her music.
twitter.com
1 min
4/5/2026
Amazon has reportedly bought large quantities of printed books, including rare and foreign-language titles, for scanning into AI training data. A 404 Media investigation tracked a rare book with a tracking device to Amazon’s VGT3 warehouse in Las Vegas, Nevada. Employees at the site reportedly said shipments of books are received there, their bindings are cut off to speed scanning, and the physical copies are destroyed. A bookseller involved in the tracked shipment said rare books can carry historical, intellectual, sentimental, and monetary value beyond their text. The Internet Archive, founded by Brewster Kahle in 1996, digitizes books while retaining their physical bindings. Its scanning centers use operators known as scribe operators to capture books one page at a time. The organization says it does not destroy books by cutting off their bindings and instead adapts machines and software for nondestructive digitization. The Internet Archive contrasts its process with Google Books, launched in 2004, which it says used industrial-scale destructive scanning to increase processing speed.
observationalepidemiology.blogspot.com
2 min
8/21/2026
Universal Music Group is engaged in a legal dispute with AI music platforms Suno and Udio, claiming they used copyrighted music for training without authorization. Major record labels argue that these AI systems were developed using significant amounts of copyrighted material.
pocketables.com
8 min
4/10/2026
Top AI models can generate near-verbatim copies of bestselling novels, indicating that they memorize more training data than previously understood. This memorization capability raises legal concerns regarding copyright and the implications for AI developers.
arstechnica.com
1 min
2/23/2026
"Better to beg forgiveness" emphasizes caution in interactions with unapproachable entities. The content includes various topics such as digital economy legislation, copyright issues, and profiles of notable figures.
pluralistic.net
18 min
7/31/2026
NEW: Musician Murphy Campbell says she isn’t making money on YouTube because an AI company is cloning her music and filing copyright claims against her own videos “An entity called Timeless Sounds IR uploaded AI-generated versions of my songs to all major streaming platforms... They used a distributor, which I just discovered, and that distributor’s name is Vydia. They used Vydia to upload all the...
twitter.com
1 min
4/5/2026
The GitHub repository for Alignment Whack-a-Mole provides the data preprocessing pipeline, finetuning scripts, memorization evaluation code, and analysis scripts related to the study of large language models' verbatim recall of copyrighted books. Example files include excerpts and generations from "The Road" by Cormac McCarthy, but full book content is not included.
github.com
4 min
4/30/2026
An AI company is generating songs that replicate the voice and style of a musical artist, leading to the artist filing a copyright claim against the AI for using her voice on her original songs. As a result, the original artist is no longer earning revenue from her music.
twitter.com
1 min
4/5/2026
No more articles to load