
theregister.com
August 18, 2026
1 min read
52/100
Summary
Google bought more than 100 million emails, 30 million recorded phone calls, and data from Microsoft Teams, Oracle, and SAP for $10 million at an auction involving crashed airline Spirit. The purchase includes a large collection of airline-related communications and enterprise-system data. The source text does not specify how Google plans to use the data or identify the auction date.
What the discussion said
The thread barely dwells on Spirit as an airline. It treats the bankruptcy auction as a vivid example of how conversational, workplace, customer-support, and operational records can become raw material for AI once a company collapses. The scale of the archive—calls, chats, emails, tickets, files, marketing addresses, purchase histories, and flight records—made commenters uneasy because it turns ordinary interactions into a transferable corporate asset. Most readers reject the reassurance that the material was de-identified. They argue that removing obvious names does not make a sprawling, cross-linked archive anonymous: stable identifiers, unusual purchases, work conversations, and outside data can make people recognizable later. Several also draw a hard consent distinction: agreeing to call recording for human quality assurance is not agreeing to supply training material for a third party’s LLM. The fact that Google selects and pays the intermediary tasked with stripping identifiers deepened suspicion, even though the legal process appears to require an independent de-identification step. A smaller counterpoint says the value is obvious: real data from a large, messy business can teach models about operations in ways polished public datasets cannot. That possibility did little to calm the prevailing fear that these systems will sharpen automated insurance, hiring, and customer-service decisions while making corporate surveillance more normal.
Where opinion split
The sharpest dispute is whether de-identification makes this a legitimate AI-data purchase. Defenders point to a formal third-party stripping identifiers before Google receives the archive and argue that authentic large-company operations data is unusually valuable for training; critics reply that linked datasets can re-identify people, so consent for a recorded support call never meaningfully became consent for model training.
Community Sentiment
Positives
Concerns