Emails, chat records, project documents, work orders—these used to be just digital remnants awaiting clean-up after a company shut down. Now, they are being reassessed, packaged, sold, and fed into an AI company's training pipeline.
An undertaker preserves the deceased's final dignity. The postmortem dignity of an enterprise is to prove that what they left behind still holds value.
On August 17, at a bankruptcy asset auction, Google bid $10 million to acquire all of Spirit Airlines' corporate data. Another bidder, Mercor, bid $7.5 million, falling short by $2.5 million.
The auctioned items were divided into three parts. The first part consisted of approximately 100 million employee emails. The second part included 500 million Microsoft Teams messages, totaling 600 million. The third part comprised calendars, spreadsheets, financial databases, project files, operational records, and a batch of internal software. Passenger profiles, frequent flyer information, and the like were not included.
600 million messages—if one person speaks 100 sentences a day, they would have to speak non-stop for over 16,000 years.
Calculating, each message was sold for 1.67 cents. Americans refer to a 1-cent coin as a penny, often not bothering to pick it up if it falls on the ground. In other words, a Spirit employee's sentence in Teams was worth one and a half pennies.
Going back to 1980, Spirit Airlines was born in Detroit, born from a trucking company called Charter One, which later transitioned to aviation. In 1992, it was renamed Spirit, meaning soul. Over the following 34 years, it set the benchmark for the entire industry to replicate the model of ultra-low-cost carriers in the U.S.
It once had a solid foundation—205 all-Airbus fleet, approximately 300 flights a day, and a projected 2024 revenue of around $5 billion. However, it faced a net loss of about $1.2 billion and carried nearly $9 billion in debt when it filed for bankruptcy.
On a late night in May 2026, the company announced it was ceasing operations. The next day, around 17,000 employees found out they were unemployed through the news.
Proceedings are still ongoing, and the transactions are awaiting approval from the bankruptcy court. Spirit is undergoing a liquidation-type closure under Chapter 11 bankruptcy, with no bankruptcy trustee taking over. The company is still supervised by the court and is selling off its remaining assets piece by piece. The sensitive data involving employee emails, Teams messages, etc., must first be handed over to an independent entity for processing. This entity was selected by Google, with Google also bearing the costs.
Furthermore, scouring through public reports, the bankrupt company selling internal data to an AI company is unprecedented. This deal is most likely the first of its kind in history.
The long-standing U.S. tech publication Gizmodo titled this transaction as follows: "Spirit is Dead, But Its Data Will Haunt Google's Servers for Generations."
A New Business Venture
There have always been people who handle the assets of defunct companies. Lawyers, liquidators, auction houses—doing this for decades. Aircraft, furniture, trademarks, patents—all sellable assets have long been sold off.
What's truly new is that starting this year, even a company's internal employee data has been put on the table.
This emerging business has two underlying reasons.
First, the number of defunct companies has increased.
In the first quarter of 2024, the failure rate of U.S. startups surged by 58% year-on-year, and the number of active VC firms decreased by 62% from its peak. The money hasn't decreased; it's just flowing more concentratedly towards AI. In 2024, U.S. AI startups raised a record $97 billion in funding. The capital market still has money, but it's becoming more reluctant to spend outside of AI.
As a result, a group of companies that could have continued to survive on funding are now hitting a wall earlier. In August 2024, the fintech company Tally, backed by a16z, announced its closure. Having raised a total of $172 million, with a peak valuation of $855 million, reaching Series D, it still couldn't secure the next round of funding.
Second, data has become more expensive.
The consumption of data for large-scale model training has reached unprecedented levels. The training data for GPT-4 consists of around 130 trillion tokens. For comparison, Google Books scanned over 40 million books in over forty years, which amounts to about 40 trillion tokens. In other words, the amount of data used for one GPT-4 training is equivalent to over three times the content of Google Books.
Epoch AI ran some numbers and concluded that high-quality language data from books, news, and Wikis will be depleted around 2026. The high-quality text available on the public internet is quickly running out, with synthetic data accounting for an increasing share. Next, AI companies will have to look beyond the public internet for new data. Years-worth of internal company emails, chat records, and work documents have now come into view.
As for the AI training dataset market, research institutions predict it could reach $9.7 billion by 2030. Taking into account various licenses, the total size of the pie is estimated to reach $67.5 billion.
On one side, more and more companies are closing down, leaving behind a large amount of internally generated data that was previously not priced; on the other side, AI companies have an increasing demand for data beyond the public internet. With both sides happening simultaneously, it is the first time that internal corporate data has the conditions for scalable transactions.
What truly adds value to this type of data is the development of Enterprise Agents. Gartner predicts that by 2026, 40% of enterprise applications will embed task-specific AI Agents, a figure that was less than 5% the previous year.
The training material required for Enterprise Agents is not exactly the same as that for large-scale models. While public web pages can provide knowledge, language, and final content, it is challenging to replicate a company's actual work processes. Communication and collaboration involve many real details, such as how a requirement is proposed, how several people discuss it, how tasks are assigned, how issues are addressed, and ultimately how delivery is made.
These processes are extensively recorded in a company's internal emails, chat logs, tickets, and project documents.
When such data begins to have clear buyers and use cases, things that were previously directly deleted when a company shut down now have standalone transaction value.
Body Snatcher
In the past, this job was done by liquidation lawyers, but now three new types of people have emerged.
The first is called a dissolution service provider, represented by SimpleClosure.
This company does only one thing: helping startups die gracefully. In 2023, the company just started, raising $1.5 million in pre-seed funding, and in May 2025, it secured $15 million in Series A funding, led by TTV Capital. Even Carta, which handles equity management and corporate affairs for many US startups, shut down its own closure service, invested in SimpleClosure, and handed over this part of customer demand to it.
By October 2025, SimpleClosure had already conducted funerals for over a thousand companies. Crunchbase gave it a nickname, "A Better Way To Fail," a kind of better way of failing. While American entrepreneurs love to say "fail fast," SimpleClosure argues that fast failure is not enough; it must also be dignified. The company even has a pricing calculator on its website – enter your company's situation, and it will calculate how much it costs to die once.
A funeral is never a waste. In April 2026, SimpleClosure launched the Asset Hub, specifically designed to handle intangible assets left behind after a company shuts down. In addition to brands, software, and customer lists, for the first time, items such as Slack messages, emails, and Jira tickets – internal work data – were explicitly put on the shelf.
This indicates that before Spirit, the market had already begun to attempt to price the internal data of dead companies, although at that time it was still startups, small transactions, and private dealings.
There is a ready-made case. When transcription and captioning company cielo24 shut down, they sold off Slack messages, internal emails, and Jira tickets accumulated over the past 13 years through SimpleClosure. CEO Shanna Johnson later told Forbes that this batch of data eventually sold for hundreds of thousands of dollars.
For a company that has already decided to close its doors, this was originally a batch of data that needed to be cleaned up, but in the end, it became an asset that could be recovered during liquidation.
For the data sales stage, SimpleClosure handed it over to Protege. This is a data exchange market specializing in AI training data licensing. In January 2026, they just received a $30 million investment led by a16z, and the founder is Bobby Samuels. Protege's initial entry point was medical imaging, where within 30 days, they gathered millions of images for pre-training for buyers.
Now, Protege is starting to apply this data licensing and trading ability to internal communication data of closing companies. SimpleClosure is responsible for company closure and asset organization, while Protege is responsible for finding buyers, completing data licensing, and transactions.
The second type of participant is bankruptcy courts and liquidation lawyers. For decades, they have been counting planes, tables, chairs, trademarks, and patents. Now, the list is beginning to include emails, Slack messages, and other internal data.
According to U.S. bankruptcy law, this data can be included as intangible assets in bankruptcy estates and sold under court supervision. Relevant law firms have also begun to establish specialized teams to handle data preservation, discovery, and organization in bankruptcy cases. Redgrave LLP has such a restructuring and discovery business.
The third type is e-discovery service providers responsible for technical execution. Companies like KLDiscovery, Epiq, and Consilio are usually involved in collecting, organizing, hosting, and reviewing enterprise data. The content in emails, Teams, and SharePoint needs to be exported, archived, and organized by them before being packaged into data that can enter the transaction process.
This industry itself already has a mature set of billing methods. EDRM regularly publishes pricing surveys, with common billing units including data collection per GB, hosting per GB per month, and document review fees.
Spirit’s 600 million messages eventually turned into an auction item, relying on this type of foundational work. However, compared to court documents and auction bids, this part rarely appears in public reports. The specifics of who handles it and how it is handled are usually not visible to the outside world.
Autopsy Checklist
For a corporate Agent, what it needs to learn is not just knowledge and standard answers, but also judgment, collaboration, error correction, and execution processes in the real work environment.
The final product tells the model what it has become, while internal records tell it how it was made.
This change has already been reflected in Agent training data. In the past, a single-line of code training sample might only consist of a few hundred tokens, involving modifying a few lines of code; now, an Agent training sample often needs to include the entire process of requirement understanding, file locating, code modification, and test verification.
Training data is transitioning from a single answer to full task execution records. For the Agent, the end result is of course important, but the judgments, operations, and feedback left during task completion are more valuable for training.
Compared to companies operating normally, the data of a shutting-down business is more likely to enter the transaction process. While the company is still operating, selling internal communications would involve trade secrets, employee privacy, non-compete risks, and customer relationships, and legal and management teams are usually very cautious. After entering liquidation, the company's main goal changes to recovering as many remaining assets as possible to secure more value for creditors.
That's why the same batch of internal data, which was difficult to sell while the company was alive, may be revalued during the shutdown phase.
The value of internal corporate data has long been recognized by the industry. Salesforce has always regarded the enterprise communications accumulated in Slack as important data assets, and Microsoft CEO Satya Nadella has emphasized multiple times that when enterprises use AI, a truly valuable part is their proprietary data and work context.
In the past, this data primarily served the company itself. Now, as AI companies begin to actively seek enterprise internal data, they have, for the first time, more clearly defined external buyers.
The Art of Pricing
While this business is still in its early stages, some reference prices have emerged in the market.
SimpleClosure and Protege handle data from closing startups, with individual transactions typically ranging from $10,000 to $100,000. Mercor offers quotes based on the chat logs and emails of acquired startup employees, with the highest reaching $300,000.
Spirit has taken the pricing to the tens of millions of dollars. Mercor bid $7.5 million, and Google ultimately offered $10 million, competing for around 600 million internal communications and other corporate data. Based solely on these 600 million messages, the average price per message is approximately 1.67 cents.
This unit price is not high. In 2024, Reuters reported that Photobucket negotiated licensing deals for about 13 billion photos and videos with an AI company, with prices around 5 cents to $1 per photo and over $1 per video. In the B2B data market, a single contact's information can be sold for a few cents to a few dollars depending on completeness and accuracy.
However, these prices are not yet enough to form a unified standard. The value of internal data from closing companies is currently mainly a matter of individual negotiation. Factors such as data volume, industry, time span, completeness, uniqueness, and what the buyer intends to do with it will all affect the final price. Spirit's $10 million bid appears more like one of the few publicly visible large-scale examples at present.
When compared to established data licensing markets, the gap becomes even more apparent. Reddit licensed user posts and comments to Google for approximately $60 million annually; News Corp's content licensing agreement with OpenAI is around $250 million for five years; xAI's partnership with Telegram amounts to $300 million; and Apple's purchase of Shutterstock image licenses falls between $25 million and $50 million.
These markets already have established buyers, licensing mechanisms, and pricing experiences. Transactions of internal data from closing companies are just getting started, with no clear rules yet on which data is most valuable or whether valuation should be done per item, by capacity, or as a whole.
Looking at Spirit's $10 million within the company's own scale is a different story. In 2024, Spirit's annual revenue was close to $5 billion, averaging around $13.7 million per day. The amount Google spent to acquire this data is less than what it makes in a day during normal operations.
For a bankruptcy liquidation, this is just a small recovery from the remaining assets; for an AI company, it is acquiring a set of long-unseen data containing the real operational processes of the business.
Cleansing and Handover
After the data transaction, it cannot be handed over directly to the buyer. Before the formal handover, it usually needs to go through several steps such as export, de-identification, organizing and packaging, court approval, and final delivery.
The first step is export. Slack's corporate data is usually exported as a ZIP file, including JSON files organized by channel, member information, and attachments.
Microsoft 365, on the other hand, can export emails and Teams messages through an eDiscovery tool. The Spirit transaction involves approximately 600 million internal communications, a large amount of data that requires processing in batches in practice. The specifics of whether this is carried out by an eDiscovery service provider or the buyer's engineering team have not been disclosed in publicly available documents.
Next is de-identification, which involves removing as much information as possible that can be traced back to specific employees. Common methods include identifying and masking personal information such as names, phone numbers, and addresses, or replacing sensitive fields with identifiers that do not directly correspond to individuals. In some statistical and training scenarios, additional random perturbation is applied to further reduce the possibility of re-identifying individuals.
However, de-identification does not equate to absolute anonymity. Even if names and emails have been removed, as long as there are enough professional, temporal, locational, or behavioral features retained in the text, there is still a possibility of re-identifying individuals through other information.
Therefore, the party responsible for this step in the Spirit transaction is crucial. According to the current transaction arrangement, a third-party independent entity will handle the de-identification process, chosen by Google and at Google's expense. Google has also committed not to exploit this dataset to re-identify individuals.
Data sales by bankrupt companies are not unprecedented. Over the past two decades, from Toysmart and Borders to RadioShack and 23andMe, cases have emerged involving the handling or sale of customer data during bankruptcy. Transactions of this nature involving consumer privacy usually face stricter scrutiny from courts, regulatory bodies, and state governments.
What sets Spirit apart is that this sale does not revolve around passenger lists but rather around employee emails, Teams messages, and other internal work data. Existing bankruptcy procedures have not established mature rules for dealing with this type of data as they have for consumer information.
Controversy has already arisen. Spirit's flight attendants union has raised objections to the transaction, leading the court to postpone approval. What was initially a batch of corporate data intended to be sold as bankruptcy assets has now become entangled in employee rights and the boundaries of AI usage.
If the transaction is ultimately approved, the data will enter the final settlement stage. However, there is little public disclosure about the data's journey from the finalized dataset to Google's systems. It is currently unknown how the data is transmitted, whether further cleansing will occur, and in what form it will ultimately enter product development or model training.
At this point, the program truly completes the transformation of data from bankruptcy assets to AI assets.
Buyer
The most clear-cut buyers of this type of data at present are model companies like Google and training data service providers like Mercor.
Google's official statement regarding this Spirit transaction is that the data will be used for product improvement and AI development. For Google, the value of this dataset lies in its documentation of a large enterprise's real operational processes, such as scheduling, collaboration, internal communication, project advancement, issue resolution, and management decisions.
These details are hard to obtain from public web pages. Particularly for corporate agents, beyond just the final outcome, what is more important is how tasks are progressed and completed within a real organization. The internal records left by Spirit happen to contain a significant amount of such process data.
Another bidder, Mercor, provides further insight into where this business is heading.
Founded in 2023, Mercor's three founders—Brendan Foody, Adarsh Hiremath, and Surya Midha—have been debate team partners since high school. The company initially focused on AI recruitment, using models to help companies screen and interview candidates. By September 2024, Mercor had evaluated around 300,000 job seekers, reaching a valuation of $250 million.
Subsequently, the company's focus shifted towards AI training data. In November 2025, when the three founders were only 22 years old, as the company's valuation rose, they became some of the youngest self-made billionaires globally. In the first half of 2026, Mercor's revenue exceeded $614 million, with about 90% coming from top AI labs like OpenAI. By July, the company was seeking a valuation of around $20 billion and had acquired Deeptune, a company specialized in building training environments for AI agents.
Mercor primarily acquires training data through two main paths.
One is by directly purchasing internal company records. They have made offers to acquired or closed-down startups to buy employee chat logs and emails, with individual companies receiving offers of up to around $300,000.
Another one is hiring people with real work experience. In October 2025, TechCrunch reported that an AI lab recruited former employees through Mercor, allowing them to convert their work experience into training tasks and feedback at an hourly rate of about $200. The CEO of Mercor stated that the company pays out over $1.5 million daily to individuals participating in AI training.
On one side, acquiring the company's retained work records, and on the other, acquiring the work experience held by employees. Mercor's core assets are essentially work processes from the real world.
This also explains why it appeared at Spirit's auction. For Mercor, 6 billion pieces of internal communication were another, larger-scale source of training data.
This type of business also comes with risks. In 2025, Scale AI sued Mercor, alleging that former employees took trade secrets; subsequently, the company experienced training-related information leaks and partnership suspensions. Data is both Mercor's business and its most critical asset to defend.
Lastly, there is a numerical contrast. Spirit operated under this name for 34 years, while Mercor, bidding on its internal data, had three founders who were only 22 years old at the time.
Broken Bench
The transaction for Spirit is not yet finalized, the court has not signed off, and Google has not acquired that set of data. There are still union objections, hearings, and many procedures to go through.
But this auction has already made something that was seldom discussed before more concrete.
In the past, when a company closed, many things did indeed disappear. Financial statements remained, trademarks remained, patents remained. However, a company's day-to-day experience usually did not. How a department conducts meetings, how a manager makes decisions, how dozens of people coordinate after a delay, why a process eventually changed to its current state – these things are rarely formally documented.
As a company dissolves, people leave, email addresses deactivate, chat groups close – they also disappear.
So, a company has always been a peculiar organization.
It can exist for decades, accumulating the experiences of tens of thousands of people, but what can truly be inherited is often only a very thin slice of it. The next company will have to hire again, make mistakes again, and learn many things all over again.
The potential AI is changing is exactly this. If emails, meetings, tickets, code changes, and internal discussions can truly be organized into training data, then a company's past experiences that could only be passed down by talent have, for the first time, found another way to be preserved.
Where this will ultimately lead is still difficult to determine. Perhaps in the future, companies will proactively save this data, perhaps employee contracts will be rewritten, perhaps bankruptcy laws will add new restrictions, perhaps companies will separately value even internal work records during financing and M&A.
Spirit has raised this issue early.
The word bankruptcy has a widely circulated etymology. Medieval Italian merchants conducted business sitting behind a bench. When their debts exceeded their assets, the bench was publicly destroyed. banca rotta, broken bench.
For centuries, when the bench broke, it was all over.
Spirit's bench has already broken. The 600 million messages it left behind are being repriced.
One message at one and a half pennies.
Original Article Link