Original title: (After a company goes bankrupt, how to make money by selling employee data)
Original author: 动察Beating
Emails, chat logs, project documents, tickets—until now, they were just digital heirlooms left behind after the company shut down, waiting to be cleared away. Now, they’re being revalued, packaged, sold, and fed into AI companies’ training pipelines.
A mortician preserves the deceased’s last dignity. For a company, that dignity after death is proving that what it left behind still has value.
On August 17, at a bankruptcy-asset auction, Google bid $10 million and bought all of Spirit Airlines’ corporate data. Another bidder—called Mercor—offered $7.5 million, but lost by $2.5 million.
The lot is split into three parts. First, about 100 million employee emails. Second, 500 million Microsoft Teams messages. Together, those two parts make 600 million messages. Third, calendars, spreadsheets, financial databases, project files, operational records, plus a batch of internal software. Passenger profiles and frequent-flyer information are not included.
600 million messages—if one person says 100 sentences a day, they would have to keep talking for more than 16,000 years.
That works out to 1.67 cents per message. Americans call a 1-cent coin a penny, something so small you would not even bother picking it up off the ground. In other words, what a Spirit employee said in Teams was worth one and a half pennies.
Going back to 1980, Spirit Airlines was born in Detroit, originally a trucking company called Charter One before switching to airplanes. In 1992 it changed its name to Spirit, meaning soul. Over the next 34 years, it became the benchmark for ultra-low-cost airlines that the entire industry copied.
It once had a strong foundation: 205 Airbus aircraft, about 300 flights a day, and revenue of about $5 billion in 2024. But that same year it posted a net loss of about $1.2 billion, and when it filed for bankruptcy it was carrying about $9 billion in debt.
One night in May 2026, the company announced it was grounding its fleet. The next day, about 17,000 employees learned from the news that they had lost their jobs.
The matter is still going through procedure, and the deal must wait for bankruptcy judge approval. Spirit is following a Chapter 11 liquidation-style shutdown, with no bankruptcy trustee taking over; the company is still under court supervision and is selling off its remaining assets one by one. Sensitive data such as employee emails and Teams messages must first be handled by an independent entity, which removes identifying information such as names and email addresses; the entity is selected by Google, and Google also bears the cost.
In addition, after combing through public reporting, no precedent could be found of a bankrupt company selling internal data to an AI company. This deal is very likely the first in history.
The title used by the veteran U.S. tech outlet Gizmodo for this deal was that Spirit is dead, but its data will haunt Google’s servers for generations.
A new business
There have always been people who clean up after dead companies. Lawyers, liquidators, auction houses have been doing it for decades. Airplanes, desks, chairs, trademarks, patents—anything worth selling has long since been sold.
What is truly new is that, starting this year, even companies’ internal employee data has been put on the market.
There are two reasons this business has emerged.
The first thing is that more dead companies are appearing.
In the first quarter of 2024, the failure rate of U.S. startups rose 58% year over year, while the number of active venture capital firms was down 62% from its peak.
The money has not actually become less; it is just flowing increasingly into AI. In 2024, U.S. AI startups took in $97 billion in funding, setting a historic record. Capital markets still have money, but they are increasingly unwilling to spend it on anything outside AI.
The result is that some companies that could otherwise have continued living on financing started hitting the wall earlier. In August 2024, fintech company Tally, backed by a16z, announced it was shutting down. It had raised a total of $172 million, reached a peak valuation of $855 million, and made it to Series D, but still could not secure the next round.
The second thing is that data has become more expensive.
Large-model training has already consumed data at an extremely staggering scale. GPT-4’s training data was about 13 trillion tokens. For comparison, Google Books scanned about 40 million books over more than 40 years, which works out to about 4 trillion tokens. In other words, one GPT-4 training run used an amount of data equivalent to more than three Google Books.

Epoch AI did the math: high-quality language data—books, news, Wikipedia—will be exhausted around 2026. The supply of high-quality text from the public internet that can be used for training is rapidly running dry, and the share of synthetic data is rising. Next, if AI companies want new data, they will have to look more and more beyond the public internet. Emails, chat logs, and work documents accumulated over years inside companies have thus entered the field of view.
And the market for AI training datasets is projected to reach $9.7 billion by 2030. Including various forms of licensing, the total market size is estimated at $67.5 billion.
On one side, more and more companies are entering shutdown and liquidation, leaving behind large amounts of internal data that had never previously been priced; on the other side, AI companies’ demand for data beyond the public internet is growing larger. With both sides appearing at the same time, internal corporate data has for the first time become eligible for large-scale trading.
What really makes this kind of data more valuable is the development of enterprise agents. Gartner predicts that by 2026, 40% of enterprise applications will have task-oriented AI agents built in, whereas in the previous year that figure was still below 5%.
The training materials enterprise agents need are not exactly the same as those for ordinary large models. Public webpages can provide knowledge, language, and finished content, but they are hard to use to reconstruct a company’s real work process. Communication and collaboration, when broken down into concrete tasks, involve many real details: how a request is raised, how several people discuss it, how tasks are assigned, how changes are made when problems arise, and how it is finally delivered.
These processes are largely stored in companies’ internal emails, chat logs, tickets, and project documents.
When this kind of data starts to have clear buyers and clear uses, things that would originally be deleted outright when a company shuts down gain separate value as tradable assets.
undertaker
In the past, liquidation lawyers did this work; now there are three more kinds of people involved.
The first is called a dissolution service provider, represented by SimpleClosure.

This company does one thing: help startups die with dignity. When it started in 2023, it raised $1.5 million in a pre-seed round, then raised another $15 million Series A in May 2025, led by TTV Capital. Even Carta, which manages equity and corporate affairs for many U.S. startups, shut down its own shutdown service, invested instead in SimpleClosure, and handed that customer demand over to it.
By October 2025, SimpleClosure had already handled the funerals of more than a thousand companies. Crunchbase gave it a name, “A Better Way To Fail,” a better way to fail. American founders like to say fail fast. SimpleClosure says fast is not enough; it also has to be dignified. Its website even has a pricing calculator: enter company details and it will calculate how much it costs to die once.
The funerals were not done for nothing. In April 2026, SimpleClosure launched Asset Hub, dedicated to handling intangible assets left behind after a company shuts down. Beyond brands, software, and customer lists, Slack logs, emails, and Jira tickets—internal work data—were being clearly put on the market for the first time.
This shows that before Spirit, the market had already begun experimenting with pricing the internal data of dead companies, except that at the time it was still startups, small transactions, and private matchmaking.
There is an existing case. When transcription and captioning company cielo24 shut down, it sold 13 years’ worth of accumulated Slack messages, internal emails, and Jira tickets through SimpleClosure. CEO Shanna Johnson later told Forbes that the data ultimately sold for hundreds of thousands of dollars.
For a company that has already decided to shut down, this was originally a batch of data that needed to be cleaned up, but in the end it became an asset that could be recovered during liquidation.
SimpleClosure handed the data-selling part to Protege. Protege is a data exchange marketplace focused on AI training data licensing. In January 2026 it raised $30 million led by a16z, and its founder is Bobby Samuels. Protege first focused on medical imaging, and within 30 days it assembled millions of images for the buyer to use for pretraining.
Now, Protege is beginning to apply this data-licensing and trading capability to the internal communications data of shut down companies. SimpleClosure handles company shutdowns and asset sorting, while Protege is responsible for finding buyers and completing data licensing and transactions.
The second type of participant is the bankruptcy court and liquidation lawyers. For decades, they inventoried airplanes, desks and chairs, trademarks, and patents. Now, email addresses, Slack logs, and other internal data are starting to appear on the list.
Under U.S. bankruptcy law, this data can be included in the bankruptcy estate as intangible assets and sold under court supervision. Law firms are also beginning to set up specialized teams to handle data preservation, forensics, and organization in bankruptcy cases. Redgrave LLP has such restructuring and forensic work.
The third type is the e-discovery service providers responsible for technical execution. Companies like KLDiscovery, Epiq, and Consilio usually handle the collection, organization, hosting, and review of corporate data. Content in email, Teams, and SharePoint must first be exported and archived by them, and then organized into a data package that can enter the transaction process.
This industry already has a mature pricing system. EDRM regularly publishes pricing surveys, and common billing units include data collection per GB, hosting per GB per month, and document review fees.
Spirit’s 600 million messages were ultimately turned into a lot through this kind of basic work. But compared with court filings and auction bids, this part rarely appears in public reporting, and who handled it and how they handled it is usually invisible from the outside.
autopsy list
For enterprise agents, they need to learn not only knowledge and standard answers, but also judgment, collaboration, error correction, and execution in a real working environment.
Finished content tells the model what was ultimately made; internal records tell it how it was actually made.
This shift is already reflected in agent training data. In the past, one code training example might have contained only a few hundred tokens, with the content being a few lines of code changes; now, one agent training example often needs to include the full process of understanding the request, locating files, modifying code, and verifying tests.
Training data is starting to shift from single answers to complete task execution records. For agents, the final outcome is of course important, but what is more valuable for training is the judgments, actions, and feedback left during the process of completing the task.
Compared with companies still operating normally, data from shut down companies is easier to move into a transaction process. While a company is still operating, selling internal communications involves trade secrets, employee privacy, non-compete risks, and customer relationships, so legal and management teams are usually very cautious. Once liquidation begins, the company’s main goal becomes recovering as much remaining asset value as possible for creditors.
This is also why the same batch of internal data is hard to sell while a company is still alive, but once it enters shutdown it may be revalued as an asset.
The value of internal corporate data has in fact been recognized by the industry for a long time. Salesforce has long treated the business communications accumulated in Slack as an important data asset, and Microsoft CEO Satya Nadella has also repeatedly emphasized that when enterprises use AI, the truly valuable part is their proprietary data and work context.
In the past, this data mainly served the company itself. Now, as AI companies actively seek internal corporate data, they have, for the first time, a clearer outside buyer.
The art of pricing
Although this business has only just begun to take shape, prices have already emerged in the market that can serve as reference points.
Data from shut down startups handled by SimpleClosure and Protege usually trades for between $10,000 and $100,000 per deal. Mercor’s offers for employee chat logs and emails from acquired startups have gone as high as $300,000.

Spirit suddenly pushed the price into the tens of millions of dollars. Mercor bid $7.5 million, and Google ultimately bid $10 million, competing for about 600 million internal communications and other corporate data. Calculated by those 600 million messages alone, the average comes to about 1.67 cents per message.
This unit price is not high. Reuters reported in 2024 that Photobucket had approached AI companies with licensing talks over roughly 13 billion photos and videos, pricing photos at about 5 cents to $1 each and videos at more than $1 each. In the B2B data market, a single contact record can also sell for a few cents to a few dollars depending on completeness and accuracy.
But these prices are still not enough to form a unified standard. How much internal data from a shut down company is worth is still mainly negotiated case by case. Data volume, industry, time span, completeness, uniqueness, and what the buyer plans to use it for all affect the final offer. Spirit’s $10 million is more like one of the few large publicly visible samples at present.
The gap becomes even more obvious when compared with a mature data-licensing market. Reddit licensed user posts and comments to Google for about $60 million a year; News Corp’s content licensing deal with OpenAI was about $250 million over five years; xAI’s deal with Telegram reached $300 million; and Apple’s purchase of Shutterstock image rights was quoted at between $25 million and $50 million.
These markets already have stable buyers, licensing methods, and pricing experience. Trading internal data from shut down companies has only just begun; which data is most valuable and whether it should be valued per message, by volume, or as a whole package still has no clear rules.
Spirit’s $10 million is another matter when viewed against the company’s own scale. In 2024, Spirit’s full-year revenue was close to $5 billion, averaging about $13.7 million per day. The money Google paid for this data was less than one day of normal operating revenue.
For bankruptcy liquidation, this is just one recovered item among the remaining assets; for AI companies, what they buy is a batch of data that has long been unpublished and contains real company operating processes.
clean-up and handover
After the data is sold, it still cannot be handed directly to the buyer. Before formal transfer, it usually goes through export, de-identification, packaging, court approval, and final delivery.
The first step is export. Slack enterprise data is typically exported as a ZIP file, including JSON files organized by channel, member information, and attachments.
Microsoft 365 can export emails and Teams messages through forensic tools. Spirit involved about 600 million internal communications, a very large volume of data, so in practice it would need to be processed in batches. Whether execution is carried out by an e-discovery provider or by the buyer’s engineering team has not been disclosed in public documents.
The next step is de-identification, meaning removing as much information as possible that could point to specific employees. Common methods include identifying and obscuring personal information such as names, phone numbers, and addresses, or replacing sensitive fields with IDs that cannot directly be matched to an individual. In some statistical and training scenarios, random noise is also added to further reduce the chance of re-identifying individuals.
But de-identification does not mean absolute anonymity. Even if names and email addresses have been removed, as long as the text retains enough occupational, time, location, or behavioral clues, there is still a possibility of re-identifying a person through other information.
So in Spirit’s deal, who is responsible for this step matters a lot. Under the current transaction arrangement, the data will be de-identified by an independent third party, but that entity is selected by Google and the cost is also borne by Google. Google has also promised not to use the data to re-identify individuals.
The sale of data by bankrupt companies is not that new. Over the past two decades, cases involving the handling or sale of customer data during bankruptcy have occurred at companies ranging from Toysmart, Borders, and RadioShack to 23andMe. These transactions involve consumer privacy and are usually subject to stricter review by courts, regulators, and state governments.
The difference with Spirit is that the core of the sale is not the passenger list, but employee emails, Teams messages, and other internal work data. Existing bankruptcy procedures have not developed mature rules for handling this kind of data the way they have for consumer information.
Controversy has already emerged. Spirit’s flight attendants’ union objected to the deal, and the court therefore delayed approval. What began as a batch of corporate data being sold as bankruptcy assets has started to implicate employee rights and the boundaries of AI use.
If the transaction is ultimately approved, the data will enter the final delivery stage. But very little has been disclosed in public documents about the path from the organized data package to Google’s systems. How the data will be transferred, whether it will undergo further cleaning, and in what final form it will enter product development or model training are all still unconfirmed.
Only when the process reaches this point does the data truly complete the transformation from bankruptcy asset to AI asset.
buyer
The clearest buyers for this kind of data right now are model companies like Google, and training-data providers like Mercor.
Google’s official explanation for the Spirit deal was product improvement and AI development. For Google, the value of this data lies in the fact that it records the real operating process of a large company, such as scheduling, collaboration, internal communication, project execution, problem handling, and management decisions.
These contents are hard to obtain from public webpages. Especially for enterprise agents, beyond the final result, what matters more is how tasks are actually driven forward and completed inside a real organization. Spirit’s internal records happen to contain a large amount of this kind of process data.
Another bidder, Mercor, better shows where this business is heading.
Mercor was founded in 2023, and the three founders, Brendan Foody, Adarsh Hiremath, and Surya Midha, had been debate teammates since high school. The company originally did AI recruiting, using models to help enterprises screen and interview candidates. By September 2024, Mercor had assessed about 300,000 job seekers and reached a valuation of $250 million.
After that, the company’s focus gradually shifted to AI training data. In November 2025, all three founders were only 22, and as the company’s valuation rose, they became one of the world’s youngest self-made billionaire groups at the time. In the first half of 2026, Mercor’s revenue exceeded $614 million, with about 90% coming from leading AI labs such as OpenAI. By July, the company was seeking a valuation of about $20 billion and had acquired Deeptune, a company dedicated to building training environments for AI agents.
Mercor has two main paths for acquiring training data.
One path is directly purchasing internal company records. It has offered bids to startups that were acquired or shut down, buying employee chat logs and emails, with the highest offer from a single company at about $300,000.

Another path is hiring people with real work experience. A TechCrunch report in October 2025 said AI labs recruited former employees of companies through Mercor, asking them to turn their work experience into training tasks and feedback, at an hourly rate of about $200. Mercor’s CEO once said the company pays more than $1.5 million per day to people participating in AI training.
Buying the work records left behind by companies on one hand, and buying the work experience held by employees on the other. Mercor’s core assets are, in essence, real-world work processes.
This also explains why it appeared at Spirit’s auction. For Mercor, 600 million internal communications are another, much larger source of training data.
This kind of business also comes with risks. In 2025, Scale AI sued Mercor, accusing former employees of taking trade secrets; after that, the company also experienced leaks of training-related information and paused partnerships. Data is both Mercor’s business and the asset it most needs to defend.
Finally, there is a numerical contrast. Spirit operated under that name for 34 years, while the three founders of Mercor, which participated in bidding for its internal data, were all only 22 at the time.
broken bench
Spirit’s deal has not been completed yet, the court has not signed off, and Google has not received the data. There are still union objections, hearings, and many procedures left to go through.
But this auction has made something that was rarely discussed before much more concrete.

In the past, when a company shut down, many things really did disappear. Financial statements stayed, trademarks stayed, patents stayed. But a company’s day-to-day experience usually did not. How a department held meetings, how a manager made decisions, how dozens of people coordinated after a delay, why a process was eventually changed into what it is today—these things were rarely formally recorded.
When a company dissolves, people leave, email accounts are deactivated, and chat groups are closed, and then they disappear with them.
That is why companies have always been such strange organizations.
It can survive for decades and accumulate the experience of tens of thousands of people, but what is truly inherited is often only a very thin slice of it. The next company still has to hire again, make the same mistakes again, and learn many things all over again.
What AI may change is exactly this. If emails, meetings, tickets, code changes, and internal discussions can really be organized into training data, then the experience a company once could only carry away through people will, for the first time, have another way to be preserved.
Where this will ultimately go is still hard to judge. Perhaps in the future companies will proactively preserve this data, perhaps employee contracts will be rewritten, perhaps bankruptcy law will add new restrictions, or perhaps internal work records will even be separately valued during financing and M&A.
Spirit brought this issue into the open ahead of time.
There is a widely circulated etymological explanation for the word bankruptcy. In medieval Italy, merchants did business sitting behind benches. When they became insolvent, the bench was publicly smashed. Banca rotta, the broken bench.
For hundreds of years, when the bench broke, that was the end of it.
Spirit’s bench has already broken. Its 600 million messages are being repriced.
One and a half pennies per message.
original link
