Emails, chat records, project documents, and support tickets—these were once just digital legacies left behind to be cleaned up after the company shut down. Now, they’re being revalued, packaged, sold, and fed into AI companies’ training pipelines.

An undertaker preserves a person’s final dignity. After a business dies, dignity means proving that what you left behind still has value.

On August 17, at a bankruptcy asset auction, Google bid $10 million and bought all of Spirit Airlines’ corporate data. Another bidder, called Mercor, bid $7.5 million but lost by $2.5 million.

The lots are divided into three parts. The first part is about 100 million employees’ email messages. The second part is 500 million Microsoft Teams messages. Together, that totals 600 million. The third part includes calendars, spreadsheets, financial databases, project documents, operating records, and a batch of internal software. Passenger files, frequent-flyer information, and the like are not included.

600 million—if one person says 100 sentences a day, they’d have to keep talking for 16,000 years.

Do the math and each message sells for 1.67 cents. Americans call a one-cent coin a penny, and if it falls on the ground, nobody even bothers to pick it up. So a single sentence an employee of Spirit says in Teams is worth one and a half pennies.

Time goes back to 1980. Spirit Airlines was born in Detroit. Its predecessor was a trucking company called Charter One, before switching to flying airplanes. In 1992 it was renamed Spirit—meaning “spirit.” After that, for 34 years, it turned the model of ultra-low-cost aviation in the U.S. into a benchmark for the whole industry to imitate.

Its foundation used to be strong: 205 aircraft all of one model, Airbus A320; about 300 flights a day; and 2024 revenue of about $5 billion. But that same year it had a net loss of about $1.2 billion. When it filed for bankruptcy, it was carrying about $9 billion in debt.

One late night in May 2026, the company announced it would stop flights. On the second day, about 17,000 employees learned from the news that they had been laid off.

The process is still underway, and the transaction will have to wait for approval from the bankruptcy judge. Spirit is following the liquidation shutdown under Chapter 11. There is no bankruptcy trustee to take over; under court supervision, the company sells off its remaining assets item by item. For sensitive data such as employee emails and Teams messages, it must first be sent to an independent organization for processing—deleting information like names and email addresses that can identify individuals. That organization is selected by Google, and Google also bears the costs.

Also, after combing through public reports, there hasn’t been a precedent where a bankrupt company sold internal data to an AI company. This deal is very likely the first of its kind in history.

The title Gizmodo—an established U.S. tech media outlet—gave to this deal is: “Spirit is dead, but its data will linger in Google’s servers, floating around for generations.”

A new business

People who collect the dead have always existed. Lawyers, liquidators, auction houses—they’ve been doing it for decades. Airplanes, desks and chairs, trademarks, patents—things that can be sold were sold long ago.

What’s truly new started this year: even internal employee data is being put on shelves.

This business has emerged for two reasons behind it.

First, more companies are dying.

In the first quarter of 2024, the failure rate of U.S. startups rose 58% year over year. The number of active risk-investment institutions was 62% lower than at its peak.

The money hasn’t really gotten less—it’s just becoming increasingly concentrated flowing into AI. In 2024, U.S. AI startups took away $97 billion in funding, setting a historical record. The capital markets still have money, but are increasingly unwilling to spend it on anything beyond AI.

The result is that a set of companies that could originally keep surviving on financing start running into the wall even earlier. In August 2024, Tally, a fintech company backed by a16z, announced a shutdown. It raised $172 million in total, with a peak valuation of $855 million. It had already reached the Series D round, but still couldn’t get the next infusion of money.

Second: the data has become more valuable.

The consumption of data for large-model training has reached an extremely staggering scale. GPT-4’s training data is about 13 trillion tokens. For comparison, Google Books scanned about 40 million books over more than 40 years, which comes out to roughly 4 trillion tokens. That means the amount of data used for a single training run of GPT-4 is equivalent to more than three Google Books.

Image source: Dongcha Beating

Epoch AI has done the math: high-quality language data—books, news, and Wikipedia—will be exhausted around 2026. High-quality text that can be used to train models is rapidly running out on the public internet, and the share of synthetic data is increasing. Next, the new data that AI companies need can only go further beyond the public internet. Internal emails, chat records, and work documents that have accumulated for years inside enterprises are suddenly in view.

As for the market for AI training datasets, research institutions predict it can reach $9.7 billion by 2030. Taking all kinds of licensing into account, the estimated total market size of the “plate” could reach $67.5 billion.

On one side, more and more companies are shutting down and liquidating, leaving huge amounts of internal data from the past that has never been priced. On the other, AI companies’ demand for data outside the public internet is growing. When both sides show up at the same time, internal enterprise data finally gets the conditions to be traded at scale for the first time.

What truly makes this kind of data more valuable is the development of enterprise Agents. Gartner predicts that by 2026, 40% of enterprise applications will embed task-oriented AI Agents, whereas the year before that figure was still under 5%.

The training materials enterprise Agents need aren’t exactly the same as those for ordinary large models. Public webpages can provide knowledge, language, and finished content, but it’s hard to reconstruct a company’s real work process. Communication and collaboration, down to specific details—how a requirement is proposed, how a few people discuss it, how tasks are assigned, how it’s modified when problems arise, and finally how it’s delivered.

These processes store vast amounts of emails, chat records, help tickets, and project documents inside companies.

Once this kind of data starts to have clear buyers and use cases, what used to be directly deleted when a company shut down also gains standalone value for trading.

The collectors of the dead

People used to do this job were liquidation lawyers; now there are three more kinds of people.

The first type is a dissolution service provider, represented by SimpleClosure.

Image source: Dongcha Beating

This company does one thing: helping startups die with dignity. When the company just started in 2023, its pre-seed round raised $1.5 million, and in May 2025 it raised $15 million in an A round led by TTV Capital. Even Carta—also a company that manages equity and corporate affairs for lots of U.S. startups—shut down its own shutdown service, pivoting to invest in SimpleClosure and handing off those customer needs to it.

By October 2025, SimpleClosure had already handled the funerals of more than 1,000 companies. Crunchbase gave it a name: “A Better Way To Fail,” a better way to fail. U.S. startup founders love to say fail fast. SimpleClosure says fast failure isn’t enough—you also have to fail with dignity. Even on its official website, there’s a pricing calculator: input the company’s situation and you can estimate how much it will cost to die.

A funeral isn’t whitewashed either. In April 2026, SimpleClosure launched Asset Hub, dedicated to handling intangible assets left behind after companies shut down. Beyond brands, software, and customer lists, for the first time, internal work data like Slack records, emails, and Jira tickets was explicitly put on shelves.

This shows that before Spirit, the market had already started trying to price internal data from dead companies, though back then it was still mostly startups, small deals, and private matchmaking.

There’s a ready-made case. When the transcription and captioning company Cielo24 shut down, it sold 13 years’ worth of accumulated Slack messages, internal emails, and Jira tickets through SimpleClosure. Later, CEO Shanna Johnson told Forbes that this batch of data ultimately sold for hundreds of thousands of dollars.

For a company that has already decided to close its doors, this was originally a batch of data that needed to be cleaned up, but in the end it became an asset that could be recovered during liquidation.

In the data-seller step, SimpleClosure hands it to Protege. Protege is a data marketplace specializing in licensing training data for AI. In January 2026, it just raised $30 million with a16z leading the investment. Its founder is Bobby Samuels. Protege’s initial entry point was medical imaging: within 30 days, it assembled millions of images for buyers to use in pretraining.

Now, Protege is starting to apply this data licensing and trading capability to internal communications data from shut-down companies. SimpleClosure handles company shutdown and asset organization; Protege handles finding buyers, completing data licensing, and executing the transaction.

The second type of participant is bankruptcy courts and liquidation lawyers. For decades, they’ve been inventorying airplanes, desks and chairs, trademarks, and patents. Now, the list is starting to include more email inboxes, Slack records, and other internal data.

Under U.S. bankruptcy law, this data can be treated as intangible assets to be included in bankruptcy estate property and sold under court supervision. Relevant law firms have also started setting up dedicated teams to handle data preservation, evidence collection, and organization in bankruptcy cases. Redgrave LLP has such reorganization and evidence-collection business.

The third type is e-discovery service providers responsible for technical execution. Companies like KLDiscovery, Epiq, and Consilio have long been collecting, organizing, hosting, and reviewing enterprise data. Content in inboxes, Teams, and SharePoint needs to be exported and archived first, then organized into data packages that can enter the trading process.

This industry itself already has a set of established ways to charge. EDRM releases pricing surveys regularly. Common billing units include data collection per GB, hosting per GB per month, and document review fees.

In the end, Spirit’s 600 million messages can be turned into an auction item thanks to this kind of basic work. It’s just that compared with court documents and auction quotes, this portion rarely appears in public reporting—who exactly handles it and how it’s handled is usually invisible to outsiders.

An autopsy checklist

For an enterprise Agent, what it needs to learn isn’t just knowledge and standard answers—it also includes judgment, collaboration, error correction, and execution in real work environments.

The final product tells the model what it was ultimately made into; internal records tell it exactly how that happened.

This kind of change is reflected in Agent training data. In the past, a single code training sample might only be a few hundred tokens, with content that modifies a few lines of code. Now, a single Agent training sample often needs to include the full process: understanding the request, locating documents, modifying code, and verifying through testing.

Training data is starting to shift from single-answer responses to complete task-execution records. For Agents, the end result is obviously important, but more valuable for training is the judgment, operations, and feedback left behind during the process of completing the task.

Compared with normally operating companies, data from shut-down enterprises is easier to enter the transaction process. When a company is still operating, selling internal communications will involve trade secrets, employee privacy, non-compete risks, and customer relationships; legal and management teams are usually very cautious. After entering liquidation, the company’s main goal becomes recovering remaining assets as much as possible, so as to extract more value for creditors.

That’s also why, with the same batch of internal data, it’s difficult to sell while a company is still alive, but during the shutdown/liquidation stage it might be re-evaluated as an asset.

The value of internal enterprise data has been recognized by the industry for a long time. Salesforce has long treated enterprise communications accumulated in Slack as an important data asset. Microsoft CEO Satya Nadella has also emphasized multiple times that when enterprises use AI, a truly valuable part is their proprietary data and work context.

In the past, this kind of data mainly served the company itself. Now, as AI companies begin actively looking for internal enterprise data, they’ve gained a clearer external buyer for the first time.

The art of pricing

Although this business is only just forming, there are already some reference prices on the market.

Shutdown-startup-company data handled by SimpleClosure and Protege: for each deal, it’s usually between $10,000 and $100,000. Mercor’s employee chat logs and email-generated quotes for acquired startups can go as high as $300,000.

Image source: Dongcha Beating

Spirit pushed the price up into the tens of millions of dollars range in one go. Mercor bid $7.5 million; Google ultimately bid $10 million. They were competing for about 600 million internal communications and other enterprise data. Based on only those 600 million messages, the average is about 1.67 cents per message.

This unit price isn’t high. Reuters reported in 2024 that Photobucket once negotiated licenses with AI companies using about 13 billion photos and videos. The photo price was about $0.05 to $1 per photo, while video was more than $1 per clip. In the B2B data market, a contact’s information can also be sold for anywhere from a few cents to a few dollars depending on completeness and accuracy.

But these prices aren’t enough to form a unified standard. How much internal data from shut-down enterprises is worth is still mostly a one-off negotiation. The data volume, industry, time span, completeness, uniqueness, and what the buyer plans to do with it all affect the final bid. Spirit’s $10 million is more like one of the few large samples currently visible.

Compared with the already mature data licensing markets, the gap is even more obvious. Reddit licenses users’ posts and comments to Google for about $60 million per year. News Corp and OpenAI’s content licensing agreement is for about $250 million over five years. The deal between xAI and Telegram amounts to $300 million. Apple purchases Shutterstock image licenses, with quotes between $25 million and $50 million.

These markets already have stable buyers, licensing methods, and pricing experience. Trading internal enterprise data from shut-down companies is only just getting started. What data is most valuable, should it be priced per item, per capacity, or as a whole package—so far there are no clear rules.

Putting Spirit’s $10 million into the company’s own scale is a different story. In 2024, Spirit’s full-year revenue was close to $5 billion—about $13.7 million per day on average. The money Google paid to buy this batch of data is less than what it earns in a single normal day.

For bankruptcy liquidation, this is just one recovered item among remaining assets; but for AI companies, what they buy is a batch of data that has long been unavailable publicly and includes the real process of how companies operate.

Cleanup and handover

After the data is成交, it still can’t be directly handed over to the buyer. Before formal transfer, there are usually several steps: export, de-identification, organizing and packaging, court approval, and the final closing.

The first step is export. Slack’s enterprise data is typically exported as a ZIP file, including JSON files organized by channel, member information, and attachments.

Microsoft 365 can export emails and Teams messages through forensic tools. Spirit involves around 600 million internal communications—huge in volume—so in practice it needs to be handled in batches. Whether it’s executed by the e-discovery service provider or the buyer’s engineering team, public documents currently haven’t disclosed it.

Next comes de-identification, meaning deleting as much information as possible that points to specific employees. Common methods include identifying and masking personal information such as names, phone numbers, and addresses, or replacing sensitive fields with codes that can’t directly map to individuals. In some statistical and training scenarios, additional random perturbations may be added to further reduce the possibility of re-identifying individuals.

But de-identification doesn’t mean absolute anonymity. Even if names and email addresses have been removed, as long as the text retains enough professional, time, location, or behavioral characteristics, it may still be possible to re-identify individuals through other information.

So in Spirit’s deal, who is responsible for this step matters. Under the current transaction arrangement, the data will be handled for de-identification by an independent third party. But the organization is selected by Google, and Google also pays the costs. Google also promises it will not use this batch of data to re-identify individuals.

Selling data by bankrupt companies isn’t that new. Over the past 20-plus years, there have been cases where customer data was handled or sold during bankruptcy, from Toysmart, Borders, and RadioShack to 23andMe. These transactions involve consumer privacy and are typically subject to stricter scrutiny by courts, regulators, and state governments.

Spirit’s difference is that the core of what’s being sold this time isn’t a passenger list, but employee emails, Teams messages, and other internal work data. Existing bankruptcy procedures haven’t formed mature rules for handling this kind of data the way they have for consumer information.

Controversy has already appeared. Spirit’s flight attendants’ union objected to the transaction, so the court delayed approval. A batch of corporate data that was originally being sold as bankruptcy assets began to involve employees’ rights and the boundaries of AI use.

If the deal is ultimately approved, only then will the data enter the final closing stage. But from a prepared data package to Google’s systems, publicly disclosed documents provide very little. How data is transmitted, whether it will continue to be cleaned, and in what form it will ultimately enter product development or model training cannot be confirmed right now.

Once the process reaches this point, the data truly completes its transformation from a bankruptcy asset into an AI asset.

Buyer

At the moment, the clearest buyers of this kind of data are model companies like Google and training-data service providers like Mercor.

Google’s official explanation for its deal with Spirit is that it’s for product improvement and AI development. For Google, the value of this batch of data is that it records the real operational process of a large enterprise—such as scheduling, collaboration, internal communication, project progress, issue handling, and management decisions.

It’s hard to obtain this kind of content from public webpages. Especially for enterprise Agents, beyond the end result, what matters more is how the task is actually advanced and completed within a real organization. The internal records left behind by Spirit happen to include a large amount of such process data.

Another bidder, Mercor, even more clearly shows what direction this business is developing.

Mercor was founded in 2023. Its three founders—Brendan Foody, Adarsh Hiremath, and Surya Midha—have been debate team teammates since high school. The company started with AI hiring, using models to help businesses screen and interview candidates. By September 2024, Mercor had evaluated about 300,000 job seekers, reaching a valuation of $250 million.

After that, the company’s focus gradually shifted toward AI training data. In November 2025, all three founders were still only 22 years old. As the company’s valuation rose, they became among the youngest self-made billionaires worldwide at the time. In the first half of 2026, Mercor’s revenue exceeded $614 million, about 90% of which came from top AI labs such as OpenAI. By July, the company was seeking about a $20 billion valuation and acquired Deeptune, a company specializing in building training environments for AI Agents.

Mercor’s acquisition of training data mainly has two pathways.

One route is to purchase company internal records directly. It once quoted acquired or shut-down startups for buying employees’ chat logs and emails. In a single company, the maximum quote is about $300,000.

Image source: Dongcha Beating

The other route is hiring people with actual work experience. TechCrunch reported in October 2025 that AI labs recruit former employees of enterprises through Mercor, turning their work experience into training tasks and feedback. The hourly rate is about $200. Mercor’s CEO has said that the company pays more than $1.5 million per day to those participating in AI training.

On one side, you buy work records left behind by companies; on the other, you buy the work experience employees possess. Mercor’s core assets, in essence, are real-world work processes.

This also explains why it appears in Spirit’s auction listing. For Mercor, 600 million internal communications are another kind of training-data source—on a much larger scale.

These kinds of businesses also come with risks. In 2025, Scale AI sued Mercor, accusing former employees of taking trade secrets. Since then, the company has also faced leaks of training-related information and pauses in collaboration. Data is both Mercor’s business and its most asset that needs defending.

Finally, there’s a contrast in numbers. Spirit operated under this name for 34 years, while Mercor—bidding with its internal data—had all three founders who were only 22 years old at the time.

A broken bench

Spirit’s deal isn’t finished yet. The court hasn’t signed off, and Google hasn’t received the data batch. There are still union objections, hearings, and many procedures to go through.

But this auction has made something that used to be rarely discussed become concrete.

Image source: Dongcha Beating

In the past, once a company shut down, many things really would disappear. Financial statements remain, trademarks remain, patents remain. But a company’s everyday experience usually doesn’t get preserved. How a department meets, how a manager makes judgments, how dozens of people coordinate after a delay happens, why a process ultimately becomes something like today—few of these things are formally recorded.

When a company dissolves, people leave, inboxes are shut down, and chat groups are closed—they disappear along with it.

So a company is always a rather strange kind of organization.

It can live for decades and accumulate the experience of tens of thousands of people, but what’s truly passed down is usually only a very thin slice. The next company will still have to hire again, make mistakes again, and learn a lot of things over again.

This may be what AI changes. If emails, meetings, help tickets, code modifications, and internal discussions really can be organized into training data, then for the first time, a company’s experience that people previously carried out of the door has another way to be preserved.

Where this will ultimately go is still hard to judge. Maybe in the future companies will proactively preserve this data; maybe employee contracts will be rewritten; maybe bankruptcy law will add new restrictions; maybe when companies seek funding and do M&A, even internal work records will be valued separately.

Spirit had this problem laid out in advance.

Bankruptcy has a widely circulated etymology explanation. In medieval Italy, a merchant sat behind a bench to do business. When debts exceed assets, the bench is smashed publicly. Banca rotta—the broken bench.

For hundreds of years, when the bench breaks, the matter is over.

Spirit’s bench is already broken. The 600 million messages it left behind are being repriced.

One and a half pennies.

  • This article is republished with authorization from: (BlockBeats)

  • Original title: (After a company shuts down, how to make money by selling employee data)

  • Original author: Sleepy, Dongcha Beating

“After a company goes under, how do you make money by selling employee data?” The story of how Google drops a trillion dollars to buy 600 million messages—this article was first published in “Crypto City.”