If You Don’t Own the Data, You Don’t Own Anything.


Ask a biotech CFO where the enterprise value sits, and the answer usually points to the pipeline. Ask a sophisticated acquirer in a serious diligence process, and the answer is increasingly different. The pipeline is where the value is expressed. The data infrastructure is where it’s stored. 

This distinction used to be theoretical. It isn’t anymore. In AI-enabled therapeutics, the proprietary datasets, the training workflows, the annotation systems, and the computational feedback loops that continuously improve a discovery platform are frequently more strategically significant than any individual asset they’ve generated. They are the engine. The molecules are what the engine has produced so far. 

Most companies have not thought carefully about whether they actually own that engine — all of it, cleanly, in a way that survives commercial pressure. 

The Data Layer Is Where the Hidden Exposure Lives

Data rights in AI discovery are genuinely complex, and the complexity tends to be underestimated for a straightforward reason: it’s invisible during the scientific work. A researcher using a licensed dataset to train a model isn’t thinking about downstream commercialization restrictions. A team integrating a third-party annotation tool isn’t reviewing sublicensing provisions. The science moves forward and the rights questions accumulate quietly underneath it. 

The restrictions that matter most typically involve one of four things: whether the company can commercialize therapeutics derived from a dataset; whether it retains full ownership of model improvements trained on licensed data; whether it can sublicense relevant rights in a partnership or acquisition; and whether its freedom to continue building on the platform is conditional on a relationship it may not want forever. 

Any one of those can constrain enterprise value significantly. A company whose lead asset derives from a platform with unresolved commercialization restrictions on its training data has a problem that no amount of patent prosecution will fix. 

The Question Investors Now Ask

It’s no longer enough to show that a molecule is patentable. Sophisticated capital increasingly wants to understand whether the platform that generated it can be scaled, licensed, and commercialized without material restrictions flowing from data and model relationships. That’s a different question and many companies aren’t prepared to answer it. 


How the Exposure Accumulates Without Anyone Noticing

The pattern is consistent across companies that discover data rights problems late. It rarely starts with a bad decision. It starts with a series of good decisions made quickly, without the full rights picture in view. 

A collaboration begins with a clear scope and informal trust. The scope expands because the science is productive and the relationship is working. A model gets retrained on a combination of proprietary and externally sourced data because it improves performance and the provenance question seems manageable. A vendor relationship that started as a tool subscription gradually becomes load-bearing infrastructure. By the time someone looks carefully at the full picture, the dependencies are real and the documentation is thin. 

None of this happens through negligence. It happens through operational momentum. The problem is that operational momentum doesn’t stop when diligence starts, but the tolerance for ambiguity does. 


Why This Shows Up as a Valuation Problem, Not a Legal Problem

This is worth being direct about: the commercial consequence of weak data rights governance isn’t usually a lawsuit. It’s a discount. Acquirers and partners who identify material uncertainty around data ownership, who can’t get clean answers about what the company actually controls, across its full discovery infrastructure, apply that uncertainty to the valuation. Sometimes they restructure the deal to shift risk. Sometimes they walk away from a conversation that was otherwise going well. 

The company usually experiences this as a negotiation that went sideways, or a deal that lost momentum, or a partner who seemed interested and then became difficult. What actually happened is that the data layer didn’t hold up to scrutiny. 

  • Licensing discussions that were expected to be straightforward become complicated when a counterparty identifies restricted training data in the platform’s history. 

  • Acquisition conversations that reached term sheet stage stall when the acquirer’s technical diligence reveals model dependencies with ambiguous ownership. 

  • Fundraising rounds that should close efficiently slow down when institutional investors start asking questions the team hasn’t prepared clean answers for. 


The Companies That Get This Right Treat Data Governance as an Asset

There is a version of data governance that is a legal compliance function, something the team does because counsel said it needed to happen. That version is better than nothing. 

There is another version that is a strategic asset. Companies that have done the work, that can produce clear provenance for every significant data source, that have structured their collaboration agreements with downstream commercialization explicitly in mind, that have documented what they own and how they own it, those companies have a material advantage in any serious commercial conversation. They close faster. They negotiate from strength. They don’t lose value on the table to ambiguity that could have been resolved years earlier at a fraction of the cost. 

“The data layer is where the real asset sits. A company that has built an extraordinary discovery platform on a foundation of ambiguous data rights hasn’t built what it thinks it’s built.” 

The pipeline is the expression of value. The data infrastructure is the source of it. Knowing who owns the source, precisely, defensibly, in a way that scales is not a legal detail. It is the strategic question underneath every other strategic question in AI therapeutics. 

 
Next
Next

You Didn’t Build a Moat. You Built a Moment.