When investors, licensees, or acquirers value a biotech patent, most of the attention goes to the obvious things: claim scope, freedom-to-operate, remaining patent term, clinical-stage progress and the competitive landscape. One component gets far less scrutiny than it deserves, given how much damage it can do – the sequence listing. For any patent disclosing nucleotide or amino acid sequences, the sequence listing isn’t a clerical attachment; it’s part of the legal disclosure itself and errors in it can quietly undercut the value of an otherwise strong patent long before anyone notices.
What a Sequence Listing Actually Is
A sequence listing is a structured, standardized disclosure of every nucleotide or amino acid sequence referenced in a patent application – identified by “SEQ ID NO:” numbers, with annotations describing what each sequence represents (a coding region, a promoter, a binding domain, an organism of origin and so on). It’s filed alongside the specification and is treated as part of the application’s disclosure, meaning it’s subject to the same requirements for accuracy, support and enablement as the rest of the patent text.
In the United States, sequence listings are governed by 37 C.F.R. §§ 1.821–1.825 and internationally by WIPO standards. This area went through a significant procedural shift in 2022: WIPO Standard ST.26 replaced the long-standing ST.25 format as the global standard, moving from a plain-text disclosure to a mandatory structured XML file with expanded annotation requirements. Applications filed on or before June 30, 2022 continue under ST.25 for the life of that application; applications filed after that date must use ST.26. Critically, an application can’t switch formats mid-prosecution – an ST.25 listing can’t later be converted to ST.26 even if amendments are filed after the transition date. For companies with active patent families that straddle 2022, this means a single portfolio may legitimately contain sequence listings in two different formats, each with its own compliance rules, which is itself a diligence trap if not tracked carefully.
Why Accuracy Matters More Than It Looks Like It Should
A typo in a sequence listing sounds like a minor drafting error. In practice, it can be far more consequential than errors almost anywhere else in a patent application, for a few structural reasons:
Sequences are claim-defining, not just descriptive. In many biotech patents, the claims themselves recite a sequence directly, or recite percent identity/similarity to a disclosed sequence (e.g., “at least 95% identical to SEQ ID NO: 3”). If the sequence listing contains an error, the claim’s actual scope may not match what the inventors intended to claim, what was actually reduced to practice, or what’s supported elsewhere in the specification.
Inconsistency between the listing and the specification is a written-description and enablement problem. If the sequence recited in the claims doesn’t match the sequence described in the body of the application – even by one residue – an examiner and later an accused infringer or petitioner in an inter partes review, has a ready-made argument that the claim lacks written description support or wasn’t enabled as filed. This isn’t a formality; it can be case-dispositive in litigation.
Errors can affect priority claims. For applications that claim priority to an earlier filing, a mismatch between the sequence disclosed in the priority document and the sequence in the later application can jeopardize the priority date itself. Losing priority can expose the claims to intervening prior art that would otherwise have been excluded – sometimes enough to invalidate the patent entirely.
Formatting non-compliance causes prosecution delay and cost, but substantive errors cause something worse. A formatting problem (wrong file type, missing annotations, non-compliance with the current WIPO standard) typically triggers an office action and a fixable delay. A substantive sequence error – the wrong sequence entirely, a transcription mistake, an incorrect SEQ ID NO cross-reference between the listing and the claims – can survive prosecution unnoticed and only surface later, during licensing diligence, litigation, or an invalidity challenge, which is precisely when it’s most expensive to discover.
How This Translates Into Valuation Risk
Patent valuation for biotech assets typically uses some combination of an income approach (discounted cash flows attributable to the patent-protected product or platform), a market approach (comparable licensing or transaction benchmarks) and a cost approach (replacement or development cost), often wrapped in a risk-adjusted NPV framework that accounts for the probability of technical and regulatory success. Sequence listing accuracy doesn’t show up as its own line item in these models, but it directly affects several of the inputs that do:
Claim validity risk (and therefore the discount rate or probability-of-success weighting). A patent whose core claims are vulnerable to a written-description or enablement challenge because of a sequence discrepancy carries meaningfully higher invalidity risk than one without that exposure. Sophisticated valuators build this into the risk adjustment applied to projected cash flows – a patent asset with a latent sequence defect should be discounted more heavily than the headline claims might suggest, even before any challenge has actually been filed.
Freedom-to-operate and blocking-patent analysis. Buyers and licensees rely on the sequence listing to determine exactly what’s claimed and what isn’t, in order to assess whether a target product falls inside or outside the claims. An inaccurate listing can cause a freedom-to-operate opinion to be built on the wrong sequence, which is a problem that doesn’t surface until someone actually compares the granted claims to the biological reality of the product – sometimes post-close.
Diligence cost and deal friction. In M&A and licensing diligence, sequence listings are increasingly checked against the underlying research data, not just against USPTO formatting requirements. Discovering an error during diligence doesn’t just create a fix-it task; it raises broader questions about the rigor of the applicant’s original filing process, which can prompt buyers to look harder at the rest of the portfolio, extend diligence timelines, or price in a broader risk discount than the specific error alone would justify.
Licensing and royalty stacking disputes. Where a license is defined by reference to specific SEQ ID NOs (common in platform-technology and biologics licensing), an inaccurate or ambiguous listing can create genuine disputes over what’s actually licensed – whether a given product falls within scope and therefore whether royalties are owed at all. These disputes are expensive regardless of outcome and the mere possibility of one is itself a valuation discount factor for licensors negotiating deal terms.
Portfolio-level exposure in platform companies. For companies whose value is concentrated in a sequence-based platform (antibody libraries, gene-editing constructs, mRNA sequences) rather than a single molecule, sequence listing errors are not isolated to one patent – the same core sequence often appears across a family of applications. An error introduced early and propagated across continuations, divisionals and foreign counterparts can compound into a portfolio-wide exposure rather than a single-patent one, which is exactly the kind of systemic risk valuators and acquirers try hardest to find and price during diligence.
Where Errors Actually Come From
Understanding the common failure points helps explain why this risk is more persistent than it should be:
- Manual transcription errors when sequence data moves from lab notebooks, sequencing output files, or spreadsheets into the sequence listing document.
- Inconsistent SEQ ID NO cross-referencing between the listing, the specification and the claims – especially after amendments during prosecution that renumber or add sequences without updating every reference.
- Format-standard confusion, particularly around the ST.25/ST.26 transition, where legacy ST.25 annotations or conventions get carried into an ST.26 filing incorrectly, or a conversion is done without re-verifying the underlying sequence data itself (not just the formatting).
- Incomplete annotation of features like coding regions, mutations, or binding sites, which can leave the disclosure ambiguous about what the sequence actually represents even if the raw sequence string itself is correct.
- Divergence across a patent family as continuations and foreign filings are prepared by different teams or outside counsel over time, without a single source of truth for the sequence data.
Practical Implications for Patent Holders and Acquirers
For companies building or maintaining a biotech patent portfolio, sequence listing accuracy is worth treating as a distinct diligence and quality-control category, not an afterthought handled entirely by outside filing agents:
- Verify sequences against original research data (not just against a previous filing) before submission and again before any conversion between ST.25 and ST.26.
- Maintain a single authoritative source for each sequence used across a patent family, so continuations and foreign counterparts don’t drift from the original.
- Cross-check SEQ ID NO references between the listing, the specification and the claims after every amendment during prosecution.
- Use validation tools to catch formatting and structural errors before filing, but don’t treat a passed validation as confirmation that the underlying sequence data itself is correct – validators check compliance with the standard, not scientific accuracy.
For investors, licensees and acquirers evaluating a biotech patent or portfolio, sequence listing review is worth building into technical diligence alongside claim-scope and prior-art analysis:
- Independently verify that claimed sequences match the sequence listing and the underlying data referenced in the specification, rather than relying solely on the applicant’s representations.
- Check whether a target portfolio spans the ST.25/ST.26 transition and confirm each application used the correct standard for its filing date.
- Trace core sequences across the full patent family to check for drift or inconsistency introduced in continuations or foreign filings.
- Factor confirmed or suspected sequence discrepancies explicitly into the risk-adjustment or discount rate used in the valuation model, rather than treating them as a pure legal formality separate from the financial analysis.
The Bottom Line
Sequence listing accuracy sits at an unusual intersection: it’s simultaneously a technical compliance requirement, a substantive legal disclosure issue and – because of how directly it can affect claim validity, priority and licensing scope – a genuine driver of financial risk. Patents built on sequences are only as strong as the accuracy of the sequences disclosed and that accuracy rarely gets the same diligence attention as claim charts or prior art searches. For anyone valuing, acquiring, or licensing a biotech patent asset, treating the sequence listing as a routine formality rather than a substantive risk factor is one of the more avoidable ways to overpay for – or overestimate the strength of – an otherwise valuable innovation.
