Introduction

In biotechnology and life sciences patent practice, sequence listings are far more than a formal filing requirement. For applications that claim nucleic acids, amino acid sequences, proteins, peptides, antibodies, or other sequence-defined biological subject matter, the accuracy and completeness of the sequence listing can directly affect how broadly an invention can be claimed and defended.

A sequence listing identifies biological sequences in a standardized format and provides a technical reference point for the disclosure and claims. An incorrect, incomplete, or inconsistent sequence can therefore create problems that extend well beyond formatting. It may raise questions about written-description support, enablement, claim construction, priority, amendment options and ultimately the enforceable scope of the patent.

For patent applicants, the practical lesson is straightforward: sequence-listing quality should be treated as part of substantive patent strategy, not merely as a filing-office task.

What Is a Sequence Listing?

A sequence listing is a standardized disclosure of nucleotide and/or amino acid sequences associated with a patent application. In international patent practice, WIPO’s ST.26 standard governs sequence listings for applications filed on or after July 1, 2022, using XML as the required electronic format under the applicable PCT framework.

The purpose is to provide a consistent, machine-readable representation of biological sequences so that patent offices, applicants and the public can identify and process sequence information reliably.

For applicants, however, the sequence listing has an additional significance: it may contain the precise sequence information needed to understand and support sequence-based claims.

Why Accuracy Matters to Claim Scope

Patent claims define the legal boundaries of the invention. When a claim depends on a particular nucleotide or amino acid sequence, an error in the corresponding disclosure can create uncertainty about exactly what the applicant has disclosed and what the claims cover.

Consider a simplified claim directed to:

“An isolated nucleic acid comprising the nucleotide sequence of SEQ ID NO: 1.”

If SEQ ID NO: 1 contains an incorrect nucleotide, the problem may not be confined to the sequence listing. The error could affect the identity of the claimed molecule itself.

Similarly, a claim directed to:

“a polypeptide comprising the amino acid sequence of SEQ ID NO: 2”

may depend critically on whether SEQ ID NO: 2 accurately represents the intended protein.

The narrower or more sequence-specific the claim, the greater the potential impact of an inaccurate sequence.

1. Sequence Errors Can Create Written-Description Problems

Under U.S. patent law, 35 U.S.C. §112(a) requires the specification to contain a written description of the invention sufficient to demonstrate that the inventor possessed the claimed invention as of the relevant filing date.

Sequence-based inventions can make this requirement particularly important.

If the application identifies one sequence but the claims ultimately rely on another, the applicant may have to establish that the claimed subject matter was actually disclosed in the application as filed.

For example, suppose an inventor intended to disclose a protein containing a particular amino acid substitution, but the sequence listing accidentally omits that substitution. If the claims later depend on the corrected sequence, the applicant may face a question about whether the corrected sequence was adequately supported in the original disclosure.

This is why sequence verification should occur before filing, rather than being treated as a proofreading exercise after prosecution begins.

2. Sequence Identity Can Affect the Breadth of a Claim

Sequence-defined claims can be drafted at different levels of breadth.

Examples include claims directed to:

The reference sequence often serves as an anchor for determining the scope of the claim.

If the reference sequence is incorrect, the resulting claim may not cover the biological subject matter the applicant intended.

For example, a claim requiring a protein having “at least 95% sequence identity to SEQ ID NO: 10” depends not only on the percentage threshold but also on the accuracy of SEQ ID NO: 10.

A one-residue or one-nucleotide error can therefore have consequences that are disproportionate to the apparent size of the mistake.

3. Sequence Listings Must Match the Specification and Claims

One of the most important quality-control principles is consistency.

The following should be checked against one another:

Specification → Sequence Listing → Claims → Figures → Experimental Data

Potential inconsistencies include:

A sequence appearing in an experimental example should also be reconciled with the corresponding sequence identifier used elsewhere in the application.

4. Sequence Numbering Errors Can Have Legal Consequences

Sequence identifiers are intended to provide an unambiguous way of referring to sequences.

An application might state:

“The antibody heavy chain is shown in SEQ ID NO: 12.”

If the specification, sequence listing and claim set use different identifiers for the same sequence, the resulting ambiguity can complicate prosecution and potentially create problems in interpreting the claims.

For this reason, sequence numbering should be generated and maintained systematically rather than manually whenever possible.

The sequence identifier should be checked every time a sequence is referenced in:

5. Sequence Listing Errors Can Affect Priority

Priority is another reason sequence accuracy matters.

For a later application or continuation-in-part, the relevant question may include whether the earlier application adequately disclosed the sequence-based subject matter.

If the earlier filing contains an incorrect sequence and the later filing introduces a corrected version, the applicant may face an issue concerning whether the corrected sequence is entitled to the earlier filing date.

This can become particularly significant where intervening prior art exists.

Accordingly, when preparing a continuation, divisional, or continuation-in-part application, practitioners should compare the sequence disclosure across the relevant family members rather than automatically carrying forward the sequence listing.

6. ST.26 Has Changed Sequence-Listing Practice

For many modern patent applications, sequence-listing preparation must account for WIPO ST.26 rather than the earlier ST.25 standard.

ST.26 uses XML and establishes standardized rules for representing nucleotide and amino acid sequences. WIPO provides an online Sequence Listing Validator and related tools for applicants preparing ST.26-compliant listings.

This technical transition creates another reason for careful quality control.

A sequence can be biologically correct but still incorrectly represented in the required format. Conversely, a technically valid XML file does not guarantee that the underlying biological sequences are correct.

These are two different quality-control questions:

  1. Is the sequence biologically accurate?
  2. Is the sequence listing formally compliant?

Both need to be addressed.

7. Sequence Listing Accuracy and Enablement

Sequence accuracy can also intersect with enablement.

For biotechnology inventions, the specification generally needs to provide sufficient information to enable a skilled person to make and use the claimed invention without undue experimentation.

If a sequence is central to the claimed biological activity but the disclosed sequence is erroneous, the applicant may encounter questions about whether the disclosure actually enables the claimed subject matter.

This is especially relevant where the claimed biological function depends on a precise sequence, structural motif, binding region, or mutation.

The problem is not necessarily that every sequence must be experimentally verified. Rather, the disclosure should accurately represent the invention and provide an adequate technical foundation for the claims being pursued.

8. Genus Claims Require Special Attention

Sequence accuracy becomes even more important when an application seeks broad genus protection.

For example, a claim might cover:

variants having at least 90% sequence identity to a reference protein and retaining a specified biological activity.

The reference sequence, identity calculation and functional limitation all contribute to the scope of the claim.

An inaccurate reference sequence can therefore affect the universe of sequences potentially falling within the claim.

Applicants should also consider whether the specification provides sufficient support across the breadth of the claimed genus. A single reference sequence does not necessarily establish possession or enablement of every possible sequence variant encompassed by a broad percentage-identity claim.

9. Functional Claims Still Depend on Accurate Sequence Disclosure

A common misconception is that sequence accuracy matters only for claims that expressly recite a sequence.

In practice, sequence information may also support claims directed to biological function.

For example, a claim might cover:

Even if the claim does not simply say “SEQ ID NO: X,” sequence information may provide important support for identifying and characterizing the claimed biological entity.

10. Sequence Data Should Be Verified at Multiple Stages

A robust workflow should include multiple verification points.

Before Filing

Verify:

During Prosecution

When claims are amended, verify that:

Before Grant

Perform a final comparison among:

The objective is to ensure that the granted claims correspond to the sequence information actually intended to be protected.

11. Common Sequence-Listing Mistakes

Several errors occur repeatedly in sequence-based patent applications.

Manual Transcription

Copying sequences manually between laboratory records, spreadsheets, specifications and patent documents increases the risk of transcription errors.

Incorrect Sequence Identifiers

A single numbering error can cause a claim to point to the wrong biological sequence.

Missing Variants

If specific variants are important to the invention, they should be disclosed and identified appropriately rather than assumed to be covered by a generic description.

Inconsistent Mutations

A mutation described in the specification should correspond exactly to the sequence listing and experimental data.

Incorrect Translation

For nucleotide sequences encoding proteins, the translated amino acid sequence should be checked against the intended protein.

Failure to Update Related Applications

A correction made in one application may not automatically correct a related family member.

A Practical Sequence-Listing Quality-Control Checklist

Before filing a biotechnology patent application, consider performing the following checks:

How Accuracy Protects Claim Strategy

Sequence-listing accuracy ultimately supports three related objectives:

1. Defining the Intended Invention

The sequence listing should accurately identify what the inventor actually developed.

2. Supporting the Desired Claim Scope

The specification and sequence disclosure should provide an appropriate foundation for both narrow and, where justified, broader sequence-based claims.

3. Preserving Flexibility During Prosecution

A well-prepared sequence disclosure gives the applicant a stronger foundation for responding to examination issues, developing fallback positions and amending claims without unnecessarily creating support problems.

This is particularly valuable in biotechnology, where claim scope may evolve substantially during prosecution as the applicant responds to novelty, inventive-step/obviousness, written-description, enablement, or other objections.

Conclusion

Sequence listing accuracy is not simply a technical filing requirement. In biotechnology patent practice, it can directly influence how an invention is disclosed, claimed, amended and ultimately protected.

An incorrect sequence identifier, nucleotide, amino acid, mutation, or sequence length can create problems that extend into written description, enablement, priority, claim construction and enforcement.

The best approach is to integrate sequence verification into the substantive patent-drafting workflow. Patent counsel, scientists, patent paralegals and sequence specialists should work from controlled source data and verify the final sequence listing against the specification and claims before filing.

For sequence-based inventions, the precision of the sequence disclosure can become the precision of the patent’s claim boundary. Investing in sequence-listing accuracy at the beginning of the patent process can therefore help prevent costly problems later

Leave a Reply

Your email address will not be published. Required fields are marked *