Patent applications involving biotechnology, molecular biology, antibodies, proteins, nucleic acids, and genetic engineering frequently describe sequences by reference to homology, identity, similarity, consensus sequences, or sequence variants rather than by setting out a single exact sequence. These disclosures can raise two separate issues: substantive patent disclosure requirements and formal sequence-listing requirements. For U.S. applications filed on or after July 1, 2022, the USPTO’s sequence-listing requirements are based on WIPO Standard ST.26 and generally require qualifying nucleotide and amino-acid sequences to be presented in a computer-readable XML Sequence Listing. The challenge for patent practitioners is determining which sequences must be specifically disclosed and listed, how consensus or variant sequences should be represented, and how to preserve adequate written description and enablement while avoiding unnecessary or inconsistent sequence disclosures.
1. Homology and Consensus Sequences in Patent Applications
Biotechnology patents commonly describe an invention using language such as:
- “a polypeptide having at least 90% sequence identity to SEQ ID NO: 1”;
- “a nucleic acid encoding a protein having at least 95% identity”;
- “a consensus sequence”;
- “a sequence variant”;
- “a conservative substitution variant”;
- “a fragment of SEQ ID NO: 1”;
- “a sequence comprising residues X through Y”;
- “a sequence having one or more substitutions”; or
- “a sequence selected from the group consisting of…”
These formulations can be useful for obtaining commercially meaningful claim scope. But broad sequence language creates a disclosure question: what sequence information actually forms part of the application’s disclosure, and what must be included in the formal sequence listing?
The answer depends on the particular sequence, how it is disclosed, the applicable filing date, and the relevant USPTO rules.
2. ST.26 Is the Current U.S. Framework for New Applications
For U.S. patent applications filed on or after July 1, 2022, qualifying nucleotide and amino-acid sequence disclosures are governed by 37 C.F.R. §§ 1.831–1.835 and WIPO Standard ST.26.
The USPTO explains that applications filed on or after July 1, 2022 must use an XML-based Sequence Listing conforming to ST.26. Applications filed before that date generally remain subject to the older ST.25 framework, with special rules for PCT national-phase applications.
The filing date is therefore a critical first checkpoint when reviewing an existing portfolio or preparing a new application.
Practical rule
Do not assume that a sequence-listing format used in an older family can simply be carried forward into a new application.
The applicable standard must be determined from the relevant filing or international filing date.
3. What Triggers a Sequence Listing?
Under 37 C.F.R. § 1.831, a patent application filed on or after July 1, 2022 that discloses qualifying nucleotide or amino-acid sequences by enumeration of their residues must contain a separate computer-readable Sequence Listing XML.
The USPTO describes the sequence listing as part of the application’s disclosure. Importantly, the sequence-listing requirement is not limited to sequences that are ultimately claimed.
This has an important drafting consequence:
A sequence can matter for sequence-listing purposes even if the applicant does not intend to place that particular sequence in a claim.
Patent teams should therefore review the entire specification, not merely the claims, when determining whether sequence-listing obligations have been triggered.
4. Exact Sequences Versus Homology-Based Definitions
An exact sequence is comparatively straightforward.
For example:
SEQ ID NO: 1: ATGCC…
The sequence is specifically enumerated and can generally be assigned a sequence identifier in the Sequence Listing XML.
A homology-based definition is different:
“A nucleic acid having at least 90% sequence identity to SEQ ID NO: 1.”
The claim or specification may define an entire class of sequences without enumerating every possible sequence falling within that class.
The formal listing requirement should not be confused with the substantive question of whether the disclosure adequately supports that broad genus.
These are separate inquiries.
Formal question
Does the disclosed sequence information fall within the regulatory definition requiring inclusion in the Sequence Listing XML?
Substantive question
Does the specification adequately describe and enable the claimed sequence genus?
A compliant sequence listing does not, by itself, establish adequate written description or enablement.
5. Consensus Sequences Require Special Attention
A consensus sequence is generally a sequence constructed to represent common residues or characteristics across multiple sequences.
For example, an applicant may identify several naturally occurring protein sequences and disclose a consensus sequence representing conserved positions.
From a patent-drafting perspective, the practitioner should distinguish among:
- The individual sequences used to generate the consensus.
- The resulting consensus sequence.
- The method used to calculate the consensus.
- Variants permitted at individual positions.
- The functional properties attributed to the consensus or variants.
Simply stating that a sequence is a “consensus sequence” may leave important technical questions unanswered.
A stronger disclosure may explain the underlying sequence set, alignment methodology, position numbering, residue-selection criteria, and permitted variation.
6. Consensus Does Not Automatically Mean “Every Possible Variant”
Suppose a specification identifies a consensus protein and states that variants having at least 90% identity are included.
That statement potentially encompasses a very large number of sequences.
The patent drafter should consider whether the specification provides enough information for a skilled person to understand the claimed genus and practice it without undue experimentation.
Relevant disclosure may include:
- Representative sequences.
- Conserved residues.
- Variable positions.
- Functional assays.
- Structural information.
- Sequence alignments.
- Mutation data.
- Examples of active variants.
- Definitions of acceptable substitutions.
- Correlation between sequence variation and biological activity.
This is particularly important when the claims are drafted around functional characteristics rather than merely sequence identity.
7. Sequence Identity Should Be Defined Carefully
Terms such as “homology,” “identity,” and “similarity” are sometimes used interchangeably in patent specifications, even though they can have different technical meanings.
A specification should, where appropriate, define:
- The sequence-comparison algorithm.
- The software or methodology.
- The parameters used.
- Whether gaps are permitted.
- Whether identity is calculated over the entire sequence or a particular region.
- How insertions and deletions are treated.
- How ambiguous residues are handled.
For example, “at least 90% sequence identity” can produce different results depending on the comparison methodology.
A carefully drafted definition can reduce ambiguity during prosecution and later claim construction.
8. Sequence Listings and Figures Are Not the Same Thing
The USPTO specifically addresses sequences presented in drawing figures. For applications subject to ST.26, qualifying nucleotide or amino-acid sequences must conform to the sequence rules concerning how they are presented and described in the Sequence Listing XML.
This means that putting a sequence into a figure does not necessarily eliminate the need to address it in the Sequence Listing XML.
For example, a patent drawing might contain:
1 ATGCCGAT…
If that sequence meets the applicable definition, the practitioner should not assume that the figure alone satisfies the sequence-listing requirement.
The sequence should be evaluated under the ST.26 rules independently.
9. Sequence Identifiers
Under the ST.26 framework, each qualifying disclosed nucleotide or amino-acid sequence is separately represented and assigned a sequence identifier. The USPTO’s MPEP explains that each disclosed qualifying sequence must appear separately in the Sequence Listing XML.
This makes consistent sequence identification important throughout the application.
A typical specification may therefore refer to:
- SEQ ID NO: 1;
- SEQ ID NO: 2;
- SEQ ID NO: 3; and so forth.
The same identifiers should be used consistently across:
- The specification.
- Claims.
- Tables.
- Figures.
- Sequence alignments.
- Examples.
- Expert materials, where applicable.
Inconsistent identifiers can create avoidable prosecution and litigation problems.
10. Partial and Fragmentary Sequences
Biotechnology applications frequently disclose partial sequences.
Examples include:
- Protein fragments.
- Antibody variable regions.
- Peptide fragments.
- Gene fragments.
- Conserved domains.
- Sequence regions separated by unknown portions.
ST.26 contains specific rules concerning sequences with defined regions separated by gaps of unknown or undisclosed numbers of residues. The USPTO explains that such regions may need to be represented as separate sequences rather than as one continuous sequence.
This can become important when a specification describes something like:
residues 1–50 … residues 101–150
without disclosing the intervening sequence.
The drafting team should not automatically concatenate the known portions into a single artificial sequence.
11. The Relationship Between Listing Requirements and Written Description
The sequence listing is a formal disclosure mechanism. It does not solve every substantive disclosure problem.
For U.S. patentability purposes, a broad claim to sequence variants can raise written-description issues if the specification does not adequately demonstrate possession of the claimed genus.
For example, a claim covering:
“a polypeptide having at least 80% sequence identity to SEQ ID NO: 1 and retaining enzymatic activity”
may encompass an enormous number of possible sequences.
The patent specification should be evaluated for whether it provides adequate support for the claimed scope.
Potential supporting material may include:
- Multiple representative sequences.
- Sequence alignments.
- Conserved motifs.
- Functional examples.
- Mutational studies.
- Structural models.
- Tables identifying substitutions.
- Experimental data demonstrating retained activity.
The appropriate level of disclosure depends heavily on the technology and claim scope.
12. Enablement Considerations
Enablement presents a related but distinct issue.
A broad sequence claim may cover many theoretical variants. The practitioner should consider whether the specification enables the skilled artisan to make and use the full scope without undue experimentation.
For functional sequence claims, particularly those involving biological activity, relevant considerations may include:
- How many variants are covered?
- How predictable is the technology?
- Are sequence-function relationships understood?
- Are there representative working examples?
- Are critical residues identified?
- Can the skilled artisan predict which substitutions will preserve function?
- Does testing every candidate require substantial experimentation?
The sequence listing itself does not answer these questions.
13. Priority and New Matter
Sequence drafting also has important consequences in continuation, divisional, and international filing strategies.
A later application may seek claims directed to a sequence variant that was not adequately disclosed in the earlier priority application.
The question is not merely whether a sequence listing exists. Counsel must determine whether the earlier application actually provides adequate support for the later claim.
This is especially important for:
- New sequence variants.
- New consensus sequences.
- Expanded identity ranges.
- Newly identified conserved regions.
- Newly discovered functional substitutions.
- Additional antibody sequences.
- Newly generated engineered sequences.
A sequence added later cannot necessarily obtain the earlier application’s priority date simply because it is scientifically related to a previously disclosed sequence.
14. ST.25 Versus ST.26: Portfolio Review
Companies with large biotechnology patent portfolios may contain both ST.25 and ST.26 applications.
The USPTO summarizes the basic distinction as follows:
| Filing situation | Applicable standard |
| Application/international filing before July 1, 2022 | Generally ST.25 |
| Application/international filing on or after July 1, 2022 | ST.26 |
| New U.S. application claiming priority to older application | ST.26 may still be required |
| PCT national phase | International filing date is important |
The USPTO specifically notes that a later application claiming priority to an earlier application containing an ST.25 listing is nevertheless required to provide an ST.26-compliant XML listing when the applicable filing date is on or after July 1, 2022.
This is an important portfolio-management issue.
15. Preparing the ST.26 XML
The USPTO recommends using WIPO Sequence, a software tool designed to help applicants generate compliant ST.26 sequence listings.
An ST.26 Sequence Listing XML contains general application information and sequence data. It must conform to the applicable XML structure and WIPO ST.26 requirements.
The file is encoded in UTF-8 and must meet specified XML and naming requirements under 37 C.F.R. § 1.834.
For U.S. applications, the USPTO also explains that the Sequence Listing XML is submitted as a separate part of the disclosure and requires an incorporation-by-reference statement identifying the XML file, creation date, and file size.
16. A Practical Disclosure Workflow
A reliable workflow can be divided into several stages.
Stage 1: Identify all sequences
Review the entire application for:
- Nucleotide sequences.
- Amino-acid sequences.
- Partial sequences.
- Consensus sequences.
- Mutant sequences.
- Variant sequences.
- Sequences appearing in figures.
- Sequences appearing in tables.
- Sequences embedded in examples.
Stage 2: Classify each sequence
Determine whether each sequence is:
- Exact.
- Partial.
- Consensus.
- Variant.
- Artificial.
- Naturally occurring.
- Defined by a formula or residue alternatives.
- Defined only by a percentage identity.
Stage 3: Determine ST.26 applicability
Confirm the relevant filing or international filing date and determine whether ST.25 or ST.26 applies.
Stage 4: Prepare the sequence data
Use appropriate sequence-management software and verify that each qualifying sequence is represented correctly.
Stage 5: Cross-check the specification
Every sequence identifier should correspond correctly to the sequence described in the specification.
Stage 6: Review substantive support
Separately evaluate written-description and enablement support for any broad sequence genus.
Stage 7: Validate the final listing
Run the appropriate validation tools and perform a human review of the resulting XML.
17. Common Drafting Mistakes
Treating “90% identity” as self-explanatory
The comparison methodology may materially affect the scope.
Assuming only claimed sequences must be listed
The sequence rules can apply to qualifying sequences disclosed in the specification even if they are not claimed.
Copying an old ST.25 listing into a new filing
New applications subject to the current rules require ST.26 XML.
Listing an artificial sequence that was never actually disclosed
A sequence listing should faithfully reflect the application’s disclosure rather than introduce new sequence information.
Ignoring sequences in figures
A sequence appearing in a figure should be evaluated under the applicable sequence rules.
Assuming sequence-listing compliance establishes written description
Formal compliance and substantive support are separate questions.
Failing to review sequence identifiers
A single numbering error can propagate through the specification, claims, examples, and prosecution history.
18. Litigation and Post-Grant Implications
Sequence disclosures can become particularly important during patent litigation, post-grant proceedings, licensing disputes, and validity challenges.
Counsel may need to determine:
- What sequence was actually disclosed on the filing date?
- Which sequence identifier corresponds to which biological molecule?
- Whether a later-added variant was supported by the original disclosure.
- Whether a claim’s percentage-identity language encompasses the accused sequence.
- What algorithm was used to calculate sequence identity.
- Whether a consensus sequence represents actual disclosed sequences.
- Whether a figure contains a sequence that was separately listed.
- Whether amendments introduced new sequence matter.
For this reason, the original application file, sequence listing, XML file, prosecution history, and any amendments should be preserved together.
19. Quality-Control Checklist
Before filing a biotechnology patent application containing sequence information, the patent team should confirm:
- The relevant filing date has been identified.
- ST.25 versus ST.26 applicability has been determined.
- All qualifying nucleotide and amino-acid sequences have been identified.
- Sequences appearing in figures and tables have been reviewed.
- Consensus sequences have been evaluated separately from their source sequences.
- Variant and partial sequences have been reviewed.
- Sequence identity methodology is defined where appropriate.
- Sequence identifiers are consistent throughout the application.
- The Sequence Listing XML has been generated in the required format.
- The XML has been validated.
- The specification contains the required incorporation-by-reference statement where applicable.
- Written-description support has been evaluated for broad sequence claims.
- Enablement has been evaluated for the full claim scope.
- Priority support has been reviewed for every important sequence genus.
- The final XML has been archived with the filed application materials.
Conclusion
Homology, consensus, and variant-sequence drafting requires coordination between substantive patent strategy and formal sequence-listing compliance. For new U.S. applications subject to the current rules, qualifying nucleotide and amino-acid sequences must generally be handled through an ST.26-compliant XML Sequence Listing. But formal listing is only one part of the analysis. Patent practitioners should separately evaluate whether the specification adequately supports the desired scope of sequence identity, homology, consensus, and functional variants. They should also ensure that sequence identifiers, figures, examples, claims, and the XML listing remain internally consistent. The most reliable approach is to treat sequence disclosure as a coordinated process: identify every sequence, classify it correctly, prepare the appropriate listing, validate the XML, and independently assess written description, enablement, and priority support. That approach can reduce prosecution problems and provide a much stronger evidentiary foundation if the resulting patent later becomes the subject of licensing negotiations, validity challenges, or patent litigation.
