A recombinant protein can be the heart of a biotech patent. But when that protein is defined by an amino acid sequence, getting the sequence listing wrong can create an avoidable prosecution problem. For U.S. patent applications, sequence disclosures are subject to detailed formatting and filing requirements. Since July 1, 2022, qualifying applications have generally been required to use WIPO Standard ST.26 and an XML-based Sequence Listing rather than the older ST.25 text format.
For biotech companies, patent counsel and R&D teams, the practical takeaway is straightforward:
The sequence is part of the invention – and the sequence listing needs to be treated as part of the patent disclosure.
ST.26 Is the Starting Point
For U.S. applications filed on or after July 1, 2022, qualifying nucleotide and amino acid sequence disclosures must generally be submitted in ST.26 XML format. Applications filed before that date generally remain subject to the older ST.25 requirements.
For U.S. national-stage applications, the relevant date is the international filing date, not the date the application enters the U.S. national stage.
That distinction matters.
A biotech company may have an older priority application containing an ST.25 sequence listing and later file a U.S. application after July 1, 2022. The existence of the earlier ST.25 listing does not automatically allow the later U.S. application to use ST.25. The USPTO states that applications subject to the post-July 1, 2022 rules must use ST.26.
Which Recombinant Protein Sequences Need to Be Listed?
This is where careful review matters.
Under 37 CFR § 1.831, the relevant sequence disclosures generally include:
Amino acid sequences
An unbranched sequence – or a linear region of a branched sequence – containing four or more specifically defined amino acids that form a single peptide backbone falls within the sequence-listing requirements.
Nucleotide sequences
An unbranched sequence – or a linear region of a branched sequence – containing 10 or more specifically defined nucleotides can fall within the requirements, including certain nucleotide analogs and modified nucleotides.
Sequences below those thresholds should not be included in the Sequence Listing XML merely because they appear somewhere in the application. The USPTO specifically states that sequences with fewer than 10 specifically defined nucleotides or fewer than four specifically defined amino acids must not be included.
What does “enumeration” mean?
A sequence does not necessarily have to appear as a conventional one-letter sequence string to qualify.
Under the USPTO’s explanation of ST.26, “enumeration of residues” includes listing residues in order using names, abbreviations, symbols, structures, or certain shorthand formulas.
That makes the sequence review more important than simply searching the specification for strings of letters.
Recombinant Protein Applications: What Should You Watch?
A recombinant protein application can contain far more sequence information than the inventors initially realize.
For example, a specification might disclose:
- A wild-type protein sequence
- A preferred engineered sequence
- Multiple substitution variants
- Truncated versions
- Signal-peptide variants
- Fusion constructs
- Linker sequences
- Nucleic acids encoding the proteins
- Polynucleotide variants
- Control or comparison sequences
Each needs to be assessed against the ST.26 rules.
A useful pre-filing exercise is to create a sequence inventory.
| Sequence category | Example | Compliance question |
| Full-length protein | Recombinant enzyme | Does it meet the amino-acid threshold? |
| Protein variant | Mutant with substitutions | Is the complete sequence disclosed? |
| Protein fragment | Binding domain | Does it meet the threshold? |
| Fusion protein | Protein + Fc domain | How should the sequence features be annotated? |
| Encoding DNA | Coding sequence | Does it meet the nucleotide threshold? |
| Modified protein | Engineered or modified amino acids | Are the modifications properly represented? |
| Comparative sequence | Reference protein | Is it actually disclosed by enumeration? |
The safest approach is to review the entire disclosure, not just the sequences intended for the claims.
Sequence IDs Matter
One of the most important practical requirements is consistent use of sequence identifiers.
If the description or claims discuss a sequence that appears in the Sequence Listing XML, the application should refer to it using its sequence identifier, such as:
SEQ ID NO: 1
The USPTO expressly requires this type of reference even when the sequence is also reproduced elsewhere in the description or claims.
This creates a clean connection between:
Claim language → specification → sequence listing → actual sequence
That connection becomes particularly important in recombinant protein claims.
For example, instead of repeatedly reproducing a long protein sequence in multiple parts of the application, a claim may define the protein by reference to the appropriate sequence identifier, together with other legally relevant limitations.
The ST.26 XML File Is More Than a Sequence Dump
An ST.26 Sequence Listing is not simply a FASTA file converted into XML.
It contains structured information about the sequences and their associated annotations.
The USPTO describes the ST.26 listing as having two broad components:
- General information, including bibliographic information associated with the application; and
- Sequence data, including the sequences and their associated feature information.
Feature tables and qualifiers can therefore matter just as much as the sequence itself.
This is particularly relevant for recombinant proteins containing:
- Modified residues
- Signal peptides
- Mature protein regions
- Cleavage sites
- Fusion regions
- Engineered features
- Other sequence-specific annotations
The sequence should not be treated as an isolated string of amino-acid letters.
Its biological features and annotations are part of the compliance exercise.
Use WIPO Sequence – But Don’t Blindly Trust It
The USPTO recommends WIPO Sequence for preparing and validating ST.26 XML sequence listings.
The software can help applicants:
- Create ST.26 sequence listings
- Enter sequence information
- Add feature keys and qualifiers
- Validate project data
- Generate XML
- Validate an existing XML file
- Import certain sequence formats
- Produce human-readable versions for review
WIPO Sequence is not mandatory. Applicants may use other software, provided the resulting XML complies with ST.26 and the USPTO rules.
But there is an important caveat:
Passing the software validator does not mean the application is automatically compliant in every respect.
The USPTO notes that the validator checks most, but not all, ST.26 requirements. Applicants remain responsible for reviewing the sequence listing and its content.
That makes a final attorney-and-scientist review essential.
The File Format Matters
For applications subject to ST.26, the Sequence Listing XML must satisfy specific technical requirements.
Among other things, the USPTO states that the listing must be:
- A single XML file
- XML 1.0 format
- Encoded using Unicode UTF-8
- Valid according to the ST.26 Document Type Definition, or DTD
This is why simply attaching a Word document, PDF, spreadsheet, or conventional text sequence listing is not an adequate substitute.
The USPTO specifically requires the XML to conform to the ST.26 DTD.
Don’t Forget the Incorporation-by-Reference Statement
One easily overlooked filing detail is the incorporation-by-reference statement.
When the Sequence Listing XML is submitted as a separate file, the specification needs an appropriate statement identifying:
- The XML file name
- The date of creation
- The file size in bytes
The USPTO provides this requirement under 37 CFR § 1.834 and related provisions.
This is a small administrative detail with a big practical lesson:
Sequence-listing compliance is not just about creating the XML. The XML and the specification need to be properly connected.
Watch the File Size
Large recombinant protein or biologics applications can generate substantial sequence data.
The USPTO currently identifies a 100 MB upload limit for Sequence Listing XML files submitted through Patent Center. XML files exceeding that limit must be submitted using the permitted alternative medium, such as a read-only optical disc.
This should be checked before filing day.
A last-minute discovery that the sequence file cannot be uploaded is exactly the kind of avoidable administrative problem a filing checklist should prevent.
What Happens If the Sequence Listing Is Missing or Defective?
A defective sequence listing does not necessarily mean the application is doomed.
But it can create prosecution work – and potentially serious disclosure concerns.
The USPTO explains that when a required Sequence Listing XML is missing or defective, the Office may issue a notification requiring compliance. An examiner can also require a compliant or replacement Sequence Listing XML if sequences disclosed in the application are missing from the listing.
For an amendment adding or replacing sequence information, additional requirements can apply.
A replacement listing may need to include:
- The complete replacement XML
- An updated incorporation-by-reference statement
- Identification of additions, deletions, or replacements
- Support for the amended sequence information in the application as originally filed
- A statement that the replacement listing contains no new matter
That last point is critical.
Fixing a formatting problem should not accidentally become an attempt to introduce new sequence disclosure after filing.
The Biggest Mistake: Treating Sequence Compliance as an IT Task
Sequence-listing preparation sits at the intersection of science, patent drafting and filing procedure.
A good workflow therefore involves three levels of review.
Scientist review
Confirm that:
- The sequences are biologically correct
- Variants are correctly represented
- Modifications are accurately identified
- Sequence annotations make scientific sense
Patent counsel review
Confirm that:
- All qualifying sequences have been captured
- Sequence IDs are used consistently
- Claims and specification references align
- The listing is supported by the application
- No unintended new matter has been introduced
Technical validation
Confirm that:
- The XML is well formed
- The ST.26 DTD requirements are satisfied
- Required fields are present
- The file validates properly
- The correct version of the software and standard is being used
No single reviewer should be expected to catch every problem.
8 Practical Tips for Better Sequence-Listing Compliance
1. Start Early
Do not wait until the night before filing to build a sequence listing.
Complex biologics applications can contain hundreds – or thousands – of sequences.
2. Build a Master Sequence Inventory
Maintain a controlled list of every sequence disclosed by the inventors.
Track:
- Sequence identifier
- Biological description
- Sequence type
- Source
- Location in the application
- Claim relevance
- ST.26 status
3. Reconcile the Listing Against the Application
Compare the final XML against:
- Specification
- Claims
- Drawings
- Examples
- Tables
- Inventor-provided sequence files
The objective is simple:
No qualifying disclosed sequence should be accidentally omitted.
4. Use One Controlled Source of Truth
Avoid having one sequence in a lab spreadsheet, another in a Word document and a third in the patent software.
Version-control the underlying sequences before generating the final XML.
5. Review Modified Residues Carefully
Engineered proteins frequently contain substitutions or modifications.
Make sure those changes are correctly represented and appropriately annotated under ST.26.
6. Validate More Than Once
Run validation before finalizing the listing – and again immediately before filing if the file has been changed.
7. Check Sequence IDs Everywhere
A single incorrect SEQ ID NO. can create confusion between the claims, specification, examples and actual sequence.
8. Preserve the Filing Version
Keep a controlled copy of exactly what was filed, including:
- XML file
- Creation date
- File size
- Validation report
- Application version
- Supporting sequence files
- Final specification and claims
That record can become invaluable during prosecution or later enforcement.
A Pre-Filing Sequence Listing Checklist
Before filing a U.S. recombinant protein application, ask:
Applicability
- Is the application subject to ST.26?
- Was the relevant filing/international filing date on or after July 1, 2022?
- Have all qualifying nucleotide and amino acid sequences been identified?
Sequence content
- Are qualifying amino acid sequences included?
- Are qualifying nucleotide sequences included?
- Are sequences below the applicable thresholds excluded?
- Are modified residues correctly represented?
- Are relevant sequence features properly annotated?
Application consistency
- Are SEQ ID NO. references consistent?
- Do the claims point to the correct sequences?
- Does the specification match the XML?
- Does the XML contain only information supported by the application?
XML compliance
- Is the file in XML format?
- Is it UTF-8 encoded?
- Does it comply with the ST.26 DTD?
- Has it been validated using current WIPO Sequence tools or another compliant validator?
- Is the file within the Patent Center upload limit?
Filing formalities
- Is the incorporation-by-reference statement included?
- Does it identify the correct file name?
- Does it state the creation date?
- Does it state the file size in bytes?
- Has the final filed XML been archived?
Final Takeaway
For recombinant protein patents, sequence-listing compliance is not a minor filing formality.
It is part of building a reliable patent disclosure.
The transition to ST.26 XML has standardized how qualifying biological sequences are presented to the USPTO, but it has also made sequence preparation more structured and technically demanding.
The best approach is to treat sequence preparation as an integrated part of patent drafting:
Identify → inventory → annotate → validate → reconcile → file → preserve.
Start with the science. Check the patent disclosure. Validate the XML. Then check everything again. For biotech companies protecting recombinant proteins, antibodies, enzymes and other biologics, that discipline can prevent a surprisingly simple sequence-listing error from becoming a much bigger prosecution problem.
