1. Why This Topic Deserves Its Own Guide
Prime editing and base editing patents are among the most sequence-dense filings in biotechnology. A single application can disclose:
- Base editor and prime editor fusion protein constructs (napDNAbp + deaminase, or napDNAbp + reverse transcriptase, often with fused UGI or MMR-evading domains)
- Hundreds to millions of guide RNA and pegRNA variants, frequently generated algorithmically against reference databases such as ClinVar
- Chemically modified nucleotides at defined positions (2′-O-methyl, 2′-fluoro, phosphorothioate, phosphonoacetate linkages, and similar)
- Non-standard structural elements unique to prime editing — spacers, gRNA scaffolds/cores, RT templates, primer-binding sites, homology arms, and linker sequences joining these regions
- Second-strand nicking guide RNAs, twinPE constructs, and other multi-component systems
Because these elements don’t map neatly onto the “plain” DNA/RNA/protein sequences that WIPO Standard ST.26 was originally designed around, prime and base editing filings are unusually prone to sequence listing errors — errors that can trigger formality objections, delay a filing date, or (in the worst case) create prosecution history estoppel issues if listings must be amended later.
This guide walks through what ST.26 requires, how it applies specifically to prime/base editing constructs, and where practitioners most commonly go wrong.
Not legal advice. This guide is general educational information about a technical filing standard. It isn’t a substitute for advice from a registered patent practitioner or your local IP office, and rules can vary by jurisdiction and change over time — always confirm current requirements with the relevant office before filing.
2. The Regulatory Baseline: WIPO ST.26
WIPO member states agreed that all sequence listings submitted as part of a patent application at the national, regional, and international level must comply with WIPO Standard ST.26. Applications filed on or after July 1, 2022 that disclose amino acid and nucleotide sequences must contain an ST.26 XML-compliant sequence listing, while listings for applications filed before that date generally remain governed by the older ST.25 standard.
Key mechanical differences from the legacy ST.25 regime:
| Feature | ST.25 (legacy) | ST.26 (current) |
| File format | Plain text (.txt) | Structured XML conforming to a defined DTD/schema |
| Non-natural sequence types | Not well supported | Explicitly supported (D-amino acids, branched sequences, nucleotide analogs) |
| Minimum sequence length | 4+ amino acids / 10+ nucleotides (looser enforcement) | 10 or more specifically defined nucleotides, or 4 or more specifically defined amino acids, for linear/unbranched sequences and linear regions of branched sequences |
| Validation | Manual/inconsistent across offices | A dedicated WIPO Sequence Validator web service that offices use to check compliance, alongside the WIPO Sequence desktop authoring tool |
| Database alignment | Not INSDC-compatible | Designed for compatibility with INSDC databases (GenBank, EMBL, DDBJ) |
The USPTO has kept pace with WIPO’s revisions: it adopted version 1.7 of ST.26 (approved by WIPO in December 2023), which improves technical terminology and descriptions, with the corresponding final rule effective July 1, 2024. Practitioners should always confirm which minor version of ST.26 a given office currently enforces, since validator behavior can shift slightly between versions.
Continuations and divisionals: the ST.26 obligation follows the filing date of the specific application, not the parent’s filing date. A continuation or divisional filed on or after the applicable implementation date must use an ST.26-compliant listing even if the parent case was filed and prosecuted entirely under ST.25 — this is a frequent trap for older CRISPR/base-editing families now spawning continuation practice.
3. What Counts as a “Sequence” in a Prime/Base Editing Application
Under ST.26, a sequence must be listed separately if it is an unbranched nucleotide sequence of 10+ specifically defined residues, or a polypeptide of 4+ specifically defined amino acids (with linear regions of branched molecules treated the same way). In practice, that threshold sweeps in nearly every functional element of a base or prime editing system:
- Editor/effector proteins: Cas9/Cas12 nickase or nuclease-dead backbones, cytidine or adenine deaminase domains (e.g., APOBEC-, TadA-derived), uracil glycosylase inhibitor (UGI) domains, reverse transcriptase domains, and the full-length fusion constructs joining them with linkers
- Guide RNA architectures: the spacer, scaffold/core, and — unique to pegRNAs — the 3′ extension arm broken into primer-binding site, RT/edit template, and homology arm sub-regions
- Second-strand nicking gRNAs used in nicking-based prime editing strategies and twinPE
- Modifier regions: optional 5′ and 3′ end modifiers, transcriptional termination signals appended to pegRNAs
- Target and edited genomic loci shown pre- and post-editing, including flanking homology sequences
- Linker sequences joining protein domains or joining a standard guide RNA to an RNA extension — note that a linker composed of amino acids (used in some fusion constructs) is measured against the 4-amino-acid threshold, while a nucleotide linker is measured against the 10-nucleotide threshold
Because prime editing patents often disclose pegRNA sequences designed programmatically against public variant databases (ClinVar-scale panels, for example), a single family can generate an extraordinarily large sequence listing — sometimes well over a million SEQ ID NOs — which raises distinct filing-logistics issues addressed in Section 6.
4. Feature Table Annotation: Where Most Errors Occur
ST.26’s XML schema requires each sequence to carry structured feature table annotations describing functional regions, not just a bare string of letters. For gene-editing constructs, common annotation pitfalls include:
- Under-annotating fusion proteins. Each functional domain (nickase, deaminase, UGI, RT, linker) should generally be annotated as a discrete feature with defined start/end coordinates — a single “misc_feature” spanning the entire fusion protein is typically insufficient and invites an examiner objection.
- Mislabeling pegRNA sub-regions. The spacer, scaffold, primer-binding site, RT template, and homology arm each have distinct biological function and should be separately annotated rather than lumped into one “misc_RNA” feature — both to satisfy ST.26 and because clear annotation supports later claim construction.
- Omitting modified-residue annotations. Chemically modified nucleotides — 2′-O-methyl-3′-phosphorothioate, 2′-O-methyl-3′-phosphonoacetate, 2′-fluoro, 2′-MOE, and related internucleotide linkage modifications commonly placed near the ends of a pegRNA — must be captured using ST.26’s modified-residue conventions, not silently represented as if they were standard ribonucleotides. Sequences containing modifications that ST.26 cannot represent as a residue-level annotation may need to be described in the specification text rather than (or in addition to) the sequence listing — confirm treatment with the relevant office.
- Inconsistent SEQ ID NO cross-referencing. In applications with combinatorial pegRNA libraries, it’s common for the specification, claims, and XML listing to drift out of sync after late-stage claim amendments. A final SEQ ID NO reconciliation pass against the XML file should be standard practice before filing.
- D-amino acids and non-natural residues — relevant where editor constructs include engineered or non-canonical residues — must use ST.26’s specific representation rules for non-standard monomers rather than the nearest standard-residue placeholder.
5. Practical Compliance Workflow
- Author in WIPO Sequence (or an equivalent ST.26-validated authoring tool) rather than retrofitting an ST.25 text listing. WIPO Sequence is the free desktop tool WIPO provides for authoring and validating ST.26-compliant listings.
- Annotate at the domain/region level as you draft, not as a post-hoc cleanup step — retrofitting feature tables onto a large pegRNA library after the specification is finalized is where most errors and delays originate.
- Run the built-in validator before submission. The desktop tool and the office-side WIPO Sequence Validator apply largely the same rules, though the desktop tool may flag as an “Error” something the Validator flags only as a “Warning” — so a clean desktop validation is a strong (not perfect) predictor of a clean office-side check.
- Reconcile the XML listing against claims and specification immediately before filing — particularly SEQ ID NO numbering, since claim amendments during drafting are a common source of drift.
- Check continuation/divisional filing dates independently — don’t assume a family’s ST.25 status carries forward.
- Confirm the applicable minor version of ST.26 with each target office, since versions (e.g., 1.7) can introduce terminology or validation changes that affect how a listing is checked.
- For very large combinatorial libraries, engage the target office early about file-size handling, submission mechanics, and whether representative subsets plus a general disclosure of the generation algorithm are acceptable in lieu of listing every computationally generated variant — practice on this point is still evolving and varies by office.
6. Jurisdiction Notes
- USPTO: Rules at 37 CFR 1.831–1.839 incorporate ST.26 by reference; the USPTO adopted ST.26 version 1.7 with a final rule effective July 1, 2024, and provides PatentPractice@uspto.gov as a point of contact for sequence-listing questions.
- EPO, JPO, KIPO, and other major offices enforce ST.26 for applications filed on or after the July 1, 2022 “big bang” date, with sequence listings needing to comply worldwide from that filing date forward.
- PCT applications: ST.26 applies at the international phase for applications filed on or after the implementation date, and national-phase entries generally inherit the international-phase listing’s compliance status — but always verify against the specific national office’s current rules, since implementation details and validator versions can differ.
7. Quick Pre-Filing Checklist
- [ ] Sequence listing is a valid ST.26 XML file (not a .txt or ST.25-format holdover)
- [ ] Every protein/nucleotide element at or above the 4-aa / 10-nt threshold is captured as its own sequence entry
- [ ] Fusion protein domains (nickase, deaminase, UGI, RT, linkers) are separately annotated with coordinates
- [ ] pegRNA sub-regions (spacer, scaffold, PBS, RT template, homology arm, end modifiers) are separately annotated
- [ ] Chemically modified nucleotides/linkages are represented using ST.26 modified-residue conventions
- [ ] SEQ ID NOs in the specification and claims match the XML listing exactly, post any late amendments
- [ ] File validated with zero errors in WIPO Sequence (or the office’s equivalent tool) immediately before submission
- [ ] Continuation/divisional filing dates checked independently for ST.26 applicability
- [ ] Target office’s currently enforced ST.26 minor version confirmed
- [ ] Large combinatorial library filing strategy (full listing vs. representative subset) cleared with the target office in advance, where volume is a concern
