Patents covering biologics – therapeutic proteins, antibodies, peptides and their derivatives – carry an obligation that small-molecule patents don’t: a formal sequence listing. When the claimed invention is a PEGylated protein, that requirement collides with a structural reality of PEGylation itself – polyethylene glycol is not an amino acid, has no place in a standard sequence alphabet and often attaches at a site, or across a distribution of sites, that a linear sequence string cannot fully describe. Getting the sequence listing right and understanding what it can and cannot capture, is a distinct technical exercise that trips up applicants who are otherwise experienced with biologics patenting.
This article covers what sequence listings require, how PEGylation is represented within them, where the listing’s limits are and the drafting practices that keep a PEGylated-protein application from stalling in formalities review.
Why Sequence Listings Exist and Who Requires Them
Any patent application that discloses a nucleotide or amino acid sequence of a specified minimum length must include a sequence listing – a standardized, machine-readable file enumerating each sequence, annotated with defined feature keys. The purpose is to let sequences be searched, compared and deposited in public databases (such as those maintained by the USPTO, EPO and WIPO, which feed into repositories like GenBank) in a consistent format across jurisdictions.
As of the current international standard, sequence listings are prepared in WIPO Standard ST.26 format (XML-based), which replaced the older ST.25 (text-based) standard. ST.26 is now mandatory for international (PCT) applications and has been adopted by the USPTO, EPO and most major patent offices for applications filed after their respective transition dates. Applicants working from older templates or precedent applications drafted under ST.25 need to confirm their filing complies with the office’s current standard – the two formats are not interchangeable and filing under the wrong one is a common source of formalities objections.
For a PEGylated biologic, the sequence listing requirement applies to the underlying protein or peptide sequence – the amino acid chain – regardless of whether that chain is later modified with PEG. A 25-mer therapeutic peptide that is PEGylated at its N-terminus still requires a sequence listing entry for the 25 amino acids; the PEG itself is not a sequence and is not separately listed as a sequence, but its presence and attachment point must be captured elsewhere in the listing’s annotation and in the specification.
PEGylation and the Amino Acid Alphabet Problem
Sequence listings are built around the standard sets of nucleotide and amino acid symbols – the 20 canonical amino acids (plus a small set of ambiguity codes and recognized non-canonical residues). PEG is a synthetic polymer, not an amino acid and has no letter in that alphabet. This creates three recurring representational challenges for PEGylated proteins:
1. PEG cannot be written into the sequence itself. A drafter cannot insert a symbol into the amino acid string to represent a PEG chain the way one might represent, say, a selenocysteine with its recognized code. The sequence entry represents only the polypeptide backbone.
2. The attachment point must be captured through feature annotation, not the sequence string. ST.26 (like ST.25 before it) supports feature tables – structured annotations attached to a sequence that describe modifications at specific residue positions, using a defined feature key (commonly a “modified residue” or “MOD_RES”-type feature, depending on the current schema) along with free-text qualifiers describing the nature of the modification. This is where PEGylation is documented: for example, an annotation noting that the modification at a given residue position is a covalent attachment of a polyethylene glycol moiety, along with descriptive detail (linear or branched PEG, approximate molecular weight, linkage chemistry) placed in the qualifier text or cross-referenced to the specification.
3. Non-specific or heterogeneous PEGylation resists clean positional annotation. Many PEGylation methods don’t produce a single, homogeneous product. Random PEGylation via lysine-reactive chemistry, for instance, can attach PEG at any of several lysine residues (and the N-terminal amine) across a population of molecules, producing a mixture of mono- and poly-PEGylated species with heterogeneous attachment sites. A feature-table annotation tied to one residue position doesn’t naturally describe “PEG attaches at any of positions 12, 45, 67, or the N-terminus, in an undetermined distribution.” Drafters typically handle this by:
- Annotating the most representative or intended primary site if one is known or preferred, while
- Describing the actual heterogeneity, distribution and characterization data (e.g., peptide mapping results showing relative occupancy at each site) in the specification’s written description rather than relying on the sequence listing to carry that burden and
- Where multiple discrete, well-characterized species are separately claimed (e.g., a site-specifically PEGylated variant at position 40 versus a randomly PEGylated mixture), treating them as distinct entries or distinct examples with their own supporting characterization.
Site-Specific vs. Random PEGylation: Drafting Implications
The claiming and disclosure strategy differs meaningfully depending on which PEGylation approach is used and the sequence listing should track that strategy rather than fight it.
Site-specific PEGylation (e.g., via an engineered cysteine, an unnatural amino acid incorporated for click chemistry, N-terminal transamination, or enzymatic conjugation at a defined tag) produces a homogeneous, well-defined product. This is the more favorable case for sequence listing purposes: the modified residue’s position is known and fixed, so a feature annotation at that specific position accurately and completely describes the attachment site. It also tends to support cleaner, more defensible claims, since the claimed molecule is a single defined species rather than a statistical population – a distinction that matters both for enablement/written description support and for later infringement analysis (a competitor’s differently-distributed random PEGylation product may or may not fall within a claim drafted around a specific site).
Random or non-specific PEGylation produces a mixture and applicants should not overstate positional specificity in either the sequence listing or the claims when the underlying chemistry doesn’t support it. Overly specific positional claiming for what is actually a heterogeneous product can create a mismatch between what’s claimed and what’s enabled, inviting a written description or enablement challenge. The more defensible approach is usually to claim the mixture by its defining characteristics (average PEG loading, molecular weight range, chromatographic or mass spectrometric profile, functional properties) while using the specification – not the sequence listing feature table – to carry the detailed characterization data.
Representing PEG Molecular Weight, Branching and Linkage Chemistry
Because the sequence listing format has no native field for polymer chemistry, PEG-specific structural details are conveyed as descriptive text within the feature qualifier and, more fully, in the specification body. Details that are commonly documented include:
- Molecular weight of the PEG moiety (e.g., 20 kDa, 40 kDa), since PEG size materially affects pharmacokinetics and is often a claim element in its own right
- Architecture – linear versus branched PEG and for branched PEG, the branching pattern
- Linkage chemistry – the specific conjugation chemistry (e.g., maleimide-thiol, NHS-ester amine coupling, reductive amination at the N-terminus, click chemistry via an incorporated azide or alkyne handle), since this determines both the attachment site’s chemical feasibility and the stability/reversibility of the linkage
- Linker structure, if the PEG is attached through an intervening cleavable or non-cleavable linker rather than directly
None of this displaces the need for a clear structural description and, typically, structural diagrams in the specification itself – the sequence listing’s feature annotation is a pointer and a formal record, not a substitute for a full written description of the conjugate’s chemistry.
Fusion Proteins, Linkers and PEGylated Multi-Domain Constructs
Many modern biologics combine PEGylation with other engineering – Fc fusions, scFv or nanobody formats, multi-domain constructs with peptide linkers between functional domains. For sequence listing purposes:
- Each distinct polypeptide chain generally requires its own sequence entry, even if the chains are covalently or non-covalently associated in the final product (e.g., the heavy and light chains of an antibody-based construct, or separate polypeptide components of a multi-chain fusion).
- Peptide linkers between domains are part of the amino acid sequence itself and are written directly into the sequence string (they are ordinary amino acids), not treated as a separate modification the way PEG is.
- Where PEGylation is applied to only one domain or chain of a multi-chain construct, the feature annotation for the PEG modification is attached to the specific sequence entry for that chain, not to the construct as a whole.
- If an unnatural amino acid is incorporated for site-specific conjugation (a common approach for precise PEGylation), that residue’s presence in the sequence needs correct handling – many unnatural amino acids don’t have a standard one-letter code and the applicant should confirm the current ST.26-permitted approach for representing them (typically an “Xaa” placeholder in the sequence combined with a feature annotation identifying the specific residue and its modification), rather than inventing a nonstandard symbol.
Common Pitfalls in Practice
Treating the sequence listing as optional or an afterthought. Because PEG dominates the physical and functional identity of a PEGylated biologic, applicants sometimes underinvest in the listing on the theory that “the interesting part is the PEG, not the protein.” But a defective or incomplete sequence listing is a formalities problem that can delay prosecution regardless of how well-characterized the PEGylation chemistry is described elsewhere.
Filing under the wrong standard. Continuing to use ST.25 conventions or templates after an office has transitioned to ST.26 (or vice versa, for older pending applications) creates format-compliance rejections that are avoidable with an early formalities check.
Overclaiming positional specificity for heterogeneous products. As discussed above, annotating a single fixed attachment site for what is actually a randomly PEGylated mixture creates a disclosure/claim mismatch.
Failing to separately characterize distinct PEGylated species when several are disclosed. If an application discloses and claims both a mono-PEGylated and a di-PEGylated variant, or PEGylation at two different alternative sites, each distinct, well-characterized species generally warrants its own clear treatment – supporting data and where sequence differences exist, its own sequence entry – rather than being folded together under a single ambiguous description.
Inconsistency between the sequence listing, the specification’s chemical description and the claims. Examiners routinely cross-check these three sources. A PEG attachment site described at one residue position in the specification’s prose but annotated at a different position in the sequence listing feature table is a red flag that invites objection and can undermine the credibility of the application’s characterization data more broadly.
Ignoring jurisdictional formatting differences. While ST.26 has substantially harmonized sequence listing format across major offices, applicants filing in multiple jurisdictions should confirm there are no office-specific supplementary requirements (for example, certain offices’ expectations around how modified or non-standard residues should be flagged) before relying on a listing prepared for one jurisdiction as sufficient for all.
Practical Recommendations
- Prepare the sequence listing early, alongside the chemistry work, not as a final-stage formalities task – the positional data needed for accurate feature annotation should come directly from the analytical characterization (peptide mapping, LC-MS, Edman degradation, or similar) rather than being reconstructed from memory at filing time.
- Use ST.26-compliant software tools provided or endorsed by the relevant patent office to generate the listing file, rather than hand-building the XML, to minimize format-compliance errors.
- Align terminology across the specification, sequence listing and claims for how the PEGylation site, PEG size and linkage chemistry are described.
- Distinguish clearly, in both the specification and the claim strategy, between site-specific and heterogeneous PEGylation products and support each with characterization data proportionate to the specificity being claimed.
- Confirm current format requirements before filing, since sequence listing standards and office-specific implementation details have changed materially in recent years and continue to be refined.
A Note on Scope
This article provides general, educational information about sequence listing practice for PEGylated biologics and is not legal advice. Sequence listing requirements, permitted feature annotations and office-specific formalities can change and the correct approach for a specific molecule, claim strategy and filing jurisdiction should be confirmed with a patent attorney or agent experienced in biologics prosecution and with the patent office’s current sequence listing guidance.
