Circular RNA, commonly abbreviated as circRNA, has become an important area of biotechnology research and patent activity. As companies and research organizations develop circRNA-based therapeutics, delivery systems, manufacturing methods and other applications, patent applications increasingly contain detailed nucleotide sequence information.
For patent applicants working in this field, describing the biological invention is only part of the challenge. Where nucleotide sequences must be included in a patent sequence listing, those sequences also need to be presented in the required standardized format.
Since July 1, 2022, WIPO Standard ST.26 has governed sequence listings for qualifying nucleotide and amino acid sequence disclosures in PCT applications and has been implemented through corresponding national and regional rules. In the United States, for example, applications with the applicable filing dates and qualifying sequence disclosures generally require a Sequence Listing XML conforming to ST.26 and 37 CFR 1.831–1.834.
For circRNA applications, the circular nature of the molecule introduces an additional issue: how should a circular nucleotide sequence be represented in the sequence listing?
What Is circRNA?
Circular RNA is a type of RNA molecule whose nucleotide chain forms a covalently closed continuous loop rather than having the conventional free 5′ and 3′ ends of a linear RNA molecule.
Patent applications involving circRNA may describe sequences associated with:
- Circular RNA molecules
- Coding sequences
- Regulatory sequences
- Internal ribosome entry sites
- Translation elements
- Circularization elements
- Intronic sequences
- Ribozymes
- Back-splice junctions
- RNA expression constructs
- RNA manufacturing processes
- Delivery systems
- Therapeutic applications
A single application may therefore contain numerous nucleotide sequences that need to be evaluated for sequence-listing purposes.
When Is a Sequence Listing Required?
ST.26 establishes requirements for nucleotide and amino acid sequence disclosures that meet its criteria. In the United States, 37 CFR 1.831 requires a Sequence Listing XML when qualifying nucleotide or amino acid sequences are disclosed by enumeration of their residues.
The important point is that applicants should not assume that only the principal circRNA sequence matters.
Depending on what is disclosed and how it is presented, other sequences in the application may also need to be considered.
Potential examples include:
- Full-length circRNA sequences
- Coding regions
- Primer sequences
- Polynucleotide variants
- Mutant sequences
- Control sequences
- Specific sequence fragments
- Sequences used in experimental examples
Each disclosed sequence should be reviewed against the applicable ST.26 requirements rather than being included or excluded based solely on its role in the invention.
The Special Issue With Circular Nucleotide Sequences
ST.26 specifically addresses circular nucleotide sequences.
Unlike a linear sequence, a circular RNA does not have an inherent beginning or end. However, a sequence listing requires the residues to be represented in a defined order and assigned positions.
ST.26 therefore requires the applicant to choose which nucleotide will be designated as residue position 1 for a circular nucleotide sequence. The numbering then proceeds continuously in the 5′ to 3′ direction, with the final residue corresponding to the total number of nucleotides.
This means that the sequence listing does not reproduce the circular molecule as a literal loop.
Instead, the circular sequence is represented as a linear sequence for sequence-listing purposes, together with information identifying its circular configuration.
How Is a circRNA Sequence Represented?
A simplified example illustrates the concept.
Imagine a circRNA containing 500 nucleotides.
Because the molecule is circular, any nucleotide could theoretically be selected as the starting position. Once the applicant selects a nucleotide as position 1, the sequence is written continuously until position 500.
The sequence listing therefore presents:
Position 1 → Position 2 → Position 3 → … → Position 500
The listing then provides the appropriate feature information to indicate that the final nucleotide connects back to the first nucleotide.
WIPO’s ST.26 guidance gives a specific example for a circular nucleotide sequence. It explains that the applicant chooses residue position 1 and describes the circular connection using the misc_feature feature key with a location such as 212^1, together with a note indicating that the molecule is circular.
The exact implementation should be prepared according to the current ST.26 requirements and the relevant filing-office rules.
Why Choosing the Correct Starting Point Matters
Because a circular molecule has no natural first nucleotide, selecting position 1 is fundamentally a representation decision.
Changing the selected starting nucleotide can change the way the sequence appears in the listing even though the underlying circular molecule is the same.
For this reason, the sequence listing should be prepared consistently with the way the invention is described in the specification and figures.
A useful internal practice is to establish a clear reference point for each circular sequence before preparing the XML file.
This can help prevent confusion when comparing:
- The sequence listing
- Sequence identifiers in the specification
- Figures
- Experimental data
- Sequence-analysis results
- Claims referring to particular regions
The Circular Junction Requires Particular Attention
One of the defining characteristics of circRNA is the junction where the end of the linearized representation connects back to its beginning.
That junction may be biologically important, particularly when the application discusses:
- Back-splicing
- Circularization
- Translation across the junction
- Junction-specific detection
- Therapeutic activity
- Sequence modifications
- Circularization efficiency
From a sequence-listing perspective, however, the junction also needs to be represented according to the applicable ST.26 feature conventions.
This is an area where simply copying a FASTA sequence into an XML file may not be sufficient.
The biological sequence and the standardized patent representation are related, but they are not necessarily identical in presentation.
ST.26 Uses a Standardized XML Format
For applications subject to ST.26, the sequence listing is not simply a text document containing sequences.
The listing is an XML file containing standardized sequence and associated information.
The USPTO explains that a Sequence Listing XML must be presented as a single XML 1.0 file encoded using Unicode UTF-8 and must conform to the ST.26 Document Type Definition (DTD).
The file contains both general application information and sequence data.
The sequence data can include information such as:
- Sequence identifiers
- Nucleotide or amino acid sequences
- Feature information
- Qualifiers
- Organism information where applicable
- Other required sequence-associated data
Using a structured format makes the biological information more consistent and searchable across patent systems.
WIPO Sequence and Validation
WIPO provides WIPO Sequence, a software tool designed to help applicants create and validate ST.26-compliant sequence listings. WIPO also provides the WIPO Sequence Validator for checking sequence-listing compliance.
WIPO currently lists WIPO Sequence as a tool for preparing nucleotide and amino acid sequence listings for national and international applications and its current software versions are updated over time.
Applicants do not necessarily have to create a listing using WIPO Sequence. The USPTO states that other software may be used, provided the resulting XML complies with the applicable ST.26 requirements. However, WIPO Sequence is highly recommended by the USPTO.
Common Compliance Challenges in circRNA Applications
CircRNA applications can contain complicated sequence disclosures, making manual preparation particularly error-prone.
1. Treating a Circular Sequence Like an Ordinary Linear Sequence
A circRNA should not simply be treated as a conventional linear molecule without considering its circular configuration.
The sequence representation needs to account for the connection between the final and first residues.
2. Failing to Establish a Consistent Position 1
Because a circular sequence has no inherent starting point, the applicant needs to select one.
That choice should be documented and applied consistently.
3. Omitting Relevant Feature Information
The nucleotide sequence alone may not communicate all information required by ST.26.
Feature keys and qualifiers may be necessary to describe particular structural or functional aspects of the sequence.
4. Confusing the CircRNA With Its DNA Template
CircRNA inventions frequently involve DNA templates, plasmids, expression constructs, or other nucleic acid intermediates.
The RNA sequence and the corresponding DNA sequence should not automatically be treated as the same sequence for sequence-listing purposes.
Each disclosed sequence needs to be evaluated under the applicable rules.
5. Missing Short or Partial Sequences
Patent applications may contain many sequence fragments in examples, figures, primers, probes, or experimental descriptions.
The sequence-listing requirements should be reviewed systematically rather than focusing only on the primary circRNA sequence.
6. Incorrect XML Structure
Even if the underlying biological sequences are correct, errors in the XML structure, encoding, required fields, feature information, or DTD compliance can create filing problems.
Sequence Listing Compliance Is More Than an XML Check
A technically valid XML file is important, but compliance begins before the XML is generated.
The sequence information should first be reviewed against the patent application itself.
For example, a sequence-listing review may ask:
- What nucleotide and amino acid sequences are disclosed?
- Which sequences meet the applicable inclusion criteria?
- Which sequences are circular?
- What nucleotide is being designated as position 1?
- What features need to be annotated?
- Are modified nucleotides involved?
- Are sequences represented consistently throughout the application?
- Does every required sequence have its own sequence identifier?
- Does the final XML pass the applicable validation checks?
This broader review helps reduce the risk of treating sequence-listing preparation as a purely technical file-conversion exercise.
Reference Numbers and Sequence Identifiers
Each sequence included in an ST.26 listing receives a sequence identification number.
ST.26 requires sequence identifiers to begin with 1 and increase consecutively. Each qualifying sequence is assigned a separate identifier, including a sequence that is identical to a region of a longer sequence.
For a complex circRNA application, careful sequence-ID management is therefore important.
The patent specification may refer to sequences as:
- SEQ ID NO: 1
- SEQ ID NO: 2
- SEQ ID NO: 3
- And so forth.
These identifiers should correspond correctly to the sequences presented in the XML listing.
A sequence-ID mismatch can create unnecessary confusion during prosecution and later review.
Modified Nucleotides in circRNA
Some circRNA technologies involve modified nucleotides or other engineered sequence features.
ST.26 specifies how nucleotide symbols are to be used and provides additional requirements for certain modifications. For example, the standard explains that the symbol t is interpreted as thymine in DNA and uracil in RNA, while uracil in DNA or thymine in RNA is treated as a modified nucleotide requiring further description in the feature table.
Accordingly, applicants should not assume that a sequence can be represented correctly simply by copying the characters from an experimental sequence file.
The biological identity and modification status of the sequence should be reviewed before finalizing the listing.
International Patent Applications and ST.26
ST.26 is particularly important for applicants pursuing international protection.
WIPO states that the standard establishes uniform requirements for nucleotide and amino acid sequence disclosures in patent applications at the international, regional and national levels.
For organizations developing circRNA technologies internationally, preparing sequence data in an ST.26-compliant manner from the beginning can reduce duplication and inconsistencies when preparing corresponding applications.
However, applicants should still confirm the specific filing requirements and procedures of each relevant patent office.
U.S. Filing Considerations
In the United States, the USPTO’s sequence rules under 37 CFR 1.831–1.835 apply to qualifying applications with the applicable filing dates.
The USPTO states that a required Sequence Listing XML can generally be submitted electronically through Patent Center, subject to the applicable file-size restrictions. The USPTO currently identifies a 100 MB limit for electronic submission; larger XML files are subject to alternative submission procedures described by the Office.
For a U.S. application, applicants should also pay attention to the requirements concerning the sequence-listing incorporation-by-reference statement and related filing procedures.
Because filing requirements can change, the current USPTO rules and guidance should be checked for the application being prepared.
A Practical circRNA Sequence Listing Checklist
Before filing a patent application containing circRNA sequences, consider the following checklist:
Sequence identification
- Identify every disclosed nucleotide and amino acid sequence.
- Determine which sequences meet the applicable sequence-listing criteria.
- Assign sequence identifiers sequentially.
- Confirm that sequence identifiers correspond to the specification.
Circular sequence representation
- Identify each circular nucleotide sequence.
- Select residue position 1.
- Number residues continuously.
- Represent the circular connection using the appropriate ST.26 feature information.
- Confirm that the circular configuration is properly described.
Sequence features
- Review coding regions.
- Review regulatory elements.
- Review junctions and other relevant features.
- Check modified nucleotides.
- Confirm appropriate feature keys and qualifiers.
XML compliance
- Generate the listing in the required XML format.
- Confirm UTF-8 encoding.
- Confirm DTD compliance.
- Validate the XML.
- Review warnings and errors.
- Confirm that the final XML corresponds to the application as filed.
Why Professional Patent Sequence Listing Services Can Help
Preparing a sequence listing for a conventional biotechnology application can already involve significant technical detail. CircRNA applications can add another layer of complexity because circular configuration, junctions, engineered sequences and associated DNA constructs may all need careful review.
A professional patent sequence listing service can assist with tasks such as:
- Identifying sequences that may require listing
- Converting sequence data into the appropriate format
- Preparing ST.26 XML listings
- Assigning and checking sequence identifiers
- Annotating sequence features
- Handling circular nucleotide representations
- Reviewing sequence consistency
- Validating XML files
- Supporting international patent filing workflows
The objective is not merely to generate an XML file. It is to create a sequence listing that accurately represents the application’s disclosed sequence information and satisfies the applicable formatting requirements.
Final Thoughts
CircRNA patent applications present unique sequence-listing considerations because the underlying molecule is circular while the patent sequence listing uses a defined linear representation.
Under WIPO ST.26, an applicant selects a nucleotide as residue position 1, numbers the sequence continuously and uses the appropriate feature information to identify the circular configuration.
For biotechnology applicants, getting these details right is important. A sequence listing is part of the patent disclosure, not simply an administrative attachment. The USPTO specifically explains that the content of a required Sequence Listing XML forms part of the disclosure of the invention under U.S. sequence-listing rules.
As circRNA research and commercialization continue to expand, accurate sequence representation will remain an important part of patent preparation.
Whether an application contains one circular RNA sequence or a large collection of engineered nucleic acids, professional patent sequence listing services can help applicants organize, prepare and validate complex sequence information before filing.
