The WIPO Sequence software is an official desktop application developed by the World Intellectual Property Organization to help applicants prepare patent sequence listings that comply with WIPO Standard ST.26. It is widely used in biotechnology and life sciences patent drafting where nucleotide and amino acid sequences must be submitted in a structured XML format.
Unlike general document tools, this software is designed specifically to enforce strict biological and legal formatting rules. Once understood, its workflow becomes predictable and highly structured.
What the Software Actually Does
At its core, WIPO Sequence software converts raw biological data into a standardized, machine-readable patent format.
| Function | Purpose |
| Sequence management | Stores DNA, RNA, and amino acid sequences |
| Annotation system | Marks functional regions like genes or CDS |
| Validation engine | Detects ST.26 compliance errors |
| XML generator | Produces final patent-ready sequence listing |
This ensures that patent offices across different countries receive uniform and legally compliant data.
Before You Start Using the Software
Preparation is important because most errors happen before data entry begins.
You should have:
- Clean sequence data (FASTA or raw format)
- Correct molecule classification (DNA, RNA, or protein)
- Verified organism names in scientific format
- Clear understanding of sequence boundaries
- Patent application reference details
A small preparation mistake can lead to validation failures later, so this step matters more than most users expect.
Creating a New Project
When you open the software, everything begins with a project workspace. Each project represents a single patent application.
You create a project by entering a project name and optionally a description. Once created, the system organizes all sequences, annotations, and metadata inside that container.
Key idea:
A single project = one patent sequence listing
Importing or Adding Sequence Data
You can either import existing files or enter sequences manually.
| Method | Best Use Case |
| FASTA import | Large datasets or lab outputs |
| ST.26 XML import | Editing existing compliant files |
| Manual entry | Small or custom sequences |
During import, the software automatically parses sequences and assigns internal identifiers.
Creating a Sequence Manually
Each sequence must be defined clearly before it can be used.
You typically provide:
- Sequence name (auto-generated or custom)
- Molecule type selection
- Raw biological sequence input
Supported molecule types include:
| Type | Example |
| DNA | ATGCGTACGT |
| RNA | AUGCGUACGU |
| Amino acid | MKTFFVLLL |
The software automatically flags invalid characters to prevent formatting issues.
Adding Organism Information
Every sequence must be linked to a biological source. This is not optional under ST.26 rules.
Common examples include:
- Homo sapiens
- Escherichia coli
- Arabidopsis thaliana
Without this information, validation will fail.
Adding Features and Functional Annotations
Features explain what different parts of the sequence do biologically.
Common feature types:
- Gene regions
- Coding sequences (CDS)
- Promoters
- Binding sites
- Signal peptides
Feature definition includes:
| Parameter | Description |
| Start position | Beginning of feature in sequence |
| End position | Ending position |
| Feature type | Biological function |
| Qualifiers | Additional metadata |
These annotations are essential because they link raw sequence data to functional biological meaning.
Understanding Sequence Validation
Validation is one of the most important steps in the entire process.
When you run validation, the software checks for:
- Missing organism names
- Invalid nucleotide or amino acid characters
- Incorrect feature boundaries
- Structural inconsistencies
- ST.26 rule violations
Validation outcome types:
| Status | Meaning |
| Error | Must be fixed before export |
| Warning | Recommended correction |
| Pass | Ready for export |
A project cannot be legally submitted unless it passes validation.
Editing and Organizing Sequences
Once sequences are added, they can be rearranged or modified at any time. This includes:
- Changing sequence order
- Editing sequence content
- Updating organism details
- Adjusting feature annotations
This flexibility is useful when refining patent drafts or making attorney-requested changes.
Generating the Final XML Sequence Listing
After successful validation, the software generates a final ST.26 XML file.
This file:
- Contains all sequences in structured format
- Includes metadata and annotations
- Follows strict WIPO schema rules
- Is accepted by international patent offices
Final output workflow:
| Step | Action |
| Validate | Ensure compliance |
| Review | Check final structure |
| Export | Generate XML file |
| Submit | Include in patent filing |
Common Mistakes First-Time Users Make
Many errors occur during early use. The most common include:
- Using incorrect nucleotide codes
- Forgetting organism names
- Incorrect feature coordinates
- Uploading improperly formatted FASTA files
- Skipping validation before export
Even small mistakes can cause full rejection of the sequence listing.
Practical Workflow Summary
The complete process can be visualized as a simple pipeline:
| Stage | Description |
| Project creation | Initialize patent workspace |
| Sequence input | Add or import biological data |
| Annotation | Define functional regions |
| Validation | Check compliance |
| Export | Generate ST.26 XML |
This structure ensures consistency across all patent applications.
Final Thoughts
WIPO Sequence software is highly structured by design. While it may seem complex at first, it actually follows a clear logic that mirrors how patent offices evaluate biological data.
Once users understand the relationship between sequences, features, validation rules, and XML output, the entire system becomes efficient and reliable for international patent filing.
If you want, I can also create a visual workflow diagram, a ST.26 cheat sheet table, or a real sample sequence listing project for practice.
