Buckets:
MuSProt Dataset Documentation
MuSProt (Multistate Protein Database) is a million-scale multimodal database for multistate proteins, designed to support programmable protein design and AI model development. It links experimentally observed conformational states of identical protein sequences from the PDB, organizes them into state clusters and transition relationships, and enriches each record with structural similarity, experimental context, state-specific function rankings and transition fidelity labels.
At a glance
| Count | |
|---|---|
Conformational observations (node rows, one per PDB chain) |
863,643 |
| Sequences (identical SEQRES observed ≥ 2 times) | 102,790 |
| Sequences with ≥ 2 conformational states | 24,673 |
| State clusters | 162,601 |
Transition triplets (edge rows, ordered observation pairs) |
113,581,394 |
| Transitions between distinct states | 25,530,666 |
Concepts
Observation
One protein chain (pdb_id, auth_asym_id) in one PDB entry. Observations whose full deposited sequence (SEQRES) is 100% identical share a sequence_id. Only sequences observed at least twice are included.
State cluster
Within each sequence, two observations are linked when their pair similarity is ≥ 0.95. The connected components of that graph are the conformational states, numbered 0, 1, 2, … per sequence (state_id).
Transition
Every ordered pair (A → B) of observations of the same sequence is an edge row. Both directions are stored; they share all structural fields and differ only in the sign of the energy differences (B − A).
Structural similarity
Pairs are superposed with US-align on a fixed residue correspondence (identical SEQRES numbering), and similarity is computed over the aligned core so that disordered or unmodelled termini do not penalise otherwise identical conformations:
tm_aln = (TM_A · L_A + TM_B · L_B) / (2 · aligned_length)
pair_similarity = 0.8 · tm_aln + 0.2 · exp(−(RMSD / 5 Å)²)
coverage = aligned_length / max(L_A, L_B)
coverage is reported separately as the completeness axis.
Fidelity labels
pair_fidelity (from pair_similarity) and state_fidelity (from state_similarity) bin how far a transition moves:
| Label | Similarity | Meaning |
|---|---|---|
identical |
≥ 0.95 | Same conformation |
low |
0.70 – 0.95 | Subtle rearrangement |
medium |
0.50 – 0.70 | Clear conformational change |
high |
≤ 0.50 | Large-scale transition |
| empty | — | Not computable (the two chains share no resolved residues) |
Download
The full dataset is a single SQLite database file (MuSProt.db, ~37 GB) with two tables, node and edge. Download it with the Download button at the top of the site.
Atomic coordinates are not included: fetch mmCIF files from the RCSB PDB or PDBe using pdb_id and auth_asym_id.
Database Schema
Table: node
One row per observation (PDB chain).
| Column | Type | Description |
|---|---|---|
sequence_id |
TEXT | Sequence identifier (e.g. SEQ049914), shared by all chains with an identical SEQRES |
uniprot_id |
TEXT | UniProt accession(s) from SIFTS; ;-joined for chimeric chains; empty if none |
pdb_id |
TEXT | 4-character PDB entry ID (lowercase) |
auth_asym_id |
TEXT | Author chain identifier (case-sensitive: d and D are different chains) |
sequence |
TEXT | Full deposited (canonical) sequence, including unmodelled residues |
sequence_length |
INT | Length of sequence |
modeled_sequence |
TEXT | Residues present in the coordinates, concatenated (gaps closed up) |
modeled_sequence_length |
INT | Length of modeled_sequence |
resolved_label_seq |
TEXT | Resolved SEQRES positions as ranges (label_seq_id, 1-based), e.g. 5-173,187-353 |
n_resolved_aa |
INT | Number of resolved standard amino acids |
resolved_coverage |
FLOAT | n_resolved_aa / sequence_length |
binders |
TEXT | Ligands, metals and cofactors on this chain (component IDs, ;-joined; water and common buffers excluded) |
binding_status |
TEXT | holo if the chain has any binder, otherwise apo |
experimental_method |
TEXT | e.g. X-RAY DIFFRACTION, ELECTRON MICROSCOPY, SOLUTION NMR |
resolution |
FLOAT | Resolution in Å; empty when not applicable (e.g. NMR) |
pH |
FLOAT | pH of the crystallisation / sample condition |
temp_K |
FLOAT | Temperature in Kelvin |
experimental_details |
TEXT | Free-text crystallisation / sample-preparation details |
chain_composition |
TEXT | Entry composition: monomeric, homomeric or heteromeric |
non_protein_polymer_binding |
TEXT | Bound nucleic acid: DNA, RNA, DNA/RNA, Other polymers or None |
initial_release_date |
TEXT | Initial PDB release date (YYYY-MM-DD) |
cath_id |
TEXT | CATH domain IDs for the chain (e.g. 3vgkA01;3vgkA02); empty if unassigned |
cath_superfamily |
TEXT | CATH superfamily code(s) (e.g. 3.30.420.40) |
pfam_id |
TEXT | Pfam family accession(s), ;-joined (e.g. PF00069); empty if none |
ecod_id |
TEXT | ECOD domain ID(s) for the chain, ;-joined (e.g. e3vgkA1;e3vgkA2) |
ecod_fid |
TEXT | ECOD family ID(s) x.h.t.f (e.g. 2484.1.1.8); the first two levels (x.h) are the homology group |
state_id |
TEXT | Conformational state of this observation within its sequence (0, 1, …) |
Rosetta |
FLOAT | Rosetta total score |
FoldX |
FLOAT | FoldX total energy |
EvoEF2 |
FLOAT | EvoEF2 total energy |
RW |
FLOAT | RW statistical potential |
RW+ |
FLOAT | RW+ statistical potential |
n_seqres |
INT | SEQRES length used for energy scoring |
n_resolved |
INT | Resolved residues used for scoring |
n_scored |
INT | Residues in the consensus set scored by Rosetta / EvoEF2 / RW / RW+ |
scored_coverage |
FLOAT | n_scored / n_seqres |
foldx_n_scored |
INT | Residues scored by FoldX |
foldx_coverage |
FLOAT | foldx_n_scored / n_seqres |
ranked_functions |
TEXT | JSON list of functional annotations, ranked for this observation |
matching_scores |
TEXT | JSON list of matching scores, index-aligned with ranked_functions |
Energies are computed on a consensus residue set shared by all observations of a sequence, so they are comparable within a sequence. Lower is more favourable.
Table: edge
One row per ordered pair of observations (A → B) of the same sequence.
| Column | Type | Description |
|---|---|---|
pdb_id_A |
TEXT | PDB ID of observation A |
auth_asym_id_A |
TEXT | Chain of observation A |
pdb_id_B |
TEXT | PDB ID of observation B |
auth_asym_id_B |
TEXT | Chain of observation B |
tm_aln |
FLOAT | TM-score normalised by the aligned length (conformation only) |
RMSD |
FLOAT | Cα RMSD over the aligned residues (Å) |
coverage |
FLOAT | aligned_length / max(L_A, L_B) |
pair_similarity |
FLOAT | 0.8 · tm_aln + 0.2 · exp(−(RMSD/5)²) |
delta_Rosetta |
FLOAT | Rosetta(B) − Rosetta(A) |
delta_FoldX |
FLOAT | FoldX(B) − FoldX(A) |
delta_EvoEF2 |
FLOAT | EvoEF2(B) − EvoEF2(A) |
delta_RW |
FLOAT | RW(B) − RW(A) |
delta_RW+ |
FLOAT | RW+(B) − RW+(A) |
state_id_A |
TEXT | State of A |
state_id_B |
TEXT | State of B |
state_similarity |
FLOAT | Mean pair_similarity over all observation pairs between the two states (within-state when state_id_A = state_id_B) |
state_fidelity |
TEXT | Fidelity label of state_similarity |
pair_fidelity |
TEXT | Fidelity label of pair_similarity |
n_shared_resolved |
INT | Residues resolved in both A and B |
delta_order_valid |
TEXT | True when resolved-residue changes are comparable (same method, not both NMR, overlapping models) |
order_invalid_reason |
TEXT | valid, cross_method, nmr, no_shared_residues or missing_data |
n_ordered |
INT | Residues resolved in B but not in A (ordered on going A → B) |
n_disordered |
INT | Residues resolved in A but not in B |
ordered_ranges |
TEXT | SEQRES positions of the ordered residues, e.g. 115-123,187-189 |
disordered_ranges |
TEXT | SEQRES positions of the disordered residues |
longest_ordered_segment |
INT | Longest contiguous ordered stretch |
longest_disordered_segment |
INT | Longest contiguous disordered stretch |
n_terminal_ordered |
INT | Ordered residues in segments touching either terminus |
n_internal_ordered |
INT | Ordered residues in internal segments |
order_evidence |
TEXT | Evidence grade for the ordering: high, medium, low or not_applicable (see below) |
ordering_with_ligand_change |
TEXT | True when the ordering co-occurs with a change in bound ligands |
sequence_id |
TEXT | Sequence shared by A and B |
When A and B share no resolved residues, tm_aln, pair_similarity and pair_fidelity are empty and RMSD / coverage are stored as 0.0.
Order / disorder labels
These columns describe the observed resolved state (residues with or without modelled coordinates), not intrinsic disorder: a residue can lack density because it is flexible, or because of crystal contacts, occupancy or modelling choices. They are only computed when both structures use the same experimental method and at least one is not NMR (NMR models build every residue); otherwise delta_order_valid is False and the counts are empty.
order_evidence grades an ordering event from its length, position and the resolution of both structures:
| Grade | Rule |
|---|---|
high |
an internal ordered segment of ≥ 8 residues, both structures better than 2.5 Å |
medium |
otherwise, if not low |
low |
longest ordered segment ≤ 2 residues, purely terminal, or either structure worse than 3.5 Å |
not_applicable |
nothing was ordered |
Knowledge graph
MuSProt is organised as a hierarchical graph. In the outer layer, each sequence is a node, and two sequences are linked when their chains carry the same set of ECOD homology groups (x.h of ecod_fid; 88.1% of sequences have an assignment). In the inner layer, every observation of a sequence is connected to every other by an edge row; contracting each state cluster to one node gives the quotient graph over states, whose edges carry state_similarity.
Indexes
Both tables are indexed for per-chain lookup: node(pdb_id COLLATE NOCASE, auth_asym_id) and edge(pdb_id_A COLLATE NOCASE, auth_asym_id_A). Use the same collation in queries to hit the index:
SELECT * FROM edge WHERE pdb_id_A = '3vgk' COLLATE NOCASE AND auth_asym_id_A = 'A';
Usage Examples
Python (sqlite3 + pandas)
import sqlite3
import pandas as pd
conn = sqlite3.connect("MuSProt.db")
# One observation
node = pd.read_sql(
"SELECT * FROM node WHERE pdb_id = ? COLLATE NOCASE AND auth_asym_id = ?",
conn, params=("3vgk", "A"),
)
# Every transition from that observation
edges = pd.read_sql(
"SELECT * FROM edge WHERE pdb_id_A = ? COLLATE NOCASE AND auth_asym_id_A = ?",
conn, params=("3vgk", "A"),
)
Transitions between distinct states with a large change
df = pd.read_sql("""
SELECT pdb_id_A, auth_asym_id_A, pdb_id_B, auth_asym_id_B,
CAST(pair_similarity AS REAL) AS similarity,
CAST(RMSD AS REAL) AS rmsd,
CAST(delta_Rosetta AS REAL) AS d_rosetta
FROM edge
WHERE pdb_id_A = ? COLLATE NOCASE AND auth_asym_id_A = ?
AND state_id_A <> state_id_B
AND pair_fidelity IN ('medium', 'high')
ORDER BY similarity
""", conn, params=("1anf", "A"))
Apo / holo pairs for one UniProt accession
df = pd.read_sql("""
SELECT e.pdb_id_A, e.auth_asym_id_A, e.pdb_id_B, e.auth_asym_id_B,
e.pair_similarity, e.pair_fidelity, nB.binders
FROM node nA
JOIN edge e ON e.pdb_id_A = nA.pdb_id COLLATE NOCASE AND e.auth_asym_id_A = nA.auth_asym_id
JOIN node nB ON nB.pdb_id = e.pdb_id_B COLLATE NOCASE AND nB.auth_asym_id = e.auth_asym_id_B
WHERE nA.uniprot_id = 'P0AEX9'
AND nA.binding_status = 'apo' AND nB.binding_status = 'holo'
""", conn)
Scanning node by uniprot_id or sequence_id reads the whole table; for repeated queries create an index first (CREATE INDEX idx_node_seq ON node(sequence_id)).
Notes
- All columns are stored as
TEXT; cast numeric fields withCAST(col AS REAL)/CAST(col AS INTEGER). ranked_functionsandmatching_scoresare JSON strings; parse withjson.loads().edgeis directional: A → B and B → A are separate rows with identical structural fields and opposite energy deltas.
Xet Storage Details
- Size:
- 12.8 kB
- Xet hash:
- 1e3c128d22d6959b0a0dd215e4d9e0c807e744aadd214eace51cd2087a2d4fe3
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.