Overview
AlphaFold 3 requires multiple genetic and structural databases for generating multiple sequence alignments (MSAs) and structural templates. These databases enable the model to leverage evolutionary information.Required Databases
AlphaFold 3 uses the following databases:Protein Databases
BFD Small
Modified BFD (Big Fantastic Database)Clustered protein sequences for fast MSA generationVersion: 2022-09-28
MGnify
Metagenomic sequencesProtein sequences from metagenomics studiesVersion: 2022_05
UniProt
Universal Protein ResourceComprehensive protein sequence databaseVersion: 2021_04
UniRef90
UniProt Reference Clusters90% identity clustered UniProt sequencesVersion: 2022_05
RNA Databases
NT-RNA
Nucleotide RNAClustered RNA sequences from NCBIVersion: 2023_02_23
RFam
RNA familiesRNA sequence families databaseVersion: 14_9
RNACentral
RNA sequence databaseComprehensive RNA sequence collectionVersion: 21_0
Structural Databases
PDB mmCIF
Protein Data Bank structures~200,000 structures in mmCIF formatVersion: 2022-09-28
PDB Seqres
PDB sequencesSequence database for template searchVersion: 2022-09-28
Quick Installation
Automated Download Script
AlphaFold 3 provides a download script that fetches all required databases:path
default:"$HOME/public_databases"
Target directory for databases. Must NOT be inside AlphaFold 3 repository.
Prerequisites
Running in Screen/Tmux
Download takes ~45 minutes on fast connections. Use
screen or tmux for long-running processes.Manual Installation
If you prefer manual download or need specific versions:Protein Databases
- BFD Small
- MGnify
- UniProt
- UniRef90
RNA Databases
- NT-RNA
- RFam
- RNACentral
Structural Databases
- PDB mmCIF
- PDB Seqres
Directory Structure
After installation, your database directory should look like:Storage Optimization
Using SSD for Performance
Copying to SSD
Using RAM Disk (Maximum Performance)
RAM disk contents are lost on reboot. Only for temporary high-performance scenarios.
Partial SSD Setup
Use SSD for frequently accessed databases, HDD for others:Database Sharding
For high-throughput environments with many CPU cores:Why Shard?
Sharding enables parallel genetic search across many CPU cores, dramatically reducing wall-clock time.
- Utilize 32+ core systems effectively
- Reduce genetic search time by 10-50×
- Maximize disk I/O parallelization
Sharding Process
1
Install seqkit
2
Shuffle Sequences
3
Split into Shards
uniref90_shuffled.fa.split/uniref90_shuffled.part_001.fa, etc.4
Rename with Padding
5
Count Sequences/Bases
Using Sharded Databases
shard_count
Specifies 128 shards with pattern
uniref90.fasta-XXXXX-of-00128integer
Total sequence count across all shards (for e-value scaling)
integer
Maximum shards to process in parallel
Recommended Shard Counts
For consistent performance, aim for equal shard sizes (~0.5-2 GB per shard).
Permissions and Access
Setting Permissions
Docker Mounts
Use
:ro (read-only) suffix for safety.Singularity Binds
Verifying Installation
Check Files Exist
Test with AlphaFold
Database Updates
AlphaFold 3 uses specific database versions from the paper. Newer versions may work but are not officially supported. If you must update:- Download new version to separate directory
- Test with known inputs
- Compare results to original databases
- Update
--db_dirflags
Troubleshooting
Download Interrupted
Corrupted Files
Insufficient Space
Permission Errors
MSA Tools Can’t Find Databases
Database Licenses
All databases are available under permissive licenses:- BFD: CC BY 4.0
- MGnify: CC0 1.0
- PDB: CC0 1.0
- UniProt/UniRef: CC BY 4.0
- NT-RNA: Modified (see paper)
- RFam: CC0 1.0
- RNACentral: CC0 1.0
Next Steps
Performance
Optimize database access and search speed
Running Docker
Use databases with AlphaFold 3