The Two Frontiers Project
Data Sharing
Environmental genomics data is a public good. We publish ours openly, and we built infrastructure to help others do the same.
7,000+
Cryopreserved samples
12+
Global expeditions
4
Continents
6
Data types
Our open data commitment
All research datasets generated by The Two Frontiers Project are published openly unless there is a specific reason not to — partner confidentiality, ethical constraints on location data, or active patent applications. The default is open.
We publish to NCBI and Zenodo, and we're happy to work directly with anyone who needs richer metadata or wants to share their own datasets under a common standard.
We believe the environmental genomics field moves faster when data doesn't sit in silos. If you have data you want to make useful, we can help with that too.
Raw Sequences
Unprocessed sequencing files (FASTQ) with full metadata, deposited in NCBI.
Processed Datasets
Assembled metagenomes, taxonomic profiles, functional annotations, and abundance tables — analysis-ready.
Environmental Metadata
GPS coordinates, collection dates, environmental parameters, sample handling notes, and protocol versions.
Analysis Pipelines
Version-controlled bioinformatics pipelines with full documentation, containerized for reproducibility.
Culture Collections
Isolated and characterized environmental strains, with genomes, phenotypic data, and availability status.
Longitudinal Records
Time-series datasets from reef, soil, and urban monitoring programs spanning multiple years.