The Two Frontiers Project

The Two Frontiers Project

Data Sharing

Environmental genomics data is a public good. We publish ours openly, and we built infrastructure to help others do the same.

7,000+

Cryopreserved samples

12+

Global expeditions

4

Continents

6

Data types

Our open data commitment

All research datasets generated by The Two Frontiers Project are published openly unless there is a specific reason not to — partner confidentiality, ethical constraints on location data, or active patent applications. The default is open.

We publish to NCBI and Zenodo, and we're happy to work directly with anyone who needs richer metadata or wants to share their own datasets under a common standard.

We believe the environmental genomics field moves faster when data doesn't sit in silos. If you have data you want to make useful, we can help with that too.

Raw Sequences

Unprocessed sequencing files (FASTQ) with full metadata, deposited in NCBI.

Processed Datasets

Assembled metagenomes, taxonomic profiles, functional annotations, and abundance tables — analysis-ready.

Environmental Metadata

GPS coordinates, collection dates, environmental parameters, sample handling notes, and protocol versions.

Analysis Pipelines

Version-controlled bioinformatics pipelines with full documentation, containerized for reproducibility.

Culture Collections

Isolated and characterized environmental strains, with genomes, phenotypic data, and availability status.

Longitudinal Records

Time-series datasets from reef, soil, and urban monitoring programs spanning multiple years.