The data

Large-scale RNA-seq needs a common language.

Public sequencing archives contain valuable biological signal. Reomics is working to connect that signal to consistent descriptions of samples, subjects, studies, and experimental context.

Inputs

Many records. One coherent view.

The proposed curation layer brings together information that currently lives across multiple archives and publications.

SRA RunInfoBioSample attributesGEO sample recordsrecount3 metadataLinked publications
The challenge

Heterogeneous descriptions

Study authors use different field names, free text, units, and levels of detail. A sample may have useful context spread across several records.

The response

Auditable harmonization

Parsing, extraction, ontology mapping, consistency checks, and sampled human review create tables that can be searched and corrected.

Scientific lineage

Experience with data at scale.

Published recount3 results show the scale of the processing foundation on which the team can draw.

recount3
316K+

human SRA runs in recount3

Published recount3 scope

recount3
763K+

human and mouse runs processed

Published recount3 scope

recount3
990 TB

compressed reads processed

Published recount3 scope

Historical recount3 figures, not Reomics platform totals. Read the 2021 Genome Biology paper.

The number of distinct human participants represented in public RNA-seq cannot be inferred reliably from run counts alone. Resolving that uncertainty is part of the metadata work.
Work with Reomics

Find the right data for your question.

We welcome conversations with biotechnology and pharmaceutical teams, researchers, and prospective collaborators.

Start a conversation