Experimental data is shared in a variety of formats. Choosing the right format extends the reusability and impact of your research, since others need to be able to open and interpret your files.
This page explains the difference between raw, processed, and results data, and provides a reference table of accepted formats for common data types.
Raw Data, Processed Data, and Results
From a reusability standpoint, data (raw or processed) is generally more valuable to future users than results, since data can support new analyses that a results file cannot.
-
Data refers to raw or partially processed information from a single sample.
-
Results are generally post-analysis information from an aggregate of samples, or figures intended for a manuscript.
For example, for RNA-seq data: raw data would be the FASTQ files, processed data would be aligned reads (.bam) or gene counts, and a differential expression analysis or volcano plot would be considered results.
This distinction is well-defined for many data types, but less so for certain assays — for example, "results" may be an acceptable format for assays that don't lend themselves to re-analysis, such as western blots. If you're unsure how your data fits, reach out to us through your Service Desk ticket.
What Data Should Be Shared?
We generally ask that raw or minimally processed data be shared whenever possible, since this is what allows secondary users to ask their own questions of the data. As a rule of thumb: if you weren't the one running this experiment, would you still want the raw data to combine it with other datasets or explore your own questions? If yes, it's a good candidate for sharing in raw or processed form. If a summary figure would suffice, results-level data alone may be sufficient.
Reference Table: Accepted Data Types and Formats
Many common experimental data types are listed below, but if you're generating a data type not included here, reach out to us through your Service Desk ticket — we're happy to help you figure out an appropriate format.
|
Data Type |
Levels |
Format |
|---|---|---|
|
Whole genome/exome sequencing |
raw and/or processed |
raw: FASTQ, unaligned BAM, CRAM; processed: VCF, PLINK (.bed/.bim/.fam), .ped/.map |
|
SNP microarray |
raw and processed |
raw: CEL, IDAT, TSV; processed: TSV, VCF, PLINK |
|
RNA sequencing (bulk) |
raw and/or processed |
raw: FASTQ, unaligned BAM, CRAM; processed: counts matrices or quantification files |
|
RNA sequencing (single-cell) |
raw and processed |
raw: FASTQ; processed: h5ad/hdf5 (following cellxgene format) |
|
Gene expression microarray |
raw and processed |
raw: CEL, IDAT, TSV; processed: TSV (normalized values) |
|
ATAC/methylation sequencing |
raw and/or semi-processed |
raw: FASTQ, unaligned BAM, CRAM; semi-processed: aligned BAM |
|
Bisulfite sequencing |
raw and/or processed |
raw: FASTQ, unaligned BAM, CRAM; processed: BED, BEDGRAPH |
|
Proteomics (LC-MS) |
raw and processed |
raw: mzML; processed: protein/peptide intensities (CSV/TSV) |
|
Western blot |
processed |
densitometry output (CSV/TSV) |
|
Metabolomics (LC-MS) |
raw and processed |
raw: mzML or vendor format; processed: metabolite intensities (CSV/TSV) |
|
16S rRNA / shotgun metagenomics |
raw and processed |
raw: FASTQ; processed: FASTA, BAM/SAM, BIOM, TSV/CSV |
|
Structured clinical/phenotype data |
processed |
CSV/TSV or XML with metadata for each variable |
|
MRI or other radiological imaging |
raw |
DICOM |
|
Immunohistochemistry/immunofluorescence |
raw |
OME-TIFF preferred, or another bio-formats compatible format |
|
In vivo tumor growth / plate-based assays |
raw and/or processed |
CSV/TSV |
|
Electrophysiology |
raw and/or processed |
NWB preferred where possible; other formats accepted (.abf, .mat, .h5/.hdf5, .csv/.tsv) |
Sharing Results and Analysis Data
To share analysis or results data specifically, see Contribute Analysis and Results Data.
Metadata Requirements
To share your data with the AD Knowledge Portal, we require annotations as described in Assemble Metadata.