AD Portal Docs

Format your Data

Experimental data is shared in a variety of formats. Choosing the right format extends the reusability and impact of your research, since others need to be able to open and interpret your files.

This page explains the difference between raw, processed, and results data, and provides a reference table of accepted formats for common data types.


Raw Data, Processed Data, and Results

From a reusability standpoint, data (raw or processed) is generally more valuable to future users than results, since data can support new analyses that a results file cannot.

  • Data refers to raw or partially processed information from a single sample.

  • Results are generally post-analysis information from an aggregate of samples, or figures intended for a manuscript.

For example, for RNA-seq data: raw data would be the FASTQ files, processed data would be aligned reads (.bam) or gene counts, and a differential expression analysis or volcano plot would be considered results.

This distinction is well-defined for many data types, but less so for certain assays — for example, "results" may be an acceptable format for assays that don't lend themselves to re-analysis, such as western blots. If you're unsure how your data fits, reach out to us through your Service Desk ticket.

What Data Should Be Shared?

We generally ask that raw or minimally processed data be shared whenever possible, since this is what allows secondary users to ask their own questions of the data. As a rule of thumb: if you weren't the one running this experiment, would you still want the raw data to combine it with other datasets or explore your own questions? If yes, it's a good candidate for sharing in raw or processed form. If a summary figure would suffice, results-level data alone may be sufficient.

Reference Table: Accepted Data Types and Formats

Many common experimental data types are listed below, but if you're generating a data type not included here, reach out to us through your Service Desk ticket — we're happy to help you figure out an appropriate format.

Data Type

Levels

Format

Whole genome/exome sequencing

raw and/or processed

raw: FASTQ, unaligned BAM, CRAM; processed: VCF, PLINK (.bed/.bim/.fam), .ped/.map

SNP microarray

raw and processed

raw: CEL, IDAT, TSV; processed: TSV, VCF, PLINK

RNA sequencing (bulk)

raw and/or processed

raw: FASTQ, unaligned BAM, CRAM; processed: counts matrices or quantification files

RNA sequencing (single-cell)

raw and processed

raw: FASTQ; processed: h5ad/hdf5 (following cellxgene format)

Gene expression microarray

raw and processed

raw: CEL, IDAT, TSV; processed: TSV (normalized values)

ATAC/methylation sequencing

raw and/or semi-processed

raw: FASTQ, unaligned BAM, CRAM; semi-processed: aligned BAM

Bisulfite sequencing

raw and/or processed

raw: FASTQ, unaligned BAM, CRAM; processed: BED, BEDGRAPH

Proteomics (LC-MS)

raw and processed

raw: mzML; processed: protein/peptide intensities (CSV/TSV)

Western blot

processed

densitometry output (CSV/TSV)

Metabolomics (LC-MS)

raw and processed

raw: mzML or vendor format; processed: metabolite intensities (CSV/TSV)

16S rRNA / shotgun metagenomics

raw and processed

raw: FASTQ; processed: FASTA, BAM/SAM, BIOM, TSV/CSV

Structured clinical/phenotype data

processed

CSV/TSV or XML with metadata for each variable

MRI or other radiological imaging

raw

DICOM

Immunohistochemistry/immunofluorescence

raw

OME-TIFF preferred, or another bio-formats compatible format

In vivo tumor growth / plate-based assays

raw and/or processed

CSV/TSV

Electrophysiology

raw and/or processed

NWB preferred where possible; other formats accepted (.abf, .mat, .h5/.hdf5, .csv/.tsv)

Sharing Results and Analysis Data

To share analysis or results data specifically, see Contribute Analysis and Results Data.

Metadata Requirements

To share your data with the AD Knowledge Portal, we require annotations as described in Assemble Metadata.


Last updated: