Skip to content

Dataset

Attribute: Dataset File Formats

Valid Value Description Ontology
AVI A multimedia container format for audio video interleaved (AVI) files that allows for synchronous audio-with-video playback. NCIT:C190162
BAI An index file associated with a binary alignment map (BAM) file, a compressed binary version of a sequence alignment/map (SAM) file, used to enable rapid random access to alignment records. NCIT:C190163
BAM A binary representation of a sequence alignment map (SAM), a compact and indexable representation of nucleotide sequence alignments, compressed using the BGZF (Blocked GNU Zip Format) library. NCIT:C153249
BED BED (Browser Extensible Data) format is a tab-delimited text format for describing genomic regions or features, such as gene models or annotation tracks, typically displayed in a genome browser. EDAM:format_3003
CDS CDS (coding sequence) refers to a nucleotide sequence file, typically in FASTA-like text format, that contains only the protein-coding portion of a gene or transcript, excluding untranslated regions. Not available
CHP Format of an Affymetrix data file containing computed (normalized) expression, genotyping, or resequencing results for individual probes on a microarray, generated by Affymetrix analysis software. EDAM:format_1644
COOL cool is a sparse, compressed binary file format built on HDF5 for storing genomic contact matrices, such as Hi-C interaction data, developed as part of the Cooler/cooltools software ecosystem. Not available
CSV A text file format that separates data elements with commas. NCIT:C182456
DAE DAE (Digital Asset Exchange) is an XML-based file format defined by the COLLADA specification for exchanging 3D graphics assets, including geometry, materials, and animation, between different software applications. Not available
DB DB is a generic file extension denoting a database file used to store structured, queryable data, most commonly associated with lightweight database engines such as SQLite. EDAM:format_3621
DCC A legacy, tab-delimited text file format used by data coordinating centers (e.g., the TCGA/TARGET Data Coordinating Center) to distribute processed microarray and other genomic data to downstream users. Not available
DS_Store .DS_Store is a hidden metadata file automatically generated by the macOS Finder to store folder view settings and icon positions; it is not a data format and typically appears as an incidental artifact rather than an intended data file. Not available
FASTA A text-based format for representing nucleotide or peptide sequences, in which base pairs or amino acids are represented using single-letter codes, with each sequence preceded by a single-line description beginning with a '>' character. EDAM:format_1929
FASTQ FASTQ format is a text-based format for storing both a biological sequence (usually nucleotide sequence) and its corresponding quality scores. EFO:0004155
FCS FCS (Flow Cytometry Standard) is a binary file format standardized by the International Society for Advancement of Cytometry for storing list-mode flow cytometry data, including per-event fluorescence and light-scatter measurements plus instrument and acquisition metadata. OBI:0000327
FIG FIG is a vector graphics file format associated with drawing and plotting tools, such as Xfig or MATLAB saved figures, used to store editable diagrams, plots, or illustrations. Not available
FREQ FREQ is a generic, tool-specific text file format used to report allele, genotype, or variant frequency data, with column structure varying by the producing software. Not available
GCG GCG is a legacy sequence file format originally used by the Genetics Computer Group (GCG) software suite, now part of the Wisconsin Package, for representing annotated nucleotide or protein sequences. EDAM:format_1935
GCT A tab-delimited, matrix file text format that describes a gene expression dataset; the columns of the matrix correspond to profiles/samples, the rows correspond to genes, and the values of the cells correspond to an expression measurement. NCIT:C123891
GCTx GCTx is a binary, HDF5-based extension of the GCT format used to store very large gene-expression or profiling matrices, developed for the Broad Institute's Connectivity Map (CMap)/LINCS projects. Not available
GFF3 A tab-separated format for sequence data. It uses one line per feature, each containing 9 columns of data, plus optional track definition lines. NCIT:C190175
GTF A tab-delimited text format based on the general feature format (GFF) but which also contains some additional information for each gene. NCIT:C190177
GZIP Format A file format consisting of a 10-byte header containing a magic number, a version number, and a timestamp, a Deflate-compressed body, and an 8-byte footer containing a checksum and the length of the original uncompressed data. NCIT:C80220
HDF HDF is the name of a set of file formats and libraries designed to store and organize large amounts of numerical data, originally developed at the National Center for Supercomputing Applications; it is supported by many platforms including Java, MATLAB, Python, and R. EDAM:format_3873
HDF5 A hierarchical, filesystem-like data format that can store metadata in the form of user-defined, named attributes, which are attached to groups and datasets, and representations of images and tables built up using datasets, groups and attributes. NCIT:C184763
HTML HTML (HyperText Markup Language) is a markup format used to structure and present content, including text, images, and hyperlinks, for display in web browsers. EDAM:format_2331
IDAT A proprietary, encrypted, compressed electronic file format from Illumina Inc. for storing genome-wide profiling data. The data contains a summary of the intensity data generated from each probe used during a sequencing or array run. NCIT:C184762
JPG JPEG (Joint Photographic Experts Group) is a lossy compressed raster image file format commonly used for photographic and other continuous-tone images. EDAM:format_3579
JSON A common open standard file format and data interchange format that uses human-readable text consisting of attribute-value pairs and arrays to store and transmit data. NCIT:C184769
LIF LIF (Leica Image File) is a proprietary container file format produced by Leica microscopy acquisition software that stores multi-dimensional, multi-channel microscope images along with associated metadata. Not available
MAP MAP files are plain-text files that associate genomic marker positions with identifiers, most commonly the PLINK .map format used in genetics to record chromosome and position information for markers paired with a corresponding genotype file. EDAM:format_3285
MAT A proprietary, binary data container format used by MATLAB software to store workspace variables. NCIT:C190178
MATLAB script The file format for MATLAB scripts or functions, stored as plain text with a .m extension and executed within the MATLAB software environment. EDAM:format_4007
MSF A proprietary mass-spectrometry format used by Thermo Scientific's ProteomeDiscoverer software; it corresponds to an SQLite database that stores peptide and protein identification results. EDAM:format_3702
MTX A format for the storage of numerical or pattern matrices in a dense (array format) or sparse (coordinate format) representation. NCIT:C201789
PDF An open file format created and controlled by Adobe Systems for representing two-dimensional documents in a device-independent and resolution-independent fixed-layout format that encapsulates text, fonts, images, and vector graphics. NCIT:C63805
PNG An extensible file format for lossless data compression of images. NCIT:C85437
PZFX An XML-based file format used by GraphPad Prism software to store analysis projects, including data tables, graphs, and statistical results. NCIT:C224011
Python Script A plain text file that stores python code. NCIT:C190184
R File Format R is a tool for data manipulation and analysis, both a programming language and a working environment; the format is used for writing and executing scripts, typically for statistical computation and graphics, and for storing complete R workspaces or selected R objects. NCIT:C190186
RAW A proprietary file format developed by Thermo Scientific to capture mass spectrometry data produced by Thermo mass spectrometers. NCIT:C190190
RDS A native binary file format used to save and load single R objects to and from a file. NCIT:C209895
ROUT A .Rout file is the plain-text output log generated when an R script is executed in batch mode (via R CMD BATCH), capturing console output, computed results, and any warnings or errors. Not available
RPROJ An .Rproj file is a plain-text configuration file created by RStudio that records project-specific settings, such as the working directory and build options, for an R project. Not available
RTF A file format that contains various formatting elements and enables exchanges of text files across different word processing platforms and applications. NCIT:C209896
SGI SGI (Silicon Graphics Image) is a raster image file format originally developed by Silicon Graphics, Inc. for storing bitmap graphics, supporting both uncompressed and run-length-encoded pixel data. Not available
SRA SRA archive format is the archive format used for input to, and export from, the NCBI Sequence Read Archive, storing raw next-generation sequencing reads and associated metadata. EDAM:format_3698
STAT STAT denotes a plain-text file containing summary statistics or quality-control metrics generated by an analysis tool, such as alignment or variant-calling statistics; its exact content and structure vary by the producing software. Not available
TAR Format A file format generated by the Unix command-line utility tar (tape archive). It is used for collecting many files into one archive file. NCIT:C190189
TDF TDF (Tiled Data File) is a binary, indexed file format used by the Integrative Genomics Viewer (IGV) to store precomputed, multi-resolution genomic data, such as coverage or feature density, for fast visualization across zoom levels. Not available
TIFF A versatile bitmap image format that supports numerous data compression schemes, widely used for high-quality raster images, including microscopy and pathology images. EDAM:format_3591
TSV A file format where each line in the file contains a single piece of data and where each field or value in a line of data is separated from the next by a tab character. NCIT:C164049
TXT A data format consisting of readable textual material maintained as a sequential file. NCIT:C85873
VCF A text-based electronic file used for storing gene sequence variation data. The first text section is composed of a header containing the metadata and keywords used in the file, followed by the body of the file which is tab-separated into eight mandatory data columns for each sample, plus optional columns for other sample-related data.NCIT:C172216
WIG Wiggle format (WIG) of a sequence annotation track that consists of a value for each sequence position, typically displayed in a genome browser. EDAM:format_3005
XML XML (Extensible Markup Language) is a flexible, tag-based text format for representing structured, hierarchical data in a manner that is both human-readable and machine-parsable. EDAM:format_2332
ZIP An archive file format that supports lossless data compression. It is used to compress one or more files together into a single location. NCIT:C190192
bed12 A BED file variant in which each feature is described using all twelve standard BED columns, with no non-standard columns allowed. EDAM:format_3586
bedgraph A format that allows display of continuous-valued data in track format. This display type is useful for probability scores and transcriptome data. NCIT:C190165
cel A type of file that stores the results of intensity calculations of pixel values for DNA microarray image analysis, including intensity values, standard deviation, pixel counts, and outlier/exclusion flags for each feature on the probe array. NCIT:C191737
cloupe The .cloupe file format is a proprietary binary format produced by 10x Genomics Cell Ranger pipelines for visualizing and exploring single-cell sequencing data in the Loupe Browser desktop application. Not available
docx A file type that uses Open XML Document format to store data for Microsoft Word and other compatible programs. NCIT:C190171
mzIdentML The exchange format for peptides and proteins identified from mass spectra, standardised by HUPO PSI-PI; it can be used for outputs of proteomics search engines. EDAM:format_3247
mzXML A common file format for proteomics mass spectrometric data developed at the Seattle Proteome Center/Institute for Systems Biology. EDAM:format_3654
pptx A visual design format that includes layout, fonts, colors, and aspect ratios for slide-based presentations. PPTX is an XML-based format that is the default for modern Microsoft PowerPoint files. NCIT:C224009
rcc Reporter Code Count (RCC) is a data file output by the NanoString nCounter Digital Analyzer, containing gene/sample information, probe information, and probe counts. EDAM:format_3580
xls A proprietary binary file format used to store spreadsheet data in Microsoft Excel. NCIT:C190191
xlsx A proprietary file format developed by Microsoft that allows the user to save a spreadsheet created in Excel in an open XML format; the file can be read and opened by other spreadsheet-compatible applications. NCIT:C164050
MGF Mascot Generic Format (MGF) is a text-based file format that encodes multiple tandem mass spectrometry (MS/MS) spectra, with m/z and intensity pairs, in a single file. EDAM:format_3651
BIGWIG A format for large sequence annotation tracks that consist of a value for each sequence position, similar to the textual wiggle (WIG) format, but indexed and compressed for efficient random access. NCIT:C190167
H5AD H5AD is an HDF5-based file format used by the AnnData Python library (and Scanpy/Seurat-compatible tools) to store annotated data matrices, such as single-cell gene-expression counts, together with associated observation and variable (cell/gene) metadata. EDAM:format_3590
H5 An HDF5 file appears to the user as a directed graph whose nodes are higher-level HDF5 objects, such as groups, datasets, and named datatypes, used to store and organize large amounts of numerical and hierarchical data. EDAM:format_3590
SF SF commonly denotes the tab-delimited transcript-quantification output file (e.g., quant.sf) produced by RNA-seq pseudoalignment/quantification tools such as Salmon, reporting per-transcript abundance estimates (TPM and read counts). Not available
PKL Format used by the Python pickle module for serializing and de-serializing a Python object structure to and from a byte stream. EDAM:format_4002
BPM BPM (BeadPool Manifest) is a proprietary Illumina binary file format that contains the design and manifest information, such as probe sequences, bead addresses, and annotations, for a genotyping or methylation microarray. Not available
Unspecified Indicates that the file format of the associated dataset or file has not been specified by the data submitter. NCIT:C38046
Pending AnnotationIndicates that curatorial annotation of the file's format has not yet been completed and is pending further review. NCIT:C53470
maf A tab-delimited text file containing mutation information aggregated across a set of variant call files (VCF). NCIT:C172215
CLS CLS is a plain-text categorical class (phenotype) label file format used by tools such as GSEA to specify sample group or class assignments corresponding to columns in an expression dataset. Not available
SCN SCN is a proprietary, TIFF-based whole-slide image file format produced by Leica Biosystems digital pathology scanners, encoding multi-resolution, multi-channel brightfield or fluorescence slide images. Not available
SVS SVS is a proprietary, pyramidal TIFF-based whole-slide image file format produced by Aperio (Leica Biosystems) digital slide scanners, used to store high-resolution histopathology images. NCIT:C162692