| AVI | A multimedia container format for audio video interleaved (AVI) files that allows for synchronous audio-with-video playback. | NCIT:C190162 |
| BAI | An index file associated with a binary alignment map (BAM) file, a compressed binary version of a sequence alignment/map (SAM) file, used to enable rapid random access to alignment records. | NCIT:C190163 |
| BAM | A binary representation of a sequence alignment map (SAM), a compact and indexable representation of nucleotide sequence alignments, compressed using the BGZF (Blocked GNU Zip Format) library. | NCIT:C153249 |
| BED | BED (Browser Extensible Data) format is a tab-delimited text format for describing genomic regions or features, such as gene models or annotation tracks, typically displayed in a genome browser. | EDAM:format_3003 |
| CDS | CDS (coding sequence) refers to a nucleotide sequence file, typically in FASTA-like text format, that contains only the protein-coding portion of a gene or transcript, excluding untranslated regions. | Not available |
| CHP | Format of an Affymetrix data file containing computed (normalized) expression, genotyping, or resequencing results for individual probes on a microarray, generated by Affymetrix analysis software. | EDAM:format_1644 |
| COOL | cool is a sparse, compressed binary file format built on HDF5 for storing genomic contact matrices, such as Hi-C interaction data, developed as part of the Cooler/cooltools software ecosystem. | Not available |
| CSV | A text file format that separates data elements with commas. | NCIT:C182456 |
| DAE | DAE (Digital Asset Exchange) is an XML-based file format defined by the COLLADA specification for exchanging 3D graphics assets, including geometry, materials, and animation, between different software applications. | Not available |
| DB | DB is a generic file extension denoting a database file used to store structured, queryable data, most commonly associated with lightweight database engines such as SQLite. | EDAM:format_3621 |
| DCC | A legacy, tab-delimited text file format used by data coordinating centers (e.g., the TCGA/TARGET Data Coordinating Center) to distribute processed microarray and other genomic data to downstream users. | Not available |
| DS_Store | .DS_Store is a hidden metadata file automatically generated by the macOS Finder to store folder view settings and icon positions; it is not a data format and typically appears as an incidental artifact rather than an intended data file. | Not available |
| FASTA | A text-based format for representing nucleotide or peptide sequences, in which base pairs or amino acids are represented using single-letter codes, with each sequence preceded by a single-line description beginning with a '>' character. | EDAM:format_1929 |
| FASTQ | FASTQ format is a text-based format for storing both a biological sequence (usually nucleotide sequence) and its corresponding quality scores. | EFO:0004155 |
| FCS | FCS (Flow Cytometry Standard) is a binary file format standardized by the International Society for Advancement of Cytometry for storing list-mode flow cytometry data, including per-event fluorescence and light-scatter measurements plus instrument and acquisition metadata. | OBI:0000327 |
| FIG | FIG is a vector graphics file format associated with drawing and plotting tools, such as Xfig or MATLAB saved figures, used to store editable diagrams, plots, or illustrations. | Not available |
| FREQ | FREQ is a generic, tool-specific text file format used to report allele, genotype, or variant frequency data, with column structure varying by the producing software. | Not available |
| GCG | GCG is a legacy sequence file format originally used by the Genetics Computer Group (GCG) software suite, now part of the Wisconsin Package, for representing annotated nucleotide or protein sequences. | EDAM:format_1935 |
| GCT | A tab-delimited, matrix file text format that describes a gene expression dataset; the columns of the matrix correspond to profiles/samples, the rows correspond to genes, and the values of the cells correspond to an expression measurement. | NCIT:C123891 |
| GCTx | GCTx is a binary, HDF5-based extension of the GCT format used to store very large gene-expression or profiling matrices, developed for the Broad Institute's Connectivity Map (CMap)/LINCS projects. | Not available |
| GFF3 | A tab-separated format for sequence data. It uses one line per feature, each containing 9 columns of data, plus optional track definition lines. | NCIT:C190175 |
| GTF | A tab-delimited text format based on the general feature format (GFF) but which also contains some additional information for each gene. | NCIT:C190177 |
| GZIP Format | A file format consisting of a 10-byte header containing a magic number, a version number, and a timestamp, a Deflate-compressed body, and an 8-byte footer containing a checksum and the length of the original uncompressed data. | NCIT:C80220 |
| HDF | HDF is the name of a set of file formats and libraries designed to store and organize large amounts of numerical data, originally developed at the National Center for Supercomputing Applications; it is supported by many platforms including Java, MATLAB, Python, and R. | EDAM:format_3873 |
| HDF5 | A hierarchical, filesystem-like data format that can store metadata in the form of user-defined, named attributes, which are attached to groups and datasets, and representations of images and tables built up using datasets, groups and attributes. | NCIT:C184763 |
| HTML | HTML (HyperText Markup Language) is a markup format used to structure and present content, including text, images, and hyperlinks, for display in web browsers. | EDAM:format_2331 |
| IDAT | A proprietary, encrypted, compressed electronic file format from Illumina Inc. for storing genome-wide profiling data. The data contains a summary of the intensity data generated from each probe used during a sequencing or array run. | NCIT:C184762 |
| JPG | JPEG (Joint Photographic Experts Group) is a lossy compressed raster image file format commonly used for photographic and other continuous-tone images. | EDAM:format_3579 |
| JSON | A common open standard file format and data interchange format that uses human-readable text consisting of attribute-value pairs and arrays to store and transmit data. | NCIT:C184769 |
| LIF | LIF (Leica Image File) is a proprietary container file format produced by Leica microscopy acquisition software that stores multi-dimensional, multi-channel microscope images along with associated metadata. | Not available |
| MAP | MAP files are plain-text files that associate genomic marker positions with identifiers, most commonly the PLINK .map format used in genetics to record chromosome and position information for markers paired with a corresponding genotype file. | EDAM:format_3285 |
| MAT | A proprietary, binary data container format used by MATLAB software to store workspace variables. | NCIT:C190178 |
| MATLAB script | The file format for MATLAB scripts or functions, stored as plain text with a .m extension and executed within the MATLAB software environment. | EDAM:format_4007 |
| MSF | A proprietary mass-spectrometry format used by Thermo Scientific's ProteomeDiscoverer software; it corresponds to an SQLite database that stores peptide and protein identification results. | EDAM:format_3702 |
| MTX | A format for the storage of numerical or pattern matrices in a dense (array format) or sparse (coordinate format) representation. | NCIT:C201789 |
| PDF | An open file format created and controlled by Adobe Systems for representing two-dimensional documents in a device-independent and resolution-independent fixed-layout format that encapsulates text, fonts, images, and vector graphics. | NCIT:C63805 |
| PNG | An extensible file format for lossless data compression of images. | NCIT:C85437 |
| PZFX | An XML-based file format used by GraphPad Prism software to store analysis projects, including data tables, graphs, and statistical results. | NCIT:C224011 |
| Python Script | A plain text file that stores python code. | NCIT:C190184 |
| R File Format | R is a tool for data manipulation and analysis, both a programming language and a working environment; the format is used for writing and executing scripts, typically for statistical computation and graphics, and for storing complete R workspaces or selected R objects. | NCIT:C190186 |
| RAW | A proprietary file format developed by Thermo Scientific to capture mass spectrometry data produced by Thermo mass spectrometers. | NCIT:C190190 |
| RDS | A native binary file format used to save and load single R objects to and from a file. | NCIT:C209895 |
| ROUT | A .Rout file is the plain-text output log generated when an R script is executed in batch mode (via R CMD BATCH), capturing console output, computed results, and any warnings or errors. | Not available |
| RPROJ | An .Rproj file is a plain-text configuration file created by RStudio that records project-specific settings, such as the working directory and build options, for an R project. | Not available |
| RTF | A file format that contains various formatting elements and enables exchanges of text files across different word processing platforms and applications. | NCIT:C209896 |
| SGI | SGI (Silicon Graphics Image) is a raster image file format originally developed by Silicon Graphics, Inc. for storing bitmap graphics, supporting both uncompressed and run-length-encoded pixel data. | Not available |
| SRA | SRA archive format is the archive format used for input to, and export from, the NCBI Sequence Read Archive, storing raw next-generation sequencing reads and associated metadata. | EDAM:format_3698 |
| STAT | STAT denotes a plain-text file containing summary statistics or quality-control metrics generated by an analysis tool, such as alignment or variant-calling statistics; its exact content and structure vary by the producing software. | Not available |
| TAR Format | A file format generated by the Unix command-line utility tar (tape archive). It is used for collecting many files into one archive file. | NCIT:C190189 |
| TDF | TDF (Tiled Data File) is a binary, indexed file format used by the Integrative Genomics Viewer (IGV) to store precomputed, multi-resolution genomic data, such as coverage or feature density, for fast visualization across zoom levels. | Not available |
| TIFF | A versatile bitmap image format that supports numerous data compression schemes, widely used for high-quality raster images, including microscopy and pathology images. | EDAM:format_3591 |
| TSV | A file format where each line in the file contains a single piece of data and where each field or value in a line of data is separated from the next by a tab character. | NCIT:C164049 |
| TXT | A data format consisting of readable textual material maintained as a sequential file. | NCIT:C85873 |
| VCF | A text-based electronic file used for storing gene sequence variation data. The first text section is composed of a header containing the metadata and keywords used in the file, followed by the body of the file which is tab-separated into eight mandatory data columns for each sample, plus optional columns for other sample-related data. | NCIT:C172216 |
| WIG | Wiggle format (WIG) of a sequence annotation track that consists of a value for each sequence position, typically displayed in a genome browser. | EDAM:format_3005 |
| XML | XML (Extensible Markup Language) is a flexible, tag-based text format for representing structured, hierarchical data in a manner that is both human-readable and machine-parsable. | EDAM:format_2332 |
| ZIP | An archive file format that supports lossless data compression. It is used to compress one or more files together into a single location. | NCIT:C190192 |
| bed12 | A BED file variant in which each feature is described using all twelve standard BED columns, with no non-standard columns allowed. | EDAM:format_3586 |
| bedgraph | A format that allows display of continuous-valued data in track format. This display type is useful for probability scores and transcriptome data. | NCIT:C190165 |
| cel | A type of file that stores the results of intensity calculations of pixel values for DNA microarray image analysis, including intensity values, standard deviation, pixel counts, and outlier/exclusion flags for each feature on the probe array. | NCIT:C191737 |
| cloupe | The .cloupe file format is a proprietary binary format produced by 10x Genomics Cell Ranger pipelines for visualizing and exploring single-cell sequencing data in the Loupe Browser desktop application. | Not available |
| docx | A file type that uses Open XML Document format to store data for Microsoft Word and other compatible programs. | NCIT:C190171 |
| mzIdentML | The exchange format for peptides and proteins identified from mass spectra, standardised by HUPO PSI-PI; it can be used for outputs of proteomics search engines. | EDAM:format_3247 |
| mzXML | A common file format for proteomics mass spectrometric data developed at the Seattle Proteome Center/Institute for Systems Biology. | EDAM:format_3654 |
| pptx | A visual design format that includes layout, fonts, colors, and aspect ratios for slide-based presentations. PPTX is an XML-based format that is the default for modern Microsoft PowerPoint files. | NCIT:C224009 |
| rcc | Reporter Code Count (RCC) is a data file output by the NanoString nCounter Digital Analyzer, containing gene/sample information, probe information, and probe counts. | EDAM:format_3580 |
| xls | A proprietary binary file format used to store spreadsheet data in Microsoft Excel. | NCIT:C190191 |
| xlsx | A proprietary file format developed by Microsoft that allows the user to save a spreadsheet created in Excel in an open XML format; the file can be read and opened by other spreadsheet-compatible applications. | NCIT:C164050 |
| MGF | Mascot Generic Format (MGF) is a text-based file format that encodes multiple tandem mass spectrometry (MS/MS) spectra, with m/z and intensity pairs, in a single file. | EDAM:format_3651 |
| BIGWIG | A format for large sequence annotation tracks that consist of a value for each sequence position, similar to the textual wiggle (WIG) format, but indexed and compressed for efficient random access. | NCIT:C190167 |
| H5AD | H5AD is an HDF5-based file format used by the AnnData Python library (and Scanpy/Seurat-compatible tools) to store annotated data matrices, such as single-cell gene-expression counts, together with associated observation and variable (cell/gene) metadata. | EDAM:format_3590 |
| H5 | An HDF5 file appears to the user as a directed graph whose nodes are higher-level HDF5 objects, such as groups, datasets, and named datatypes, used to store and organize large amounts of numerical and hierarchical data. | EDAM:format_3590 |
| SF | SF commonly denotes the tab-delimited transcript-quantification output file (e.g., quant.sf) produced by RNA-seq pseudoalignment/quantification tools such as Salmon, reporting per-transcript abundance estimates (TPM and read counts). | Not available |
| PKL | Format used by the Python pickle module for serializing and de-serializing a Python object structure to and from a byte stream. | EDAM:format_4002 |
| BPM | BPM (BeadPool Manifest) is a proprietary Illumina binary file format that contains the design and manifest information, such as probe sequences, bead addresses, and annotations, for a genotyping or methylation microarray. | Not available |
| Unspecified | Indicates that the file format of the associated dataset or file has not been specified by the data submitter. | NCIT:C38046 |
| Pending Annotation | Indicates that curatorial annotation of the file's format has not yet been completed and is pending further review. | NCIT:C53470 |
| maf | A tab-delimited text file containing mutation information aggregated across a set of variant call files (VCF). | NCIT:C172215 |
| CLS | CLS is a plain-text categorical class (phenotype) label file format used by tools such as GSEA to specify sample group or class assignments corresponding to columns in an expression dataset. | Not available |
| SCN | SCN is a proprietary, TIFF-based whole-slide image file format produced by Leica Biosystems digital pathology scanners, encoding multi-resolution, multi-channel brightfield or fluorescence slide images. | Not available |
| SVS | SVS is a proprietary, pyramidal TIFF-based whole-slide image file format produced by Aperio (Leica Biosystems) digital slide scanners, used to store high-resolution histopathology images. | NCIT:C162692 |