Data Catalog
A Data Catalog entry captures the schema.org/Bioschemas-flavored dataset-cataloging metadata that Synapse's own Data Catalog feature stores directly on a Dataset entity — fields like measurementTechnique, license, accessType, conditionsOfAccess, funder, and includedInDataCatalog. It describes the same real-world dataset as a portal Dataset entry, but as a second, complementary metadata facet rather than a duplicate record: DataCatalog_id always equals the corresponding Dataset entry's own dataset ID.
This model applies only to datasets backed by a real Synapse Dataset entity (i.e., hosted or indexed directly on Synapse, as opposed to externally hosted datasets described only through a portal Dataset entry).
Where Data Catalog Metadata Comes From¶
Unlike the other models on this site, Data Catalog fields are not submitted through this repo's manifest/template process. They are populated directly on the Synapse Dataset entity itself, through Synapse's own Data Catalog UI and metadata assistant, at the time the dataset is created or curated on Synapse.
The Cancer Complexity Knowledge Portal (CCKP)'s knowledge-graph pipeline reads these annotations directly from the Dataset entity and merges them onto the same subject as the entity's portal Dataset metadata, so the two facets end up describing one unified dataset in the CCKP knowledge graph.
Who Maintains Data Catalog Annotations?¶
Whoever manages the corresponding Synapse Dataset entity — typically the same Principal Investigators, Data Managers, and Research Staff who host or index the dataset's files on Synapse. Keeping these annotations current on the Dataset entity (rather than in a separate submission) ensures the CCKP knowledge graph reflects accurate licensing, access, and provenance information for the dataset.
Full Field Reference¶
Below is the full field reference table with attributes and their descriptions.
| Attribute | Description | Required | Column Type | Format | Regex Pattern | Standard Terms | CDE | Examples |
|---|---|---|---|---|---|---|---|---|
| DataCatalog_id | Unique ID associated with a Data Catalog entry, used as a key for record linking and updates. | True | string | None | ^syn\d{7,8}$ | None | ||
| GrantView Key | Unique GrantView_id foreign key(s) that group the resource with other components, as part of the same grant-associated collection. Please provide multiple values as a comma-separated list. | False | string | None | (CA\d{6}|Affiliated/Non-Grant Associated) | None | ||
| Study Key | The unique Study_id foreign keys associated with the resource, found in the grant Study information. Used to group the resource with other components. Please provide multiple values as a comma-separated list. | False | string | None | None | None | ||
| DatasetView Key | Unique DatasetView_id foreign key(s) that link metadata entries as part of the same collection. Please provide multiple values as a comma-separated list. | False | string | None | None | None | ||
| dataCatalogStudyId | The Synapse ID of the Project associated with this dataset's grant. Populated as a native Synapse annotation directly on the Dataset entity; distinct from this repo's own Study_id/Study Key. | False | string | None | ^syn\d{7,8}$ | None | ||
| portal | The Sage Bionetworks portal that curated this record (e.g. CCKP). Distinguishes provenance if this vocabulary is ever shared across multiple Sage portals. | False | string | None | None | None | ||
| community | The Sage Bionetworks project or initiative with which the dataset is associated. | False | string | None | None | None | ||
| description | A text description of the dataset. Can be identifical to Dataset Description. | False | string | None | None | None | ||
| dataCatalogContributor | Organization(s), insitution(s), or person(s) that generated and/or deposited the data in Synapse. Can be identical to Grant Institution Name. | False | string_list | None | None | None | ||
| keywords | Keywords associated with the dataset. Can be identical to Publication Keywords. | True | string_list | None | None | None | ||
| individualCount | The number of participants or samples associated with the files in the dataset. Can be identical to Study Number of Participants or DSP Number of Participants | False | number | None | None | None | ||
| link | The url where the dataset is displayed, typically on a public-facing data portal. Can be identical to Dataset Url. | False | string | uri | None | None | ||
| croissant_s3_file_object | A Synapse url where the croissant file associated with this dataset is stored. | False | string | uri | None | None | ||
| accessType | Access type for the dataset. | False | string | None | None | View | ||
| alternateName | An altername name that can be used for search and discovery improvement. | False | string | None | None | None | ||
| conditionsOfAccess | Additional requirements a user may need outside of Data Use Modifiers. This could include additional registration, updating profile information, joining a Synapse Team, or using specific authentication methods like 2FA or RAS. Omit property if not applicable/unknown. | False | string | None | None | None | ||
| countryOfOrigin | Origin of individuals from which data were generated. Omit if not applicable/unknown. | False | string_list | None | None | View | ||
| creator | Organization or person that is creator of the dataset. Default is the PI of the project and/or the user who created all files in the dataset. | True | string_list | None | None | None | ||
| authors | Author list for the dataset's associated publication. Distinct from creator, which captures the person(s)/organization(s) that generated or deposited the data itself. | False | string_list | None | None | None | ||
| dataCatalogDataUseModifiers | List of data use ontology (DUO) terms that are true for dataset, which describes the allowable scope and terms for data use. Omit property if not applicable/unknown. | False | string_list | None | ^(DUO:\d{7}|DUOPlus\d+|Pending Annotation)$ | None | ||
| datePublished | Date data were published/available on Synapse. | False | number | None | None | None | ||
| funder | The entity(ies) responsible for funding the data contributors. | False | string_list | None | None | View | ||
| includedInDataCatalog | Link(s) to known data catalog(s) the dataset is included in. | False | string | None | None | None | ||
| dataCatalogLicense | The terms and conditions underwhich the data are allowed to be accessed and used. Unless information for license is clear, this should default to UNKNOWN. | False | string | None | None | View | ||
| measurementTechnique | Assay used to generate original data. Omit if not applicable (e.g. for curated dataset such as a list compounds from a database or text extracted from Wikipedia). Can be identical to Dataset Assay. | False | string_list | None | None | View | ||
| dataType | Broad data-type category for the dataset (e.g. gene expression, chromatin activity, genomic variants). Aligned with modules/shared/mc2_iconTag_map_3-4-25.csv's label taxonomy. | False | string_list | None | None | View | ||
| Species | The species of origin corresponding to the model, or from which the biospecimen originates. | True | string_list | None | None | View | ||
| subject | Applicable subject term(s) for dataset cataloging - in practice, a free-text model-system/specimen descriptor (e.g. C57BL/6, NOD SCID), not a Library of Congress Subject Heading. | False | string_list | None | None | None | ||
| title | The name of the dataset. | True | string | None | None | None | ||
| citation | Academic articles that are recommended by the data provider to be cited in addition to the dataset doi itself. Can be identical to DatasetView.PublicationView Key. | False | string | None | None | None | ||
| doi | Digital Object Identifier (DOI) for citation and identification. Can be identical to Dataset Doi | False | string | uri | None | None | ||
| visualizeDataOn | Links to external visualization platforms. | False | string | uri | None | None | ||
| yearProcessed | Year the data were processed (used with specific series only). | False | number | None | None | None | ||
| specimenCount | Number of unique specimens in the dataset. Can be identical to DSP Number of Samples. | False | number | None | None | None | ||
| series | Data series designation (e.g., NF-OSI Processed Data). | False | string | None | None | None | ||
| downloadType | Hosting type for the dataset's files, as annotated directly on the Synapse Dataset entity (distinct from the CCKP portal's own DatasetView.downloadType column). | False | string | None | None | View | ||
| externalRepositoryUri | External repository identifier(s) for the dataset, as CURIE-shaped references (e.g. geo:GSE12345, sra:SRP12345, bioproject:PRJNA12345). | False | string_list | None | None | None | ||
| ageGroup | Age group of the individual(s)/subject(s) the dataset's data were generated from (e.g. Adult). Free text pending a confident controlled vocabulary - too few real values observed to seed one yet. | False | string | None | None | None | ||
| diseaseFocus | Disease area the dataset focuses on, as annotated directly on the Synapse Dataset entity. Not populated in current practice. | False | string | None | None | None | ||
| manifestation | Tumor type or other clinical manifestation associated with the dataset. Can be identical to Dataset Tumor Type. | False | string | None | None | View |