Skip to content

Model and identifiers

This page covers the LinkML schemas, class alignment, the "Key"-to-"Ref" join convention, and how every IRI is minted. See Overview for context, and the kg-pipeline README for scripts and flags this page doesn't repeat.

LinkML schemas

Schema Source Notes
schema/mc2_model.linkml.yaml This repo's mc2.model.csv + modules/mapping.yaml Generated by a vendored converter (Stage 0 in Build and validation)
schema/cckp_portal.linkml.yaml Hand-authored Imports mc2_model.linkml's shared enums; its 5 classes match the live Synapse table shapes

cckp_portal.linkml.yaml declares a schema.org and Biolink mapping per class (see the class-alignment design).

Class schema.org Biolink
Dataset schema:Dataset (exact) biolink:Dataset (exact)
Publication schema:ScholarlyArticle (exact) biolink:Publication (exact)
Tool schema:SoftwareApplication (exact) biolink:InformationContentEntity (close)
Grant schema:MonetaryGrant (exact) biolink:AdministrativeEntity (close)
EducationalResource schema:LearningResource (close) biolink:InformationContentEntity (close)

Biolink is materialized as a second rdf:type on every instance and checked on every build; the schema.org alignment stays a class-level claim only (real schema: predicates come from the DataCatalog layer).

Foreign keys: "Key" in the model, "Ref" in the graph

Every foreign-key attribute in the model is named "<Entity> Key" (Dataset Key, Grant Key, ...), defined once in modules/shared/annotationProperty.csv; a Key value is a plain string naming another record, not an embedded object. Ref edges are a separate thing: cckp_portal.linkml.yaml marks specific portal-schema columns — never a model "Key" field — with a cckp_join annotation, and build_triples.py resolves those into a Ref edge. mc2_model.linkml.yaml carries no cckp_join at all, so a model "Key" field (like File View's Biospecimen Key) stays a plain literal unless a separate stage resolves it, as link_sagebrain.py does by grouping rows on the literal value.

Portal column Joins to Ref edge
grantNumber Grant.grantNumber cckp:grantNumberRef
pubMedId Publication.pubMedId cckp:pubMedIdRef
dataset/datasets Dataset.datasetAlias cckp:datasetRef

A Dataset's grant number matched to its Grant, next to an unmatched value that stays as plain text.

Identifiers and namespaces

One function mints every instance IRI, so every script that points at an existing row does it the same way.

A decision tree from a row's id string to which of three IRI forms it gets.

A row with a real Synapse id gets Synapse's own canonical IRI — the same base governanceDUO and sagebrain-model use, so the graphs join with no translation. A row with no Synapse id gets a placeholder IRI this pipeline owns.

Dataset and Grant use a declared identifier slot (datasetId, grantId); the other three classes fall back through a per-class field, and finally to a synthetic id, "synthetic-" + sha1(...)[:16], hashed from whatever fallback fields are populated:

Class Falls back to Last resort, hashed
Publication pubMedId publicationTitle + doi
Tool toolName description + downloadUrl
EducationalResource internalIdentifier/alias title

A CV value confirmed to have no real ontology term instead gets a provisional IRI, .../cckp-portal/terms/{field}/{slug}, rather than a dead-end literal or a guess.

Namespace Base IRI Used for
cckp: https://w3id.org/mc2-center/cckp-portal/ Every class/predicate this pipeline mints
(no prefix) https://w3id.org/mc2-center/cckp-portal/data/ Locally-minted instance IRIs
(no prefix) https://w3id.org/mc2-center/cckp-portal/terms/ Provisional placeholder concepts
(no prefix) https://www.synapse.org/Synapse: Canonical Synapse-entity IRIs
mc2: https://w3id.org/mc2-center/mc2-model/ MC2 model schema terms
biolink: https://w3id.org/biolink/vocab/ Dual typing, private-layer stub nodes
schema: https://schema.org/ Class alignment, real DataCatalog predicates
sagecdm: https://sage-bionetworks.github.io/SageCommonDataModel/ SCDM federation nodes
sagebrain: https://w3id.org/synapse/sagebrain# Private-layer links into sagebrain-model
(ontology bases) NCIT/EDAM/ROR/SNOMED/... Resolved ontology and registry terms
shape: https://w3id.org/mc2-center/cckp-portal/shapes# SHACL shape names
gov: https://w3id.org/synapse/governance# governanceDUO's own namespace — never asserted here

See Federation and sagebrain-model for why kg-pipeline never mints or asserts anything in gov:.

A worked example

Real, trimmed output for one Dataset row, syn20826574 (kg-pipeline/data/rdf/Dataset.ttl, from a live build):

@prefix cckp: <https://w3id.org/mc2-center/cckp-portal/> .
@prefix biolink: <https://w3id.org/biolink/vocab/> .

<https://www.synapse.org/Synapse:syn20826574> a biolink:Dataset,
        cckp:Dataset ;
    cckp:assay "CRISPR" ;
    cckp:consortium "CSBC" ;
    cckp:datasetAlias "Genetic Interactions of Chromatin-Related Genes" ;
    cckp:datasetId "syn20826574" ;
    cckp:doi "https://doi.org/10.1038/nmeth.4286" ;
    cckp:doiIri <https://doi.org/10.1038/nmeth.4286> ;
    cckp:fileFormats "Pending Annotation" ;
    cckp:fileFormatsTerm <http://purl.obolibrary.org/obo/NCIT_C53470> ;
    cckp:grantNumber "CA209891" ;
    cckp:grantNumberRef <https://www.synapse.org/Synapse:syn10140998> ;
    cckp:pubMedId "28481362" ;
    cckp:pubMedIdIri <https://pubmed.ncbi.nlm.nih.gov/28481362> ;
    cckp:pubMedIdRef <https://w3id.org/mc2-center/cckp-portal/data/Publication/28481362> ;
    cckp:species "Human" ;
    cckp:speciesTerm <http://purl.obolibrary.org/obo/NCIT_C14225> ;
    cckp:tumorType "Not Applicable" ;
    cckp:tumorTypeTerm <http://purl.obolibrary.org/obo/NCIT_C48660> ;
    cckp:version 1 .

The subject is the dataset's own Synapse IRI, not a second identifier. doi/pubMedId are raw portal values; doiIri/pubMedIdIri are resolvable IRIs templated from their shape.

tumorTypeTerm/speciesTerm/fileFormatsTerm are the real NCIT concept each value resolved to. grantNumberRef/pubMedIdRef are the two resolved joins, one Synapse-canonical and one locally minted.