Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
50 changes: 43 additions & 7 deletions context/0.0.1/sourcelume.jsonld
Original file line number Diff line number Diff line change
Expand Up @@ -2,15 +2,51 @@
"@context": {
"@version": 1.1,
"sl": "https://sourcelume.apache.org/ns/0.0.1#",
"dct": "http://purl.org/dc/terms/",
"schema": "https://schema.org/",
"spdx": "http://spdx.org/rdf/terms#",
"xsd": "http://www.w3.org/2001/XMLSchema#",

"id": "@id",
"type": "@type",

"name": "sl:name",
"description": "sl:description",
"version": "sl:version",
"createdAt": {
"@id": "sl:createdAt",
"@type": "http://www.w3.org/2001/XMLSchema#dateTime"
"ProvenanceRecord": "sl:ProvenanceRecord",

"identifier": "dct:identifier",
"name": "schema:name",
"version": "schema:version",
"license": {
"@id": "dct:license",
"@type": "@id"
},
"licenseScope": "sl:licenseScope",
"licenseNote": "sl:licenseNote",
"licenseCategory": "sl:licenseCategory",
"licenseHistory": "sl:licenseHistory",
"creator": "schema:creator",
"role": "sl:role",
"created": {
"@id": "dct:created",
"@type": "xsd:dateTime"
},
"added": {
"@id": "schema:datePublished",
"@type": "xsd:dateTime"
},
"contentCreated": {
"@id": "schema:dateCreated",
"@type": "xsd:dateTime"
},
"origin": "sl:origin",
"custodyChain": "sl:custodyChain",
"agent": {
"@id": "schema:agent",
"@type": "@id"
},
"action": "sl:action",
"startTime": {
"@id": "schema:startTime",
"@type": "xsd:dateTime"
}
}
}
}
26 changes: 26 additions & 0 deletions mappings/croissant-crosswalk.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
# Crosswalk: Sourcelume ↔ MLCommons Croissant

> **Status: draft.** Field-level mappings below are a starting proposal, not confirmed against
> the current Croissant release. Discuss on `[email protected]` before treating any row
> as final.

Croissant (`http://mlcommons.org/croissant/`) is a schema.org-based JSON-LD vocabulary for
describing ML datasets. Sourcelume records provenance/custody/licensing specifically, so most
Croissant terms (`distribution`, `recordSet`, `field`) have no Sourcelume equivalent — the overlap
is concentrated in dataset identity and licensing metadata, both of which Croissant itself borrows
from schema.org.

| Sourcelume field | Croissant / schema.org term | Notes |
|---|---|---|
| `identifier` | `schema:identifier` (on `cr:Dataset`) | Both typically hold a DOI or persistent URL. |
| `name` | `schema:name` | Direct match. |
| `license` | `schema:license` | Direct match — both expect an IRI or SPDX identifier. |
| `creator` | `schema:creator` | Same schema.org property, but Sourcelume's `creator` is an array of role-tagged (`originator`/`curator`/`distributor`) parties as of 0.0.1, rather than Croissant's single value. |
| `origin` | *(no direct equivalent)* | Croissant doesn't model a provenance narrative; stays Sourcelume-specific. |
| `custodyChain` | *(no direct equivalent)* | Custody-chain modeling is Sourcelume's core addition over Croissant. |

## Open questions

- Should a Sourcelume record be embeddable as a property of a Croissant `cr:Dataset` (e.g. a
`sourcelume:provenance` extension property), or should the two stay as separate, cross-linked
documents? Needs dev-list discussion.
16 changes: 16 additions & 0 deletions mappings/otdi-crosswalk.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
# Crosswalk: Sourcelume ↔ Open Trusted Data Initiative (OTDI)

> **Status: draft, lowest confidence of the three crosswalks.** OTDI's public vocabulary is less
> stable/documented than Croissant or SPDX as of this writing — treat every row here as a
> placeholder pending confirmation from the AI Alliance's published OTDI artifacts.

| Sourcelume field | OTDI concept (tentative) | Notes |
|---|---|---|
| `identifier` | dataset identifier | Needs confirmation of OTDI's canonical ID property name. |
| `origin` | provenance narrative / data source description | OTDI is reported to emphasize consent and collection-method disclosure more heavily than Sourcelume's 0.0.1 free-text `origin` field does. |
| `custodyChain` | chain-of-custody / handling event | Likely the strongest conceptual overlap between the two specs, but needs a real OTDI schema reference to map field-by-field. |

## Open questions

- Get a concrete OTDI schema/example on the dev list so this crosswalk can move past
conceptual-only mapping.
20 changes: 20 additions & 0 deletions mappings/spdx-ai-profile-crosswalk.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
# Crosswalk: Sourcelume ↔ SPDX AI Profile

> **Status: draft.** Needs review against the current SPDX 3.0 AI profile release before being
> treated as authoritative.

The SPDX AI Profile extends SPDX 3.0 with AI-specific classes (`ai:AIPackage`,
`ai:DatasetPackage`) and already has strong primitives for license expression, which Sourcelume
should reuse rather than duplicate.

| Sourcelume field | SPDX AI Profile term | Notes |
|---|---|---|
| `identifier` | `spdx:SpdxId` / external identifier on `DatasetPackage` | SPDX favors its own ID scheme; mapping needs a stable rule (e.g. always carry the original DOI as an `ExternalIdentifier`). |
| `license` | `spdx:licenseConcluded` / `spdx:licenseDeclared` | SPDX distinguishes *declared* vs. *concluded* license — Sourcelume's single `license` field is closer to `licenseDeclared` since Sourcelume doesn't adjudicate accuracy. |
| `creator` | `spdx:suppliedBy` / `spdx:originatedBy` | SPDX splits "who supplied this copy" from "who originated the data" as distinct relationships; Sourcelume's `creator` is an array of role-tagged parties (`originator`/`curator`/`distributor`) as of 0.0.1, which maps more directly now but still isn't a 1:1 term match. |
| `custodyChain` | *(no direct equivalent)* | Closest SPDX concept is a `Relationship` between elements, but there's no first-class chain-of-custody event type. |

## Open questions

- Should `license` split into `licenseDeclared`/`licenseConcluded` in 0.0.2 to match SPDX more
closely, or is the ambiguity intentional (Sourcelume publishes claims, not adjudications)?
4 changes: 3 additions & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,9 @@ dependencies = []

[project.optional-dependencies]
dev = [
"jsonschema",
# [format] pulls in rfc3987, needed so jsonschema actually enforces
# "format": "uri" instead of silently treating it as a no-op annotation.
"jsonschema[format]",
"pyld",
"rdflib",
"pyshacl",
Expand Down
147 changes: 134 additions & 13 deletions schema/0.0.1/sourcelume.schema.json
Original file line number Diff line number Diff line change
@@ -1,36 +1,157 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://sourcelume.apache.org/schema/0.0.1/sourcelume.schema.json",
"title": "Apache Sourcelume Record",
"description": "Minimal skeleton schema for an Apache Sourcelume record. Fields and constraints will be filled in as the specification develops.",
"title": "Apache Sourcelume Provenance Record (0.0.1)",
"description": "Minimal viable structural schema for a Sourcelume ProvenanceRecord: dataset identity, one or more role-tagged creators, one license claim, and an ordered custody chain. No signature block yet.",
"type": "object",
"required": [
"id",
"type",
"identifier",
"name",
"license",
"creator",
"created",
"added",
"contentCreated",
"origin",
"custodyChain"
],
"properties": {
"@context": {},
"id": {
"type": "string",
"description": "Unique identifier (IRI) for this record."
"format": "uri"
},
"type": {
"const": "ProvenanceRecord"
},
"identifier": {
"type": "string",
"description": "The record type."
"minLength": 1,
"description": "Identifier for the dataset being described (e.g. a DOI or persistent URL), not for this record."
},
"name": {
"type": "string",
"description": "Human-readable name for the record."
"minLength": 1
},
"version": {
"type": "string"
},
"description": {
"license": {
"type": "string",
"description": "Free-text description of the record."
"format": "uri",
"description": "IRI of the claimed license (e.g. an SPDX license page URL)."
},
"version": {
"licenseScope": {
"type": "string",
"minLength": 1,
"description": "Optional. Jurisdiction or scope in which the claimed `license` applies, for cases where a license (especially a public-domain claim) is jurisdiction-scoped rather than global. For example: 'US' for US judicial-work-product public domain, or 'jurisdictions with a life+70 copyright term where the author died before 1954' for a copyright-expiration public-domain claim. Omit for licenses that apply globally (e.g. CC0, MIT)."
},
"licenseNote": {
"type": "string",
"minLength": 1,
"description": "Optional. Free-text qualifier for the claimed `license` IRI, for cases where the IRI alone is materially misleading. The classic case is a `license` IRI that reflects a processing artifact rather than an intended selection (e.g. a code subset that filters to a copyleft license by a bug, so the license IRI reads copyleft even though the intent was permissively-licensed code). Use this field to flag the discrepancy so an auditor does not take the bare IRI at face value. Omit when the IRI faithfully represents the licensing situation."
},
"licenseCategory": {
"type": "string",
"enum": [
"open-source",
"open-data",
"public-domain",
"agreement-supplied",
"use-restricted",
"risk-based",
"other"
],
"description": "Optional. Category of the claimed `license`, so the license's licensing *model* is checkable at a glance rather than only by reading the IRI. The single `license` IRI cannot, by itself, distinguish a standard permissive open-source license (e.g. Apache 2.0, MIT) from a use-restricted license (open in IP grants but restricting specific harmful uses) from a risk-based license (tiered access by risk band). Categories: 'open-source' (standard permissive OSS, e.g. Apache 2.0), 'open-data' (standard permissive data license, e.g. ODC-By, CC0), 'public-domain' (CC Public Domain Mark or equivalent), 'agreement-supplied' (data shared under a bilateral agreement whose terms do not permit public sharing), 'use-restricted' (open IP grants but with use-based restrictions), 'risk-based' (tiered access by risk band), 'other'. Omit when the IRI is a well-known standard license whose category is obvious."
},
"licenseHistory": {
"type": "array",
"minItems": 1,
"description": "Optional. Record of prior licenses the dataset was distributed under, earliest first, when the dataset's own dataset-level license changed over time. The current license remains in the `license` field; this array records the superseded ones with their effective dates. Each entry is an object with `iri` (the prior license IRI) and `effectiveDate` (xsd:dateTime when that license became effective). Omit for datasets whose license has been stable since release. Example: a dataset released 2023-08 under a risk-based license, later switched to ODC-By, would carry `license` = ODC-By and `licenseHistory` = [{ iri: <prior-risk-based-license-IRI>, effectiveDate: 2023-08-01 }].",
"items": {
"type": "object",
"required": ["iri", "effectiveDate"],
"properties": {
"iri": {
"type": "string",
"format": "uri"
},
"effectiveDate": {
"type": "string",
"format": "date-time"
}
},
"additionalProperties": true
}
},
"creator": {
"type": "array",
"minItems": 1,
"items": {
"type": "object",
"required": ["type", "name"],
"properties": {
"type": {
"enum": ["schema:Organization", "schema:Person"]
},
"name": {
"type": "string",
"minLength": 1
},
"role": {
"type": "string",
"enum": ["originator", "curator", "distributor"],
"description": "Optional. Distinguishes who created the underlying data from who curated or redistributed it. Omit if the record has a single, undifferentiated creator."
}
},
"additionalProperties": true
}
},
"created": {
"type": "string",
"description": "Version of the record content, e.g. following semver."
"format": "date-time"
},
"createdAt": {
"added": {
"type": "string",
"format": "date-time",
"description": "Creation timestamp, ISO 8601."
"description": "When this dataset was added to the collection being described (e.g. when a source was incorporated into an aggregator corpus). Mirrors the `added` field commonly found in dataset datasheets."
},
"contentCreated": {
"type": "string",
"format": "date-time",
"description": "When the dataset's underlying documents/content were originally created (e.g. the historical date range of the texts). Mirrors the `created` field commonly found in dataset datasheets. Use the start of the range if a range applies."
},
"origin": {
"type": "string",
"minLength": 1,
"description": "Free-text description of where the dataset's underlying data came from."
},
"custodyChain": {
"type": "array",
"minItems": 1,
"description": "Ordered chain of custody events, earliest first. A single-hop dataset has exactly one entry.",
"items": {
"type": "object",
"required": ["agent", "action", "startTime"],
"properties": {
"agent": {
"type": "string",
"format": "uri"
},
"action": {
"type": "string",
"minLength": 1
},
"startTime": {
"type": "string",
"format": "date-time"
}
},
"additionalProperties": true
}
}
},
"required": ["id", "type"],
"additionalProperties": true
}
}
Loading
Loading