Skip to content

Docs: explain where the data comes from and how sources are arbitrated - #43

Merged
lambda2 merged 1 commit into
masterfrom
docs/understanding-the-data
Sep 5, 2026
Merged

Docs: explain where the data comes from and how sources are arbitrated#43
lambda2 merged 1 commit into
masterfrom
docs/understanding-the-data

Conversation

@lambda2

@lambda2 lambda2 commented Sep 5, 2026

Copy link
Copy Markdown
Member

A page for the reader who wants to judge whether to trust a value, without reading any code.

Why a second page on this

Data provenance documents the facts endpoint — it answers where did this value come from for someone holding an API token and reading JSON. It does not answer what a botanist or a data consumer asks first: what are these sources, why does one beat another, and how much of this can I rely on?

This is the editorial companion, and it sits first in the Advanced section so it is read before the field reference.

What it covers

  • The sources, what each one actually contributes, and roughly how much of the dataset it covers — including that Baseflor is ~6 400 species of western European flora, not a global layer
  • Why Wikipedia never feeds a numeric field: an article quoting a height is quoting a source, and we would rather cite that source
  • The four principles behind the ranking — human review over automated import, measurement over assertion, regional flora over global aggregator within its region, nomenclatural authority over computed consensus
  • That the ranking is an editorial judgement, stated plainly, including its known limitation: it is global where authority is often field-specific
  • What the tricky fields mean, in particular that ecological indicators describe habitat, not tolerance, cover only the temperate European flora, and that 0 is a value there rather than a blank — the exact confusion behind trefle-api#90
  • Why a range never becomes an average: a midpoint would be our arithmetic, not anyone's observation
  • An honest account of coverage: nomenclature is in good shape, botanical traits are largely empty, and the page says so — "the database is not wrong on these fields; it is empty, which is a different problem and, in our view, a more honest one"
  • What a low completeness figure means, so it is not read as unreliability

Verification

Build green, new page generated, both internal links resolve (onBrokenLinks defaults to throw in Docusaurus 3, so a bad link would have failed the build).

The provenance page documents the facts endpoint, which answers 'where
did this value come from' for a developer holding an API token. It does
not answer the question a botanist or a data consumer asks first: what
are these sources, why does one win over another, and how much of this
can I actually rely on.

This is the editorial companion. No code, no JSON: the sources and what
each contributes, the four principles behind the ranking (and the plain
statement that it is an editorial judgement applied globally rather than
field by field), what the less obvious fields mean, and an honest account
of coverage — nomenclature is solid, botanical traits are largely empty,
and that is a gap rather than an inaccuracy.

Says explicitly that ecological indicators describe habitat rather than
tolerance, that they only cover the temperate European flora, and that 0
is a value there and not a blank — the confusion behind trefle-api#90.
Also states why a range never becomes an average, and why Wikipedia never
feeds a numeric field.
@lambda2
lambda2 merged commit 180e538 into master Sep 5, 2026
1 check passed
@lambda2
lambda2 deleted the docs/understanding-the-data branch September 5, 2026 06:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant