Standards

The provenance standard: every row says where it came from.

AI search is full of confident numbers with no sourcing behind them. Citabld publishes the opposite standard: every subquery in every library carries a label describing how it was produced, and that label survives into scoring, reporting, and the build sheet a client signs.

Definition

A provenance label is a required field on every subquery row in a Citabld library that records how the row was produced: MEASURED from an engine surface, DETERMINISTIC from a reproducible rule, EXTRAPOLATED from a comparable engine, or MODELED from behavioral profiles. Labels prevent projections from being read as observations.

The four labels

MEASURED

Captured directly from an engine surface that exposes its own retrieval, such as Perplexity's displayed search steps or a citation panel returned with the answer.

The only label that supports a claim about what an engine actually did.

DETERMINISTIC

Produced by a documented, reproducible rule rather than an observation, such as expanding a query template across a known product line or geography set.

Safe to act on, because anyone can regenerate the same rows from the same inputs.

EXTRAPOLATED

Derived from measured behavior on one engine and projected onto a similar engine with published architectural similarity.

Directionally reliable, weaker than measured, and always flagged with the source engine.

MODELED

Generated from behavioral profiles when the engine exposes no retrieval trace at all, which is the case for Gemini and Google's AI surfaces.

Useful for coverage planning, never presented as evidence of engine behavior.

The four labels

A number without provenance is a decoration.

MEASURED

The only label that supports a claim about what an engine actually did.

DETERMINISTIC

Safe to act on, because anyone can regenerate the same rows from the same inputs.

EXTRAPOLATED

Directionally reliable, weaker than measured, and always flagged with the source engine.

MODELED

Useful for coverage planning, never presented as evidence of engine behavior.

Why the standard exists

The AI visibility market grew faster than its evidence base. Tools report share of voice across engines that publish no retrieval data, and agencies present synthesized query lists as observed engine behavior. Both are guesses wearing the clothing of measurement. When a client cannot tell which rows were observed and which were generated, they cannot allocate budget rationally, and they cannot tell whether a result came from the work or from noise.

Labeling fixes that at the row level rather than in a disclaimer nobody reads. A theme built on twelve measured rows and three modeled rows is a different investment than a theme built on fifteen modeled rows, even when both look identical in a coverage chart. Our scoring treats them differently, our reports state the mix, and our recommendations say plainly when a decision rests on a model.

What is a provenance label?

A provenance label is a required field on every subquery row in a Citabld library that records how that row came to exist: MEASURED, DETERMINISTIC, EXTRAPOLATED, or MODELED. The label travels with the row into scoring, into the build sheet, and into every deliverable, so no reader can mistake a projection for an observation.

Why can't every subquery be measured?

Because most engines do not publish their internal decomposition. Perplexity shows its search steps, and some surfaces return citation sets that reveal what was retrieved, but Google's AI Mode and AI Overviews, Gemini, and most assistant surfaces expose no query logs to anyone, including their own advertisers. Any vendor selling you measured Google fan-out data is selling you a model with the label removed.

How does provenance change what gets built?

Priority arithmetic weights measured rows above modeled ones, so production capital flows first to themes with observable evidence behind them. Modeled rows still earn coverage, because a theme with no content cannot be cited under any engine, but they never justify a claim in a report.

See it applied

Every row in your library will say where it came from.