Method / Step 02 · Updated August 2026

One head query does not have one fan-out. It has a space of them.

The library is the core artifact of AI Query Architecture and the deliverable clients keep. For every head query we generate the subquery space across nine engine behavior profiles and three to six persona cells, then cluster the rows into the 30 to 60 themes that actually decide retrieval. Generation runs to saturation, not to a target count, because padded sets are noise sold as thoroughness.

Definition

A subquery library is a database of the smaller questions AI engines generate from a brand's main buyer questions. Citabld builds one across nine engines and every buyer persona, labels each row with how it was obtained, and groups the rows into scored themes. The library serves as both the content plan and the measurement list.

Why nine engines instead of one list?

Because every engine decomposes differently, and a single generic subquery list is wrong on eight of them. Each engine is parameterized from measured behavior rather than described in prose, so a row generated for Copilot looks nothing like a row generated for ChatGPT.

enginemeasured behaviorwhat it demands
chatgpthigh lexical drift, near-total run variance, memory residentbroad coverage across phrasings the user never typed
perplexitynear-verbatim fragments, highly stableclassic keyword discipline still transfers; our calibration anchor
copilotcompresses questions into short keyword stringskeyword-dense content wins, persona effects weakest
claudehands back its search index rankings largely untouchedBrave visibility is the most direct citation lever that exists
google ai modelargest fan-outs, eight to twelve subqueries standardfragment-level ranking coverage, never single-run readings
ai overviewssimilar conclusions to AI Mode, largely different URLsscored as a separate surface, never merged into 'Google'
geminideepest personalization substrate, behavioralfull persona matrix
grokiterative loop, generates follow-ups from what it findssecond-order content, and X discourse as a source
deepseekretrieves nothing by defaultclean parametric probe; never enters citation metrics

What is a persona cell?

A persona cell is one validated buyer profile that the library generates fragments for. We model six dimensions for business buyers: role and seniority, company size, geography, vertical, technical sophistication, and constraint posture, chosen because they are the dimensions that measurably survive engine rewrites. Location survives almost always. Price constraints frequently do not.

Validation before generation

Cells are validated against closed-won deals or mined reviews before they drive generation, so the matrix reflects your actual buyers rather than a marketing persona deck.

Committees and veto paths

We map the committee, because the champion's ChatGPT can be convinced while the CFO's Copilot un-convinces. Veto paths get their own cells.

Step 02 · Library

One list of subqueries is a guess. Nine engine profiles and your persona cells are a map.

Every row carries its provenance label, so you always know which fragments were observed and which were modeled.

What is the Citabld Decomposition Calculus?

The Citabld Decomposition Calculus (CDC) is our named framework for deriving a head query's subquery space instead of guessing at it. The insight it is built on: engines do not generate subqueries from the question, they generate them from the answer. CDC parses the head query into a semantic frame, blueprints what a complete answer must contain, then applies a closed set of expansion operators to produce the fragments, conditioned per engine and per persona. Because the operators are content-free machinery and the query supplies the content, the same calculus handles any head query in any industry, and it is falsifiable: every fragment captured from a live engine either traces back to an operator or grows the algebra.

Derived, not recalled

There is no pattern library being pasted in. Each library row is the output of a specific operator applied to your query's frame, which is why the same head query provably yields different spaces for different buyers.

Modeled honestly

CDC also models what engines destroy: price constraints mutate, brand sets compress, geography survives. The library covers the post-attrition forms, not just the pretty originals.

What does a provenance label mean?

Every row carries one of four labels describing how it was obtained. Three major surfaces expose no query logs to anyone, so their rows are always MODELED and always say so. This is not a disclaimer at the bottom of a report. It is the spine of the product, because you cannot make good decisions on data that hides what it is.

labelmeaning
MEASUREDcaptured from a live engine
DETERMINISTICfollows from a documented pipeline, such as Claude's Brave inversion
EXTRAPOLATEDderived from patent taxonomy or vendor description
MODELEDgenerated from our engine profiles, calibrated against the observable engines

What do you actually receive?

A working database, not a PDF. Every row is a fragment with its engine, persona cell, journey stage, provenance label, and theme cluster. Every theme carries its scores. The same file is the content plan for production and the tracked-prompt inventory for measurement, which is why the library keeps paying after the engagement.

  • /500 or more subquery rows, generated to saturation across your head queries.
  • /30 to 60 theme clusters, each scored for coverage, winnability, and priority.
  • /Every row labeled MEASURED, DETERMINISTIC, EXTRAPOLATED, or MODELED.
  • /Journey stage tags that determine which metric is honest for each theme.
  • /Yours to execute with any team, including one that is not us.
Next · Score

See this phase run on one of your own head queries, free.

Six questions, one modeled fan-out across seven engines, and one finished piece of content for the subquery you pick.