One head query does not have one fan-out. It has a space of them.
The library is the core artifact of AI Query Architecture and the deliverable clients keep. For every head query we generate the subquery space across nine engine behavior profiles and three to six persona cells, then cluster the rows into the 30 to 60 themes that actually decide retrieval. Generation runs to saturation, not to a target count, because padded sets are noise sold as thoroughness.
A subquery library is a database of the smaller questions AI engines generate from a brand's main buyer questions. Citabld builds one across nine engines and every buyer persona, labels each row with how it was obtained, and groups the rows into scored themes. The library serves as both the content plan and the measurement list.
Why nine engines instead of one list?
Because every engine decomposes differently, and a single generic subquery list is wrong on eight of them. Each engine is parameterized from measured behavior rather than described in prose, so a row generated for Copilot looks nothing like a row generated for ChatGPT.
| engine | measured behavior | what it demands |
|---|---|---|
| chatgpt | high lexical drift, near-total run variance, memory resident | broad coverage across phrasings the user never typed |
| perplexity | near-verbatim fragments, highly stable | classic keyword discipline still transfers; our calibration anchor |
| copilot | compresses questions into short keyword strings | keyword-dense content wins, persona effects weakest |
| claude | hands back its search index rankings largely untouched | Brave visibility is the most direct citation lever that exists |
| google ai mode | largest fan-outs, eight to twelve subqueries standard | fragment-level ranking coverage, never single-run readings |
| ai overviews | similar conclusions to AI Mode, largely different URLs | scored as a separate surface, never merged into 'Google' |
| gemini | deepest personalization substrate, behavioral | full persona matrix |
| grok | iterative loop, generates follow-ups from what it finds | second-order content, and X discourse as a source |
| deepseek | retrieves nothing by default | clean parametric probe; never enters citation metrics |
What is a persona cell?
A persona cell is one validated buyer profile that the library generates fragments for. We model six dimensions for business buyers: role and seniority, company size, geography, vertical, technical sophistication, and constraint posture, chosen because they are the dimensions that measurably survive engine rewrites. Location survives almost always. Price constraints frequently do not.
Validation before generation
Cells are validated against closed-won deals or mined reviews before they drive generation, so the matrix reflects your actual buyers rather than a marketing persona deck.
Committees and veto paths
We map the committee, because the champion's ChatGPT can be convinced while the CFO's Copilot un-convinces. Veto paths get their own cells.
One list of subqueries is a guess. Nine engine profiles and your persona cells are a map.
Every row carries its provenance label, so you always know which fragments were observed and which were modeled.
What is the Citabld Decomposition Calculus?
The Citabld Decomposition Calculus (CDC) is our named framework for deriving a head query's subquery space instead of guessing at it. The insight it is built on: engines do not generate subqueries from the question, they generate them from the answer. CDC parses the head query into a semantic frame, blueprints what a complete answer must contain, then applies a closed set of expansion operators to produce the fragments, conditioned per engine and per persona. Because the operators are content-free machinery and the query supplies the content, the same calculus handles any head query in any industry, and it is falsifiable: every fragment captured from a live engine either traces back to an operator or grows the algebra.
Derived, not recalled
There is no pattern library being pasted in. Each library row is the output of a specific operator applied to your query's frame, which is why the same head query provably yields different spaces for different buyers.
Modeled honestly
CDC also models what engines destroy: price constraints mutate, brand sets compress, geography survives. The library covers the post-attrition forms, not just the pretty originals.
What does a provenance label mean?
Every row carries one of four labels describing how it was obtained. Three major surfaces expose no query logs to anyone, so their rows are always MODELED and always say so. This is not a disclaimer at the bottom of a report. It is the spine of the product, because you cannot make good decisions on data that hides what it is.
| label | meaning |
|---|---|
| MEASURED | captured from a live engine |
| DETERMINISTIC | follows from a documented pipeline, such as Claude's Brave inversion |
| EXTRAPOLATED | derived from patent taxonomy or vendor description |
| MODELED | generated from our engine profiles, calibrated against the observable engines |
What do you actually receive?
A working database, not a PDF. Every row is a fragment with its engine, persona cell, journey stage, provenance label, and theme cluster. Every theme carries its scores. The same file is the content plan for production and the tracked-prompt inventory for measurement, which is why the library keeps paying after the engagement.
- /500 or more subquery rows, generated to saturation across your head queries.
- /30 to 60 theme clusters, each scored for coverage, winnability, and priority.
- /Every row labeled MEASURED, DETERMINISTIC, EXTRAPOLATED, or MODELED.
- /Journey stage tags that determine which metric is honest for each theme.
- /Yours to execute with any team, including one that is not us.
See this phase run on one of your own head queries, free.
Six questions, one modeled fan-out across seven engines, and one finished piece of content for the subquery you pick.