Skip to content
Mindflow Marketing — home
Book a call Free Visibility Check
Results Pricing
Get my free Visibility Check Book a call
THE METHOD

How we work and measure: the Entity Integration Framework.

Five stages, run in order, measured in public. In owner language: name it, build it, get found, measure it, prove it. This page shows the machinery underneath.

Jump to the X/12 measurement protocol, published in full

FIVE STAGES, IN ORDER

What is the Entity Integration Framework?

The Entity Integration Framework is the five-stage method Mindflow Marketing runs on every engagement, in order: Identity, Architecture, Retrieval, Measurement, Outcome. Stage four, Measurement, runs first and last, because it sets the baseline and then proves the outcome. Every stage is published here so the work can be followed without a call.

Identity → Architecture → Retrieval → Measurement → Outcome

01Identity

Engines can only recommend a business they can verify. Stage one makes you unambiguous: one exact name, address, and phone everywhere; a hardened Google Business Profile; entity signals and schema that agree with each other; every conflicting listing hunted down and fixed.

Who you are, provable everywhere.
02Architecture

Your website gets structured so both people and machines can use it: pages organized around how customers actually buy, answer-first formatting engines can quote, clean technical foundations, and crawler access open where it earns you visibility.

A site built to be read and cited.
03Retrieval

This is where you become the answer: citable content on the pages that matter, presence on the surfaces engines actually consult, reviews at a steady honest cadence, and the citation trail AI assistants follow when someone asks who to call.

Show up where the answer gets formed.
04Measurement

Everything is scored against your baseline: map-grid position city by city and the rankings that matter for your trade in the monthly report, and Share of Answer quarterly — the 12 real buying questions, scored X/12, never a percentage. The Work Ledger records every completed task.

Three numbers. One ledger. No mystery.
05Outcome

Visibility only counts when the phone rings. Stage five ties the numbers to calls and jobs, prunes what isn't earning, doubles down on what is, and sets the next quarter's priorities, in writing.

More of the right calls, proven.
PUBLISHED IN FULL

What is the X/12 Protocol, and what does it measure?

The X/12 Protocol is the complete measurement protocol behind AI Share of Answer. It is published so that the number can be independently understood and reproduced: by you, by another agency, or by anyone who wants to check our work.

What is not published here: source code and automation credentials, client-confidential queries or data, internal QA checklists unrelated to the calculation, and vendor configuration that would introduce a security risk. None of those are needed to reproduce the number.

How are the twelve questions selected?

The twelve are drawn from what a buyer actually types or asks when they are close to hiring, not from keyword volume.

We build the candidate set from four sources: queries the client's own sales team reports hearing, the People Also Ask and related-question surfaces for the category, the phrasing competitors target in their own page titles, and the natural-language forms a buyer uses with an assistant, which are longer and more conversational than typed search.

From that candidate pool we select twelve that are commercial (the asker is choosing a provider, not researching a concept), answerable (a named business could legitimately appear in the response), and stable (the question will still be asked in six months). Twelve is a working compromise: enough for a percentage to mean something, few enough to run repeatedly by hand without sampling shortcuts.

How are buyer intent and commercial relevance validated?

A question only enters the set if a reasonable answer to it could include a business like the client's. “How does a heat pump work” fails that test, because it is informational, and a good answer names no companies. “Who installs heat pumps in Marietta” passes.

We validate this by running each candidate once before the set is locked and reading what comes back. If the response names no providers at all, the question is measuring something other than commercial visibility and is replaced. This validation pass is disclosed in the baseline report along with any questions that were rejected and why.

Which surfaces are tested?

Google AI Overviews, ChatGPT and Perplexity as the counted set, each scored and reported separately. A blended number that hides a zero on one platform conceals the single most useful fact in the report. Google AI Mode is spot coverage only and is not part of the counted set. Copilot is not sampled. We report Bing indexation and Bing Places listing state instead, because presence there is verifiable while a sampled score would not be. A protocol that never names what it excludes is not a protocol.

Are accounts, location and personalization controlled?

Yes, and this is where most published AI-visibility numbers quietly fall apart. Every run is executed signed out, in a fresh session with no conversation history, with browsing and memory features disabled where the platform allows it.

Location is set explicitly to the market being measured rather than inherited from the machine running the test, and the location used is recorded in the report. Where a platform does not permit location to be set deterministically, that limitation is stated against the affected surface rather than silently absorbed into the score.

When and how often do the tests run?

The baseline runs before any work begins. After that, the full protocol runs quarterly, on the same twelve questions, under the same conditions.

Quarterly rather than monthly is deliberate. These surfaces move on their own between runs for reasons unrelated to anything an agency did, and a monthly cadence produces noise a client would reasonably mistake for progress. The monthly report carries the other two numbers; the X/12 figure updates on the quarter, and is dated.

How do mentions, recommendations and citations differ?

These are three different things and collapsing them is the most common way an AI-visibility number is inflated. We score them separately:

The three outcomes scored separately under the X/12 Protocol v2.16, Mindflow Marketing, last updated 26 July 2026
Outcome What it means How it is counted
MentionThe business name appears in the response body.Reported alongside as a separate count
RecommendationThe response actively puts the business forward as an option for the asker.Counted in the headline X/12 figure
CitationThe business’s own domain is linked or named as a source underpinning the answer.Reported alongside as a separate count

A business can be cited without being recommended, and mentioned without either. The headline X/12 figure counts recommendations, because that is what a buying question is actually asking for. Mentions and citations are reported alongside as separate counts, never folded into the headline.

How is the X/12 score calculated?

X = the number of the twelve questions, on a given surface, where the business appears as a recommendation in the majority of runs. The denominator is always twelve. The result is reported per surface, never averaged across surfaces.

“Majority of runs” means at least two of three; see the next item. A business recommended in one run of three does not score; the appearance is recorded in the evidence file and noted as unstable, which is useful information but is not a point.

How are inconsistent answers and repeated runs handled?

Each question is run three times per surface, in separate sessions. These systems are non-deterministic and a single run is close to meaningless.

A question scores when the business is recommended in at least two of the three runs. Where results split, recommended once and absent twice, the question is logged as unstable and the split is shown in the report rather than resolved by rounding.

Instability is a finding in its own right: it usually means the business sits at the boundary of the model's consideration set, which is a different strategic position from being absent, and it is often the fastest thing to move.

How is evidence captured?

Every run is captured as a full-text response with a timestamp, the surface, the location setting used, and the run number. Nothing is summarised at capture time.

The evidence file is delivered with the report and belongs to the client. It exists so the number can be checked rather than trusted: a client, or a client's other agency, can read the raw responses and verify the score was counted correctly. That is the point of publishing a protocol at all.

What are the known limitations and sources of variance?

These systems change without notice. A model update can move results between runs for reasons that have nothing to do with any work performed. We do not claim otherwise, and a quarter where the number falls is reported as a number that fell, with what we know about why. Twelve questions is a sample, not a census. It describes the buying questions selected, not every possible question in the category. Personalization cannot be fully eliminated, only reduced. Signed-out, fresh-session, location-set conditions get close, and the residual variance is real. Three runs is a small sample. It is enough to distinguish stable presence from noise; it is not enough to produce a confidence interval, and we do not present one. Recommendation is a judgement call at the margin. Where a response mentions a business ambiguously, the call is made conservatively: ambiguous cases are scored as not-recommended, and flagged in the evidence file so the client can disagree.
Protocol version

X/12 Protocol v2.16 · last updated 26 July 2026. Changes to this protocol are versioned and dated. Where a change would break comparability with a client’s earlier baseline, the prior version is used for that client until the next baseline is re-run, and the report says which version produced the number.

MEASURE. FIX. BUILD. PROVE.

What does the framework look like in owner language?

The five stages map onto four owner words. Measure is your baseline (stage 4 runs first and last). Fix is Identity and the urgent parts of Architecture. Build is Architecture and Retrieval. Prove is Measurement and Outcome: the Mindflow Monthly, three numbers, and the Work Ledger. Same machine, plain words.

Rankings are never guaranteed. Anything we could not trace to a primary source is absent from this page, not estimated. The audit runs the same six layers described on the pricing page, and Share of Answer is scored quarterly.

Start with the free Visibility Check See public pricing
THE FRAMEWORK

The five stages, in order

Stage 2, Architecture →
Your website gets structured so both people and machines can use it, organised around how customers actually buy.…
Stage 1, Identity →
Engines can only recommend a business they can verify. Identity is the stage that makes you unambiguous.…
Stage 4, Measurement →
Everything is scored against your baseline, three numbers, one ledger, and a protocol published in full.…
Stage 5, Outcome →
Visibility only counts when the phone rings. Stage five ties the numbers to calls and jobs, and decides what happens n…
Stage 3, Retrieval →
This is where you become the answer, present on the surfaces engines actually consult when someone asks who to call.…
Book a call Free Visibility Check