Editor’s note from Mike: Today’s piece is a collab with our friend Tom Critchlow, forward deployed engineer at Alephic and founder of AI Search Leaders. It’s written in the singular first-person because we effectively became one brain on it.
As a reminder, Tom is taking the stage at BRXND NYC 2026 on November 5 to provide a state of AI brand visibility you won’t want to miss. We look forward to seeing you at The Times Center in 76 days. If you haven’t bought a ticket already, the button below beckons.
The AI Search Maturity Matrix
I spoke with a company recently that spends $500k / year on prompt tracking. They track thousands of prompts and there’s an executive-wide email that goes out every week with a “citation rate” line chart that wiggles.
Meanwhile, skyrocketing DataforSEO spend on Ramp suggests that more nimble companies are moving away from enterprise GEO tools and building their own custom solutions in house.
But so what? Prompt tracking isn’t a strategy.
How your brand is perceived by AI models feels (correctly!) existential to many CEOs and CMOs. A core job of the CMO in this moment is presenting a coherent strategy for how to evolve a brand in a world where LLMs increasingly mediate brand preference.
The challenge of building an actual AI search strategy is that you need a way to allocate resources beyond the time horizon of the models, a strategy can’t fluctuate every time a new model drops. The real task, then, is to build AI Search as an organizational capability rather than a dashboard. That’s where the concept of a maturity matrix comes in.
The Five Levels of AI Search Maturity
To say this another way, the core question marketing leaders should be asking is “how can we evolve our AI search maturity?”
AI Search maturity is the organizational ability to turn noisy model outputs into foundational insights about the information ecosystem, then turn that into durable market influence.
To that end, we’ve built an organizational maturity model to help you understand how sophisticated you are - and how you might get better. The model has three dimensions: measurement capability, executive knowledge and strategic posture. Companies are rarely advancing equally across all dimensions.
Hopefully this helps you have a conversation with your leadership team and think about not just problem solving but capacity building.
Level 1: Instrumented - “how visible are we in AI search?”
This is where everyone was 12-18 months ago. You’ve loaded up a set of prompts into a prompt tracking engine and can answer basic questions like “are we more or less visible than our competitors”?
Basic prompt tracking has a number of key problems and challenges though:
Citations ≠ Mentions ≠ Recommendations - AI search is qualitatively different than traditional search and knowing your ranking position is far less important than knowing HOW you’re recommended. A brand can appear as a citation, but not appear at all in the recommendation or worse, become an ANTI-recommendation.
You need to aggressively segment your prompts - not all prompts are created equal. Branded vs unbranded, aided vs unaided prompts, different stages of the buyer journey etc. You need to be careful about how you setup your prompt types.
I think more or less every executive team knows AI search is important at this point but at this level you’re still in a kind of vague “change is coming” mindset and have trouble really pinpointing how the new world is different from the old world.
Level 2: Contextual - “which customer moments and sources actually matter?”
Two of the most important characteristics of AI search in my mind is that you can use AI search for more of the customer journey than you could regular search, and that it’s not just ranking answers but actually giving you qualitative advice.
So your setup has to grapple with this reality with:
Prompt segmentation that maps to your customer journey. The good news is that most brands already have a reasonable notion of their customer journey from brand marketing - so the first step is to map your prompt tracking to that journey and to use the same kind of language the brand team uses (aided vs unaided awareness, likelihood to recommend etc)
Qualitative analysis of the prompt responses. Beyond just knowing how often you show up you need to be probing the responses for qualitative insights - are we being recommended strongly? Which gaps exist where we’re being cited but not recommended?
At this stage your leadership should be able to articulate gaps - a difference in how the brand shows up in AI search vs regular search. You should be able to point to areas of weakness and strength.
Level 3: Operational - “Can our organization act on what we learn?”
Once you know where the gaps are (level 2) the next level is acting on them. Most AI search problems are consensus problems across the internet. Do we have a reputation and brand awareness that’s strong and repeated everywhere?
The strategic work becomes coordination. Brand, content, product data, PR, community and technical teams organize around the information gaps the company has found. Some gaps require a new page while others might require better product data, a clearer position, stronger third-party evidence or fixing the underlying customer experience.
Citation rate and mention rate varies tremendously in AI search but measuring consensus is a little more long-term. To do that your measurement stack has to be able to probe the qualitative prompt responses not just for sentiment but for coherence: are our brand attributes correct everywhere? Is it mentioning the brand values that we want to repeat everywhere?
It also starts mapping the information ecosystem producing the answers. Is the model drawing from the company website, a retailer feed, Reddit threads, review sites, old press coverage, a product database or a competitor comparison page? Where does the ecosystem agree with the company’s intended positioning, and where does it tell a different story?
In my experience very few companies are actually at level 3. Most are at level 2, aspiring to become level 3.
Level 4: Exploratory - “Are we paying attention to how AI search is changing?”
Despite all of the hype about AI search, I still don’t think people appreciate how much better the models have gotten. AI has crossed the chasm of building a truly superior search experiences for everyday consumer use cases than ten blue links. Recently on vacation I had ChatGPT using skills, operating on scheduled actions and providing a level of recommendation and usefulness that regular search can’t touch. I used Google Maps a bunch but I can’t recall using Google once.
Getting cited and mentioned is one part of AI search today but I think you should be paying attention to the frontier too. Click based analytics is effectively dead vaporware. Brands can be discovered and recommended without a single site visit. Agents can autonomously research and purchase products.
No one is ready for this world but smart brands should be investing in exploration. Some areas that I might invest in here today:
WebMCP:- While agentic commerce protocols still feel early, clunky and not widely used, the idea of giving every website a structure, agent-native access point will really upend a lot of things. It feels like we’ll see some major innovation here in the concept of a webMCP before the end of the year. (one to watch: Cloudflare can now give every site a webmcp endpoint)
Model memory vs agentic web search - beyond mention/citation rate, it’s important to understand how the models understand your brand and your products in their training data vs what they learn from agentic web search. Open source models can provide raw thinking traces and can shine a light on various brand attributes associated with your brand and allow you to probe model memory more directly
I only know a small handful of people operating at this level.
Level 5: Experimental - “Can we prove that our interventions work?”
Level 5 is the crown jewel. Can you actually test and iterate on AI search as a field? This is harder than it sounds given the volatility in the space.
There are two promising pathways I’ve seen here:
First, using a platform like SearchPilot (disclaimer: my brother runs it) you can actually construct A/B tests for site changes to understand what moves the needle. True for regular SEO and increasingly for AI search too.
Secondly - you can build simulators that enable you to actually feed different context into the models and see how they react. This allows you to estimate the impact of changing owned or influenced properties and seeing how it changes the likelihood to recommend.
We’ve been playing with this approach to build simulation environments and see how changing descriptive language changes the likelihood to recommend and whether brand / product features influence the models.
Here’s a glimpse at the output - we took a real hotel and ran hundreds of simulated AI recommendation queries against four real competitors, changing only the wording of the listing each time. Reordering the copy to lead with what business travelers care about took the model’s recommendation rate from 16% to 78%, while removing the brand names barely moved it: the words, not the brand, were driving the choice.
More on this in later newsletters but get in touch if you want to explore this for your brand!
Where is your brand on the maturity matrix? Where is this framework wrong? How would you extend it? We’d love to hear from you.
If you have any questions, please be in touch. As always, thanks for reading.
— Mike










