Synthetic Isn't a Dirty Word
Synthetic prompts are the only way to measure AEO performance consistently, not a shortcut around real consumer intent.

Bluefish Team
Share
There's a lot of noise in the agentic marketing space right now. The discourse is moving fast and rewards confidence over accuracy. Vendors make claims they can't substantiate with methodologies that sound rigorous but collapse under scrutiny.
This year, OpenAI and Anthropic extended memory features to users on their free tiers, enabling AI assistants to build profiles for anyone with a ChatGPT or Claude account. With these profiles, AI systems deliver customized answers that vary wildly from user to user, even when given the same prompt.
Other AI monitoring tools use a fixed set of popular prompts to measure a brand’s AI visibility, but in a channel where responses vary by user, this approach doesn’t give marketers a ton of insight to invest in. Data collected through a faulty methodology reveals nothing meaningful about why a brand is winning or which audiences matter.
This is why Bluefish generates synthetic prompts, tens of thousands of representative simulations of real-world consumer intent, to capture the true depth and variety of AI responses.
In AI, “synthetic” is sometimes used as a pejorative, with the implication that prompts designed by an agentic marketing platform are less valid than prompts typed by a human consumer. The argument is that real measurement requires watching real queries in real time, and while that logic may seem to make sense on the surface, it fundamentally misunderstands what credible measurement requires and how AI thinks.
Rigorous measurement requires a highly defined, controlled baseline, so that changes to that baseline can be accurately attributed. The question worth asking about any AI measurement approach is whether the baseline is credible, stable, transparent, and grounded in real-world demand, which can only be accomplished with synthetic prompts.
Unlike traditional search, AI doesn't react to syntax or keyword matching. Instead, AI responds to intent and context. When a consumer asks, "what's the best sunscreen" or "recommend a sunscreen for sensitive skin," the model attempts to understand the intent behind the message, rather than focusing on the prompt’s exact phrasing. Bluefish builds prompt configurations around those underlying intents, which are the real signals driving AI recommendations.
Those configurations are built from a three-layer foundation:
Proprietary AI behavioral data: The topics, categories, and contexts that AI models surface in recommendations across ChatGPT, Gemini, Claude, and others. This data represents what AI actually thinks about a category.
Google Search demand: Every topic and category is validated against real consumer search volume, including Google AI Overviews and AI Mode.
The brand itself: Customer teams review and approve every prompt set against their product priorities, target audiences, and competitive landscape, before configurations are finalized.
The result is a baseline built in partnership with customers to be credible, transparent, grounded in real-world demand, and measured consistently over time.
That consistency is what makes the data meaningful. When a measurement system even slightly changes its prompt inputs, isolating what drove any change to the results becomes impossible. Did a brand’s visibility improve because their content got better, or maybe it was because the AI model updated? Rotating sets of prompts cannot answer these questions. With Bluefish, only one variable changes: the brand's actual performance.
Bluefish designs synthetic prompts that are controlled, reproducible, and help brands look inside the “mind” of AI to produce data with real impact.


