Kojable research · Persona conditioning · AI search queries
Do Persona Prompts Change How AI Searches? Evidence from 1,937 Grounded Search Queries
Across 149 matched business scenarios, five persona conditions produced strongly differentiated grounded search queries. Query differences tracked answer differences, but publisher-level source selection remained unmeasurable.
Key finding
When the same underlying business scenario was answered under five different persona conditions, Gemini reported substantially different grounded search-query formulations.
Across 149 complete matched scenarios, exact normalized query sets were almost entirely different across persona pairs. Token-level divergence averaged 0.682, and TF-IDF lexical divergence averaged 0.887.
The persona patterns visible in the final answers were also visible in the search queries: SEO Manager queries were much more technical, Demand Generation queries much more commercial, and CMO queries much more governance-oriented.
Qualification: This study measures recorded search-query formulation in a historical grounded Gemini run. It does not establish that persona instructions caused the query differences, that different publishers were retrieved, or that different searches were better searches.
- Query formulation
- Persona conditioning
- 1,937 usable queries
- 12 min read
At a glance
Research snapshot
| Study element | Detail |
|---|---|
| Research question | When the same scenario is framed for different personas, does Gemini formulate different grounded search queries? |
| Parent study | 150 matched scenarios × 5 persona conditions |
| Primary complete scenarios | 149 |
| Substantive/query-valid responses | 749 |
| Recorded usable query entries | 1,937 |
| Primary persona-pair rows | 1,490 |
| Exact query-set divergence | 0.9991 |
| Token divergence | 0.6822 |
| TF-IDF lexical divergence | 0.8870 |
| Source records audited | 9,799 |
| Reliable publisher identities | 0% |
| Primary decision | Query-only differentiation |
| Major limitation | Persona condition remained confounded with collection position |
Direct answer
Do persona prompts change how Gemini searches?
Answer
In this historical matched experiment, yes at the recorded query-formulation level.
The same neutral business scenario, presented under different persona conditions, usually produced different search queries.
Literal query overlap was almost nonexistent, and the differences remained substantial when measured using token and lexical representations rather than exact-string matching.
But the study does not establish the full retrieval pathway.
The stored source URLs were unresolved grounding wrappers, so we could not determine whether the different queries led to different publishers or source pools.
The publication-safe conclusion is:
Persona-associated differentiation appeared in the grounded search queries Gemini reported, not only in the final answer text.
Research question
Why this question matters
Grounded AI systems do more than generate prose.
When search is enabled, the model may formulate one or more queries before or during answer construction.
That creates an observable layer between:
prompt
and:
final answer.
For AI visibility, AEO and GEO research, this layer matters because two prompts that lead to different search formulations may expose different information needs—even when they originate from the same underlying business question.
But query telemetry is easy to overinterpret.
A different query does not automatically mean:
- A different publisher was retrieved
- A different source influenced the answer
- The model performed deeper research
- The query was better
- The final answer was more accurate
Our earlier research on Gemini query fan-out asked a different question:
When similar prompts are phrased differently, does Gemini search for similar information?
That study found a strong relationship between prompt similarity and query-set similarity.
This study holds the underlying scenario constant and changes the persona condition.
The question is narrower:
Does the decision context supplied by the persona correspond to a different recorded search formulation?
Study design
Study design
The parent experiment began with 150 neutral business scenarios.
Each scenario was rendered under five persona-conditioning bundles:
- CMO
- VP of Product Marketing
- SEO Manager
- Demand Generation Director
- Founder/CEO
Gemini 3.6 Flash answered with Google Search grounding enabled.
A later quality audit excluded one malformed response, leaving 749 substantive responses.
Primary matched query comparisons used the 149 complete scenarios with all five persona conditions present.
Each complete scenario creates ten unordered persona pairs, giving:
149 × 10 = 1,490 primary pair comparisons.
Across the substantive responses, the grounded output contained 1,937 usable search-query entries.
All 1,937 passed the query-quality screen.
There were:
- No within-response duplicate query entries
- Only nine normalized queries shared across persona conditions within matched scenarios
The analysis compared query sets at several levels because exact-string comparison alone can create a misleading ceiling.
Finding 1
The same scenario almost never produced the same literal search queries across personas
The first comparison used normalized exact query sets.
The metric is:
1 − query-set Jaccard similarity
where:
- 0 means identical query sets
- 1 means no exact normalized query overlap
Mean divergence was:
0.9991
with a 95% scenario-bootstrap interval of:
0.9978–0.9999.
That means exact normalized search-query overlap across persona conditions was almost nonexistent.
But this result needs careful interpretation.
Exact query strings are brittle.
Two queries can use different strings while expressing closely related information needs.
So exact divergence establishes that literal queries were different.
It does not tell us how different their content was.
Finding 2
The query differences remained substantial after moving beyond exact strings
To measure degree rather than literal identity, we used two graded representations.
Token-set divergence
Mean divergence:
0.6822
95% CI:
0.6710–0.6933
TF-IDF lexical divergence
Mean divergence:
0.8870
95% CI:
0.8798–0.8938
| Query representation | Mean divergence | 95% CI |
|---|---|---|
| Exact normalized query sets | 0.9991 | 0.9978–0.9999 |
| Token sets | 0.6822 | 0.6710–0.6933 |
| TF-IDF lexical | 0.8870 | 0.8798–0.8938 |
This matters because the result is not simply:
“The strings were different.”
The recorded queries also differed substantially in their vocabulary and formulation.
Finding 3
Different persona conditions produced different search-language signatures
The persona differences visible in final answers were also visible upstream in the query language.
Using a fixed vocabulary codebook defined before the retrieval outcomes were inspected, we compared technical, commercial and governance-oriented language in the recorded queries.
| Persona condition | Mean queries | Mean query words | Technical / 1k words | Commercial / 1k | Governance / 1k |
|---|---|---|---|---|---|
| CMO | 2.68 | 8.44 | 9.09 | 30.28 | 35.90 |
| Demand Generation Director | 2.46 | 8.45 | 6.22 | 75.64 | 3.52 |
| Founder/CEO | 2.33 | 8.53 | 9.54 | 34.98 | 10.02 |
| SEO Manager | 2.90 | 8.18 | 74.92 | 17.77 | 1.06 |
| VP of Product Marketing | 2.56 | 8.23 | 8.83 | 33.16 | 4.32 |
SEO Manager queries were much more technical
SEO Manager recorded roughly 74.9 technical terms per 1,000 query words, compared with single-digit rates for the other four conditions.
Demand Generation queries were much more commercial
Demand Generation recorded approximately 75.6 commercial terms per 1,000 query words.
CMO queries were much more governance-oriented
CMO recorded approximately 35.9 governance terms per 1,000 query words.
These are not measurements of how real SEO managers, demand-generation leaders or CMOs search.
They are properties of the model's recorded query formulations under the historical persona-conditioned prompt bundles.
Pair distances
Which persona pairs searched most differently?
Using lexical divergence:
| Persona pair | Lexical query divergence |
|---|---|
| Demand Generation Director / SEO Manager | 0.9281 |
| SEO Manager / VP Product Marketing | 0.9252 |
| CMO / SEO Manager | 0.9182 |
| Founder/CEO / SEO Manager | 0.9162 |
| Demand Generation Director / VP Product Marketing | 0.8710 |
| CMO / VP Product Marketing | 0.8684 |
| CMO / Demand Generation Director | 0.8662 |
| Demand Generation Director / Founder/CEO | 0.8655 |
| Founder/CEO / VP Product Marketing | 0.8612 |
| CMO / Founder/CEO | 0.8505 |
The strongest observed lexical query divergence was:
Demand Generation Director / SEO Manager
The weakest was:
CMO / Founder/CEO
That mirrors the broader persona programme.
SEO Manager tended to occupy the most distinctive region of the response and query space.
CMO and Founder/CEO were repeatedly among the closest conditions.
Again, “closest” and “farthest” describe representational distance—not quality or real-role similarity.
Finding 4
More different search language was associated with more different answers
The parent analysis had already created a persona-residual answer distance after masking direct persona labels and profile phrases.
We joined those answer distances to the query divergences for the same:
scenario × persona pair
This let us ask:
When two persona conditions formulate more different searches, do they also tend to produce more different final answers?
Exact query divergence was almost saturated.
Its within-scenario association with residual answer distance was:
r = −0.012
95% CI:
−0.043 to 0.016
That is essentially no association.
The graded query measures were different.
Token divergence ↔ residual answer distance
r = 0.514
95% CI:
0.459–0.566
Lexical query divergence ↔ residual answer distance
r = 0.495
95% CI:
0.436–0.553
| Query metric | Association with residual answer distance |
|---|---|
| Exact query-set divergence | −0.012 |
| Token divergence | 0.514 |
| Lexical divergence | 0.495 |
The relationship also appeared descriptively across the ten persona-pair profiles.
Pair-profile Spearman correlations were:
- Token divergence: 0.939
- Lexical divergence: 0.818
In other words, persona pairs that were more differentiated in query language also tended to be more differentiated in their final residual answers.
This is an association.
It is not evidence that the query difference caused the answer difference.
Search and answer generation occurred inside the same grounded call.
Topic stability
The pattern remained high across all five topic clusters
Lexical query divergence remained high across all five clusters:
| Topic cluster | Lexical query divergence |
|---|---|
| AI optimisation | 0.8941 |
| Answer Alignment | 0.9093 |
| Content authority | 0.8644 |
| Conversion / pipeline | 0.8659 |
| Search visibility | 0.9018 |
Range:
0.864–0.909
The exact magnitude varied, but no topic cluster showed a collapse in persona-associated query differentiation.
This does not make the effect causal.
It does show that the observed query differentiation was not confined to one subject area.
Source boundary
What we could not establish: different queries do not prove different sources
This is the most important boundary in the study.
The historical grounded responses contained:
9,799 source records.
Every stored source URI was classified as an unresolved grounding wrapper.
Reliable normalized publisher/domain coverage was:
0%.
We predefined a source-normalization gate before interpreting source differences.
The source gate therefore failed.
We did not calculate:
- Publisher Jaccard overlap
- Domain divergence
- Persona-specific publisher preferences
- Query-to-source association
- Source-to-answer association
This means the result is:
query-only differentiation
not:
retrieval-pathway differentiation.
And source non-identification does not mean the personas converged on the same evidence.
It means the stored data cannot answer that question.
Interpretation boundary
Why query fan-out should be treated as telemetry, not hidden reasoning
The query layer is useful because it exposes part of the observable search process.
But it should not be confused with the model's hidden reasoning.
A recorded query tells us something the system searched for or reported searching for.
It does not prove:
- Why the model chose that query
- What internal representation generated it
- Which retrieved document mattered most
- Whether the search was necessary
- Whether a later query was causally conditioned on an earlier result
Query telemetry can reveal:
- Search-plan variation
- Information-need framing
- Provider differences
- Persona-associated formulation differences
It cannot, by itself, establish hidden reasoning or GEO influence.
AEO research
What this means for AI visibility and AEO research
The practical implication is not:
“Track every query Gemini generates.”
The more useful conclusion is:
The query layer can reveal variation that the final answer alone may hide.
If two persona conditions ask different grounded search questions before producing an answer, then monitoring only the final response can miss an important part of the observable process.
For research teams, this suggests a layered evaluation model:
- Prompt layer
What did we ask?
- Query layer
What information needs did the grounded system formulate?
- Source layer
Which evidence did it retrieve or cite?
- Answer layer
What did it finally say?
The current study observes layers 1, 2 and 4.
Its source layer was not resolvable.
That distinction should remain explicit in any AEO or GEO workflow.
Panel design
Implications for persona-panel design
The query findings also help explain why persona panels can remain diverse even when the underlying scenario is held constant.
The five persona conditions did not simply apply different wording to the same apparent search plan.
Their recorded search formulations themselves differed.
That means persona-panel diversity can arise before answer synthesis.
But this still does not establish that every query difference is useful.
A more technical query may be more relevant, less relevant, redundant, overly narrow or simply different.
The next step for stronger research would be to connect persona-conditioned query formulation to reliably resolved source pools, claim-level evidence, factual accuracy and recommendation quality.
Those links are not identified here.
Method and reproducibility
Methodology
The analysis reuses the frozen historical V2 persona corpus.
PR1's substantive-response quality gate supplied 749 usable responses.
The primary matched query analysis used 149 complete five-persona scenarios and all ten unordered persona pairs per scenario.
That produced 1,490 primary pair rows.
Query normalization applied Unicode NFKC normalization, lowercase conversion, whitespace collapse and surrounding sentence-punctuation cleanup.
Three query-comparison families were used:
- Exact normalized query-set Jaccard divergence
- Token-set divergence
- TF-IDF lexical divergence
Scenario-level bootstrap intervals resampled whole neutral scenarios rather than treating persona-pair rows as independent observations.
For query-to-answer association, pair rows were centered within scenario and uncertainty was estimated by whole-scenario resampling.
No live search, source resolution, new model generation or external embedding API was used.
Study boundaries
Limitations
- Persona and collection position are confounded
The historical persona order was fixed within every scenario. The study therefore cannot isolate persona from collection-position effects.
- Query telemetry is incomplete
Stored search queries describe provider-reported observable query behavior. They do not expose the full internal search or reasoning process.
- The source layer is unresolved
All stored source URIs were grounding wrappers. Publisher-level source differentiation is not identified.
- Exact divergence is a ceiling metric
Exact normalized query-set divergence was almost saturated. It establishes literal difference but is weak for measuring degree.
- One response per persona/scenario
The historical experiment does not estimate within-condition generation variability.
- Different queries do not mean better queries
The study does not measure relevance, authority, factual support or business value.
FAQ
Frequently asked questions
Did different persona prompts make Gemini use different search queries?
Yes in this historical run. Exact normalized query overlap across persona conditions was almost nonexistent, and token and lexical query representations also showed substantial differentiation.
Which persona pair had the most different search language?
Demand Generation Director and SEO Manager had the highest observed lexical query divergence at 0.928.
Which pair was the closest?
CMO and Founder/CEO had the lowest observed lexical query divergence at 0.851.
Does that mean those real roles search similarly?
No. These are AI outputs under persona-conditioned prompts, not observations of real professionals.
Did different search queries lead to different sources?
The study cannot answer that. All 9,799 stored source records used unresolved grounding-wrapper URLs, so reliable publisher identity was unavailable.
Did more different queries produce more different answers?
Token and lexical query divergence were positively associated with residual answer distance within scenarios, with correlations of 0.514 and 0.495 respectively. This is an association, not causal mediation.
Is exact query overlap a good metric?
It is useful for showing whether literal queries match, but in this dataset exact divergence was almost saturated. Token and lexical measures were more informative about degree.
Does query fan-out show the model's reasoning?
No. It provides observable search-plan telemetry, not direct access to hidden reasoning.
Conclusion
Conclusion
Persona prompts changed more than the final wording in this historical grounded Gemini experiment.
Across 149 matched scenarios, different persona conditions were associated with substantially different recorded search-query formulations.
The difference was visible at three levels:
- Almost no exact normalized query overlap
- Substantial token divergence
- High lexical divergence
The query language also reflected distinct persona-associated emphases.
SEO Manager queries were more technical.
Demand Generation queries were more commercial.
CMO queries were more governance-oriented.
And persona pairs with more differentiated query language tended to produce more differentiated residual answers.
But the research stops at the boundary the data supports.
We cannot say the different queries caused the different answers.
We cannot say different publishers were retrieved.
And we cannot say more differentiated searches were better searches.
The strongest conclusion is:
In this matched grounded-AI run, persona-associated differentiation was visible in the system's recorded search-query formulation before it was visible in the final answer.
For the broader programme—including prompt-echo robustness, task persistence and panel-size analysis—see Persona Prompts Change More Than Tone: Evidence from 750 Grounded AI Responses.