Kojable research · Marketing role framing · 167 eligible responses
Do AI Role Prompts Change the Decision Lens? Financial Accountability Is the Clearest Observed Difference
An exploratory analysis of explicitly role-framed responses drawn from an 800-response marketing corpus finds that leadership-labelled prompts were associated most clearly with financial and organisational accountability—not a generic strategy-versus-execution divide.
By Piush Vaish, founder and CEO of Kojable.
Key finding
Budget, cost or investment language appeared 24.8 percentage points more often in the two leadership-labelled roles, while ROI, return or impact language showed a 19.9-point gap.
Qualification: The four role conditions were not matched counterfactuals. The 167 eligible prompts may differ in topic, template and wording, so the result is an exploratory within-corpus association rather than a causal estimate of role or seniority.
- AI role prompting
- Marketing personas
- Decision framing
- Financial accountability
- 18 min read
Study overview
Executive summary
Role prompts are often treated as a presentation layer: ask an AI system to answer as a CMO, an SEO manager or a content strategist, and the expectation is that tone and vocabulary will change while the underlying decision logic remains broadly similar.
This analysis tests a narrower question: when a marketing role is explicitly named in the prompt, what kinds of considerations become more or less prominent in the resulting answer?
Across 167 eligible responses, prompts explicitly naming a leadership role—CMO or VP Marketing—were associated with substantially more language about budget, cost, investment, ROI, return, impact, teams and resources than prompts explicitly naming Content Strategy Consultant or SEO Manager.
The data did not show an equally clear strategy-versus-execution divide. Strategic language was common across all four roles, while tactical and technical language differed only modestly at the group level.
Principal finding
Within this synthetic corpus, the two leadership-labelled roles were associated most clearly with financial and organisational-accountability language. The evidence does not support a broad leadership-versus-practitioner split in strategic, tactical or technical content.
Answer first
Direct answer
The clearest observed difference between the two leadership-labelled and two practitioner-labelled roles was financial and organisational accountability language.
Budget, cost or investment language appeared in 79.7% of leadership-framed responses and 55.0% of practitioner-framed responses: a difference of +24.8 percentage points.
ROI, return or impact language appeared in 78.7% of leadership-framed responses and 58.8% of practitioner-framed responses: a difference of +19.9 percentage points.
By contrast, tactical and technical language differed by only a few percentage points, and strategic language was common in both groups. The most defensible interpretation is that explicit role framing was associated with a different decision lens, not that the role title independently caused the differences or that one kind of answer was better.
Research snapshot
The corpus, comparison and inference boundary
| Response generator | Gemini 3 Flash Preview |
|---|---|
| Recorded run | 15 March 2026 |
| Wider response corpus | 800 successful model responses across eight marketing and growth roles |
| Directly eligible responses | 167 prompts explicitly naming exactly one target role |
| Roles and eligible samples | CMO 40; VP Marketing 47; Content Strategy Consultant 39; SEO Manager 41 |
| Unit of analysis | One synthetic model response |
| Main measures | Boundary-aware keyword-family presence, response length, recorded search-query count and grounding-source object count |
| Statistical summary | Equal-weight role means and 10,000 role-stratified bootstrap resamples |
| Inference boundary | Exploratory within-corpus comparison; not a matched causal experiment |
Analytical population
Why the analysis uses 167 responses, not all 800
The wider corpus contains 800 responses and metadata assigning 100 prompts to each of eight roles. But the role stored in metadata was not always explicitly present in the text shown to the model.
An exact-title audit found 800 responses in the wider corpus, 336 prompts explicitly naming one of the eight roles, and 167 prompts explicitly naming exactly one of the four roles examined here.
This restriction improves treatment validity, but it creates a selection boundary. Role-explicit prompts may differ in template, topic, wording, intent or other characteristics. Because the four roles were not applied to matched versions of the same prompt, composition differences may contribute to the observed role-level contrasts.
Finding 1
Financial-accountability language is the strongest observed separation
Each of the four roles received equal weight within its leadership or practitioner group, preventing roles with slightly more eligible responses from dominating the comparison.
Budget-related language appeared in 79.7% of leadership-framed responses and 55.0% of practitioner-framed responses, a 24.8-point gap. ROI, return or impact language appeared in 78.7% and 58.8%, respectively, a 19.9-point gap.
| Role | Budget / cost / investment | ROI / return / impact |
|---|---|---|
| CMO | 85.0% | 85.0% |
| VP Marketing | 74.5% | 72.3% |
| Content Strategy Consultant | 53.8% | 59.0% |
| SEO Manager | 56.1% | 58.5% |
The two leadership-labelled roles were not identical, but both showed higher prevalence of the two financial keyword families than either practitioner-labelled role.
Sensitivity analysis
The tested financial gap was not simple direct-keyword repetition
The financial comparisons were repeated after excluding prompts that already contained the corresponding financial keyword family. The observed gaps did not shrink.
- Prompts without budget, cost or investment terms
Leadership equal-role average: 78.6%. Practitioner equal-role average: 52.6%. Difference: +26.0 percentage points.
- Prompts without ROI, return or impact terms
Leadership equal-role average: 77.9%. Practitioner equal-role average: 57.2%. Difference: +20.7 percentage points.
This rules out one narrow explanation: direct repetition of the exact tested financial keywords already present in the prompt does not account for the observed gap.
The result measures financial-accountability language, not financial reasoning, P&L competence or investment quality. “Impact” and “return” can also appear outside a strictly financial context.
Finding 2
The data does not support a clean strategy-versus-execution divide
| Keyword family | Leadership | Practitioner | Difference |
|---|---|---|---|
| Budget / cost / investment | 79.7% | 55.0% | +24.8 pp |
| ROI / return / impact | 78.7% | 58.8% | +19.9 pp |
| Team / resources / staff | 78.4% | 62.5% | +16.0 pp |
| Strategy / strategic | 94.1% | 85.4% | +8.8 pp |
| Tactics / implementation | 42.7% | 46.2% | −3.5 pp |
| Technical / implementation / setup | 66.5% | 68.8% | −2.2 pp |
Within-corpus bootstrap intervals were clearly above zero for budget, financial-outcome and team/resource language. The strategy difference was smaller, with a lower interval boundary close to zero. Tactical and technical intervals crossed zero broadly.
Practitioner-framed responses were not “non-strategic”, and leadership-framed responses were not detached from implementation. The more distinctive leadership-associated layer was financial justification and organisational accountability.
Finding 3
Group labels hide meaningful role-specific differences
- CMO
The strongest financial signature: 85.0% contained budget/cost/investment language and 85.0% contained ROI/return/impact language.
- VP Marketing
Lower financial prevalence than CMO but stronger team/resource language: 89.4%, compared with 67.5% for CMO.
- Content Strategy Consultant
Strategy language appeared in every eligible response, but the role title itself makes that result difficult to interpret independently.
- SEO Manager
Responses showed comparatively high technical framing and the largest average number of grounding-source objects among the four roles.
The two leadership-labelled roles share a stronger financial-accountability signature in this corpus, but the four roles are not internally interchangeable. This is an observation about a model-and-prompt dataset, not a psychological profile of the occupations.
Finding 4
Search-query volume, grounding objects and response length do not move together
| Role group | Search queries | Grounding-source objects | Response length |
|---|---|---|---|
| Leadership-labelled | 8.10 | 2.13 | Approximately 4,528 characters |
| Practitioner-labelled | 6.70 | 2.98 | Approximately 4,815 characters |
Practitioner-framed responses returned approximately 0.86 more grounding-source objects per response and were somewhat longer, while leadership-framed responses recorded more search-query fan-out.
Evaluation principle
Search-query volume, grounding-object count and response length are separate generation behaviours. They should not be collapsed into one measure of depth or quality.
More search queries did not automatically produce more grounding objects, and more grounding objects did not automatically produce a longer answer. None of these measures establishes evidence authority, relevance, independence, recency or claim support.
Workflow implications
What this means for teams using role prompts
Treat role prompts as decision-lens assumptions
A selected role can be associated with different criteria being foregrounded, not merely a change in tone.
State required financial criteria explicitly
If cost, resources, expected return, measurement horizon or failure conditions matter, require them rather than assuming they will emerge from a practitioner label.
Avoid a strategy-versus-execution stereotype
Strategic language was widespread, while tactical and technical prevalence differed only modestly between the groups.
Test multi-role review rather than assuming it works
Practitioner drafting followed by leadership-labelled review is a plausible workflow hypothesis, but this study did not test whether it improves outcomes.
Evaluate evidence quality separately
Measure authority, relevance, independence, recency, claim-level entailment and coverage—not merely grounding-object count.
Evidence summary
What this study supports
- Financial language
The two leadership-labelled roles were associated with substantially more budget/cost/investment and ROI/return/impact language.
- Direct-keyword sensitivity
The financial gaps remained after prompts containing the corresponding exact financial keywords were excluded.
- Organisational accountability
Team/resource language was more common under the leadership-labelled roles, especially VP Marketing.
- Weak generic divide
Tactical and technical language showed little group-level separation; strategy language was common and partly exposed to title echo.
- Separate generation mechanics
Practitioner-labelled responses had more grounding objects and were longer; leadership-labelled responses generated more recorded search queries.
- Role-level heterogeneity
Specific roles differed enough that the leadership/practitioner grouping is not a complete explanation.
Claim boundary
What this study does not support
- Claims about people
The results do not show that real marketing leaders are more financially accountable than real practitioners.
- Causal seniority or role-title effects
Job seniority and the role title alone were not isolated as causes of the response differences.
- Better recommendations
More financial language does not establish a better recommendation.
- Better research or evidence
More queries do not indicate better research, and more grounding-source objects do not indicate better evidence.
- Length as quality
Longer responses are not necessarily more useful or accurate.
- Universal categories
The findings do not establish leadership and practitioner as universal or internally homogeneous persona categories.
- Generalisability
The results do not establish replication across models, versions, prompt families, industries or time periods.
Method and boundaries
Methodology and limitations
- Eligible subset
The comparison uses 167 responses selected because the prompt explicitly named exactly one target role. This improves treatment validity but may introduce selection differences.
- Group estimates
Each role received equal weight within the leadership or practitioner group. Response-level keyword-family presence supplies the principal measures.
- Bootstrap summary
Intervals use 10,000 role-stratified resamples of this fixed synthetic subset. They do not describe a population of real professionals.
- Recorded model snapshot
The analysis covers one Gemini 3 Flash Preview run recorded on 15 March 2026.
- Reproducibility disclosure
The local publication package includes the article and four figures, but no public analysis script, derived CSV tables or complete reproducibility bundle. The publication documents the supplied methods and estimates without claiming independent reproducibility from the package alone.
- No matched counterfactuals
The four roles were not applied to matched versions of the same tasks, so unmeasured prompt-composition differences may contribute.
- Small role samples
Eligible role samples range from 39 to 47 responses.
- Keyword-family measures
Term presence does not determine whether reasoning is correct, important or contextually appropriate; several dictionary terms are semantically broad or overlapping.
- Partial prompt-echo test
Excluding exact tested financial terms does not eliminate broader semantic prompt cues.
- Evidence metadata
Query and grounding-object counts do not establish authority, claim support, factual accuracy or retrieval usefulness.
- One model snapshot
The observed patterns may not reproduce in another model, version, retrieval environment or collection period.
Research agenda
What a stronger follow-up experiment should test
A confirmatory study should manipulate role framing directly rather than select role-explicit prompts after generation.
- Hold the complete prompt skeleton constant
Keep the marketing task, topic, wording and context matched across role treatments.
- Run every target role on every task
Generate CMO, VP Marketing, Content Strategy Consultant and SEO Manager conditions.
- Include a generic no-role control
Establish a neutral baseline for each matched marketing task.
- Randomise and replicate
Randomise execution order and generate each role-task condition multiple times.
- Pre-specify outcome measures
Assess financial-accountability coverage, strategic prioritisation, implementation completeness, resource and risk awareness, factual accuracy, decision usefulness and source quality.
The matched task should be the primary inferential unit. A further sensitivity analysis should separate narrow economic terms such as budget, cost, investment, ROI, revenue, margin, payback and profitability from broader terms such as impact and return.
Research lineage
How this fits the broader persona research
The finance-persona similarity work asks whether persona-conditioned outputs remain differentiated after accounting for prompt design. Its cue-ablation follow-on asks whether that residual survives removal of explicit persona/profile language. A separate matched AEO study asks how persona conditions shift search, pipeline, AEO and SEO language.
This paper uses a separate marketing/growth corpus and focuses on decision framing within explicitly named marketing-role prompts: which considerations become more prominent in the generated advice?
The financial-accountability result is the clearest observed answer in this corpus. It remains exploratory because the role conditions were not matched counterfactuals.
FAQ
Frequently asked questions
How many responses were used in the leadership-versus-practitioner comparison?
The comparison used 167 explicitly role-framed responses drawn from a wider 800-response marketing corpus. The 167 prompts explicitly named exactly one of CMO, VP Marketing, Content Strategy Consultant or SEO Manager.
Did leadership-role prompts produce more strategic answers?
Not in a clean general sense. Strategy language was common across all four roles, while tactical and technical differences were small. The strongest observed separation was financial and organisational-accountability language.
Does this prove that real marketing leaders think differently from practitioners?
No. This is an analysis of synthetic model responses, and the role conditions were not matched counterfactuals. The results do not isolate a causal effect of job seniority or describe real professionals.
Did more grounding-source objects mean better evidence?
No. Grounding-source object count measures the quantity of recorded grounding artefacts, not source authority, relevance, correctness or claim support.



