The Mechanics of AI Generated Selection

AI Search Research Note 002 · Experiments 004 and 005 · September 2026

Two controlled experiments examining why stable candidate sets can still produce unstable brand selections — and how persona construction, evidence activation and model-generated weighting shape the final recommendation.

Research Note 001 ended where Brand Selection emerged. Research Note 002 begins at that point and asks a narrower question:

Why can the same decision context repeatedly produce almost the same Candidate Set while still producing different winners?

STABLE CANDIDATE SETPERSONA CONSTRUCTIONCRITERIA WEIGHTINGEVIDENCE ACTIVATIONPERSONA BRAND FITVARIABLE BRAND SELECTION
Research Note 0022 experiments
24documented responses
173visible fan-outs
1,696raw source retrievals
264citations

→ Download the full Research Note (PDF)


Key Finding

Across Experiment 004, the Candidate Set remained highly stable, with a mean pairwise Jaccard similarity of 0.877.

But the Primary Selection did not stabilize. Fonds Finanz, JDC and Netfonds / NFS were each selected in four of twelve runs.

Experiment 005 then showed that increasing search depth substantially expanded the evidence volume without proportionally expanding the Candidate Set or producing a deterministic winner.

Candidate stability does not imply Selection stability.


Research Note 002 · Visual Summary

Five Findings at a Glance

Use the arrows, dots or swipe to move through the five observed and reconstructed layers.

Why This Question Matters

Research Note 001 documented how informational and operational questions can escalate into commercial categories, Candidate Sets and justified Brand Selection.

Research Note 002 examines what happens after several plausible providers already occupy the decision space.

→ Read Research Note 001


Experiment 004 — Candidate Set Stability

Experiment 004 ran twelve independent ChatGPT Search sessions in fresh, logged-out incognito chats. The prompt remained unchanged. Follow-ups and regenerations were excluded. Twelve runs were evaluable; one technical failure was replaced by a documented replacement run.

Stability metricResult
Evaluable runs12
Mean Candidate Set size5.75
Distinct candidates7
Core Candidates (at least 9/12)6
Invariant candidates (12/12)4
Mean pairwise Jaccard0.877
Runs with Primary Selection12/12

Candidate inclusion

CandidateInclusionsRateClass
Fonds Finanz12100%Core and invariant
BCA12100%Core and invariant
blau direkt12100%Core and invariant
JDC12100%Core and invariant
Netfonds / NFS1083%Core
VEMA1083%Core
DEMV18%One-off

The Candidate Set was highly stable. Four brands appeared in every run; only DEMV remained a one-off peripheral inclusion.


Stable Candidates, Unstable Selection

The stable Candidate Set did not produce a stable recommendation.

Primary SelectionSelectionsSelection per inclusion
Fonds Finanz433%
JDC433%
Netfonds / NFS440%
BCA00%
blau direkt00%
VEMA00%

Fonds Finanz, JDC and Netfonds / NFS each won four runs. BCA, blau direkt and VEMA remained stable candidates without becoming the Primary Selection.

Exploratory criterion coding

All-finance breadth and Investment / Wealth each appeared in seven of twelve runs as the most frequent criterion families. Banker / target-group fit appeared in five runs. Technology / integration, liability umbrella / regulation and insurance strength each appeared in three.

The criteria were more stable than their brand-specific weighting. Fonds Finanz won through breadth. Netfonds won through Investment / Wealth and Banker Fit. JDC occupied a hybrid position. This criterion coding was developed post hoc and is therefore exploratory.


Experiment 005 — Selection Evidence

Experiment 005 repeated the same decision context in twelve new ChatGPT Free runs. Six used the natural prompt. Six added an instruction to search the web and use current sources. All runs took place on 14 September 2026 in fresh chats with Memory disabled.

ConditionRunsFan-outsRaw sourcesCitations / run
A — Natural Search64049316.0
B — Explicit Current Search61331,20328.0

The explicit search instruction generated 3.3× as many visible fan-outs and 2.4× as many raw source retrievals per run. The final Candidate Set still remained mostly between five and seven companies. Additional research deepened the evidence much more than it broadened the selection space.

Primary Selection by condition

Primary SelectionNaturalExplicit Current SearchTotal
JDC246
Fonds Finanz213
Netfonds / NFS / finfire213

Natural Search produced an even 2–2–2 distribution. Under Explicit Current Search, JDC won four of six runs. This is an observed association in a small sample, not a causal effect of the search instruction.


The AI Result Graph

The experiments illuminate different parts of the same result graph. Experiment 004 measures Candidate Set stability and selection variability. Experiment 005 opens the area between prompt, search, evidence and outcome weighting.

StageFunctionEvidence status
1. Prompt signalsBank background, self-employment and desired service breadthObserved
2. Task interpretationTask framed as strategic provider selectionResponse framing
3. Constructed decision personaFormer banker and future independent intermediaryReconstructed
4. Criteria weightingBreadth, Investment, technology, liability umbrella and portfolioObserved / coded
5. Candidate eligibilityStable core set of six brandsStrongly supported
6. Evidence activationMarket sources and Capability Pages retrievedObserved
7. Evidence interpretationRetrieved material translated into claims and relationshipsCase inspection
8. Persona–Brand FitCapabilities related to the constructed personaReconstructed
9. Selection weightingOne candidate prioritized over alternativesObserved outcome
10. Primary SelectionJDC, Fonds Finanz or Netfonds wins24 runs

Two evidence paths converge. Rankings and market overviews legitimize brands as candidates. Provider and product pages document concrete capabilities. Model-side interpretation connects those capabilities to the constructed persona and produces a ranking.

No single source explains the complete selection decision.


The AI Result Graph, a functional reconstruction from prompt signals through task interpretation, decision persona, criteria weighting, candidate eligibility, evidence activation and interpretation, persona-brand fit, selection weighting and Primary Selection, with observed, supported and reconstructed stages distinguished.
Functional reconstruction based on Experiments 004 and 005. Observed, supported and reconstructed layers are distinguished explicitly; this is not a claimed universal internal model architecture.

Seven Findings

1. Candidate stability does not imply Selection stability

A mean pairwise Jaccard of 0.877 coexisted with an exact 4–4–4 winner distribution in Experiment 004.

2. Stable criteria do not determine a single winner

The same broad criterion families recurred, but their brand-specific weighting differed.

3. Search depth expands evidence more than candidate breadth

Explicit search strongly increased fan-outs and source retrievals without proportionally widening the Candidate Set.

4. Explicit search coincided with stronger JDC concentration

JDC won four of six Explicit Current Search runs. The sample supports an observed association, not a causal claim.

5. Different sources perform different jobs

Rankings and market overviews can legitimize candidates. Provider homepages establish market role and breadth. Product and platform pages support concrete capabilities. Regulatory sources structure the decision space. Ownership reporting can make dependencies and strategic risk legible. Model inference translates capabilities into personal fit and ranking.

6. Retrieval is not Selection

Some retrieved brands did not survive into the final Candidate Set or recommendation. Experiment 004 shows the same separation at another layer: BCA and blau direkt appeared in all twelve Candidate Sets but never won.

7. Evidence Interpretation is a separate risk layer

One inspected case connected blau direkt with Insurgo even though the cited Insurgo page described a migration from blau direkt to Insurgo rather than an integrated blau direkt capability. A semantically plausible association became a factually problematic relationship.

Grounding boundary: First-party pages can document capabilities, but they do not independently prove that the provider is better suited to the persona. The comparative superiority inference is produced primarily by the model.


What Companies Can Influence

ClassInfluenceable surfaceBoundary
OwnedClear product, platform, target-group and exit documentationSupports capabilities, not independent superiority
EarnedComparisons, studies, specialist media, reviews and partnershipsNot fully controllable
DistributedConsistent relationships between brand, persona, use case and criterionWorks across multiple sources and contexts
OpaqueInternal weighting, ordering and synthesisNo defensible direct intervention surface

The practically relevant unit is not the individual citation. It is the distributed result graph. A brand becomes more selection-capable when independent market legitimacy, concrete Capability Evidence and consistent Persona / Use Case associations converge.


What This Research Does Not Show

The experiments do not reveal proprietary ranking parameters or prove that a specific source causes a specific winner.

They do not show that Explicit Current Search will generally favor JDC, that the observed winner distribution will remain stable over time, or that the reconstructed decision persona corresponds to a directly observable internal model state.

The distinction remains important:

Observed: responses, Candidate Sets, Primary Selections, visible fan-outs, raw retrievals, citations and inspected evidence relationships.

Reconstructed: persona construction, criteria weighting and Persona–Brand Fit as explanatory layers connecting those observations.

Strategic inference: distributed associations may increase selection capability, but the experiments do not causally test an optimization intervention.


Methodology and Limitations

Research Note 002 synthesizes two connected experiments using the same commercially actionable core prompt.

Experiment 004: twelve independent repetitions with no prompt variation, measuring Candidate Set similarity and Primary Selection.

Experiment 005: six Natural Search and six Explicit Current Search runs, measuring search scope, evidence activation, traceability and Primary Selection.

Across both experiments, 24 responses were evaluated. Experiment 005 documented 173 visible fan-outs, 1,696 raw source retrievals and 264 citations.

The sample remains small and limited to one market, one decision persona, one commercial prompt family and a narrow collection period. Experiment 004 criterion coding was developed post hoc and is exploratory. Experiment 005's condition comparison is descriptive and does not establish causality.


The Next Research Question

Experiment 006 should vary the persona systematically while keeping the product question and search condition constant.

Possible contrasts include:

This would test whether the currently reconstructed relationship between persona, criteria weighting and Brand Selection can be made experimentally visible.

Experiment 007 should vary the evidence environment — for example with and without an independent comparison source or with deliberately changed Capability Claims.

That would begin to isolate the currently reconstructed edges of the result graph.


Closing Perspective

Research Note 001 showed how AI-mediated search can move from problem understanding to Category Association, Candidate Sets and justified Brand Selection.

Research Note 002 shows why that final step should not be treated as a simple ranking event.

A stable Candidate Set can coexist with unstable winners.

More search can activate substantially more evidence without materially expanding the candidate space.

And the final recommendation appears to depend not only on what evidence exists, but on how the system interprets the task, constructs the decision persona and weights that evidence for the inferred context.

AI selection depends on what evidence exists — but also on which persona the system constructs and how it weights that evidence for the inferred decision context.

→ Download Research Note 002 (PDF)

→ Browse all AI Search Research