The Mechanics of AI Generated Selection
Two controlled experiments examining why stable candidate sets can still produce unstable brand selections — and how persona construction, evidence activation and model-generated weighting shape the final recommendation.
Research Note 001 ended where Brand Selection emerged. Research Note 002 begins at that point and asks a narrower question:
Why can the same decision context repeatedly produce almost the same Candidate Set while still producing different winners?
→ Download the full Research Note (PDF)
Key Finding
Across Experiment 004, the Candidate Set remained highly stable, with a mean pairwise Jaccard similarity of 0.877.
But the Primary Selection did not stabilize. Fonds Finanz, JDC and Netfonds / NFS were each selected in four of twelve runs.
Experiment 005 then showed that increasing search depth substantially expanded the evidence volume without proportionally expanding the Candidate Set or producing a deterministic winner.
Candidate stability does not imply Selection stability.
Research Note 002 · Visual Summary
Five Findings at a Glance
Use the arrows, dots or swipe to move through the five observed and reconstructed layers.
Why This Question Matters
Research Note 001 documented how informational and operational questions can escalate into commercial categories, Candidate Sets and justified Brand Selection.
Research Note 002 examines what happens after several plausible providers already occupy the decision space.
Experiment 004 — Candidate Set Stability
Experiment 004 ran twelve independent ChatGPT Search sessions in fresh, logged-out incognito chats. The prompt remained unchanged. Follow-ups and regenerations were excluded. Twelve runs were evaluable; one technical failure was replaced by a documented replacement run.
| Stability metric | Result |
|---|---|
| Evaluable runs | 12 |
| Mean Candidate Set size | 5.75 |
| Distinct candidates | 7 |
| Core Candidates (at least 9/12) | 6 |
| Invariant candidates (12/12) | 4 |
| Mean pairwise Jaccard | 0.877 |
| Runs with Primary Selection | 12/12 |
Candidate inclusion
| Candidate | Inclusions | Rate | Class |
|---|---|---|---|
| Fonds Finanz | 12 | 100% | Core and invariant |
| BCA | 12 | 100% | Core and invariant |
| blau direkt | 12 | 100% | Core and invariant |
| JDC | 12 | 100% | Core and invariant |
| Netfonds / NFS | 10 | 83% | Core |
| VEMA | 10 | 83% | Core |
| DEMV | 1 | 8% | One-off |
The Candidate Set was highly stable. Four brands appeared in every run; only DEMV remained a one-off peripheral inclusion.
Stable Candidates, Unstable Selection
The stable Candidate Set did not produce a stable recommendation.
| Primary Selection | Selections | Selection per inclusion |
|---|---|---|
| Fonds Finanz | 4 | 33% |
| JDC | 4 | 33% |
| Netfonds / NFS | 4 | 40% |
| BCA | 0 | 0% |
| blau direkt | 0 | 0% |
| VEMA | 0 | 0% |
Fonds Finanz, JDC and Netfonds / NFS each won four runs. BCA, blau direkt and VEMA remained stable candidates without becoming the Primary Selection.
Exploratory criterion coding
All-finance breadth and Investment / Wealth each appeared in seven of twelve runs as the most frequent criterion families. Banker / target-group fit appeared in five runs. Technology / integration, liability umbrella / regulation and insurance strength each appeared in three.
The criteria were more stable than their brand-specific weighting. Fonds Finanz won through breadth. Netfonds won through Investment / Wealth and Banker Fit. JDC occupied a hybrid position. This criterion coding was developed post hoc and is therefore exploratory.
Experiment 005 — Selection Evidence
Experiment 005 repeated the same decision context in twelve new ChatGPT Free runs. Six used the natural prompt. Six added an instruction to search the web and use current sources. All runs took place on 14 September 2026 in fresh chats with Memory disabled.
| Condition | Runs | Fan-outs | Raw sources | Citations / run |
|---|---|---|---|---|
| A — Natural Search | 6 | 40 | 493 | 16.0 |
| B — Explicit Current Search | 6 | 133 | 1,203 | 28.0 |
The explicit search instruction generated 3.3× as many visible fan-outs and 2.4× as many raw source retrievals per run. The final Candidate Set still remained mostly between five and seven companies. Additional research deepened the evidence much more than it broadened the selection space.
Primary Selection by condition
| Primary Selection | Natural | Explicit Current Search | Total |
|---|---|---|---|
| JDC | 2 | 4 | 6 |
| Fonds Finanz | 2 | 1 | 3 |
| Netfonds / NFS / finfire | 2 | 1 | 3 |
Natural Search produced an even 2–2–2 distribution. Under Explicit Current Search, JDC won four of six runs. This is an observed association in a small sample, not a causal effect of the search instruction.
The AI Result Graph
The experiments illuminate different parts of the same result graph. Experiment 004 measures Candidate Set stability and selection variability. Experiment 005 opens the area between prompt, search, evidence and outcome weighting.
| Stage | Function | Evidence status |
|---|---|---|
| 1. Prompt signals | Bank background, self-employment and desired service breadth | Observed |
| 2. Task interpretation | Task framed as strategic provider selection | Response framing |
| 3. Constructed decision persona | Former banker and future independent intermediary | Reconstructed |
| 4. Criteria weighting | Breadth, Investment, technology, liability umbrella and portfolio | Observed / coded |
| 5. Candidate eligibility | Stable core set of six brands | Strongly supported |
| 6. Evidence activation | Market sources and Capability Pages retrieved | Observed |
| 7. Evidence interpretation | Retrieved material translated into claims and relationships | Case inspection |
| 8. Persona–Brand Fit | Capabilities related to the constructed persona | Reconstructed |
| 9. Selection weighting | One candidate prioritized over alternatives | Observed outcome |
| 10. Primary Selection | JDC, Fonds Finanz or Netfonds wins | 24 runs |
Two evidence paths converge. Rankings and market overviews legitimize brands as candidates. Provider and product pages document concrete capabilities. Model-side interpretation connects those capabilities to the constructed persona and produces a ranking.
No single source explains the complete selection decision.
Seven Findings
1. Candidate stability does not imply Selection stability
A mean pairwise Jaccard of 0.877 coexisted with an exact 4–4–4 winner distribution in Experiment 004.
2. Stable criteria do not determine a single winner
The same broad criterion families recurred, but their brand-specific weighting differed.
3. Search depth expands evidence more than candidate breadth
Explicit search strongly increased fan-outs and source retrievals without proportionally widening the Candidate Set.
4. Explicit search coincided with stronger JDC concentration
JDC won four of six Explicit Current Search runs. The sample supports an observed association, not a causal claim.
5. Different sources perform different jobs
Rankings and market overviews can legitimize candidates. Provider homepages establish market role and breadth. Product and platform pages support concrete capabilities. Regulatory sources structure the decision space. Ownership reporting can make dependencies and strategic risk legible. Model inference translates capabilities into personal fit and ranking.
6. Retrieval is not Selection
Some retrieved brands did not survive into the final Candidate Set or recommendation. Experiment 004 shows the same separation at another layer: BCA and blau direkt appeared in all twelve Candidate Sets but never won.
7. Evidence Interpretation is a separate risk layer
One inspected case connected blau direkt with Insurgo even though the cited Insurgo page described a migration from blau direkt to Insurgo rather than an integrated blau direkt capability. A semantically plausible association became a factually problematic relationship.
Grounding boundary: First-party pages can document capabilities, but they do not independently prove that the provider is better suited to the persona. The comparative superiority inference is produced primarily by the model.
What Companies Can Influence
| Class | Influenceable surface | Boundary |
|---|---|---|
| Owned | Clear product, platform, target-group and exit documentation | Supports capabilities, not independent superiority |
| Earned | Comparisons, studies, specialist media, reviews and partnerships | Not fully controllable |
| Distributed | Consistent relationships between brand, persona, use case and criterion | Works across multiple sources and contexts |
| Opaque | Internal weighting, ordering and synthesis | No defensible direct intervention surface |
The practically relevant unit is not the individual citation. It is the distributed result graph. A brand becomes more selection-capable when independent market legitimacy, concrete Capability Evidence and consistent Persona / Use Case associations converge.
What This Research Does Not Show
The experiments do not reveal proprietary ranking parameters or prove that a specific source causes a specific winner.
They do not show that Explicit Current Search will generally favor JDC, that the observed winner distribution will remain stable over time, or that the reconstructed decision persona corresponds to a directly observable internal model state.
The distinction remains important:
Observed: responses, Candidate Sets, Primary Selections, visible fan-outs, raw retrievals, citations and inspected evidence relationships.
Reconstructed: persona construction, criteria weighting and Persona–Brand Fit as explanatory layers connecting those observations.
Strategic inference: distributed associations may increase selection capability, but the experiments do not causally test an optimization intervention.
Methodology and Limitations
Research Note 002 synthesizes two connected experiments using the same commercially actionable core prompt.
Experiment 004: twelve independent repetitions with no prompt variation, measuring Candidate Set similarity and Primary Selection.
Experiment 005: six Natural Search and six Explicit Current Search runs, measuring search scope, evidence activation, traceability and Primary Selection.
Across both experiments, 24 responses were evaluated. Experiment 005 documented 173 visible fan-outs, 1,696 raw source retrievals and 264 citations.
The sample remains small and limited to one market, one decision persona, one commercial prompt family and a narrow collection period. Experiment 004 criterion coding was developed post hoc and is exploratory. Experiment 005's condition comparison is descriptive and does not establish causality.
The Next Research Question
Experiment 006 should vary the persona systematically while keeping the product question and search condition constant.
Possible contrasts include:
- insurance focus
- investment focus
- solo founder
- scalable brokerage business
This would test whether the currently reconstructed relationship between persona, criteria weighting and Brand Selection can be made experimentally visible.
Experiment 007 should vary the evidence environment — for example with and without an independent comparison source or with deliberately changed Capability Claims.
That would begin to isolate the currently reconstructed edges of the result graph.
Closing Perspective
Research Note 001 showed how AI-mediated search can move from problem understanding to Category Association, Candidate Sets and justified Brand Selection.
Research Note 002 shows why that final step should not be treated as a simple ranking event.
A stable Candidate Set can coexist with unstable winners.
More search can activate substantially more evidence without materially expanding the candidate space.
And the final recommendation appears to depend not only on what evidence exists, but on how the system interprets the task, constructs the decision persona and weights that evidence for the inferred context.
AI selection depends on what evidence exists — but also on which persona the system constructs and how it weights that evidence for the inferred decision context.




