GTM AI Index
Index Research › Repeat measurement · September 2026 Edition
Repeat measurement · September 2026 Edition

Asked the same question twice, the models moved their own first choice 66% of the time. The leader still held in 15 of 20 categories.

Twenty categories were asked again, with nothing changed: the same prompts, the same model versions, the same settings, one buyer segment, six framings, twelve models, 1,440 answers. Whatever differs between the two runs is noise, and the noise is what this page measures. That is the reason the index asks every question six ways of every model rather than once: on its own, one ask moved 66% of the time; seventy-two, read together, are what the index reports.

Model by model

For each model, the share of its questions where the top pick differed between the two runs. The pairs are one question asked in each run; a pair flips when the first choice is not the same product.
Qwen 3.7 Flash80%79 of 99
Kimi K274%81 of 110
Gemini 3.5 Flash72%68 of 94
GLM 4.7 FlashX72%63 of 88
Claude Haiku 4.571%54 of 76
MiniMax M2.568%61 of 90
DeepSeek V4 Flash67%68 of 102
Mistral Small66%65 of 98
Llama 4 Maverick64%41 of 64
GPT-5.4 mini62%48 of 77
Grok 4.1 Fast59%66 of 112
Perplexity Sonar32%23 of 72

Qwen 3.7 Flash changed its first choice most often, 80% of its 99 pairs; Perplexity Sonar least, 32% of 72. Pooled over every model and question, 66%.

A flip is the same model, the same question, a different first choice a few days apart. It says how much one answer can be trusted on its own, and nothing about why the model answered as it did.

Category by category

The leader's share of first choices in the edition run and in the repeat, and whether the leader was the same product both times. Sorted by the size of the move.
CategoryLeader in the edition runShare, run oneShare, repeatMoveLeader
CRM HubSpot CRM41%56%+15 pointsheld
AI visibility Otterly20%31%+11 pointsheld
Webinars Livestorm23%12%-10 pointschanged: Demio
Conversation intel Avoma52%44%-8 pointsheld
Attribution & MMM Dreamdata29%24%-6 pointsheld
Email marketing ActiveCampaign40%35%-5 pointschanged: HubSpot Marketing Hub
Identity Hightouch15%10%-5 pointschanged: Segment
Helpdesk Freshdesk44%40%-4 pointsheld
Sales intel Apollo.io76%72%-4 pointsheld
Intent data Bombora35%31%-4 pointsheld
SEO Surfer SEO25%22%-3 pointschanged: Semrush
Sales rooms Aligned40%36%-3 pointsheld
Support chatbots Intercom35%32%-3 pointsheld
CDP Segment48%45%-3 pointsheld
PRM PartnerPortal.io24%27%+2 pointsheld
Warehouse & ETL Hightouch42%40%-2 pointsheld
Product analytics Amplitude30%31%+2 pointschanged: Mixpanel
Marketing automation HubSpot Marketing Hub55%57%+2 pointsheld
Sales engagement Salesloft35%34%-1 pointsheld
Marketing CMS WordPress33%33%+0 pointsheld

In 15 of the 20 categories the same product led both runs. Where the leader changed, the two products were within 5 points of each other in the edition run.

The share floor

The bar a change has to clear before the index calls it a change, in the unit of the change itself.
10 pointsthe floor: the 90th percentile of the moves above
4 pointsmedian move of a leader's share on a repeat
15 pointsthe largest move, CRM

From the next edition on, a product's change in share counts as movement only when it is larger than 10 points, and a new leader is reported only when it clears the old one by more than that. Nine repeats in ten move a leader less. The floor is measured again with every edition and the method page carries the rule: how the floor is measured.

Cite this

GTM AI Recommendation Index, September 2026 Edition: repeat measurement. gtm-ai-index.com/research/repeat-measurement/. Published under CC BY 4.0. Every figure on this page is computed from the published edition and changes with it; the edition and its date are the citation.

The output is the models' output. Nothing here says the models can be steered, and nothing here is a recommendation by the index.