GTM AI Index
How the index is made

Methodology

Everything needed to replicate an edition, dispute a placement, or check a number. If a step is not written here, the index does not do it.

01

What this measures

The index records what frontier AI models say when a buyer asks them which go-to-market product to use. It is not a review site and collects no user ratings. G2 measures what buyers claim after purchase; this measures what buyers are told before purchase.

Each edition the same prompts, six per category, go to the same models with web search enabled, in a fresh session per prompt with no conversational carryover. A judge model labels every product each answer names. The labels are the permanent record. Everything published is derived from them and can be recomputed.

02

Categories

Fourteen categories so far, across four of the six go-to-market verticals, chosen for high buyer search intent and a contested vendor landscape; two carry a conflict disclosure under the conflict policy below. The roadmap on the Data page lists 73 across all six. The AI visibility category is deliberate: ranking the answer engine tracking vendors inside an AI recommendation index is a reflexive test of the method.

The taxonomy is the index's own. To keep it legible to vendors and buyers, each category is cross-referenced to where it sits in the MartechMap, the marketing technology landscape by chiefmartec and MartechTribe, in G2's category tree, and in current analyst evaluations. Those taxonomies belong to their publishers and are cited by name, not reproduced. Analyst references were checked against the publisher on 2026-09-08.

Roadmap: 14 categories covered, 15 planned, 44 open, for an addressable set of 73 across six go-to-market verticals: Marketing (7 covered of 27), Sales (3 covered of 12), Customer (1 covered of 9), Revenue operations (0 covered of 9), GTM data and infrastructure (3 covered of 12), Partner and channel (0 covered of 4). The Management domain of the landscape (7 categories) is out of scope: Team and operations tooling rather than products a buyer asks an AI to recommend for marketing. Not in the addressable set.

#CategoryBuyer phraseLandscape referenceG2Analyst coverage
01Customer data platformscustomer data platformData › Customer Data PlatformCustomer Data Platforms (CDP)Gartner Magic Quadrant for Customer Data Platforms (2026)
02Product analyticsproduct analytics platformData › Mobile & Web AnalyticsProduct Analytics
03Content management systems for marketing sitesmarketing website CMSContent and Experience › CMS & Web Experience ManagementWeb Content ManagementGartner Magic Quadrant for Digital Experience Platforms (2026)
04SEO and content optimization platformsSEO and content optimization platformContent and Experience › SEOSEO Tools, Content MarketingGartner Magic Quadrant for Content Marketing Platforms (2026, adjacent)
05Webinar and virtual event platformswebinar and virtual event platformSocial and Relationships › Events, Meetings & WebinarsWebinar Platforms, Event Management
06Sales engagement platformssales engagement platformCommerce and Sales › Sales Automation, Enablement & IntelligenceSales EngagementGartner Magic Quadrant for Revenue Action Orchestration (first edition, December 2025; succeeds the Sales Engagement Applications market)
07Conversational intelligence and call recordingconversation intelligence and call recording platformSocial and Relationships › Call Analytics & ManagementConversation IntelligenceGartner Magic Quadrant for Revenue Action Orchestration (December 2025, adjacent: conversation intelligence is one of its converged capabilities)
08Attribution and marketing mix modelingmarketing attribution and marketing mix modeling platformData › Marketing Analytics, Performance & AttributionMarketing Attribution, Marketing AnalyticsGartner Magic Quadrant for Marketing Mix Modeling Solutions (2025); Forrester Wave: Marketing Measurement and Optimization Services (Q1 2026)
09Customer support and helpdeskcustomer support and helpdesk platformSocial and Relationships › Customer Experience, Service & SuccessHelp DeskGartner Magic Quadrant for the CRM Customer Engagement Center (2025)
10Data warehouse and reverse ETL for marketingmarketing data warehouse and reverse ETL stackData › iPaaS, Cloud/Data Integration & Tag ManagementReverse ETL, Data WarehouseGartner Magic Quadrant for Cloud Database Management Systems (2025); Gartner Magic Quadrant for Data Integration Tools (2025)
11AI visibility and answer engine tracking toolsAI visibility and answer engine tracking toolContent and Experience › SEO (AEO/GEO, noted by the landscape as an accelerating area)Generative Engine Optimization
12B2B intent data providersB2B intent data providerData › Audience/Marketing Data & Data EnhancementBuyer Intent Data Providers
13Marketing automation and campaign managementmarketing automation platformContent and Experience › Marketing Automation & Campaign/Lead ManagementMarketing AutomationGartner Magic Quadrant for B2B Marketing Automation Platforms (2025); Gartner Magic Quadrant for Multichannel Marketing Hubs (2025)
14Email marketing platformsemail marketing platformContent and Experience › Email MarketingEmail MarketingForrester Wave: Email Marketing Service Providers (Q1 2026)
03

Prompts

Six framings per category. Variants one and two are semantically identical and lexically different; they exist to measure paraphrase stability. Variant six captures anti-recommendations, which carry commercial weight and which most trackers ignore. Every prompt is published verbatim.

Every prompt frames the buyer as a mid-market B2B software company, and the scale prompt says 500 people. The index does not yet ask on behalf of a small business or an enterprise, so every standing on this site is a mid-market standing. Company size, industry, question framing and who answers are the four dimensions of the buyer; today there is one combination of them, and the Data page lists the planned ones, including small-business and enterprise sizes, other industries, a job-to-be-done framing that describes the work instead of naming the category, and human buyers answering the same questions.

#TemplatePurpose
01What is the best {phrase} for a mid-market B2B software company?Direct recommendation. Paired with variant 2 for paraphrase stability.
02Which {alt_phrase} would you recommend to a mid-sized B2B software company?Semantically identical to variant 1, lexically different. Uses alt_phrase.
03What are the top {plural} and how do they differ?Comparative. Surfaces the considered set and how the model differentiates.
04What is the best {phrase} for a company with a limited budget?Constrained by budget.
05We are a 500 person company evaluating a {phrase}. What should we look at?Constrained by scale.
06Which {plural} should I avoid or be cautious about?Negative framing. Captures anti-recommendations.

Settings: temperature provider default, max output 4,000 tokens, no system prompt, fresh session per prompt, Anthropic search capped at 5 uses per answer.

04

Model set

Coverage across labs rather than depth within one: two Anthropic models spanning a generation boundary, two OpenAI models spanning one, one Google model, one challenger. Version strings are recorded exactly as returned on every call. A change to any string is a new model row and forces a noise floor recalibration for that model.

ModelVersion stringLabGenerationSearch toolEndpoint
Claude Opus 5claude-opus-5anthropiccurrentweb_search_20260209api.anthropic.com
Claude Opus 4.8claude-opus-4-8anthropicpriorweb_search_20260209api.anthropic.com
GPT-6 Astragpt-6-astraopenaicurrentweb_searchapi.openai.com
GPT-5.6 Solgpt-5.6-solopenaipriorweb_searchapi.openai.com
Gemini 3.1 Progemini-3.1-pro-previewgooglesinglegoogle_searchgenerativelanguage.googleapis.com
Switched from gemini-3.8-flash on 2026-09-08: Flash never invoked Google Search in testing, Pro grounds. Preview build.
Perplexity Sonar Prosonar-prochallengersinglenativeapi.perplexity.ai
05

Scoring

Only products the answer treats as candidates in the asked category are counted, so a warehouse named as a data source in a CDP answer is not a CDP candidate. When the answer recommends a stack, only the product filling the category's role is a first choice; the rest are alternatives.

Categories, methodologies, analyst firms and people are never counted, and nothing named only inside a cited URL is counted.

A Claude Opus 5 judge, disclosed as an Anthropic model, reads each answer with the prompt and category and returns one object per named product with the name exactly as written, a label, its position and a verbatim evidence quote. Judge is an Anthropic model and is disclosed as such. It runs at effort low with structured output so labels are deterministic given the text.

Published weights

LabelWeightMeaning
First choice+3The product the answer leads with for the asked category
Alternative+2Named as a viable option alongside the first choice
Mention+1Named without endorsement
Soft negative−2Named with a caveat that discourages the buyer
Hard negative−3Named as something to avoid
06

Derived metrics

paraphrase stability
Per model, the share of categories where the first-choice set on the direct prompt equals the set on the paraphrase.
first-choice share
Per category and product, first-choice labels across the direct, paraphrase, budget and scale prompts and all models, divided by all first-choice labels in the category. The comparative and negative prompts do not ask for a pick and are excluded.
contested
No product holds more than 40% of first choices. The top share is always published next to the label because a category just above the line is not meaningfully different from one just below it.
negative rate
Per category and product, soft plus hard negative labels divided by all labels. Separates sentiment from salience.
quadrants
Products with at least 10 labels in a category placed by first-choice share and negative rate. Leader at 30% share or more, criticized at 25% negative or more: endorsed leader, criticized default, criticized challenger, accepted challenger.
lab treatment
For a lab with products in the category set, how its own model labels those products against how every other model labels the same products, as mean label weight. Only measurable with that control group; a lone self-preference count is not published.
discontinued
A positive label on a product the catalog marks discontinued. A retrieval failure worth naming.
noise floor
Per model, the share of (model, prompt) pairs whose first-choice set differs between an edition run and its calibration repeat. Pooled across models for movement verdicts.
07

Normalization

Product names are resolved through a versioned vendor table (v2026-09-08.5: 538 vendors with aliases, 20 bare names with category-scoped readings, 9 exclusions). Matching is case-insensitive and strips a trailing parenthetical. Resolution order for a raw name in a category:

  1. category alias
  2. global alias or canonical name
  3. exclusion
  4. trailing tier words removed, then the first two steps again
  5. split on separators with every part resolving
  6. unresolved

A bare vendor name resolves to that vendor's product for the asked category when it has exactly one (Salesforce in customer support is Service Cloud). Where the vendor has no product in the asked category, the name stays unresolved and is listed in the report. This is a directional assumption: a model writing a bare vendor name may mean the platform generally rather than the in-category product. The report lists every category-scoped resolution so the assumption is visible and reversible. A combined answer produces one label per product with the same label and evidence. Only applied when every part resolves; otherwise the name stays unresolved.

Adding a vendor can change how historical raw names resolve (a bare Salesforce in attribution stays unresolved only until a Salesforce attribution product exists). The catalog therefore follows the same discipline as the aliases: every change bumps the version and regenerates the series.

Excluded by rule: Google Ads (ad platform self-reported attribution, not a vendor in the category); Meta Ads Manager (ad platform self-reported attribution, not a vendor in the category); LinkedIn Ads (ad platform self-reported attribution, not a vendor in the category); TikTok (ad platform self-reported attribution, not a vendor in the category); Amazon (ad platform self-reported attribution, not a vendor in the category); Meta Conversion Lift (ad platform lift study, not a vendor in the category); Google geo experiments/Ads lift studies (ad platform lift study, not a vendor in the category); Native Google, Meta, TikTok and Amazon attribution (ad platform self-reported attribution, not a vendor in the category); Excel (general-purpose tool, not a vendor in the category).

08

Noise floor

Before any trend is claimed, the full matrix runs twice within one week with nothing changed: same prompts, same model versions, same settings. Whatever differs between the two runs is variance, not movement, and it sets the threshold below which the index says nothing changed.

It is reported per model as a first-choice flip rate, because models differ here and the differences are themselves a finding. A model whose repeat-run flip rate is as high as its paraphrase instability is not showing a paraphrase effect; its stability number is a floor and is marked as one wherever it appears.

Status for the September 2026 Edition: the calibration repeat is scheduled for 12 September 2026. The figure shown on the index until then comes from a simulated repeat and is labeled as such.

09

Data capture

Every prompt and every full response is stored verbatim with its timestamp, the model version string as returned, whether the model invoked search, the cited source URLs it exposed, latency, token counts and cost. Cited sources reflect live retrieval only and say nothing about pretraining data.

The judge's raw labels, one per named product with the name exactly as written and a verbatim evidence quote, are the permanent record.

An edition is 504 model calls. The September 2026 Edition cost $100.77 in model calls and $14.36 in judging; 86% of answers invoked search; median latency 42 seconds.

10

Editions and cadence

A monthly edition on the same date each month, which matches the pace at which the drivers move: model version updates and shifts in the retrieval corpus. A special edition within days of a major frontier model release, published as a comparison against the prior generation. Quarterly written analysis, only once enough editions have accumulated to say something with substance.

Every edition is archived permanently at a stable URL. If the project ends, the final edition is marked final on the site.

11

Conflict policy

The author operates Gane, which is building an operating system for B2B go-to-market. Its product overlaps with marketing automation and campaign management and with email marketing, so both categories are in the index with a conflict disclosure on each page rather than an exclusion. The data in those categories is collected and scored exactly like every other category and published unmodified; if Gane is named by a model, that label is published like any other. The disclosure exists so a reader can weigh it, and the raw record exists so a reader can check it.

The judge is an Anthropic model and is disclosed as such on every page.

12

Change log

Instrument changes: the vendor table, the judge rubric, the model set. Each entry names the first edition scored under it.

InstrumentChangeFirst edition under it
Vendor table v2026-09-08.1Seeded 38 CDP-centric vendors, then 438 vendors from the September 2026 Edition normalization backlog across all twelve categories.September 2026 Edition (re-scored)
Vendor table v2026-09-08.2Category-scoped aliases for bare parent and multi-product names (Salesforce, HubSpot, Adobe, Oracle, SAP, Semrush, Ahrefs, Salesloft, Outreach, Zoom, Clari, Gong). Split rule for combined answers such as A / B and A + B. Exclusion list for ad-platform self-attribution and general-purpose tools. Nine products added from the 38 unresolved September 2026 Edition names.September 2026 Edition (re-scored)
Vendor table v2026-09-08.3status: discontinued on eleven products so a positive label on a dead product is reportable. No resolution change.September 2026 Edition (re-scored)
Vendor table v2026-09-08.4Categories 13 (marketing automation and campaign management) and 14 (email marketing) added with a conflict disclosure: 49 vendors, category-scoped readings for HubSpot, Adobe, Oracle, Zoho, Microsoft, Salesforce, SAP, Twilio and Intuit, and existing vendors extended into the two categories. Gane is in the catalog so that if a model names it, the label is published like any other.September 2026 Edition (re-scored)
Vendor table v2026-09-08.5Tier-suffix rule (a product tier counts as the product). Aliases for Salesforce Account Engagement naming and four email vendors added after the first read of categories 13 and 14.September 2026 Edition (re-scored)
Judge rubricPending: jointly named products (A / B, A + B) return one object per product. Held until the calibration pair is complete so both runs are judged by the same rubric; historical combined names are split at report time by the vendor table rule.October 2026 Edition
Model setGoogle row switched from gemini-3.8-flash to gemini-3.1-pro-preview before the first edition because Flash never invoked search in testing.September 2026 Edition

Known weaknesses

Source capture reflects live retrieval only, not the pretraining corpus.
The category taxonomy is a judgment call and vendors will dispute their placement.
Fourteen categories is not the go-to-market landscape; the roadmap lists 73.
Reading a bare vendor name as its in-category product assigns specificity the model did not write. Every such reading is listed on the category page so the assumption is visible and reversible.
Self-preference is only measurable for a lab whose own products fall inside the category set. In this edition that is Google. Nothing is claimed about Anthropic or OpenAI.
The author operates Gane, whose product overlaps with two covered categories. Both are in the index with a disclosure rather than excluded, which a reader may reasonably weigh.