What this measures
The index records what frontier AI models say when a buyer asks them which go-to-market product to use. It is not a review site and collects no user ratings. G2 measures what buyers claim after purchase; this measures what buyers are told before purchase.
Each edition the same prompts, six per category, go to the same models with web search enabled, in a fresh session per prompt with no conversational carryover. A judge model labels every product each answer names. The labels are the permanent record. Everything published is derived from them and can be recomputed.
Categories
Fourteen categories so far, across four of the six go-to-market verticals, chosen for high buyer search intent and a contested vendor landscape; two carry a conflict disclosure under the conflict policy below. The roadmap on the Data page lists 73 across all six. The AI visibility category is deliberate: ranking the answer engine tracking vendors inside an AI recommendation index is a reflexive test of the method.
The taxonomy is the index's own. To keep it legible to vendors and buyers, each category is cross-referenced to where it sits in the MartechMap, the marketing technology landscape by chiefmartec and MartechTribe, in G2's category tree, and in current analyst evaluations. Those taxonomies belong to their publishers and are cited by name, not reproduced. Analyst references were checked against the publisher on 2026-09-08.
Roadmap: 14 categories covered, 15 planned, 44 open, for an addressable set of 73 across six go-to-market verticals: Marketing (7 covered of 27), Sales (3 covered of 12), Customer (1 covered of 9), Revenue operations (0 covered of 9), GTM data and infrastructure (3 covered of 12), Partner and channel (0 covered of 4). The Management domain of the landscape (7 categories) is out of scope: Team and operations tooling rather than products a buyer asks an AI to recommend for marketing. Not in the addressable set.
| # | Category | Buyer phrase | Landscape reference | G2 | Analyst coverage |
|---|---|---|---|---|---|
| 01 | Customer data platforms | customer data platform | Data › Customer Data Platform | Customer Data Platforms (CDP) | Gartner Magic Quadrant for Customer Data Platforms (2026) |
| 02 | Product analytics | product analytics platform | Data › Mobile & Web Analytics | Product Analytics | — |
| 03 | Content management systems for marketing sites | marketing website CMS | Content and Experience › CMS & Web Experience Management | Web Content Management | Gartner Magic Quadrant for Digital Experience Platforms (2026) |
| 04 | SEO and content optimization platforms | SEO and content optimization platform | Content and Experience › SEO | SEO Tools, Content Marketing | Gartner Magic Quadrant for Content Marketing Platforms (2026, adjacent) |
| 05 | Webinar and virtual event platforms | webinar and virtual event platform | Social and Relationships › Events, Meetings & Webinars | Webinar Platforms, Event Management | — |
| 06 | Sales engagement platforms | sales engagement platform | Commerce and Sales › Sales Automation, Enablement & Intelligence | Sales Engagement | Gartner Magic Quadrant for Revenue Action Orchestration (first edition, December 2025; succeeds the Sales Engagement Applications market) |
| 07 | Conversational intelligence and call recording | conversation intelligence and call recording platform | Social and Relationships › Call Analytics & Management | Conversation Intelligence | Gartner Magic Quadrant for Revenue Action Orchestration (December 2025, adjacent: conversation intelligence is one of its converged capabilities) |
| 08 | Attribution and marketing mix modeling | marketing attribution and marketing mix modeling platform | Data › Marketing Analytics, Performance & Attribution | Marketing Attribution, Marketing Analytics | Gartner Magic Quadrant for Marketing Mix Modeling Solutions (2025); Forrester Wave: Marketing Measurement and Optimization Services (Q1 2026) |
| 09 | Customer support and helpdesk | customer support and helpdesk platform | Social and Relationships › Customer Experience, Service & Success | Help Desk | Gartner Magic Quadrant for the CRM Customer Engagement Center (2025) |
| 10 | Data warehouse and reverse ETL for marketing | marketing data warehouse and reverse ETL stack | Data › iPaaS, Cloud/Data Integration & Tag Management | Reverse ETL, Data Warehouse | Gartner Magic Quadrant for Cloud Database Management Systems (2025); Gartner Magic Quadrant for Data Integration Tools (2025) |
| 11 | AI visibility and answer engine tracking tools | AI visibility and answer engine tracking tool | Content and Experience › SEO (AEO/GEO, noted by the landscape as an accelerating area) | Generative Engine Optimization | — |
| 12 | B2B intent data providers | B2B intent data provider | Data › Audience/Marketing Data & Data Enhancement | Buyer Intent Data Providers | — |
| 13 | Marketing automation and campaign management | marketing automation platform | Content and Experience › Marketing Automation & Campaign/Lead Management | Marketing Automation | Gartner Magic Quadrant for B2B Marketing Automation Platforms (2025); Gartner Magic Quadrant for Multichannel Marketing Hubs (2025) |
| 14 | Email marketing platforms | email marketing platform | Content and Experience › Email Marketing | Email Marketing | Forrester Wave: Email Marketing Service Providers (Q1 2026) |
Prompts
Six framings per category. Variants one and two are semantically identical and lexically different; they exist to measure paraphrase stability. Variant six captures anti-recommendations, which carry commercial weight and which most trackers ignore. Every prompt is published verbatim.
Every prompt frames the buyer as a mid-market B2B software company, and the scale prompt says 500 people. The index does not yet ask on behalf of a small business or an enterprise, so every standing on this site is a mid-market standing. Company size, industry, question framing and who answers are the four dimensions of the buyer; today there is one combination of them, and the Data page lists the planned ones, including small-business and enterprise sizes, other industries, a job-to-be-done framing that describes the work instead of naming the category, and human buyers answering the same questions.
| # | Template | Purpose |
|---|---|---|
| 01 | What is the best {phrase} for a mid-market B2B software company? | Direct recommendation. Paired with variant 2 for paraphrase stability. |
| 02 | Which {alt_phrase} would you recommend to a mid-sized B2B software company? | Semantically identical to variant 1, lexically different. Uses alt_phrase. |
| 03 | What are the top {plural} and how do they differ? | Comparative. Surfaces the considered set and how the model differentiates. |
| 04 | What is the best {phrase} for a company with a limited budget? | Constrained by budget. |
| 05 | We are a 500 person company evaluating a {phrase}. What should we look at? | Constrained by scale. |
| 06 | Which {plural} should I avoid or be cautious about? | Negative framing. Captures anti-recommendations. |
Settings: temperature provider default, max output 4,000 tokens, no system prompt, fresh session per prompt, Anthropic search capped at 5 uses per answer.
Model set
Coverage across labs rather than depth within one: two Anthropic models spanning a generation boundary, two OpenAI models spanning one, one Google model, one challenger. Version strings are recorded exactly as returned on every call. A change to any string is a new model row and forces a noise floor recalibration for that model.
| Model | Version string | Lab | Generation | Search tool | Endpoint |
|---|---|---|---|---|---|
| Claude Opus 5 | claude-opus-5 | anthropic | current | web_search_20260209 | api.anthropic.com |
| Claude Opus 4.8 | claude-opus-4-8 | anthropic | prior | web_search_20260209 | api.anthropic.com |
| GPT-6 Astra | gpt-6-astra | openai | current | web_search | api.openai.com |
| GPT-5.6 Sol | gpt-5.6-sol | openai | prior | web_search | api.openai.com |
| Gemini 3.1 Pro | gemini-3.1-pro-preview | single | google_search | generativelanguage.googleapis.com Switched from gemini-3.8-flash on 2026-09-08: Flash never invoked Google Search in testing, Pro grounds. Preview build. | |
| Perplexity Sonar Pro | sonar-pro | challenger | single | native | api.perplexity.ai |
Scoring
Only products the answer treats as candidates in the asked category are counted, so a warehouse named as a data source in a CDP answer is not a CDP candidate. When the answer recommends a stack, only the product filling the category's role is a first choice; the rest are alternatives.
Categories, methodologies, analyst firms and people are never counted, and nothing named only inside a cited URL is counted.
A Claude Opus 5 judge, disclosed as an Anthropic model, reads each answer with the prompt and category and returns one object per named product with the name exactly as written, a label, its position and a verbatim evidence quote. Judge is an Anthropic model and is disclosed as such. It runs at effort low with structured output so labels are deterministic given the text.
Published weights
| Label | Weight | Meaning |
|---|---|---|
| First choice | +3 | The product the answer leads with for the asked category |
| Alternative | +2 | Named as a viable option alongside the first choice |
| Mention | +1 | Named without endorsement |
| Soft negative | −2 | Named with a caveat that discourages the buyer |
| Hard negative | −3 | Named as something to avoid |
Derived metrics
- paraphrase stability
- Per model, the share of categories where the first-choice set on the direct prompt equals the set on the paraphrase.
- first-choice share
- Per category and product, first-choice labels across the direct, paraphrase, budget and scale prompts and all models, divided by all first-choice labels in the category. The comparative and negative prompts do not ask for a pick and are excluded.
- contested
- No product holds more than 40% of first choices. The top share is always published next to the label because a category just above the line is not meaningfully different from one just below it.
- negative rate
- Per category and product, soft plus hard negative labels divided by all labels. Separates sentiment from salience.
- quadrants
- Products with at least 10 labels in a category placed by first-choice share and negative rate. Leader at 30% share or more, criticized at 25% negative or more: endorsed leader, criticized default, criticized challenger, accepted challenger.
- lab treatment
- For a lab with products in the category set, how its own model labels those products against how every other model labels the same products, as mean label weight. Only measurable with that control group; a lone self-preference count is not published.
- discontinued
- A positive label on a product the catalog marks discontinued. A retrieval failure worth naming.
- noise floor
- Per model, the share of (model, prompt) pairs whose first-choice set differs between an edition run and its calibration repeat. Pooled across models for movement verdicts.
Normalization
Product names are resolved through a versioned vendor table (v2026-09-08.5: 538 vendors with aliases, 20 bare names with category-scoped readings, 9 exclusions). Matching is case-insensitive and strips a trailing parenthetical. Resolution order for a raw name in a category:
- category alias
- global alias or canonical name
- exclusion
- trailing tier words removed, then the first two steps again
- split on separators with every part resolving
- unresolved
A bare vendor name resolves to that vendor's product for the asked category when it has exactly one (Salesforce in customer support is Service Cloud). Where the vendor has no product in the asked category, the name stays unresolved and is listed in the report. This is a directional assumption: a model writing a bare vendor name may mean the platform generally rather than the in-category product. The report lists every category-scoped resolution so the assumption is visible and reversible. A combined answer produces one label per product with the same label and evidence. Only applied when every part resolves; otherwise the name stays unresolved.
Adding a vendor can change how historical raw names resolve (a bare Salesforce in attribution stays unresolved only until a Salesforce attribution product exists). The catalog therefore follows the same discipline as the aliases: every change bumps the version and regenerates the series.
Excluded by rule: Google Ads (ad platform self-reported attribution, not a vendor in the category); Meta Ads Manager (ad platform self-reported attribution, not a vendor in the category); LinkedIn Ads (ad platform self-reported attribution, not a vendor in the category); TikTok (ad platform self-reported attribution, not a vendor in the category); Amazon (ad platform self-reported attribution, not a vendor in the category); Meta Conversion Lift (ad platform lift study, not a vendor in the category); Google geo experiments/Ads lift studies (ad platform lift study, not a vendor in the category); Native Google, Meta, TikTok and Amazon attribution (ad platform self-reported attribution, not a vendor in the category); Excel (general-purpose tool, not a vendor in the category).
Noise floor
Before any trend is claimed, the full matrix runs twice within one week with nothing changed: same prompts, same model versions, same settings. Whatever differs between the two runs is variance, not movement, and it sets the threshold below which the index says nothing changed.
It is reported per model as a first-choice flip rate, because models differ here and the differences are themselves a finding. A model whose repeat-run flip rate is as high as its paraphrase instability is not showing a paraphrase effect; its stability number is a floor and is marked as one wherever it appears.
Status for the September 2026 Edition: the calibration repeat is scheduled for 12 September 2026. The figure shown on the index until then comes from a simulated repeat and is labeled as such.
Data capture
Every prompt and every full response is stored verbatim with its timestamp, the model version string as returned, whether the model invoked search, the cited source URLs it exposed, latency, token counts and cost. Cited sources reflect live retrieval only and say nothing about pretraining data.
The judge's raw labels, one per named product with the name exactly as written and a verbatim evidence quote, are the permanent record.
An edition is 504 model calls. The September 2026 Edition cost $100.77 in model calls and $14.36 in judging; 86% of answers invoked search; median latency 42 seconds.
Editions and cadence
A monthly edition on the same date each month, which matches the pace at which the drivers move: model version updates and shifts in the retrieval corpus. A special edition within days of a major frontier model release, published as a comparison against the prior generation. Quarterly written analysis, only once enough editions have accumulated to say something with substance.
Every edition is archived permanently at a stable URL. If the project ends, the final edition is marked final on the site.
Conflict policy
The author operates Gane, which is building an operating system for B2B go-to-market. Its product overlaps with marketing automation and campaign management and with email marketing, so both categories are in the index with a conflict disclosure on each page rather than an exclusion. The data in those categories is collected and scored exactly like every other category and published unmodified; if Gane is named by a model, that label is published like any other. The disclosure exists so a reader can weigh it, and the raw record exists so a reader can check it.
The judge is an Anthropic model and is disclosed as such on every page.
Change log
Instrument changes: the vendor table, the judge rubric, the model set. Each entry names the first edition scored under it.
| Instrument | Change | First edition under it |
|---|---|---|
| Vendor table v2026-09-08.1 | Seeded 38 CDP-centric vendors, then 438 vendors from the September 2026 Edition normalization backlog across all twelve categories. | September 2026 Edition (re-scored) |
| Vendor table v2026-09-08.2 | Category-scoped aliases for bare parent and multi-product names (Salesforce, HubSpot, Adobe, Oracle, SAP, Semrush, Ahrefs, Salesloft, Outreach, Zoom, Clari, Gong). Split rule for combined answers such as A / B and A + B. Exclusion list for ad-platform self-attribution and general-purpose tools. Nine products added from the 38 unresolved September 2026 Edition names. | September 2026 Edition (re-scored) |
| Vendor table v2026-09-08.3 | status: discontinued on eleven products so a positive label on a dead product is reportable. No resolution change. | September 2026 Edition (re-scored) |
| Vendor table v2026-09-08.4 | Categories 13 (marketing automation and campaign management) and 14 (email marketing) added with a conflict disclosure: 49 vendors, category-scoped readings for HubSpot, Adobe, Oracle, Zoho, Microsoft, Salesforce, SAP, Twilio and Intuit, and existing vendors extended into the two categories. Gane is in the catalog so that if a model names it, the label is published like any other. | September 2026 Edition (re-scored) |
| Vendor table v2026-09-08.5 | Tier-suffix rule (a product tier counts as the product). Aliases for Salesforce Account Engagement naming and four email vendors added after the first read of categories 13 and 14. | September 2026 Edition (re-scored) |
| Judge rubric | Pending: jointly named products (A / B, A + B) return one object per product. Held until the calibration pair is complete so both runs are judged by the same rubric; historical combined names are split at report time by the vendor table rule. | October 2026 Edition |
| Model set | Google row switched from gemini-3.8-flash to gemini-3.1-pro-preview before the first edition because Flash never invoked search in testing. | September 2026 Edition |