The index is free to read and its whole record is free to download. It is also new, and it measures something nobody has a settled method for yet. These are the three ways to make it better that do not involve buying anything.
Every prompt, answer and judge label is downloadable on the Data page. If you run something against it and find something the index missed, send it here. Work that holds up is linked from this page with your name on it, and where it changes how a number is produced it is credited on the Method page and in the changelog.
A result someone else can rerun from the same files. A method note that measures something better than the current rubric does. An error in how a name, a share or a rank was resolved. Those get fixed and the correction is logged.
A reanalysis that happens to raise your own standing, or a finding whose data is not published so nobody can check it. Corrections to the vendor table go through claiming your page, not through here.
Every edition is 15,768 model calls across 12 models, and then every product any of them named is labeled by the judge, with the evidence quote kept. That bill arrives monthly whether or not anyone subscribes. Contribute whatever it is worth to you; there is nothing behind it, and nothing is withheld from anyone who does not.
This section comes down once subscriptions cover the running cost. It is here because they do not yet, and saying so is more honest than pretending the runs pay for themselves.
Every answer in the index comes from an API call with no memory, no history and no personalisation, in a fresh session. A person asking the same question inside a consumer app is not in that situation: the app may carry memory of them, a different system prompt, a different model version, and a profile built over months. Whether the recommendation changes because of that, and by how much, is a real question this index cannot answer from API calls alone.
Either answering the buying questions yourself, or running them through the consumer apps you already use and sending back what you got, along with which app, whether memory was on, and whether the session was fresh.
Your role, the size and industry of the company you buy for, and the memory setting. That is what makes a human answer comparable to an API one. It is published in the same form as everything else: aggregated, never as a person.
Rates depend on how much of the prompt set a round covers, and are agreed before anything starts. Panel answers are labeled as such wherever they appear and are never mixed into the model figures.