Skip to content
Why we do / don’t

Why are ChatGPT and Gemini measured on the consumer plane?

ChatGPT and Gemini are measured on the consumer plane because that is where buyers read the answer. A developer API returns a different response than the same question asked in the product interface, so an API-only reading would describe a surface your customers never see. The divergence is the point, not a footnote.

An API answer and an interface answer are produced by different retrieval and grounding paths. They can name different sources, order them differently, and attribute differently — for the same question, asked the same minute.

That is why ChatGPT and Gemini are collected from the consumer plane wherever we can, and why we state which plane every other engine came from instead of implying they are all alike. Our remaining measured engines run through official developer APIs or through a search surface, and each carries its own label.

What this means for your numbers: a consumer-plane reading is closer to what a buyer would have seen, and it is still a sample. Answers drift between runs, so a single reading is directional and a trend across runs is the thing worth acting on.

It also sets the bar for adding an engine. If we cannot collect it on a plane we can describe, we would rather not report it than report it under a label that overstates what we saw.

What this answer is based on

  • Engine registry — per-engine collection plane and provenance label
  • Provenance labelling policy

Related answers

Last updated August 3, 2026

← All answers