Ask the same question twice, get two numbers (“revenue” is three questions)

· once got three revenue numbers for one month🔁 with an ask-it-again simulator, not a lecture

Published October 4, 2026

The 60-second version

  • Different answers usually mean an ambiguous question — gross or net, IST or UTC, order or payment date — not a broken model.
  • Ask important questions three times: agreement is a confidence check, disagreement is a free warning to define the metric first.
  • The fix: save agreed definitions, make the AI state its assumptions, and ask instead of guessing — temperature zero only hides the problem.
bottom line: disagreement is a signal, not a bug

The CMO asks the AI assistant a simple question for Monday's slide: "What was our revenue last month?" The answer comes back: ₹48.2L.

Later, building the deck in a fresh chat, she asks again. ₹44.6L. A third try gives ₹48.9L. Same question, same data, same model — three different numbers. Now she trusts none of them.

Here is the twist: none of the three was a bug. Each was a reasonable reading of an ambiguous question. Gross or net of refunds? Calendar month in IST or UTC? "revenue" is three questions The model quietly picked one interpretation each time, and the pick changed between runs. The fix isn't a better model. It's turning that disagreement into a signal.

One question, three answers

Consistency check
Run 1 and run 3₹48.2LGross order value, IST month, cancelled orders excluded
Run 2 and run 5₹44.6LNet of discounts and refunds
Run 4₹48.9LGross, but the month counted in UTC
Gap between highest and lowest₹4.3L (9.6%)Bigger than most month-on-month changes
What was wrongThe question"Revenue" had no agreed definition
What fixed itOne saved definitionEvery run now returns ₹44.6L, net, IST
Five runs, three answers, zero bugs. The disagreement was the most useful thing the AI produced all day: it pointed straight at a definition nobody had written down.

Why the same question gets different answers

Three separate things make an AI analyst's answer change from one run to the next. Only one of them is "randomness".

Sampling

Language models pick each word with a bit of chance built in. Small differences early in the SQL can lead to a different query by the end.

Ambiguity

"Revenue last month" has several correct SQL translations — gross or net, IST or UTC, order date or payment date. The model must choose, and doesn't always choose the same.

Context

A fresh chat, a different earlier question, or slightly different schema notes fed to the model nudge it towards a different reading.

Press the button below to ask the same question five times. Each run shows the hidden assumption the model made and the bit of SQL that changed.

LIVE · ASK AGAIN“What was our revenue last month?”

you ask“What was our revenue last month?”

Run 1₹48.2L

Gross order value, September in IST, cancelled orders excluded.

SUM(total_price) WHERE status != 'cancelled'
₹48.2L
1 different answer in 1 run

Disagreement is free information

Most teams treat a changing answer as proof the AI is unreliable. Flip it around: if five runs agree, the question was probably clear. If they disagree, the question is ambiguous — and you've just learned that before putting a number on a slide.

This trick is called self-consistency: ask the same question a few times, compare the answers, and only trust the ones that agree.

One run

🎯 A single confident number

Fast, cheap, and unverifiable. You see ₹48.2L and have no way to know two other answers were equally likely.

  • Looks precise
  • Hides the assumption it made
  • Fails silently on ambiguous questions
Three to five runs

🧭 Agreement as a confidence check

Same answer every time? Use it. Different answers? Show the readings side by side and ask which one you meant.

  • Catches ambiguity before it reaches a slide
  • Costs a few extra model calls
  • Turns 'the AI is flaky' into 'define revenue'

Is checking worth the extra model calls? Usually by a mile:

LIVE · DRAG ITWhat does self-consistency cost?
50
3
₹2
15%
Extra model cost / month
₹6,000
Ambiguous questions flagged / month
225
Cost per question caught
₹27
CHEAP INSURANCE

Catching an ambiguous question costs less than a cup of coffee. One wrong number in a board deck costs far more.

live mathsestimates are fine — this is a what-if sandbox

Turning "temperature" to zero doesn't fix this. A model with no randomness gives the same answer every time — but it may be the same wrong reading every time, and now nothing warns you. Removing the symptom hides the ambiguity instead of resolving it.


Fix the question, not the temperature

Self-consistency finds the ambiguous questions. Definitions make them go away.

From ambiguous to answerable

Reporting Hierarchy
Tier 1
Write the definition once

Decide what 'revenue' means for your business — net or gross, which timezone, which date — and save it where the AI reads it every time. A semantic layer is the sturdy version of this.

revenue = net, IST, paid_at
Tier 2
Make the model state its reading

Every answer starts with its assumptions in plain words: 'Net revenue, September in IST, by payment date.' A wrong reading becomes visible in one glance.

assumptions first
Tier 3
Ask instead of guessing

When runs disagree, or a term has no saved definition, the AI should ask a clarifying question rather than pick one silently.

unclear → ask

The simplest durable fix is a canonical view: one place where "revenue" is computed the agreed way, which the AI is told to use instead of rebuilding the logic each time.

One agreed definition of revenue

With that view in place, "revenue last month" has one reading — and every run returns the same number for the right reason. For where definitions should live long-term, see semantic layers vs prompting; to test answers against known-good results, see benchmarking on your golden questions.

Consistency habits start with the top 10 metrics →

  • Run important questions three times — anything going into a report, a budget decision or a leadership slide. Disagreement means 'define it first'.
  • Require assumptions in every answer — net or gross, timezone, which date. If the AI can't say, it should ask.
  • Save each settled definition — every time a disagreement is resolved, write the answer into a view or semantic layer so it never comes back.

Quick gut-check

One question. If you get it, the whole post clicks. 30 seconds, promise

You ask an AI 'How many customers did we get last week?' three times and get 412, 412 and 389. What should you do?


Frequently asked questions

Why does ChatGPT or another AI give different answers to the same data question?

Three reasons: the model samples its words with some randomness, many business questions have more than one correct reading, and small context differences (a new chat, different earlier messages) push it towards different interpretations. Ambiguity is usually the biggest cause.

Should I set the AI's temperature to zero for analytics?

It makes answers repeatable, which helps, but it doesn't make them right. An ambiguous question will get the same interpretation every time, with no sign that another reading existed. Fix definitions first; use low temperature as a finishing touch.

What is self-consistency in AI?

It's asking a model the same question several times and comparing the answers. Agreement suggests the answer is stable; disagreement shows that the question or the model's reasoning is uncertain. In analytics it's a cheap way to catch ambiguous questions before they reach a report.


The summary

  • The same question can produce different AI answers because of sampling, ambiguity and context.
  • Ambiguity is the main culprit: "revenue" or "customers" often has several correct readings.
  • Ask important questions several times — disagreement is a free warning that the question needs a definition.
  • Temperature zero hides ambiguity rather than fixing it.
  • Save agreed definitions in views or a semantic layer and make the AI state its assumptions.

Takeaways for your next report

  • Different answers to the same question usually mean an ambiguous question, not a broken model.
  • Run important questions three times; if the answers disagree, define the metric before you use any of them.
  • Every answer should state its assumptions: net or gross, timezone, which date.
  • Temperature zero makes answers repeatable, not correct.
  • Each settled definition belongs in a view or semantic layer, so it never has to be argued again.
stick this on your Monday report
Chinmay Raibagkar

Chinmay Raibagkar

About author →

Founder of DataLens AI. He helps non-technical teams read their ad and database numbers with confidence — which number to trust, what to do next, and what to ignore.