Ask an AI answer engine whether a company is reliable and you receive what looks like a finished assessment. Clear wording. Confident advice. A neat list of sources.
But would the same engine give you the same answer five minutes later?
To find out, we submitted one simple reputation question—“Is Temu reliable?”—ten times each to ChatGPT, Gemini, and Perplexity. That produced 30 answers collected on 7 October 2026, all in English, in guest/default mode, with visible Web search.
The headline was remarkably stable. All 30 answers treated Temu as a real marketplace that could be used selectively for inexpensive, non-critical purchases. Not one recommended trusting it for expensive or safety-sensitive products.
Yet beneath that shared conclusion, the engines told three noticeably different stories.
That difference is the real finding—and it matters far beyond Temu.
The experiment in one minute
This was a deliberately narrow pilot designed to measure AI answer stability, not to declare whether Temu is objectively reliable.
The setup was straightforward:
- One exact prompt:
Is Temu reliable? - Three engines: ChatGPT, Gemini, and Perplexity
- Ten responses per engine
- Thirty responses in total
- One collection date: 7 October 2026
- Guest/default mode with explicit Web search
- Responses, timestamps, claims, citations, and source URLs preserved
We compared each engine only with itself. Ten responses produce 45 unique within-engine pairs, so the study contained 135 pairwise comparisons in total.
Each answer was coded for 17 possible claims or pieces of advice: legitimacy, delivery, product quality, refunds, safety, privacy, seller variability, regulatory action, complaints, payment protection, and others. We then measured how much the claim sets overlapped from one run to the next.
This is a pilot, not a universal ranking. One company, one prompt, one language, and one day cannot establish how an engine always behaves. It can, however, expose a measurement problem that SEO, GEO, and online reputation teams routinely overlook.
The surprising result: the recommendation never changed
Across all 30 responses:
- 30/30 described Temu as a legitimate marketplace rather than an outright scam.
- 30/30 described product quality as variable.
- 30/30 discussed delivery reliability or timing.
- 30/30 recommended using Temu only for cheap, non-essential, or low-risk purchases.
- 30/30 advised against relying on it for expensive or safety-sensitive products.
- 0/30 recommended avoiding the platform for every purchase.
- 0/135 within-engine pairs produced a positive-to-negative verdict flip.
If we had measured only the final recommendation, the result would have looked almost boring: the engines agreed, and they kept agreeing.
But a reputation is not just a verdict. It is the collection of facts, risks, associations, and sources used to justify that verdict.
Once we examined those elements, the apparent consensus fractured.
One verdict, three reputation narratives
ChatGPT, Perplexity, and Gemini did not construct “Temu reliability” from the same evidence.
ChatGPT told a regulatory-and-complaints story
ChatGPT was the most stable at the claim level. Its answers repeatedly returned to the same group of ideas:
- product quality varies;
- third-party sellers create inconsistency;
- consumer complaints matter;
- product safety deserves caution;
- a US enforcement action is relevant.
ChatGPT mentioned an actual enforcement action in 10 out of 10 answers and negative complaint or review evidence in 10 out of 10. It mentioned seller variability every time.
Privacy concerns, however, appeared in none of its ten answers. Neither did the most recent major EU enforcement event in our dated truth set.
High stability therefore did not mean complete coverage. It meant that ChatGPT repeated a narrow narrative very consistently—including its omissions.
Perplexity told the broadest story
Perplexity covered the widest range of issues. Its responses combined product quality, refunds, privacy, consumer reviews, enforcement, and practical shopping precautions.
It mentioned refund or support friction in all ten runs, privacy concerns in nine, and negative review evidence in all ten. It was also the only engine to surface the May 2026 EU fine repeatedly.
Its framing moved around more than the other engines. Four responses opened with a conditionally positive stance, while six used explicitly mixed or uneven language. Yet the practical recommendation never changed: limit Temu to low-risk purchases.
This distinction is easy to miss. The tone varied; the decision advice did not.
Gemini told a privacy-and-safe-shopping story
Gemini consistently emphasized privacy, payment protection, variable quality, and practical purchasing precautions.
Privacy appeared in 10 out of 10 responses, as did protected-payment advice. But Gemini did not mention negative complaint evidence in any run. It also omitted a concrete enforcement action in all ten responses.
Gemini repeatedly supplied a general one-to-two-week delivery estimate. Temu’s current French terms instead provide listing-specific estimates and an approximate delivery date before purchase, so the repeated universal timeframe was more definite than the official policy supports.
The same brand was therefore associated with different risks depending on the engine:
| Reputation theme | ChatGPT | Perplexity | Gemini |
|---|---|---|---|
| Refund or support friction | 9/10 | 10/10 | 0/10 |
| Privacy or data concern | 0/10 | 9/10 | 10/10 |
| Actual enforcement or penalty | 10/10 | 7/10 | 0/10 |
| Negative complaint or review evidence | 10/10 | 10/10 | 0/10 |
| Third-party seller variability | 10/10 | 4/10 | 5/10 |
| Protected-payment advice | 4/10 | 9/10 | 10/10 |
For online reputation management, this is more consequential than a simple positive, neutral, or negative label. A company can receive the same overall verdict while being defined by a different risk narrative in every engine.

Which engine was the most stable?
The answer depends on what “stable” means.
| Engine | Claim overlap | Lexical similarity | Source-page overlap | Pilot stability score |
|---|---|---|---|---|
| ChatGPT | 0.899 | 0.592 | 0.659 | 88.04 |
| Gemini | 0.763 | 0.558 | 0.491 | 80.89 |
| Perplexity | 0.774 | 0.491 | 0.933 | 74.11 |
ChatGPT had the highest claim stability. Perplexity had the most stable source-page set. Gemini maintained a stable overall stance and relatively consistent wording, but its precise source pages changed more.
The pilot stability score shown here combines claim overlap, lexical similarity, stance, recommendation, caveats, and contradictions. It is an exploratory communication device, not a validated industry standard. The individual measurements are more useful than the final number.
This is precisely why asking “Which engine is most consistent?” is too vague. An answer can be stable in its conclusion, unstable in its reasoning, and highly repetitive in its sources—all at once.
More links did not mean more independent evidence
The source analysis produced another warning for GEO measurement.
ChatGPT used a narrow recorded source pool: four canonical pages across the entire experiment. Its run-level sources regularly included the FTC, BBB, and Which?. This produced a stable but concentrated evidence stack.
Perplexity recorded ten unique pages in every run. Runs two through ten reused the same ten-page set. Its source overlap was therefore exceptionally high—but none of the 100 run-level page occurrences came from a regulator, court, company filing, or another domain classified as official primary evidence.
That is not the same as ten independent confirmations. It looks more like a persistent retrieval pool.
Gemini’s source field contained 101 URL mentions. After removing duplicate text fragments and canonicalizing the links, those mentions collapsed to 41 unique page occurrences within runs and only 11 unique pages across the entire dataset. The raw mentions were heavily concentrated in Reddit, Avast, and Panda Security.
This creates a practical rule for AI visibility research:
Never use raw citation count as a proxy for evidence diversity.
A serious source audit should distinguish:
- raw URL mentions;
- canonical pages;
- independent domains;
- primary versus secondary sources;
- source recurrence across repeated runs;
- whether a citation actually supports the associated claim.

Stability is not accuracy
An engine can repeat the same incomplete or outdated story perfectly.
We therefore checked selected claims against a dated official-source truth set. Two examples illustrate the difference.
First, the European Commission announced a €200 million Digital Services Act fine against Temu on 28 May 2026. Perplexity surfaced this event repeatedly. ChatGPT and Gemini did not make it part of their stable narrative. The Commission’s announcement says the risk assessment failed to examine illegal-product risks diligently and underestimated how often EU consumers could encounter illegal products.
Second, a US federal court entered an order in September 2025 resolving allegations under the INFORM Consumers Act. The order included a $2 million civil penalty and compliance measures. ChatGPT repeatedly surfaced this event and regularly included an FTC source in its recorded source set. The US Department of Justice announcement describes the settlement and allegations.
The point is not that one engine found the “right” controversy. The point is that current authoritative evidence existed, yet no engine consistently assembled the complete official record.
The same caution applies to consumer-policy claims. Temu’s France-facing policy provides a voluntary return window of up to 90 days, but it includes exclusions and shorter periods for certain electronics and large appliances. “Temu offers 90-day returns” can therefore be directionally useful and still misleading when stated without qualifications. The details are available in Temu’s current return and refund policy.
What this means for SEO, GEO, and ORM teams
One answer is a screenshot, not a measurement
A single run would have captured the broad Temu recommendation correctly in this pilot. It could still have produced the wrong conclusion about whether privacy, enforcement, complaints, seller variability, or refunds were central to the engine’s narrative.
Monitoring one answer once is not AI reputation measurement. It is anecdotal observation.
Sentiment dashboards miss the story
All three engines arrived at roughly the same practical advice. A sentiment-only dashboard would report strong agreement.
The claim-level data revealed the strategically important differences. ChatGPT associated the brand with enforcement and complaints. Gemini associated it with privacy and safe payment. Perplexity assembled a wider but largely secondary evidence pool.
For ORM, the question is not only “Is the answer positive or negative?” It is also:
- Which risks recur?
- Which facts disappear?
- Which sources define the entity?
- Does the official response appear?
- Is an old controversy written in the present tense?
- Does the recommendation change even when the tone does not?
Stable visibility can reveal a weakness
Perplexity’s nearly fixed source set looks excellent on a stability chart. From another perspective, it may indicate a narrow retrieval universe or reuse within a short collection window.
ChatGPT’s high claim stability makes its recurring story predictable. It also makes recurring omissions predictable.
Stability is valuable because it shows what users are likely to encounter repeatedly. It should never be confused with completeness, authority, or truth.
What this pilot does—and does not—prove
Thirty responses are enough to expose variation and test a methodology. They are not enough to support universal claims about the three engines.
The pilot is limited by:
- one company and one reputation question;
- one language;
- one collection date;
- sequential rather than interleaved engine collection;
- consumer interfaces with unpinned underlying models;
- incomplete documentation of browser, location, cookies, and personalization;
- a single-coder exploratory claim taxonomy;
In short, this pilot reveals a pattern worth testing—not a universal conclusion about AI answer stability.
The takeaway
In this 30-response pilot, the answer to “Is Temu reliable?” was stable at the decision level and unstable at the narrative level.
All three engines advised roughly the same thing: Temu may be suitable for cheap, low-risk purchases, but not for products where quality or safety matters.
What changed was the reputation surrounding that advice:
- ChatGPT repeated a regulatory-and-complaints narrative.
- Perplexity delivered the broadest narrative from an almost fixed pool of consumer, review, cybersecurity, media, and community sources.
- Gemini emphasized privacy and safe-shopping practices while omitting concrete enforcement and complaint evidence.
The lesson for GEO and ORM is simple:
A brand’s AI reputation is not one answer. It is a distribution of recurring claims, omissions, recommendations, and sources.
Until those elements are measured separately, “AI visibility” remains an impression—not an evidence-based metric.
Frequently asked questions
Is 30 responses enough for an AI answer-stability study?
It is enough for a clearly labelled pilot that demonstrates variation, tests a coding framework, and generates hypotheses. It is not enough to generalize across all prompts, dates, languages, locations, or model versions.
Which engine was the most stable?
ChatGPT had the highest claim-level stability in this pilot. Perplexity had the most stable recorded source-page set. The answer therefore depends on whether stability refers to claims, wording, stance, recommendation, or sources.
Did any engine call Temu a scam?
No. All 30 responses described it as a legitimate marketplace, although every response also advised caution and limited the recommendation to inexpensive, non-critical purchases.
Does a stable answer mean an accurate answer?
No. A system can repeat an incomplete, weakly sourced, or outdated narrative consistently. Stability and factual accuracy require separate evaluation.
Why should brands repeat AI prompts?
Because one run may omit a risk, citation, official response, correction, or current event that appears in other runs. Repetition reveals which elements are persistent and which are volatile.