We Asked Perplexity and ChatGPT the Same 20 Candle Questions. They Agreed on 19% of the Brands.
If you sell something, you have probably wondered which AI assistant to optimise for. The honest answer, on this evidence, is that they are not the same target.
We put 20 real buyer questions about candles and home fragrance to Perplexity and to ChatGPT, in the same browser session on the same day, and recorded every brand each one named. Across both engines the answers contained 66 distinct brands. Only 15 of those 66 were named by both. Measured as set overlap, the agreement is 0.19.
That number on its own is just a curiosity. What makes it useful is where the disagreement sits.
Naming the brand collapses the disagreement
Sixteen of our twenty questions were open-ended — “what are the best soy candles?”, “best cheap candles that smell good”. Four supplied the brands in the question itself — “Diptyque vs Voluspa: which is worth the money?”, “Is Yankee Candle worth buying?”.
Split the results that way and the pattern is not subtle:
| Question type | Questions | Brand overlap between engines |
|---|---|---|
| Open-ended (“best X for Y”) | 11 with brands | 0.00 – 0.29 |
| Brand named in the question | 4 | 1.00 on all four |
On every question where we named the brands, both engines discussed exactly those brands and nothing else. On every question where we did not, they largely named different companies.
This is worth sitting with, because it inverts a common assumption. The assistants are not disagreeing about which brands are good. When handed the same candidates they reason about them similarly — both picked Voluspa over Diptyque on value, and for much the same reason. They are disagreeing about which brands exist, or at least which ones are worth retrieving. The judgement is shared; the shortlist is not.
For a merchant, that means being “known to AI” is not one problem. It is one problem per engine, at the retrieval stage, before any question of whether your product is actually good.

The two engines are reading different things
The overlap number understates how different the answers feel, because the engines are not even producing the same kind of artefact.
Perplexity answered in prose. Typical length was 3,000–4,200 characters, structured with headings and comparison tables, and every substantive claim carried a numbered citation — to the New York Times, to Forbes, to niche review sites, and very often to Reddit.
ChatGPT free tier mostly answered with a product carousel: a row of cards, each with a brand, a product name, a price, and a one-line blurb. Its answer to “what are the best candles for gifts?” was 378 characters long in total, of which the prose portion was the two words “If you”. No citations.
The sources differ accordingly. Asked for cheap candles that smell good, Perplexity returned Yankee Candle, Public Goods, P.F. Candle Co., Otherland and Goose Creek, citing Forbes and Homes & Gardens. ChatGPT opened with “Walmart’s candle selection is surprisingly strong” and returned Better Homes & Gardens, Mainstays, Allswell and Goose Creek — three of which are Walmart house brands. One brand in common.
That reads less like two models disagreeing and more like two different indexes: one built from editorial and forum content, one built from retail product feeds.

Perplexity’s answer to a neighbouring question shows the other format: prose, a comparison table, and a citation on each claim.

Some questions have no brands to win
Four of the twenty questions produced no brand names at all, in either engine: soy vs beeswax, soy vs paraffin, “are expensive candles worth it”, and “are soy candles worth paying more for”. Both engines answered them with material properties, safety notes and price-band reasoning, and named nobody.
These are high-intent, high-volume questions. They are also, for a merchant, unwinnable in the sense that matters: there is no recommendation slot to occupy. Content aimed at them can still earn citations, but it will not put your name in the answer.
The inverse is more actionable. Scenario questions — relaxation, sleep, small apartments, pet odours, fall scents — were the most brand-dense in our set. Perplexity named eight brands for “best candles for relaxation” and six for “best candles to help you sleep”. If you are choosing where to compete, that is where the slots are.
One brand won both engines
Across all 40 answers, one name appeared more than any other: Voluspa, with 5 mentions in Perplexity and 9 in ChatGPT. It led the luxury-under-$50 question in both. In ChatGPT’s answer to that question it held six of the eight ranked picks, including the closing recommendation.

It also won the head-to-head we set up deliberately. Asked “Diptyque vs Voluspa: which is worth the money?”, both engines chose Voluspa, and both did the same arithmetic: Perplexity put Diptyque at roughly $1.30–1.60 per burn hour against Voluspa’s $0.38; ChatGPT gave Voluspa five stars for value against Diptyque’s three. Brand prestige did not survive contact with cost-per-hour.
The one head-to-head where the engines split was Trudon vs Jo Malone London. Perplexity declined to pick, calling it a tie. ChatGPT picked Trudon and wrote a “Why Trudon wins” section.
The SEO winner that AI would not recommend
While designing the questions we ran five broad probes through Google’s AI Overviews to find out which brands appear most often. One name kept surfacing in the ordinary search results and never once in an AI answer.
A brand called Lumond ranked #1 in its own “Top 10 Scented Candles 2026” listicle on its own domain, appeared at #1 in a third-party ranking, and held the surfaced top answer in a Reddit thread. On conventional search it looks like a category leader.
It was named as a recommended brand zero times — not in any of the five Google AI Overviews, not in any of Perplexity’s 20 answers, not in any of ChatGPT’s 20.
It did appear once, in a way worth noticing. Answering “Is Yankee Candle worth buying?”, Perplexity cited lumond.com twice — as a source for two claims about Yankee Candle, a competitor. The content farm succeeded at becoming a citation and failed at becoming a recommendation.
Those are different outcomes, and conflating them is easy. Being read is not being chosen.

What we would do with this
Three things follow, if this pattern holds beyond one category:
- Audit per engine, not “AI” in general. A brand strong in Perplexity may be absent from ChatGPT, and the fix is different in each case — editorial and forum presence for one, retail product-feed presence for the other.
- Chase scenario questions, not material questions. “Best candles for a small apartment” has six recommendation slots. “Soy vs beeswax” has none.
- Do not read citations as endorsements. Being quoted as a source and being named as a pick are separate outcomes, and the Lumond case shows you can have the first without the second.
Method
Twenty questions across five categories — best-for, brand comparison, price band, worth-it, and scenario. The brands used in the comparison and worth-it questions were not chosen by us: we ran five category-level probes through Google AI Overviews first and took the most frequently named brands from those AI answers, so the head-to-heads reflect what the engines were already talking about.
Both engines were queried on 2026-08-28 from one clean Chrome profile created for this test, with no prior browsing history, signed in to free-tier accounts. Perplexity was queried by URL; ChatGPT by typing each question into a fresh session. Answers were read from the rendered page and transcribed verbatim. Brand rank is order of first appearance.
Limitations, stated plainly:
- One run, one day. AI answers vary between runs. Nothing here is a stable ranking, and we have not repeated the test.
- Free tiers. Paid tiers may retrieve differently, and ChatGPT’s product-carousel behaviour in particular may be tier- or region-specific.
- One profile, one region, one language. The browser was signed in, not anonymous, on a residential connection in one country. The account’s interface language is Chinese, which localised some ChatGPT product blurbs and conversation titles; answer prose was English throughout and brand names were unaffected.
- Overlap is computed on normalised brand names. “Nest” and “Nest New York” are treated as one brand for the overlap figure, and as written in the per-question data.
- Our first pass undercounted. An initial capture truncated long answers and lost brands named late; three questions were wrongly recorded as having no brands. Every answer was re-read in full before publication. The corrected figures are the ones above.
The full dataset is published as CSV: every brand mention with rank and reason, the cross-engine comparison by question, brand presence across engines, the Lumond tracking log, and the question list.
If you repeat this in another category we would like to see what you get — particularly whether the “naming the brand collapses the disagreement” effect holds. It is the finding here we would most like to see someone else fail to reproduce.