Does naming brands influence AI answer sentiment?
Ask an AI assistant which washing machine brands are most reliable and it is free to choose. Ask whether Samsung is better than LG and you have already chosen the candidates for it. Both questions are perfectly reasonable in AI visibility tracking, but we wanted to know whether they also invite different sentiment.
The problem is that, with most AI visibility tracking tools, sentiment is often presented as a single score that describes a brand in the abstract. In reality, they describe how a brand was treated in a particular answer to a particular question.
What was tested
I took 48 questions from a UK home electronics and appliances study recently conducted for Geometriqs and ran two variants through ChatGPT with web search. The first version was the original, with no brand in the question (e.g. "For pure performance, which TV brands or specific models should I have on my radar?"). The second kept the same subject but ended by asking whether Samsung was better than LG (and, as a separate variant, vice versa). I generated four comparison answers for each original question and averaged their sentiment labels. Every answer was rated positive, balanced, neutral or negative. The purpose was to compare sentiment in open answers with sentiment in prompted comparisons.
The answer: not reliably
Of course the questions that named Samsung and LG produced answers mentioning Samsung and LG. The open questions mentioned other brands as well (or sometimes instead). A fair comparison can only be made when a brand appeared in the answer.
We scored positive as +1, balanced or neutral as 0, and negative as −1, then averaged the four head-to-head answers for each question. Neither brand showed a statistically reliable change.
Average sentiment score on matched questions
Individual labels did move. Comparing the open label with each of the four head-to-head labels on the same matched questions produced 55% agreement for Samsung and 53% for LG. On those same questions, exact repeats of a head-to-head prompt agreed 64% of the time for Samsung and 76% for LG. The differences between methods are therefore hard to separate from ordinary variation in the model's answers.
Exact sentiment label agreement
There is still a clear selection effect. Samsung and LG received no negative ratings when they appeared in the 48 open answers. In the 192 forced comparisons, 57 answers contained at least one negative rating (59 negative brand ratings in total). Of those 59 ratings, 58 concerned a brand that had not appeared in the matching open answer — 51 for LG and seven for Samsung. Only one negative rating applied to a brand the model had selected naturally. Naming a brand therefore creates opportunities for criticism that an open answer may never create.
So which approach should a brand use?
For a visibility study, the open question is the more useful starting point. It shows whether a brand earns a place in the answer without being prompted, and how the assistant talks about it once it is there. That is closer to the discovery questions people actually ask.
Direct comparisons are still useful, but they answer a narrower question: what does the model say once a shopper has already reduced the choice to these two names? That can help with competitor briefs, product pages and sales material.
From this limited experiment, I did not find any evidence to suggest that one method makes brands look better or worse. Rather, it's that the open method measures sentiment after selection, while the directed method removes selection by forcing brands into every answer. Which one is most useful depends on the question the brand is trying to answer.