Inteligência Artificial
How to appear in AI answers: what decides a citation
Branddi · Published on
There is no AI-specific optimization. Google states in its own documentation that there is no extra requirement and no special file needed to appear in AI Overviews and AI Mode: the page must be indexed and eligible to be shown with a snippet. What weighs most in the citation choice is what third-party sources say about your brand.
That last sentence is the uncomfortable part. Most of the effort spent on "showing up in AI" goes into a company's own website — and that is precisely the material generative engines lean on least when assembling an answer about a brand.
What does Google officially say about appearing in AI Overviews?
The AI Features and Your Website documentation, from Google Search Central, is shorter and blunter than the market suggests. Three statements matter:
- Eligibility is the only rule. The page must be indexed and eligible to be shown with a snippet, meeting Search's technical requirements. The documentation is explicit: "There are no additional technical requirements."
- There is no special optimization. "There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary."
- There is no new file to create. "You don't need to create new machine readable files, AI text files, or markup to appear in these features." And on markup: "There's also no special schema.org structured data that you need to add."
This conflicts directly with much of what is sold today as AI optimization services:
- What gets sold — What the documentation says
- Create an llms.txt for the model to read — No need to create machine readable files or "AI text files"
- Add AI-specific schema — There is no special structured data to add
- Optimize separately from SEO — The same SEO practices apply; these are not separate systems
- Guarantee the citation — Indexing and serving are never guaranteed
The documentation also describes the query fan-out mechanism: instead of answering the question as typed, the system issues several related searches across subtopics and data sources, then builds the response from that set. This is why the link list in an AI answer often does not match the first page of classic search — it is not the same slice of query.
Two operational recommendations appear in Google's text and are worth noting, because they are genuinely actionable: keep structured data consistent with the visible text on the page, and use nosnippet, data-nosnippet, max-snippet or noindex when the intent is to limit what gets displayed.
If there is no special optimization, what decides the citation?
Google does not detail its selection criteria. The most solid evidence available comes from two independent studies — both preprints, not yet peer-reviewed.
The first is Generative Engine Optimization: How to Dominate AI Search, by Chen, Wang, Chen and Koudas, dated 10 September 2025. The authors ran controlled experiments across multiple verticals, languages and paraphrases of the same query, on ChatGPT, Perplexity and Gemini, using Google as the comparison. The central finding: generative engines show a systematic and pronounced bias toward earned media — authoritative third-party sources — over brand-owned content and social media. Google distributes more evenly.
The second is News Source Citing Patterns in AI Search Systems, by Kai-Cheng Yang, dated 7 July 2025, which analyzed over 366,000 citations drawn from more than 24,000 conversations and 65,000 responses from models by OpenAI, Perplexity and Google. Two conclusions apply beyond the news coverage it studies: different providers cite different sources even while sharing behavioral patterns; and within the news slice — 9% of the total — low-credibility sources are rarely cited.
Together, the two studies sketch a practical picture:
- You cannot optimize once and cover every engine. Each one has its own source preference, and the studies point to divergence in domain diversity, content freshness and sensitivity to how the question is phrased.
- Your website is not the primary source. It counts, but it weighs less than what third parties have published about you.
- Source credibility acts as a filter. At least within the measured slice, the system tends to avoid poor sources — which only helps if a good source exists and is accurate.
Why does the citation depend more on third parties than on your own site?
It makes sense from the model's perspective. Content published by the brand itself is interested material: it describes the product the way the manufacturer wants it described. A system that has to answer "which is the best option" or "is this company trustworthy" tends to give more weight to whoever has no stake in the answer.
The consequence is uncomfortable to manage. You can have the best page in your industry about your product and still not be cited — or be cited through a third party that describes your brand incorrectly, out of date, or in a deliberately distorted way. The answer arrives with the same authoritative tone regardless of the quality of the source behind it. This shift in first contact was covered in the top of the funnel is out of your control and in the end of the click.
What happens when the sources about your brand are contaminated?
This is where the subject stops being content marketing and becomes brand protection.
A "third-party source" is not only press, reviews and directories. It is also the site copying your brand, the listing from an unauthorized seller, the fake profile presenting itself as an official channel, the complaint page carrying distorted information. If the engine weights third parties above owned content, and that third-party surface is polluted, the model repeats the pollution.
The scale of that exposure in Brazil shows up in Branddi's own survey of 500 Brazilians across every state, collected on 12 January 2026, with 95% confidence and a 3.3 percentage-point margin of error:
- Indicator — Result
- Use AI to research products and services — 66%
- Have already bought influenced by an AI recommendation — 54%
- Had difficulty telling whether a suggestion was impartial or sponsored — 45%
- Cite the spread of scams as a concern — 28%
The full figures are in more than half of Brazilians have bought under AI influence. The number that matters here is the 45%: almost half of people cannot assess where the recommendation they received came from. That is the scenario in which a hostile source pays off more than it would in traditional search, where the user at least sees the domain before clicking. When the purchase becomes mediated by an agent, the problem changes scale — the subject of agentic AI.
How do you audit what AI says about your brand today?
Before trying to influence the answer, you have to measure what it is. The minimum routine has five steps:
- Build a list of 20 to 30 real questions. Do not invent them: use what sales hears, what arrives through the contact form, and your branded queries in Search Console. Include the uncomfortable ones — "is [brand] trustworthy?", "is [brand] a scam?", "best alternative to [brand]".
- Run the list on each engine separately. ChatGPT, Gemini, Perplexity and AI Overviews answer differently. An aggregate result hides the problem that exists in only one of them.
- Record the sources, not just the answer. The answer changes on every run; the list of cited domains is the stable data point and the one you can actually act on.
- Classify each cited domain into four buckets: yours, legitimate third party, outdated third party, hostile or fraudulent third party.
- Repeat on a fixed cadence and compare the lists. What matters is how the composition shifts over time, not a single day's snapshot.
What do you do with what the audit finds?
Each bucket calls for a different action, with a different owner:
- Finding — Action — Owner
- None of your sources are cited — Publish extractable material and pursue mentions in authoritative third-party sources — Marketing and content
- Legitimate source with wrong or outdated data — Contact the outlet and request a documented correction — Communications
- Clone site, fake profile or improper ad — Takedown and continuous monitoring for recurrence — Brand protection and legal
- Unauthorized seller cited as a channel — Marketplace removal and channel governance review — Channel and legal
- Competitor cited on your branded query — Check for trademark misuse in advertising — Brand protection
An honest caveat: the first three rows are communications work and no tool solves them. The last two are the ground where a brand protection operation acts directly — and they are also the ones that contaminate the answer fastest, because fraudulent sources are usually produced at volume.
Frequently asked questions
Do I need an llms.txt file to appear in AI answers?
For Google, no. The AI features documentation states that there is no need to create machine readable files, "AI text files" or specific markup to appear in AI Overviews and AI Mode. No other major engine has documented llms.txt as a requirement so far.
Does schema.org markup increase the chance of being cited?
Google states there is no special structured data to add for AI features. The recommendation recorded in the documentation is different and still holds: keep the markup consistent with the visible text on the page.
Does blocking the AI crawler protect my brand?
It blocks your content, not the subject. Since engines weight third-party sources above owned content, blocking your own site tends to remove the official version from the answer and leave the rest in place. At Google, the documentation also warns that AI features are part of Search — restricting Googlebot affects both.
Does appearing in AI replace traditional SEO?
No. The eligibility Google documents starts from indexing and snippet display, which keeps SEO as a prerequisite. What changed is that it is no longer sufficient on its own.
How often should the audit be repeated?
Monthly covers most cases. During a campaign, a launch or a seasonal peak, weekly is worth it — that is when fraudulent sources are produced in greater volume and the citation mix shifts fastest.
If the audit shows clone sites, fake profiles or unauthorized sellers among the sources AI cites about your brand, the problem is no longer a content problem. Branddi's complete protection monitors and removes that kind of source across search, marketplaces, social networks and advertising — which is the third-party material models read when someone asks about you.