Snippet Controls for AI Search
You can limit how search engines and AI systems use your content through robots.txt rules for specific AI crawlers, robots meta directives such as nosnippet and max-snippet, and the data-nosnippet attribute. Each control has a different scope, and most involve a trade-off with visibility.
  • Robots.txt controls access by crawler. Meta directives control how fetched content is displayed.
  • For Google's AI features, nosnippet, max-snippet and data-nosnippet are the available controls.
  • Google-Extended opts out of Gemini training and grounding, not AI Overviews.
  • Bing uses nocache and noarchive to limit how content is used in Copilot.
  • Every restriction reduces visibility somewhere, so be clear about what you are trying to achieve.

Not everyone wants their content summarised by AI. Publishers may want to protect subscriptions, some businesses have licensing concerns, and others simply want certain sections left out. The controls exist, but they are spread across different mechanisms and do not all do what their names suggest.

Start with the goal

  • "I do not want my content used to train models." Block training crawlers in robots.txt.
  • "I do not want to appear in AI answers at all." Block AI search crawlers and apply snippet restrictions for Google and Bing.
  • "I want to be cited, but not have everything given away." Use max-snippet and data-nosnippet selectively.
  • "I want maximum AI visibility." Allow everything and check nothing is blocking by accident.

Control 1: robots.txt by user agent

This is the main control for AI companies other than Google and Microsoft search.

 User-agent: GPTBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: OAI-SearchBot
Allow: / 
  • Block training bots such as GPTBot, ClaudeBot and CCBot to stay out of training data.
  • Block search bots such as OAI-SearchBot, Claude-SearchBot and PerplexityBot to stay out of those products' answers.
  • Rules can apply to specific folders, for example allowing marketing pages and disallowing premium content.

Robots.txt is voluntary, applies from now on and does not remove content already collected.

Control 2: Google-Extended

Google-Extended is a token used in robots.txt. Disallowing it tells Google not to use your content for training Gemini models or for grounding answers in the Gemini app and related services. It has no effect on Google Search, and that includes AI Overviews and AI Mode, which are Search features and use Googlebot.

Control 3: robots meta directives

These are the only way to limit use in Google's AI search features while staying indexed.

 <meta name="robots" content="max-snippet:160, max-image-preview:large" /> 
  • nosnippet: no text snippet in results, and the content is not used as a direct input for AI Overviews or AI Mode.
  • max-snippet:[number]: caps the characters that can be shown or used.
  • noindex: removes the page from Search altogether.

They can also be sent as an X-Robots-Tag HTTP header, which is the option for PDFs and other files.

Control 4: data-nosnippet

To exclude part of a page while leaving the rest available:

 <p>Our full methodology is available to subscribers.
<span data-nosnippet>The detailed figures are 42, 17 and 93.</span></p> 

It works on span, div and section elements. This is the most precise tool available and the least used.

Control 5: Bing and Copilot

Microsoft uses existing directives to govern how content appears in Copilot answers:

  • nocache: Copilot may use only the URL, title and snippet.
  • noarchive: content is not used in Copilot answers and not used for training Microsoft's generative models.

Neither affects normal Bing rankings.

Control 6: paywall markup

If content sits behind a paywall, use the isAccessibleForFree property in structured data with a CSS selector identifying the gated section. This tells Google the hidden text is a paywall, not cloaking.

The trade-offs

  • nosnippet strips your description from standard results and rules out featured snippets, which usually cuts clicks.
  • max-snippet set too low can have a similar effect.
  • Blocking AI search crawlers removes you from citations while competitors remain.
  • Blocking training crawlers may reduce how well future models know your brand.
  • None of these controls stops a model summarising what other sites say about you.

Emerging mechanisms

Several efforts aim to give site owners finer-grained preferences, including an IETF working group on AI preferences, CDN-level content signals, licensing standards and pay-per-crawl schemes. Adoption is still developing, so treat them as additions to the established controls above.

Common mistakes

  • Using Google-Extended to try to leave AI Overviews.
  • Adding nosnippet site-wide without modelling the traffic impact.
  • Blocking a crawler in robots.txt and placing a meta directive on the same page, where it will never be read.
  • Assuming llms.txt is an opt-out mechanism. It is not.
  • Forgetting CDN-level bot rules that override everything else.

How to test

  • Check directives in the page source and response headers.
  • Use URL Inspection to confirm Google sees the directive.
  • Search for the page and confirm the snippet behaves as expected.
  • Review logs to confirm blocked crawlers have stopped and allowed ones get 200 responses.

Add your title here

This is a paragraph. Writing in paragraphs lets visitors find what they are looking for quickly and easily. Make sure the title suits the content of this text.

Contact Us Amy Time