How to Control How Much of Your Page an AI Answer Can Reproduce
Every page without a directive is published at its widest setting, and no engine reports that. Google's nosnippet now also prevents content being used as a direct input for AI Overviews, while Bing's NOARCHIVE removes the link from chat answers, not only the text.
Step One: Assign Every Revenue Page a Ceiling Before Touching Markup
Assignment comes before markup because the default is not neutral. Google states that snippets are automatically created from page content, so a page carrying no directive is published at its widest reproduction setting, chosen by nobody, much as a browser computes a page semantic model whether anyone reviews it or not. This Digital Strategy Force procedure assigns every revenue page one of five ceilings, applies the directive each engine honours, then verifies the result actually took effect.
Digital Strategy Force models those settings as the DSF Reproduction Aperture Model: five levels running from Open, where an engine may reproduce any portion at any length, to Sealed, where neither the text nor the link survives into an answer. Each level is a commercial decision before it is a technical one.
The rule governing the model is the Silent Default Principle: a page with no directive is not neutral, it sits at the widest aperture, then no engine will ever report that it does. There is no warning, no console message, no traffic event. Both halves of the usual failure are invisible for the same reason, which is that nobody chose the setting, so nobody audits it.
A page that carries no directive is not neutral. It is published at its widest setting, then no engine will ever report that it was.— DSF Answer Engineering Division
Start by listing the pages that carry revenue, then write one line per page saying what an engine should be allowed to reproduce and why. A definitional explainer whose job is to be quoted everywhere belongs at Open. A proprietary benchmark that took a quarter to produce does not. Do this in plain language first, because a ceiling assigned from commercial reasoning survives a template migration, while one assigned from a tag list does not.
| Level | Page archetype | What an engine may reproduce | Directive |
|---|---|---|---|
| Open | Definitional explainers, glossary entries | Any portion, any length | None applied |
| Metered | Premium research, thought leadership | A length-capped preview of the finding | max-snippet |
| Masked | One proprietary finding inside open prose | Everything except the carved-out region | data-nosnippet |
| Closed | Pages that must be found, never quoted | The listing, without reproduced text | nosnippet, or NOCACHE |
| Sealed | Rarely correct, usually applied by accident | Nothing, in a chat answer, not even the link | NOARCHIVE, at Bing |
One caveat belongs here rather than at the end. No single directive produces identical exposure across engines, so a level is an intent, then each engine gets its own implementation. The rest of this procedure is the translation.
Step Two: Apply the Google Set, then Accept the Coupling It Forces
Google offers three reproduction controls, then couples them to AI answers in a way that is easy to miss. The robots meta tag reference documents nosnippet as suppressing a text snippet or video preview, adding that this applies across web search, Images, Discover, plus AI Overviews and AI Mode. The same entry states it will also prevent the content from being used as a direct input for AI Overviews and AI Mode.
That sentence is the reason this is not an old SEO topic. The snippet dial is now documented as an AI answer control, in Google's own words, on its own reference page. The same page says max-snippet will limit how much of the content may be used as a direct input for those features, which makes length a genuine lever rather than a cosmetic one.
| Directive | Effect on the search snippet | Effect on AI Overviews and AI Mode | Granularity |
|---|---|---|---|
| nosnippet | No text snippet or video preview; a static thumbnail may remain | Prevents use as a direct input, per the reference page | Whole page |
| max-snippet:[n] | Caps the snippet at n characters | Limits how much may be used as a direct input | Whole page |
| data-nosnippet | Excludes only the marked region | Removes that region from what can be drawn on | div, span, section |
| noarchive | Nothing, the cached link feature no longer exists | Nothing | Retired |
Now the coupling. Google's documentation on AI features states that to be eligible to appear as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet. That is a precondition, not a threat, so read the consequence carefully: a page carrying nosnippet fails the stated eligibility test for a supporting link. Google does not say the link is removed. It says what eligibility requires, then nosnippet withholds it.
This is why the Closed level is expensive at Google in a way it is not elsewhere. Suppressing the quotation and suppressing the attribution are the same action. A team that reaches for nosnippet to stop an answer reproducing a page is also, by the same stroke, arguing itself out of being named beside that answer.
A vendor who types zero intending a small snippet has not set a tight cap. They have set the Closed level on a page that was meant to be Metered, which is exactly the class of error this audit exists to find.
Step Three: Apply the Bing Set, Where NOARCHIVE Removes the Link
Bing inverts the intuition that a cache directive is the gentler option. Microsoft states that content tagged NOARCHIVE will not be included in Bing Chat answers, then adds that it will not be linked to in the answers either. Content carrying NOCACHE may still be included, with only the URL, snippet and title displayed.
Read those two sentences beside each other. The tag that sounds like a cache-management detail is the most severe setting in the entire set, because it removes the page from the answer surface completely, including the attribution. The tag that sounds equally technical is the moderate one.
The scope guard matters as much as the reversal, so state it plainly. Microsoft assures publishers that content carrying either tag will still appear in Bing search results. Neither directive delists a page. Both govern what happens inside the chat answer, which is precisely where a demand-capture page most needs to be present.
| Tag | In search results | Text in an answer | Link in an answer | Generative captions |
|---|---|---|---|---|
| Neither tag | Present | Reproduced | Present | Generated |
| NOCACHE | Present | URL, snippet, title only | Present | Suppressed |
| NOARCHIVE | Present | None | Removed | Suppressed |
The fourth column carries a consequence most teams never connect to these tags. Microsoft states that a site preferring not to have generative captions created can achieve that simply by using the NOCACHE or NOARCHIVE tags, which means either tag also changes how ordinary Bing result descriptions are written. A directive applied for one reason quietly governs three surfaces.
Bing also honours the length dial. Its snippet controls announcement defines max-snippet as the maximum text length in characters of a snippet in search results, with zero meaning no text snippet shown, then negative one meaning no length limit. The Metered level therefore translates cleanly across both engines, which is not true of any other level in the model.
Masked translates too, as of recently. Bing shipped support for data-nosnippet in October 2025, describing it as a way to mark specific sections of a page's HTML so they do not appear in Bing search snippets or AI-generated answers, while confirming that content marked this way is still indexed normally. That last clause is what makes Masked the most underused level in the set.
Step Four: Resolve Conflicts Per Engine, Because the Rule Runs Backwards
A page carrying more than one reproduction directive is common, because directives accumulate across template changes, vendor engagements, then compliance requests. What is not common is knowing how each engine resolves the collision, and the two engines resolve it in opposite directions.
Google resolves toward severity. Its special tags documentation states that in the case of conflicting robots meta tags the more restrictive tag applies, giving the worked example that a page carrying both max-snippet:50 and nosnippet will have the nosnippet tag apply. Bing resolves the other way, stating that if content has both NOCACHE and NOARCHIVE tags, it will be treated as NOCACHE.
The practical consequence is that a single page can miss its assigned level in both directions at once, landing tighter than intended at Google then looser than intended at Bing. That is not a hypothetical. It is the ordinary outcome of leaving two directives on a page for a few years.
So record the expected outcome per engine, not per page. A ledger with one resolution column is wrong by construction. Two columns, one per engine, is the minimum honest form, then the row is only complete when both are filled.
Want a Reproduction Ceiling Set Deliberately Rather Than Inherited? Digital Strategy Force audits the directive layer page by page, so the setting on every revenue page is one somebody actually chose.
Step Five: Audit Directives That Were Equivalent the Day They Were Applied
The suppressed page is usually not a mistake. It is a correct decision whose meaning changed underneath it, which is why this step looks backwards before it looks at markup.
Bing originally documented the two tags as one thing. Its 2008 robots exclusion protocol post lists a single combined entry, NOARCHIVE and NOCACHE together, defined only as telling a search engine not to show a cached link for a page. On that definition the two were interchangeable, then choosing either was a matter of taste.
Fourteen years later they were still described that way. Bing's guidance for subscription and paywall content instructs publishers to use a noarchive robots meta tag or the equivalent nocache tag on subscription content that should not be cached. That page is still live, still recommends either tag for paywalled material, then predates the divergence by sixteen months.
A publisher who followed that paywall guidance in 2022, on the reasonable assumption that the two tags were the same, applied NOARCHIVE to gated commercial pages. In September 2023 those pages left Bing chat answers entirely, link included. Nothing on the site changed. No error appeared anywhere.
Google's side of this step is shorter but the same shape. Its reference now records that the noarchive rule is no longer used to control whether a cached link is shown, because the cached link feature no longer exists. A directive that meant something specific for two decades now means nothing at one engine, then means everything at the other.
So the audit step is mechanical. Pull the raw source and the response headers for the twenty highest-value demand-capture pages, search both for the four directive strings, then date every hit against the timeline above. Any NOARCHIVE applied before September 2023 was almost certainly chosen under the old definition, which makes it a candidate for correction rather than evidence of intent.
Step Six: Deliver Each Directive at the Layer That Matches the Asset
A directive the crawler never receives is indistinguishable from no directive at all, so the delivery layer is part of the decision rather than an implementation detail left to whoever picks up the ticket.
Microsoft has documented since 2008 that these tags can be present as meta tags in the page HTML or as X-Robots-Tag entries in the HTTP header, noting that the header form allows non-HTML resources to carry the same rules. That clause resolves the single most common gap in this whole area, which is the priced research PDF.
A PDF, a spreadsheet, or a slide deck has no HTML head, so it cannot carry a meta tag at all. Companies at this revenue tier routinely publish their most valuable material in exactly those formats, then assign it a ceiling in a spreadsheet, then never deliver that ceiling to any engine. The asset sits at Open while the register says Metered.
| Asset | Meta tag | HTTP header | Failure mode if the wrong layer is used |
|---|---|---|---|
| Standard HTML page | Yes | Yes | Either works, so the risk is duplication across both layers |
| PDF or Office document | No | Yes, required | The asset stays at Open no matter what the register says |
| Region within a page | Attribute only | No | A header applies to the whole document, overshooting the intent |
| Page blocked in robots.txt | Never read | Never read | The directive cannot be seen, so it has no effect at all |
The last row is the trap that defeats the entire procedure silently. Google states that for a page-level rule to be effective the page must not be blocked by a robots.txt file, then must otherwise be accessible to the crawler. Blocking a page in robots.txt and applying a directive to it are mutually exclusive, because the crawler that would read the directive is the one being turned away.
Step Seven: Verify on Both Engines Before Recording a Page as Done
Applying a directive and confirming its effect are separate steps, then the gap between them is measured in crawls rather than minutes. Skipping verification is how a register fills with rows that describe intentions instead of outcomes.
Start with what is actually being served. In Chrome, DevTools shows HTTP header data by selecting the request in the Network panel then opening the Headers tab. Filter to the document request so the page's own headers are isolated from every asset it loads, then read the response headers for an X-Robots-Tag entry.
Then check what the crawler received rather than what the browser rendered. Google's guidance is to use the URL Inspection tool to see the HTML that Googlebot received, which matters because a directive injected by a script may not be present in the fetched HTML at all. For teams verifying at scale, the URL Inspection API exposes fields including whether the page is blocked by robots.txt, whether indexing is blocked by a rule, then whether the fetch succeeded.
| Check | Tool | Pass condition | Confounder |
|---|---|---|---|
| Received header | DevTools Network, Headers tab | X-Robots-Tag present on the document request | Reading an asset request rather than the document |
| Crawler-visible markup | Google URL Inspection | Directive appears in the fetched HTML | Browser view source shows a script-injected tag the crawler lacks |
| Bing-side state | Bing URL Inspection | Crawl and index state reported as expected | Assuming a Google result describes the Bing outcome |
| Effect on the answer | Recrawl, then re-observe | Reproduced text matches the assigned level | Checking before recrawl, which shows the old state |
The last confounder is the one that wastes the most time. Google states plainly that it does not crawl a page immediately after a fix is published, so Search Console can continue showing an error for a page that has already been corrected until that page is crawled again. A verification run the same afternoon will report failure for work that succeeded.
What the Directive Layer Cannot Do
Four limits bound this entire procedure, then stating them is what keeps the audit honest rather than aspirational.
There is no reproduction-volume control for ChatGPT. OpenAI's crawler documentation offers per-bot allow or disallow in robots.txt with no length parameter anywhere, so the choice is binary rather than graduated, and it carries a nuance worth knowing: sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links.
There is no AI-Overviews-only opt-out at Google. The company has said it is exploring updates to its controls to let sites opt out of generative features specifically, which is exploratory language about a control that does not exist today. Anyone promising that setting is describing a future, not a configuration.
Access controls are also a separate axis from reproduction controls. Google documents that Google-Extended does not impact a site's inclusion in Google Search nor act as a ranking signal, which makes it a training lever rather than a snippet lever. Cloudflare's Content Signals Policy similarly expresses permission by purpose, distinguishing search from AI input, without offering any volume dial.
| What gets asked for | Does it exist | Nearest available control |
|---|---|---|
| Stay in Google Search, leave AI Overviews | No | max-snippet caps how much may be used as input, at the cost of a shorter snippet everywhere |
| Limit how much ChatGPT quotes | No | Per-bot allow or disallow, which is presence or absence rather than volume |
| Protect one section, publish the rest | Yes | data-nosnippet, honoured by both Google and Bing, with the page still indexed |
| Block training, keep answer visibility | Yes | Google-Extended, which does not affect Search inclusion or ranking |
The fourth limit is the largest. None of this governs what a model already absorbed during training, and none of it reaches an answer generated from parametric memory rather than a live fetch, which is a separate problem from never being selected for retrieval at all. The directive layer controls reproduction at retrieval time, which is a real and useful boundary, then it is not a recall of anything already learned.
Where a Reproduction Ceiling Pays for Itself
Run this procedure once and two findings tend to arrive together, pointing in opposite directions. Premium material sits at Open, reproduced in full inside answers that satisfy the reader completely. Demand-capture pages sit at Sealed, removed from chat answers by a tag applied years ago under a definition that no longer holds.
Neither shows up in analytics, because neither produces an event. There is no error, no warning, no traffic drop attributable to a cause. The page that is over-reproduced still gets impressions. The page that is suppressed simply never appears, and an absence generates no record anywhere a marketing team is looking.
That is what makes the Silent Default Principle worth the audit rather than a slogan. Every page on the internet carries a reproduction setting right now. On most sites nobody selected any of them, nobody has reviewed them since, then the two engines that matter read them in ways that diverge on the same markup.
The work itself is small. Assignment is a morning of commercial judgement, application is a template change, then verification is a console and a wait for recrawl. What makes it worth doing is not the difficulty but the invisibility: this is one of the few remaining places where a deliberate decision costs almost nothing, and where the absence of a decision has been quietly charging a company for years.
FAQ — Reproduction Limits
Does NOARCHIVE remove a page from Bing search results?
No. Microsoft assures publishers that content carrying either tag will still appear in Bing search results. The documented effect of NOARCHIVE is scoped to chat answers, where tagged content is neither included nor linked to. The page stays findable, then loses its place inside the answer.
What happens when a page carries both NOARCHIVE and NOCACHE?
Microsoft states that content carrying both tags is treated as NOCACHE, resolving toward the less restrictive setting. Google resolves the opposite way for its own directives, applying the more restrictive tag, so an identical page can land tighter than intended at one engine then looser at the other.
Can a page be kept out of AI Overviews while keeping a normal Google snippet?
No such control is documented today. Google states that eligibility as a supporting link requires a page to be indexed and eligible to be shown with a snippet, which ties the two together. Google has said it is exploring updates to its controls, which describes an intention rather than a setting available now.
How is a reproduction directive applied to a PDF?
Through the X-Robots-Tag HTTP response header, because a PDF has no HTML head to carry a meta tag. Microsoft has documented since 2008 that the header form allows non-HTML resources to carry these rules. A research PDF assigned a ceiling in a spreadsheet but delivered no header sits at the widest setting regardless.
Do these directives limit how much ChatGPT reproduces from a page?
No. OpenAI documents per-bot allow or disallow in robots.txt with no snippet-length or partial-extraction parameter, so the available choice is binary rather than graduated. Opting out of OAI-SearchBot means a site will not be shown in ChatGPT search answers, though it can still appear as a navigational link.
Why would a directive applied years ago suppress a page only now?
Because the meaning changed after it was applied. Bing defined NOARCHIVE and NOCACHE as one combined entry in 2008, meaning only that no cached link would be shown, then still called them equivalent in 2022 paywall guidance. The September 2023 split gave them opposite effects on chat answers without anything on the site changing.
Next Steps — Reproduction Limits
- ▶ Assign every revenue page one level in the model with a one-line commercial reason, before any markup is touched.
- ▶ Pull the raw source and response headers for the twenty highest-value demand-capture pages, then search both for all four directives.
- ▶ Date every NOARCHIVE you find against September 2023, since anything older was chosen under a definition that no longer applies.
- ▶ Move directives for PDFs and other non-HTML assets to the response header, then confirm no page carrying a directive is disallowed in robots.txt.
- ▶ Re-verify each corrected page on both engines after recrawl rather than the same afternoon, since the console reports the old state until then.
Digital Strategy Force runs this audit across the pages that carry revenue, then sets each ceiling deliberately on both engines. Talk to the answer engine optimization team.
Open this article inside an AI assistant — pre-loaded with DSF's framework as the lens.