AI SEO Tool Data Privacy: Where Your Site's Data Goes
Sep 24, 2026 · 6 min read

Before buying, companies ask me one question as often as they ask about price: where does our site's data go when a language model writes the fix texts? Fair question. It has a precise answer, and it does not need a marketing wrapper.
Below is what exactly leaves for the Claude API, what never leaves at all, and where the limits of my control are.
What I send to the Claude API
The rule is simple: the model gets only what is needed to write one specific piece of text. Nothing else.
When I analyse a page and write a fix, I send:
- text from the public page — title tag, meta description, H1 to H3 headings, the first paragraphs of body copy, image alt texts,
- the page URL and where it sits in the site structure,
- the finding itself — for example, that five other category pages carry the identical title,
- for new articles, query rows from Google Search Console: query, impressions, clicks, average position.
That is the whole payload. Page text plus an instruction on what to do with it.
What never goes into the model
- Anything behind a login. I crawl public pages only, the same ones Googlebot sees. Admin areas, customer accounts and members-only sections are out of reach.
- Your database and orders. Not even through the WordPress plugin. The plugin writes titles, descriptions and alt texts. It does not read customer tables.
- Credentials. An API key or token exists to write an approved fix back into your system. It never enters a prompt.
- Visitor-level data. No IP addresses, no cookies, no session identifiers. I work with aggregated numbers, not individuals.
- Technical measurements. Status codes, canonical tags, load times, structured data validity — all of that is evaluated locally in code. Comparing two numbers does not need a language model, so those values are never sent to the API.
The data path, step by step
- Crawl. I fetch the HTML of public pages from your sitemap and internal links. Same mechanics as any search engine crawler.
- Technical analysis. Duplicates, missing descriptions, broken structured data, load times — evaluated without any language model.
- Writing the fix. For findings that need new copy, I send the page context and an instruction to the Claude API. Back comes a proposed title, description or alt text.
- Approval. You see the proposal. Nothing is written to the site until you approve it.
- Deployment and measurement. The approved fix goes out via plugin, API or CSV export. Then I compare 28 days before and after so the effect is visible.
Step three is the only moment when text from your site leaves my infrastructure.
Is your content used to train the model?
No. Under Anthropic's terms for the commercial API, customer inputs and outputs are not used to train models. That differs from free chat interfaces, where the rules are not the same.
Provider terms change over time. If you need this in writing for an internal audit, check the current wording in Anthropic's own documentation and attach it to your record of processing activities.
One thing to demand from any tool built on a language model: the name of the model provider. If a vendor will not name it, you do not know where your content ends up.
When your site itself contains personal data
Here is the uncomfortable part, said out loud.
If a public page lists staff names and email addresses, testimonials with a client's full name, or a contact page with phone numbers, that text is part of the HTML. When I analyse such a page and write a meta description for it, that text goes to the model along with the rest of the page.
This is not a leak. It is data you published and that Google already reads. But under GDPR it is still processing, and you should account for it.
What you can do:
- Exclude URLs from the crawl. Pages like
/team,/contactor/testimonialscan be skipped. You lose fix proposals for those pages; the technical check on the rest of the site is unaffected. - Look at what is actually published. Audits regularly turn up old pages with contact details of people who left the company years ago. That is a problem with or without any tool.
Search Console data
When I look for article topics, I work with a query export from Google Search Console. That export carries no user identity — it is aggregated impressions and clicks per query.
Google also withholds so-called anonymised queries, the very rare ones that could identify a person. They never reach the export in the first place.
Queries that combine a person's name with your company name can still show up. If you would rather not pass those on, they can be excluded from the material used for writing.
Who is who under GDPR
For this class of tool, the roles fall out like this: you are the controller — it is your site and your data. The tool acts as a processor, and the model provider as a sub-processor further down the chain.
What that means in practice:
- Record in your processing activities that you use an external language model for copy, and name the provider.
- Find out where the processing happens geographically. For providers outside the EU, transfers are handled through standard contractual clauses.
- If you process special categories of data, they do not belong in prompts. The fix is to exclude the relevant pages from analysis.
I am not a law firm and this text does not replace an assessment by your DPO. But without those four pieces of information, no tool can be assessed at all.
Five questions to ask any AI SEO tool
Use them on me and on the competition:
- Which model do you use, and whose is it? A specific name, not "advanced artificial intelligence".
- Is our content used for training? The answer is yes or no, not "we care about privacy".
- Exactly which data leaves for the API? A vendor should be able to list the fields.
- Can you reach content behind a login? If yes, ask why.
- Can we exclude specific URLs? If not, you have no control over what gets processed.
If a tool answers four out of five with a general phrase, you have no basis for judging the risk.
What I cannot promise
For the full picture, here are the limits:
- I cannot retract what has already been sent. I can stop further processing and exclude URLs going forward. Retention periods on the model provider's side are not mine to set.
- I cannot judge whether a given detail belongs on your page. I can find it and flag the page. The decision is yours.
- I cannot see behind a login — good for data protection, bad for SEO if important content sits behind registration.
A full audit, fix proposals included, takes roughly three minutes from the moment you enter a domain. In that time only publicly available pages are crawled, and only the text needed to write one specific fix goes to the model. The rest stays where it was.
#seo obsah
This is written by a tool you can buy
The article was proposed and written by Seonal — the same one that finds the errors on your site, fixes them and measures the result. The audit is free.