Privacy Policy
Last updated: September 15, 2026
Traceten ("Traceten", "we", "us") provides an analytics service that detects AI-referred traffic on our customers' websites and attributes that traffic to revenue. This Privacy Policy explains what data we collect, why we collect it, how we protect it, and the rights you have over it.
This policy covers two distinct relationships:
- Customers who sign up for a Traceten account. We act as a data controller for the limited account information we collect from them.
- Visitors to websites that have installed the Traceten tracking snippet. We act as a data processor for these visitors on behalf of our customers, who are the controllers.
1. What we collect from website visitors
When the Traceten snippet runs on a customer's website, we collect behavioral signals that help classify whether a visit originated from an AI tool such as ChatGPT, Claude, Perplexity, Gemini, or Copilot.
Data we collect by default
- Page views: URL, page title, referrer, and timestamp. We also read four tracking parameters off the URL:
utm_source,utm_medium,utm_campaign, and the genericrefparameter used by partner and affiliate links.utm_sourceandutm_mediumare how an AI assistant identifies itself when it tags its outbound links; we use them to classify the visit and do not store them as separate fields.utm_campaignandrefare stored as separate fields, so our customer can group their own traffic by them, and never affect how we classify a visit. What those two contain is chosen by our customer, or by whoever links to them; we do not generate, inspect, or validate it, and we show it back only to that customer, only for their own site. - Behavioral signals: scroll depth, click events, time on page, and rough mouse-movement timing patterns. We also count copy and form interactions on the page (counts only; we never read what was copied or typed). These signals are how we tell humans apart from automated traffic. Two per-session figures serve a second purpose and are shown to you in your dashboard: a pageview count, which we derive on our servers from records we already hold rather than receiving it from the browser, and the total time your pages spent in the foreground of the visitor’s browser. On sites where every navigation is a full page load, that foreground figure was already sent to us and we now keep it instead of discarding it. On single-page apps the browser previously discarded it whenever a visitor moved between in-app views; those moves now send it, along with the other signals in this list.
- Outbound link destination: when a visitor clicks a link that takes their current tab from the customer’s site to another site, we record the destination domain name only (for example
github.com). We never record the path or the query string of the destination, and we observe nothing on the destination site. We record nothing when the link opens in a new tab or is opened with a modifier key, because in those cases the visitor has not left the page. This is how we show a site owner where their visitors go when they leave. - Browser and device: user-agent string, screen resolution, language, timezone, and a coarse device memory class (the browser's reported JavaScript heap limit, where available).
- Approximate location: country and city, derived from the IP address at our edge ingestion layer and stored alongside the visit. We use it to show site owners which regions their AI-referred traffic comes from, and as one signal in telling human visits apart from automated ones. We never store the postcode, street address, or GPS coordinates, and we never derive a location more precise than the city from any visitor. Where a map is shown, the point drawn comes from a public list of city centre coordinates, looked up as the page renders and never stored: it is the city's own point, identical for everyone there, or the country's point where the city is unknown or too small to name. See Third-party data and licences. City is kept for the same period as the visit it belongs to: 730 days for the raw event, and between 730 and 760 days for the session-level record. Our lawful basis is legitimate interests; the assessment behind that is summarised at docs.traceten.com/privacy/gdpr.
- IP address (hashed only): the raw IP is hashed with HMAC-SHA-256 under a key derived per site, at our edge ingestion layer (Cloudflare Workers). Only the resulting hash is forwarded to storage. The raw IP never reaches our analytics database, our logs, or disk. This holds for every visitor IP, without exception. The one place we ever retain a raw address is the optional AI crawler tracking feature in section 3, and only where the address matches an IP range an AI vendor publishes for its own crawlers, or a hostname that vendor's DNS forward-confirms.
Data we collect only with explicit consent
Two coarse device-capability signals are collected only after a visitor grants consent through the site's consent management platform (or an equivalent opt-in): whether WebGL is available (true/false), and a single entropy score derived from a small canvas rendering test. We use them to distinguish real browsers from automated ones. The raw canvas image never leaves the browser, and neither signal is used to identify a visitor or to link activity across sites.
Data we do not collect by default
- The analytics snippet does not collect names, email addresses, phone numbers, payment details, or any other personally identifiable information by default. Two exceptions, both under the customer’s control: an email address is handled if the customer calls
traceten.identify(), which hashes it at our edge and discards the plaintext, or if one reaches us on a revenue integration so a payment can be matched to a session, which the customer may do by sending it to our Payment API, and which also happens on its own once a Stripe, Shopify, Lemon Squeezy or Polar connection exists, because those providers put the buyer’s address on the order they send us. Paddle sends none. In every case it is hashed on consumption and only the hash is stored. (We also collect an email if you choose to join our early-access waitlist. See section 2.) - We do not read form input, password fields, or page text content.
- We do not maintain a cross-site identity graph. Cookies are first-party only and scoped to the customer’s own registered domain, never beyond it, and we never use browser or device signals to link a visitor’s activity on one customer’s site to their activity on another.
Cookies set by the snippet
_traceten_sid: first-party session cookie, 30-minute sliding window,SameSite=Lax. Used to group page views into a single session._traceten_vid: first-party visitor cookie, 1-year lifetime,SameSite=Lax. Used to identify returning visitors on the same site._traceten_cart: first-party checkout cookie, 24-hour lifetime,SameSite=Lax(so it survives the redirect to Shopify's checkout domain and back). Set only when a Shopify checkout is detected; holds the checkout's cart token so the resulting purchase can be attributed to the visit._traceten_optout: first-party opt-out preference, 30-day lifetime,SameSite=Strict. Set when you opt out of tracking. While it is present, the snippet stops writing the visitor cookie, stops collecting the higher-entropy device signals, and stops exit, click, in-app navigation and goal tracking, including goals fired from thedata-traceten-goalanddata-traceten-scrollHTML attributes, whose listeners are never installed. The initial page view is still sent. See section 9 for what opting out does and does not stop.
All of these cookies are set with the Secure flag on HTTPS sites, and all are strictly first-party: the browser sends them only to the customer’s own site, never to another site and never to Traceten. Their values do reach us, because the snippet reads them and includes the visitor ID, session ID and cart token in the events it sends to our ingest endpoint. That is the measurement the customer installed Traceten to do, and section 1 above describes it. The snippet also reads a _traceten_consent cookie, which the site’s consent management integration may set to record your consent choice for the consent-gated signals described above. It never writes a consent value; the one time it touches that cookie is to clear it when you deny consent.
By default these cookies are readable only on the exact host that set them. A customer can turn on cross-subdomain scoping per site, after which they become readable across the subdomains of that customer’s own registered domain, so a visitor who moves from www.example.com to app.example.com is recognised as one person instead of two, but this requires the customer’s explicit confirmation in their dashboard; nothing is broadened automatically when a site is created. When it is on, the cookie’s name also carries a value unique to that site, so that two Traceten sites sharing a registered domain cannot be mistaken for one another. The cookies hold pseudonymous identifiers: a random ID, or a keyed hash of an identifier the customer passed to identify(). No name, email address, postal address or password is ever written to a cookie, though on Shopify stores _traceten_cart holds the checkout token, which refers to a checkout that does contain the buyer’s details. The cookies are still never sent to any other registered domain. A customer can turn cross-subdomain scoping back off per site, after which the snippet writes each cookie only to the host that set it; a cookie already stored at the wider scope stays in the browser until it expires, which for the visitor cookie is up to a year. Opting out is broadened the same way on purpose: an opt-out recorded on one subdomain is honoured on the others.
The names change with the scope. A cookie written to one host is _traceten_vid; one broadened to the registered domain is _traceten_vid_ followed by eight characters of the customer’s site key, and the same for the session and cart cookies. The suffix is what stops two Traceten sites under one registered domain from sharing a visitor, since a broadened cookie is visible to both. _traceten_optout is deliberately left un-suffixed, so a refusal covers every site under that domain. Where this policy names a cookie, read it as the base name.
All of these cookies use SameSite=Lax except _traceten_optout, which uses SameSite=Strict. Lax is what lets a visitor’s session and visitor ID survive their first hit after clicking a link inside an AI answer (ChatGPT, Perplexity, and similar), which is a top-level cross-site navigation. The cookies are still never sent on cross-site subresource requests or inside embedded iframes. Opt-out stays Strict because it never needs to ride a cross-site navigation.
2. What we collect from customers
When a customer creates a Traceten account, we collect:
- Name and work email, supplied through our authentication provider (Clerk).
- Company name and billing address, for invoicing.
- Credentials for integrations the customer chooses to connect. For Stripe this is a restricted API key the customer creates in Stripe and pastes into Traceten; for Lemon Squeezy an API key; for Polar an organization access token; for Paddle an API key; for Shopify an OAuth access token. Every one of these is stored with the signing secret for the webhook we set up on the customer’s behalf, except Shopify, which has none. All of these are encrypted at rest with AES-256-GCM under a 256-bit key held as a deployment secret, never in source or in the database. We use a pasted key for four things and nothing else: to confirm which account it belongs to, to read the revenue events we attribute, to create the webhook endpoint on the customer’s own account with that provider, and to delete an endpoint we created. Those reads happen when a key is connected, when the customer uses the “Test connection” button, and during the one-off import of the customer’s own payment history, which is bounded to the last 90 days by two different mechanisms. Stripe and Paddle can filter by date, so we ask them to and nothing older is ever returned to us. Lemon Squeezy and Polar cannot, so we read newest-first and stop paging at the first record older than the window. That means one page (the one straddling the boundary) is retrieved and parsed in full, including the buyer email addresses on the records that turn out to be older than 90 days. Those records are discarded in memory: they are not imported, not written to any store, and not logged. No page beyond the boundary is requested. An endpoint we created is deleted when the customer disconnects, when they replace the key, when the site it is attached to is deleted, and when the account is erased. We never return a credential after it is stored. For a pasted key, the only part that can be read back is its last four characters, shown so the customer can tell which key is connected.
- Identifying details of the connected payment account, stored alongside the credential: the provider’s own account identifier, and where the provider reports them, a display name and a default currency. What is read differs by provider, so we state it exactly rather than in general: from Stripe, all three; from Lemon Squeezy, the store id, its name and its currency; from Polar, the organization id and its name, but no currency; from Paddle, none of them, because it has no endpoint that reports which account a key belongs to (there the identifier is the Seller ID the customer types in, recorded without our being able to verify it, and no display name or currency is stored at all). Where a provider scopes its data by a store, organization or seller (Lemon Squeezy, Polar, Paddle), that identifier is entered by the customer and stored in the clear beside the encrypted credential; it is not a secret. Where we do have a display name we show it in the dashboard so the customer can see which account is connected. For a business that trades under a person’s own name (a sole trader, for example), that display name is personal data, and it is deleted with the rest of the integration when the connection is removed or the account is erased.
- API keys generated by the customer. We store two one-way hashes of the same key, one Argon2id and one SHA-256, plus the first 16 characters of the key, which identify it without revealing it (never the plaintext). The two hashes exist because two systems verify the key: the account API in our backend, and the edge network that receives events. Once issued, an API key cannot be retrieved again from us. We also store the identifier of the team member who created the key, and use it for three things: to check that the person is still an administrator when the key is used to create a site or issue another credential; to record that person as the author of a chart note written with the key; and to revoke every key they created when they are removed from the customer’s organization or their user is deleted. It is the same identifier our authentication provider gives us; it is not their name or email address.
- Chart notes the customer’s team pins to days on a site’s traffic chart: the note text (up to 500 characters), the day it is pinned to, when it was created and last changed, the identifier our authentication provider gives the team member who wrote it and the one who last changed it (never a name or email address), and whether it was written in the dashboard, with an API key (and which key) or through an AI assistant the customer connected. A note is visible to everyone in the customer’s organization who can see the site, and to the customer’s API keys and connected AI assistants; when an assistant reads a note, its text goes to that assistant’s vendor. A note is deleted when a team member deletes it, when its day falls outside the 730-day analytics retention window, and when its site or the account is erased. When a team member’s user is deleted, their identifier on every note is replaced with a placeholder.
- Crawl tokens generated by the customer to enable AI crawler tracking (section 3). Each token belongs to a single site and authenticates that site's server-side crawl reports. We store a SHA-256 hash of the token, plus the first 16 characters of the token, which identify it without revealing it (never the plaintext), and it cannot be retrieved again once issued.
- Standard server logs (request timestamps, status codes, error traces).
Early-access waitlist
If you ask to join our early-access waitlist, we collect the email address you submit, with your consent, so that we can contact you about early access. We also store a salted, one-way hash of your IP address (never the raw IP) purely to rate-limit abuse of the signup form. We use the email only to email you about early access. We do not sell it, share it with advertisers, or add it to unrelated marketing. We keep it until we launch or until you ask us to remove you, whichever comes first. To be removed at any time, email privacy@traceten.com and we will delete your signup.
3. AI crawler tracking (optional, installed on the customer's server)
AI crawler tracking is a separate, optional feature. It is not part of the browser snippet and it is off unless a customer installs it. AI crawlers such as GPTBot, ClaudeBot, and PerplexityBot do not run JavaScript, so the snippet cannot see them: it needs a small package (@traceten/ai-crawl) running on the customer's own web server, plus a secret token issued for that specific site.
What this feature records are requests made by automated crawler software, not visits by people using a browser.
What the customer's server sends us
For each crawler request, the package sends:
- the site identifier and the time of the request;
- the URL that was requested and the HTTP method. The URL is sent to us in full, including any query string. We remove the query string, and any fragment, on receipt at our edge, before the record is stored or placed on any internal queue, so what we keep is the address of the page and nothing that was attached to it. The HTTP method is likewise transmitted to us but is not retained in our analytics database;
- the HTTP status the customer's server returned, where their framework exposes it;
- the User-Agent string the crawler sent, truncated to 1,024 characters;
- the IP address the request came from, as seen by the customer's server. This is the only place that address is observable: the crawler connects to the customer, not to us.
What we do with the crawler's IP address
We use it to answer a single question: was this really the crawler it claimed to be? We compare the address against the crawler IP ranges the AI vendors publish themselves. If that does not match and the claimed vendor supports it, we then run a reverse DNS lookup, which sends the address to Cloudflare's public DNS resolver to see whether it resolves back to that vendor's own hostname. That happens before we know whether the address belongs to a crawler at all, on failed attempts as well as successful ones, unless a recent identical lookup is still cached. That cache holds the address itself, in memory in the instance that made the lookup, for up to one hour. It is never written to disk and never logged. Where an address is reported, we store a one-way hash of it (HMAC-SHA-256 under a key derived per site) rather than the address.
One consequence worth stating plainly: because the lookup is the test, someone holding a valid crawl token can report an address of their choosing with a user agent claiming a vendor that supports reverse DNS, and cause that address to be sent to the resolver. Two things bound it. The crawl endpoint always requires a token, so the person doing it is an identified customer of ours rather than an anonymous stranger, and we can identify and stop them. That endpoint is also rate-limited in its own budget, separate from every other kind of traffic, which caps how fast this could be done at all.
We are deliberately not counting the cache as a third control. It is keyed on the address being looked up, so it only prevents repeat lookups of the same address; someone submitting a different address each time misses the cache every time. It reduces ordinary traffic to the resolver. It does not bound the behaviour described above.
What happens to the address itself depends on the answer:
- If the check proved the request came from the vendor's own published crawler range, or from a hostname the vendor's DNS confirms, we keep the address. It is one the AI vendor itself publishes as belonging to its crawlers, and keeping it lets us debug our own verification and notice when a vendor's published range file falls behind the machines they are actually crawling from.
- In every other case we discard the address and keep only the hash. An unverified address could belong to anyone, including a person on a home connection, so we do not keep it. We do not look up the network operator or the country for these requests either: our records have fields for both, and we write nothing into them.
This boundary is enforced by the database itself, not only by the code that writes to it: a crawl record that is not network-verified is rejected outright if it carries an address. Today we go further than that rule requires. Where the only evidence is a third-party routing database (which shows that an address sits somewhere inside a company's network, rather than that it belongs to that company's crawler fleet), we keep the hash and discard the address, even though the record would be allowed to hold it.
There is a second case where we keep less than the rule allows. Each plan includes a monthly number of crawler requests. Past it, unless the customer has turned on metered overage, we stop doing the reverse DNS lookup, we stop keeping the individual per-request records, and we do not keep the address even where the vendor's own published range matched: the hash alone is kept. Nothing is refused and no totals are lost: the counts behind the customer's charts stay complete. What is reduced is the per-request detail, and the effect on this section is that fewer addresses are retained, not more.
How long we keep it
- Individual crawl records: 90 days, on every plan. In our analytics database these are the only records of a crawl request that hold the requesting address or a hash of it.
- Daily totals: kept indefinitely. These are counts only, in three rollups: one grouped by crawler, company, category, verification result and a coarse response class (2xx, 3xx, 4xx, 5xx or unknown), one grouped by page path, company, category and verification result, and one grouped by page path, crawler, company, category and verification result. None holds an address or a hash, which is what lets us show a year-over-year trend without keeping the detail behind it.
- Closing an account, or deleting an individual site, erases all of the above for the affected sites, including the daily totals. Both take effect 30 days after the request is made. The collection of new analytics data for those sites stops as soon as the request is made, and the request can be cancelled at any point in that window.
4. AI answer-engine visibility (optional, off by default)
Customers can ask us to measure whether AI answer engines mention and cite their brand. The customer writes a short list of questions, and on a schedule we send each question to answer engines and store what came back. Unlike the rest of the service, which receives data about what happened on a customer’s website, this feature sends customer-authored text out to third parties on a schedule, so it is described separately and in full. (Two other flows send data out to a THIRD PARTY, both described in section 3: the reverse-DNS lookup used to verify a crawler’s address, and nothing else. Separately, our prompt-suggestion tool fetches the customer’s OWN pages to build candidate questions; that is a request to the customer’s own site, not a disclosure to anyone.)
The feature is not enabled. It is gated by two separate switches, both of which are off in every Traceten deployment, and neither of which is turned on by a customer. No question has ever been sent to any of the providers below from production, and none of them is processing personal data on our behalf today. They are listed in the Planned section of our sub-processors page, which is what starts the 30-day notice and objection window described there. If you want to object before the feature is enabled, email legal@traceten.com.
What is sent, and to whom
What leaves our infrastructure is the question the customer typed, together with a location signal taken from the language and region the customer set on that question: an approximate country for the three chat APIs, and a location and language code for the search query. Neither is derived from any person. No visitor data is sent, and nothing we measure about a website is sent: no visitor or session identifier, no cookie, no IP address or hash of one, no page URL, and no account or site identifier. The question itself is text the customer wrote, and where they built it with our suggestion tool it contains words taken from their own page titles, headings and navigation labels. A question is limited to 2,000 characters and is otherwise sent as written.
- Perplexity AI, Inc. receives the question as a chat message to its Sonar API, with an output-token limit and an approximate country.
- DataForSEO OU receives the question as a Google search query rather than as a chat prompt, with a location code, a language code and a fixed device and operating-system profile. DataForSEO runs that search against Google from its own infrastructure and returns the results page, so the question also reaches Google LLC as an ordinary web search, handled by Google as it handles any other search. We do not use Google's Gemini API or any grounding API, and the retention terms attached to those products do not apply here.
- Anthropic, PBC receives the question as a user message to the Anthropic API, with an output-token limit, an approximate country, and a web search tool capped at three searches per request. Anthropic’s API has no equivalent of OpenAI’s
storesetting, so there is no server-side retention flag for us to set on this path. Only for customers who buy the engine coverage add-on. - OpenAI OpCo, LLC receives the question as the input to the OpenAI Responses API, with output-token and tool-call limits, a low reasoning-effort setting, an approximate country, and a web search tool. We send
store: falseon every call, which under OpenAI’s API terms opts that call out of OpenAI’s default server-side retention of the request and the response. That is what we ask for and what our own code proves; what happens inside OpenAI’s systems is governed by their terms, not by ours. Only for customers who buy the engine coverage add-on.
We identify ourselves to all four as Traceten-Visibility/1.0 (+https://traceten.com). The questions are written by the customer, so a customer who puts personal data in one transmits it to these providers; our documentation tells them not to, and neither our retention limits nor our deletion tooling reaches inside a provider’s systems.
What we store from the answer
We store the answer verbatim, up to 100,000 characters, and the engine’s list of cited sources, up to 100 of them with titles up to 512 characters and links up to 2,048 characters. Both are unredacted prose written by a third party: we did not author it, we do not filter it, and it can name people who are neither our customers nor visitors to their sites. Alongside it we store the exact model identifier, the tool configuration, the language and region, the timestamp and how long the call took, so that any figure can be traced back to the request that produced it. A link’s query string and fragment are removed before it is stored.
From that answer we derive two narrower records: each place the customer’s brand was named, as a verbatim span of up to 500 characters with a confidence score and how it was matched; and each cited link, reduced to its address, path and domain. Daily counts are then rolled up from those.
None of these records carries a visitor identifier, a session identifier, or an IP address in any form. They describe what a model said when we asked it, not what anybody did on a website, and nothing joins them to a person’s browsing.
How long we keep it
- Answers and their source lists: 90 days. This is the limit that bounds the evidence behind any figure.
- Brand mentions and resolved citations: 365 days. They outlive the answers they came from, so a year of trend data is not limited to the shorter window of the evidence behind it.
- The questions and brand names themselves: for as long as the site exists. A question’s text is stored verbatim, as the customer wrote it, and so are the brand names, aliases and domains they asked us to look for. A question cannot be edited in place, because stored runs reference it and rewording it would retroactively change what they claim to have asked; changing the wording archives the old question and starts a new one, so past wordings persist. There is no other route by which a question is removed. This is the field most likely to contain a person’s name, because a customer may have built it from their own page headings, so it is worth stating plainly.
- Daily counts: kept indefinitely. Two sets. One holds counts per question, engine and brand: no prose, no link, no identifier of any person. The other holds counts per cited link, and keeps that link (its address, path and domain) for as long as the account exists. Those links are third-party pages an engine chose to cite, not pages anyone visited. Nothing else in the feature is kept indefinitely, and deleting the site is the only thing that clears either.
- Deleting a site, or closing an account, erases all of the above for the affected sites, including the daily counts, which have no expiry of their own and are cleared by nothing else. Both take effect 30 days after the request. Because these records carry no identifier for a person, a per-visitor deletion request cannot reach them; there is no key to look one up by. A request from someone named in a stored answer is handled by hand at privacy@traceten.com.
5. How we use the data
- To classify each visit as AI-referred or not, score its confidence, and attribute it to a source (e.g., ChatGPT, Perplexity).
- To match visits against revenue events from connected billing systems (Stripe, Shopify, Lemon Squeezy, Polar, Paddle, or any processor the customer posts to our Payment API), so customers can see what their AI traffic is worth.
- To produce aggregated analytics in the customer's Traceten dashboard.
- Where a customer has AI answer-engine visibility enabled, to send that customer's own questions to the answer engines named in section 4 and report what those engines said. No visitor data is used for this and none is sent.
- To operate, secure, and improve the service.
We do not sell visitor data. We do not share visitor data with advertisers. We do not use visitor data to train machine learning models that benefit anyone other than the customer whose site the data was collected from, except for our internal AI-classification model, which is trained on aggregated, non-identifying behavioral signals.
6. Where data is stored
- Raw behavioral event records (730 days): ClickHouse Cloud. Used for dashboard queries. Raw event records are permanently deleted after 730 days by an automatic database retention policy (TTL).
- Conversion and revenue-attribution records (1 year): ClickHouse Cloud, permanently deleted after 1 year by the same automatic retention mechanism. These records reference the pseudonymous visitor identifier, never a name or email.
- Refund records (life of the site): Postgres (Neon), encrypted at rest. One record per refund or cancellation a customer’s payment processor reported. What it holds changes once the refund has been applied. Until then, which is usually seconds but lasts longer when a refund reaches us before the order it refunds, the record holds the payment processor’s own identifiers for the payment and the refund, the amount refunded in the merchant’s own currency, that currency’s code, the date of the refund, and, on Shopify stores, the titles, quantities and per-item prices of the returned items. Those figures are there because we rebuild the refund from this record when the original order arrives. Once the refund has been applied, or settled as one whose original order never arrived, we clear the amount, the currency and the returned-item list in the same write that marks it done, leaving the payment processor’s identifiers for the payment and the refund, the date of the refund, and our own bookkeeping about how the record was processed. That remainder has no expiry date: we keep it for the life of the site because it is what stops a payment processor’s redelivered refund webhook subtracting the same money a second time. No visitor identifier, no session identifier and no customer name, email or address is held at either stage. The whole record is deleted when the site is deleted, through the process described in section 9.
- Session summaries (730 to 760 days): ClickHouse Cloud. One compact record per session, keyed by the pseudonymous visitor and session identifiers. It holds the session’s entry page, its exit page (the last page we measured in the visit), attributed source, the country and city of the first event, the device type, operating system and browser of the first event, the number of pageviews in the session, the total time your pages spent in the foreground of the visitor’s browser, and, on stores that use it, the shopping-cart token we use to match a session to an order. Deletion runs on monthly storage partitions, so a given session is removed at least 730 and at most about 760 days after it started.
- Lifetime-revenue totals (life of the account): per-visitor revenue totals derived from the records above, keyed by the pseudonymous visitor identifier. These are retained while the customer’s account is active so that lifetime-value reporting keeps working, and they are removed through the deletion process described in section 9.
- Aggregated daily rollups (life of the account): ClickHouse Cloud. Daily per-site counts (pageviews, top pages, revenue by AI source, and the country, device type and referring site of visits) with no visitor or session identifier column. Distinct-visit counts are stored as aggregate sketches derived from hashed session identifiers. These power year-over-year trend reporting and are removed through the deletion process described in section 9.
- Daily visitor counts (730 days): ClickHouse Cloud. A daily per-site count of distinct visitors, split new against returning and by AI source. Unlike the rollups above, its sketch is derived from the hashed visitor identifier rather than the session identifier, so we do not keep it for the life of the account: it expires with the raw event records it is derived from. This and top pages by country are the two aggregated rollups we time-limit.
- Top pages by country (730 days): ClickHouse Cloud. A daily per-site count of which page URLs were read from which country, with no visitor or session identifier column. This is the second of the two aggregated rollups we do not keep for the life of the account: a page paired with a country is more revealing than either on its own, so we bound how long we hold it.
- Individual AI crawler records (90 days): ClickHouse Cloud. One record per crawler request reported by a customer's server, including the requested URL and page path, the User-Agent, the crawler and its category, the response status, the verification result and its confidence, and a hash of the requesting address, plus the address itself only where it was matched to an AI vendor's published crawler range or forward-confirmed hostname, as described in section 3.
- AI crawler daily rollups (life of the account): ClickHouse Cloud. Daily per-site counts of crawler requests in three rollups (one grouped by crawler, company, category, verification result and a coarse response class; one grouped by page path, company, category and verification result; one grouped by page path, crawler, company, category and verification result), with no address and no hash. These power year-over-year crawl trend reporting and are removed through the deletion process described in section 9.
- AI answer-engine answers (90 days): ClickHouse Cloud. One record per engine call, holding the verbatim answer and the engine’s list of cited sources, plus the question’s identifier, the engine, the exact model identifier, the tool configuration, the language and region, and the timestamp. Both the answer and the source titles are unredacted third-party prose, as described in section 4.
- AI answer-engine mentions and citations (365 days): ClickHouse Cloud. Each place a tracked brand was named, as a verbatim span of up to 500 characters with its confidence score and match method, and each cited link reduced to its address, path and domain. No visitor or session identifier.
- AI answer-engine daily counts (life of the account): ClickHouse Cloud. Two rollups. Daily per-site counts per question, engine and brand, which hold counts only and no prose; and daily per-site counts per cited link, which keep that link’s address, path and domain alongside the counts. Neither holds any answer text. These power year-over-year visibility trends and are removed through the deletion process described in section 9.
- Account metadata: Postgres (Neon), encrypted at rest. This includes the questions a customer has configured for AI answer-engine visibility, and the brand names and domains they have asked us to look for, which are stored as the customer wrote them.
- Sessions and rate-limit counters: Upstash Redis, with TTLs that match the cookie lifetimes above.
Raw event records are deleted after 730 days, and conversion and revenue-attribution records after 1 year, as described above. Individual AI crawler records and AI answer-engine answers are each deleted on their own shorter schedules, given in the entries above. Only the per-visitor summaries, the aggregated daily rollups, the AI crawler daily rollups, the AI answer-engine daily counts, and what remains of a refund record after it has been applied, all described above, are kept beyond those windows. We plan to add an encrypted cold-storage archive (Amazon S3, Parquet format, SSE-KMS) for historical reporting; if we do, we will update this policy before it launches.
You select a data residency region when you create a site, and we record it against the site. Region-based storage routing is not yet in place: all behavioral data currently lands in one region regardless of the region selected, listed on our sub-processors page. Transfers are covered by the standard contractual clauses in our DPA. We will update this policy and notify affected accounts before region routing goes live.
7. Sub-processors
We use a small number of vetted sub-processors to operate the service. The current list, with each sub-processor's location and role, is published at /legal/sub-processors. We notify customers at least 30 days before adding a new sub-processor that processes personal data.
Four providers are listed there as planned for the AI answer-engine visibility feature described in section 4: Perplexity, DataForSEO, Anthropic and OpenAI. They are not processing personal data on our behalf: the feature is off, and each would begin only after the 30-day notice and objection window on that page has run. What each of them would receive is set out per provider on that page and in section 4.
8. Security
- All traffic between the snippet, our edge, and our backends is encrypted with TLS 1.2 or higher.
- Data at rest is encrypted (ClickHouse Cloud, Postgres).
- API keys are stored as two one-way hashes of the same key, Argon2id and SHA-256, and server-side crawl tokens as a SHA-256 hash. All are 256-bit random values, and none can be recovered from what we store, by us or by anyone else.
- Three kinds of secret have to be usable again, so they are stored as AES-256-GCM ciphertext under a 256-bit key held as a deployment secret (never in source or in the database) rather than hashed: credentials for connected integrations (a pasted Stripe restricted key, a Lemon Squeezy or Paddle API key, a Polar organization access token, a Shopify OAuth access token); the signing secret for the webhook endpoint we create at connect; and the signing secret for each outbound webhook a customer configures us to send. Who generates that webhook secret differs by provider: Stripe and Paddle generate it and return it when we create the endpoint, and we read it at that moment and never ask for it again; for Lemon Squeezy and Polar we generate it ourselves and send it to them. In every case the copy we hold is the one we encrypt at creation. Shopify has no per-connection signing secret. These secrets are not wrapped by a key-management service and the key is not customer-managed.
- Access to production systems requires SSO and hardware-key MFA.
- Logs are scrubbed of known secret patterns before they leave the originating service.
9. Visitor rights (GDPR, CCPA, and similar laws)
If you are a visitor to a website that uses Traceten and you want to exercise your rights of access, correction, or deletion, your first point of contact is the operator of the site you visited; they are the controller of your data. They can use the Traceten deletion process to remove your data from our systems.
You can also contact us directly at privacy@traceten.com. We will route the request to the relevant customer and confirm completion within 30 days.
Because we do not collect names, emails, or other direct identifiers from visitors by default, identifying your records may require you to provide the website you visited and, where available, your _traceten_vid cookie value. We cannot identify visitors without one of these.
Opting out. Setting the _traceten_optout cookie, or sending a DNT: 1 header, puts the snippet into opt-out mode. This stops the persistent _traceten_vid cookie (a transient in-memory ID is used instead, so visits cannot be joined across sessions), the higher-entropy device signals (WebGL support, canvas entropy, approximate JavaScript heap size, timezone offset, paste detection), the Shopify cart token, and exit, click, in-app navigation and goal tracking (goals fired from the data-traceten-goal and data-traceten-scroll HTML attributes are included: neither listener is installed in opt-out mode). It does not stop the initial page view: that is still sent, and it still carries the page URL, referrer, user agent, screen size, language, and timezone, and we still derive an approximate country and city from the request IP as we would for any page view. The _traceten_sid session cookie is also still written. If you want collection to stop outright, the site operator must gate the snippet behind their consent tool, which we enforce on our side. To remove data already collected, use the deletion route described above.
The _traceten_vid visitor identifier is stored alongside pageview events and, when a site operator records a conversion event via the Traceten server SDK, alongside that conversion record. This linkage allows customers to attribute revenue to the AI traffic that drove it, and it is required to fulfil Art. 17 deletion requests: a deletion request keyed on _traceten_vid removes the corresponding pageview events, session data, and conversion records from our systems.
10. International data transfers
Where data is transferred from the European Economic Area, the United Kingdom, or Switzerland to a country that has not received an adequacy decision, we rely on the EU Standard Contractual Clauses (Module 2: controller to processor), incorporated into our Data Processing Agreement at /legal/dpa.
11. Children
Traceten is a B2B analytics product and is not directed to children under 16. We do not knowingly collect data from children. If you believe a child's data has reached our systems, contact privacy@traceten.com and we will delete it.
12. Changes to this policy
We will update this policy as the product evolves. Material changes are announced to customers by email at least 30 days before they take effect. The "Last updated" date at the top of this page reflects the most recent change.
13. Contact
Privacy questions and data-subject requests: privacy@traceten.com
Contractual and legal questions: legal@traceten.com
Traceten, [address]