RoktCrawl crawler policy

Effective 29 September 2026.

This policy states how Rokt's storefront crawler, RoktCrawl, identifies itself, how its requests can be verified, what it accesses, how it limits its requests, what it retains, and how a site operator can opt out.

1. Scope

Rokt is an e-commerce technology company. We work with online retailers and other businesses to present relevant offers to their customers at the moment of purchase. RoktCrawl is operated by Rokt and visits the public storefronts of businesses that work with us, whether as a partner running our placement or as an advertiser whose campaigns we serve. We crawl a site because understanding how its pages are structured — its search results, product pages and category pages — is what lets our systems classify pages and better serve offers, without you having to send us that information yourself. RoktCrawl is not a general web crawler and does not index the web.

2. Identification

Every request carries the token RoktCrawl/0.1 (+https://docs.crawl.rokt.com) appended to a standard browser user-agent string. The token links to this site, and this page, https://docs.crawl.rokt.com/policy.html, is the crawler's policy. The browser and platform named in the rest of the string are the ones actually in use; the crawler does not disguise itself as another client.

3. Verification (Web Bot Auth)

Every request the crawler sends is signed with an HTTP Message Signature (RFC 9421, Ed25519) in the form the Web Bot Auth drafts describe, so a site or the bot-management service in front of it can verify cryptographically that a request came from RoktCrawl rather than from something copying its user-agent string. The Signature-Agent header names the origin that publishes the key; the public key directory at that origin is served without credentials, with the content type application/http-message-signatures-directory+json, and is itself signed with the same key. The keyid on each signature is the JWK thumbprint of the key in the directory.

Signature-Agent
https://crawl.rokt.com
Key directory
https://crawl.rokt.com/.well-known/http-message-signatures-directory

4. Robots exclusion

The crawler requests robots.txt before any other page and applies its rules strictly. It honours rules addressed to RoktCrawl and the general rules addressed to *. A disallowed address is not requested, even when the site's own search form produced it. If robots.txt cannot be read because the site withholds or blocks it, the crawler treats the site as disallowing everything and requests nothing further. Rules are read afresh on every visit.

5. What is accessed

A visit consists of: the site's robots.txt, its sitemap files, its homepage, and a browser session that submits a short list of generic product words to the site's own search box and browses a sample of the resulting search, product and category pages, including their sorting, filtering and paging controls.

The crawler does not sign in, does not create accounts, does not add items to a cart, does not enter checkout, and does not submit any form other than the search box. Addresses for cart, bag, checkout, order, account and login pages are excluded from browsing. It does not attempt to bypass rate limiting, bot detection or any other access control: a page that answers with a block or a challenge is recorded as blocked and is not retried.

6. Request rate

Requests to a single site are spaced at least 2.5 seconds apart, with up to 1.5 seconds of additional random delay, and the crawler loads one page of a site at a time. Within a page load, the browser retrieves the page's own assets as an ordinary browser would. Each visit covers a bounded number of pages, and scheduled crawls run approximately weekly per region. Requests time out after 30 seconds and follow at most ten redirects.

7. What is retained

The crawler retains, for each page it loads: the address, the response headers, a copy of the page body, a screenshot, the addresses of requests the page made, and the crawler's own log of the visit. This material is stored in Rokt's own systems and used to derive and verify the address patterns described above and to describe the storefront's technical characteristics (platform, search capabilities, structured data). Retention period: to be set.

The crawler has no account on any site and sees only what an anonymous visitor sees. It does not seek out personal data, and when it later reads search terms out of addresses, values that appear to be a person's details are withheld rather than recorded.

8. Source addresses

The crawler operates only from the fixed network addresses below, so that a site operator can verify that a request identifying itself as RoktCrawl originated from Rokt. A request carrying the RoktCrawl token from any other address did not come from Rokt. The same list is published as plain text, one address per line in CIDR form, at https://docs.crawl.rokt.com/ips.txt.

RegionAddresses
US West (Oregon)16.145.117.88/32 32.186.91.66/32 184.33.228.4/32
US East (N. Virginia)3.212.239.12/32 34.195.230.14/32 100.56.59.13/32
Europe (Ireland)54.73.34.189/32 3.248.25.71/32 34.242.239.9/32
Asia Pacific (Sydney)13.54.60.54/32 32.236.51.124/32 13.211.111.197/32

9. Opting out

A site operator may exclude the crawler at any time with a robots.txt rule:

User-agent: RoktCrawl
Disallow: /

The exclusion takes effect at the next visit. An operator may also ask Rokt to exclude a site directly using the contact below; such requests are honoured without requiring a robots rule.

10. Contact

Questions, exclusion requests and reports about the crawler's behaviour go to crawler@rokt.com. Include the site's hostname and, where available, the time and source address of the requests concerned.

11. Changes to this policy

If the crawler's identification, behaviour or source addresses change, this page is updated before the change takes effect.