How Provena identifies itself when it visits your site.
If you found this page from a request in your logs, this is the crawler behind it. This page states what it is for, how it behaves, and how to allow or block it. Provena does not hide its crawler and does not try to get around a site that says no.
Status: the Provena Extract crawler is in development and is not visiting sites yet. This page is published ahead of the first request so that site operators can decide how to treat it before it ever arrives. It will be updated before operation starts.
Preservation on request, not indexing.
- What it does
- Visits a small, defined set of public pages at the request of a Provena customer and preserves what a browser would show: the rendered page, a screenshot, the resources loaded, and technical context, each with its own integrity fingerprint.
- Why
- The customer needs a documented, verifiable record of how those pages looked at a given moment, for a legal, compliance, or investigative matter. Each visit is a one-off job with a stated scope, not a standing crawl.
- What it does not do
- It does not build a search index, does not train models, does not resell or republish content, and does not revisit sites on its own. What it preserves is delivered only to the customer who requested it.
- What it never does
- It does not disguise itself as a person, does not rotate identities or proxies to avoid detection, and does not attempt to solve CAPTCHAs or bypass access challenges. A block is a block; it is recorded as such and respected.
One name, in every request.
Every request carries the crawler’s own token in the User-Agent header, with a link back to this page:
ProvenaExtract/1.0 (+https://provena.legal/bot)
- Real browser rendering. Pages are loaded in a standard Chrome build, so the request pattern looks like a browser loading a page: HTML first, then the stylesheets, scripts, and images that page references.
- Cryptographic identification, planned. Provena intends to sign requests under the Web Bot Auth draft (HTTP Message Signatures), so that a site can verify the request came from Provena rather than from someone copying the token. The public key directory will be linked from this page when signing is in place.
- Origin addresses. Requests come from cloud infrastructure in the region chosen for the customer’s matter. The address ranges in use will be listed on this page before operation starts, for operators who maintain allow lists.
- Every visit is recorded. The time, the address, the pages requested, and the responses received become part of the customer’s custody record, so what the crawler did on your site can be reconstructed and checked later.
Slow, bounded, and polite by design.
- Follows robots.txt. Rules for ProvenaExtract and for * are honoured, including Disallow and Crawl-delay.
- One page at a time per site, with a pause between pages. There is no parallel fetching against the same host.
- Bounded scope. Each job has an explicit list of pages or an explicit limit on depth and page count, set by the customer and recorded. It does not wander across a site.
- Stops at challenges and blocks. A CAPTCHA, an interstitial, a 403, or a 429 ends the automated visit to that page. The response is preserved as what it is: evidence that the site restricted automated access at that time.
- Public pages only, unless the account holder of a restricted area supplies their own credentials and states that they are entitled to use them. That authorisation is recorded with the job.
- No forms, no actions. It reads pages. It does not submit forms, create accounts, post content, or interact with the site beyond loading what a visitor would see.
Your site decides.
To block the crawler entirely
User-agent: ProvenaExtract Disallow: /
To allow it, or part of the site
User-agent: ProvenaExtract Allow: / Disallow: /private/
Sites behind Cloudflare or a similar service can also allow or block the crawler by its User-Agent token in their bot rules. Once request signing is in place, they will be able to verify the signature as well. A general Disallow: / for * is respected like any other rule.
Questions or a request from your logs you did not expect?
Write to us with the URL, the time, and the User-Agent you saw. We answer operators directly, and a report of unexpected behaviour is treated as a defect.