What our crawler reads, and how to block it.
ProbeliAIBot fetches public web pages for Probeli AI, which shows businesses how AI engines talk about their brand.
How to recognize it
Every request ProbeliAIBot sends carries this user agent:
Mozilla/5.0 (compatible; ProbeliAIBot/1.0; +https://probeli.ai/bot)
In robots.txt, its token is probeliaibot (in any capitalization). It downloads pages without running their JavaScript, sends no cookies, never signs in or submits forms, and never requests private or internal network addresses.
What it reads, and when
ProbeliAIBot fetches a page only when a Probeli AI customer's workspace needs it:
- Pages that AI answers cite. When an answer from ChatGPT, Gemini, Google AI Mode and Perplexity cites a page, ProbeliAIBot reads that page to find the brands it names. It doesn't follow the links it finds there.
- A website's AI site audit. When a customer runs the audit for the website their workspace tracks, ProbeliAIBot reads its robots.txt, homepage, sitemaps, llms.txt and llms-full.txt, a few common API description paths (such as
/openapi.json), up to 40 pages, and a sample of the internal links on them (aHEADrequest, or a shortGET). - A homepage, when a customer enters a website. Creating a workspace, changing its website or adding a competitor loads the top of that site's homepage, up to the end of its
<head>, to name the brand.
How often
- Cited pages: one page at a time from each site, and the next page at least 1 second after the previous one (from each of our servers), or longer when your robots.txt sets a
Crawl-delay, which it follows up to 10 seconds. The spacing is between pages: the requests for one page go out right after each other (your robots.txt, when ProbeliAIBot has no recent copy of it, then the page and the redirects it follows, to your site or another). A page it has read, or found blocked, isn't requested again for 7 days, and after that only if an AI answer cites it again. When a page fails for a reason that can pass (a server error, a rate limit or a timeout), ProbeliAIBot tries again later, never sooner than aRetry-Afterheader asks and 3 tries in all, then waits until an AI answer cites the page again. - Audits: only when a customer starts one, and each workspace can start only a few a day. An audit sends at most 4 requests at a time to your site (from each of our servers) and doesn't follow
Crawl-delay. - Homepage reads: one page load, following redirects, each time a customer enters your website.
How it follows robots.txt
ProbeliAIBot reads the robots.txt group for probeliaibot, or the * group when there is none, with Allow, Disallow and the * and $ wildcards (RFC 9309).
- Cited pages: a page your robots.txt disallows is never requested, and neither is a redirect to one. A missing robots.txt allows everything; one that can't be read (a server error, a rate limit or a timeout) stops ProbeliAIBot from reading any of your pages until it can. ProbeliAIBot keeps a site's robots.txt for up to 1 day, so a change can take that long to apply.
- Audits: robots.txt is read again for every audit, and the pages, links and API description paths the audit picks on your site are only ones it allows. It checks the address it picks: when that address redirects, the audit follows the redirect without checking robots.txt again for the new address. If it disallows your homepage, or can't be read, the audit is limited to site-wide files: robots.txt, sitemaps, llms.txt and llms-full.txt.
A few requests aren't subject to robots.txt:
- robots.txt itself;
- the sitemaps, llms.txt and llms-full.txt an audit checks, and API description files that audited pages link to on other websites;
- an audit's first homepage load when robots.txt gives no answer at all (to tell whether the site is up), or when the homepage redirects to another host, whose robots.txt the audit reads right after;
- the redirects an audit follows, to another path on your site or to another website (robots.txt is checked for the address the audit picked, not for the one it redirects to);
- the homepage read when a customer enters a website.
How to block it
To block ProbeliAIBot on your whole website, add this to your robots.txt:
User-agent: probeliaibot
Disallow: /
To keep it out of part of your website, disallow only that path, for example Disallow: /private/. To slow it down, add a Crawl-delay line (in seconds) to its group.
When your robots.txt blocks it, Probeli AI can't find mentions on your pages, and the audit of your website is limited. To refuse the homepage reads too, block requests whose user agent contains ProbeliAIBot on your server or firewall. We don't publish a list of IP addresses.
Questions and requests
Write to [email protected] with questions about ProbeliAIBot, a problem with how it reads your website, or a request to delete a page from our cache. Our Privacy Policy explains what we store from the pages it reads and for how long.
Curious what AI says about your brand?
Probeli AI asks ChatGPT, Gemini, Google AI Mode and Perplexity the questions your buyers ask, every day, and shows who they recommend.
7-day free trial · No credit card
Insights with Webflow as the example brand: seven pages AI engines cite for its prompts that feature competitors but not Webflow, led by a top-priority review page on techradar.com cited by all four engines for five prompts; two are in progress and three completed, one of them automatically after the page started mentioning Webflow.