AstraVerifyBot
Last updated 2026-09-20
What it is
AstraVerifyBot is the fetcher behind the discoverability scan at astraverify.com. It runs only when a person enters your domain and presses Scan; it does not crawl the web on a schedule and it does not follow links beyond the pages listed below. Its user-agent string is: Mozilla/5.0 (compatible; AstraVerifyBot/1.0; +https://astraverify.com/bot).
What it requests
- The root page over https:// (following up to five redirects), and http:// and the www / apex alternate once each, to check redirects.
- /robots.txt, /llms.txt and the sitemap declared in robots.txt (or /sitemap.xml), each capped at 200 KB.
- Up to four pages listed in the sitemap, capped at 300 KB each, to check their title, description, canonical and structured data.
- The og:image referenced by the root page (first 64 KB only, to read its dimensions).
- One random, non-existent URL to check that missing pages return 404.
- When a user asks for it, the root page once more with a named crawler’s user-agent string (Googlebot, Bingbot, OAI-SearchBot, PerplexityBot, ClaudeBot or GPTBot) to show how your edge treats that crawler. Those requests come from AstraVerify’s servers, not from the company named.
How often
About 14 to 18 requests per fresh scan, run concurrently and finished within a few seconds. Results are cached for 15 minutes and served to everyone who opens the result link; a domain is scanned freshly at most 30 times per hour across all users. Nothing is fetched when someone merely opens a result link.
Where it comes from
Requests originate from Google Cloud (Cloud Run, us-central1). The addresses are dynamic, so there is no fixed IP list to allow; identify the bot by its user-agent string.
How to block it
Add this to your robots.txt and AstraVerifyBot will report the rule and stop fetching pages:
User-agent: AstraVerifyBot
Disallow: /