AdvizzBot

Only when the owner of a site asks us to. A visit follows one of two actions in our product: someone enters their website address to create an assistant for it, or an existing customer presses Re-read website in their account. We do not crawl the web looking for sites, we keep no discovery queue, and we do not index anything for search.

If AdvizzBot appears in your logs and you did not ask for it, the most likely explanation is that someone at your company signed up — and the second is that somebody typed your domain by mistake. Either way, write to hello@advizz.io and we will tell you which it was.

Shopping carts, checkouts, logins, admin panels, search results and feeds are skipped by address, whether or not robots.txt mentions them. Each request waits at most 30 seconds, stops at 5 MB, and follows at most 3 redirects.

A run reads a bounded number of pages — 15 when an assistant is first created, and from 10 to 250 on a re-read, depending on the customer’s plan. The number of runs is capped per account per month as well, from 3 to 200 by plan, and two runs for the same site cannot overlap. Pages are fetched one after another, never in parallel.

AdvizzBot obeys robots.txt. To keep it out of part of your site:

User-agent: AdvizzBot
Disallow: /private/

To keep it out entirely:

User-agent: AdvizzBot
Disallow: /

With that in place we stop before the first page and tell our customer what happened, in those words, with your domain named — so if they are entitled to the access they can come and ask you for it, instead of filing a bug with us.

The rules are read the way search engines read them: the group naming AdvizzBot wins over the * group; within a group the longest matching path wins, and Allow wins a tie, so a narrow Allow can open a path inside a broad Disallow; * matches any run of characters and $ anchors the end of the address; an empty Disallow: permits everything. If robots.txt is missing, unreachable or unparseable, we treat the site as open — the same as search engines do, and for the same reason: a typo in a file should not silently cost a site owner a working assistant.

Crawl-delay is not needed here and is not read: the pacing above is already lower than what it would ask for. Blocking us by User-Agent at your firewall works too, of course; robots.txt is simply faster for both sides, because we check it before we knock.

The text of the pages is turned into the knowledge base of the assistant belonging to the account that requested the read, and is used to answer that account’s visitors. It is not shared with other customers, not sold, and not used to train AI models. Details are in our Privacy Policy.

Advizz-catalogue-sync/1.0 (+https://advizz.io/bot) is a separate, much narrower job: it reads one product feed — a Shopify products.json, a WooCommerce Store API endpoint, or a CSV — whose address a customer typed into their own account. It does not follow links and does not read pages: it asks for that one address, on a nightly schedule and when the customer presses sync. If you see it against an address you did not publish for that purpose, write to us and we will stop it.

Questions, complaints, or a request to stop: hello@advizz.io. Quantum Market Hub LLC, 8 The Green, Suite R, Dover, DE 19901, United States.