If you found us in your server log

This page explains what we do, what we keep, and how to have your site left out.

What we are

webXtrend is a domain and DNS information service. We publish facts about domains: their DNS records, where they redirect to, which other domains resolve to the same IP address, which social media profiles they link to, and a few properties of their start page.

Our customers use this for market research, competitive analysis, IT security and brand protection.

How to recognise it

Our crawler is called webXtrendBot. It sends this user agent:

Mozilla/5.0 (compatible; webXtrendBot/1.1; +https://webxtrend.com/crawler)

Its requests come from addresses whose reverse DNS name ends in .webxtrend.com, and that name resolves back to the same address. A request that calls itself webXtrendBot but fails this check is not from us.

What the crawler requests

For each domain it requests /robots.txt and then the start page over HTTP or HTTPS, following the redirects the start page sends. It does not log in anywhere, it does not submit forms, it does not follow links into your site, and it does not try to reach pages that are not linked from the public internet.

Each domain is revisited periodically so that the published data stays current.

How to exclude your site

webXtrendBot reads your robots.txt and follows every rule that names webXtrendBot. To keep it away from your whole site, add this:

User-agent: webXtrendBot
Disallow: /

Rules written for all robots (User-agent: *) do not apply to webXtrendBot. They are usually meant for search engines, and following them would leave sites out of statistics they are part of. To exclude us, name us as shown above. The change takes effect on our next visit.

What happens when you exclude it

We no longer request your page, and we delete the document we had stored for it. Data we derived from the page is removed with the next update of the dataset it belongs to. Information that does not come from the page itself, such as DNS records, remains.

What we store

We store the HTML document the server returned, together with the TLS certificate and the IP addresses the domain resolved to. The document is used to derive the facts we publish, for example whether a site exists, which language it is in, and which social media profiles it links to.

What we publish

On the search result page we show the title and the meta description your page provides for exactly this purpose, plus the size of the document in bytes. We do not republish your page, your text, or your images, and our API returns no HTML at all.

Other ways to reach us

If you cannot change your robots.txt, write to info@webxtrend.com from an address at the domain in question, or from the address in its Impressum, and name the domain. We will stop fetching it and remove what we have stored for it.

If you believe that something we display concerns you as a private individual, the same address applies. You have a right to object under Art. 21 GDPR; see section 15 of our privacy policy.

Legal basis

Reading and analysing publicly reachable web pages is text and data mining within the meaning of Art. 4 of Directive (EU) 2019/790, implemented in Germany as section 44b of the Copyright Act. Where the data relates to an identifiable person, we rely on Art. 6(1)(f) GDPR; the balancing we applied is set out in section 11 of our privacy policy. We honour a reservation of rights addressed to webXtrendBot in your robots.txt, as described above.

We are webXtrend, based in Germany. Our full provider identification is in the imprint.