Skip to main content

How to block ICC Crawler

Operated by NICT. ICC-Crawler automatically crawls the Internet and collects web pages. ICC-Crawler is operated by the Universal Communication Research Institute at the National Institute of Information and Communications Technology (NICT).

nginx

In the server block, then reload.

nginx
# In the server block. 403 rather than 444: a closed connection tells the
# operator nothing, and an agent that gets a status code can log it.
if ($http_user_agent ~* "(ICC-Crawler/)") {
    return 403;
}

Apache

In .htaccess or the vhost.

Apache
BrowserMatchNoCase "ICC-Crawler/" bad_bot

<RequireAll>
    Require all granted
    Require not env bad_bot
</RequireAll>

Cloudflare

Security → WAF → Custom rules.

Cloudflare
# Security → WAF → Custom rules, action: Block
(http.user_agent contains "ICC-Crawler/")

WordPress

A child theme's functions.php, or a small plugin.

WordPress
// functions.php of a child theme, or a small plugin. Runs before WordPress
// builds the page, so a blocked agent costs one PHP process and no queries.
add_action('init', function () {
    $agent = $_SERVER['HTTP_USER_AGENT'] ?? '';

    foreach (['ICC-Crawler/'] as $needle) {
        if (stripos($agent, $needle) !== false) {
            status_header(403);
            exit('Blocked: ICC Crawler');
        }
    }
});

A rule on the user agent is a rule on a string it chose

The server rules match on a header the requester sets, which stops the agent that says who it is and not the one that copies somebody else's name. Botscope verifies identity against DNS and published address ranges, records which check decided each request, and shows you the ones that lied.

Others like it

Everything we know about ICC Crawler