How to block Library Of Congress Web Archiving
Operated by United States Library of Congress. The Library of Congress Web Archive manages, preserves, and provides access to archived web content selected by subject experts from across the Library, so that it will be available for researchers today and in the future. More information on the programme here: https://www.loc.gov/programs/web-archiving/about-this-program/ And information about crawling policy here: https://www.loc.gov/programs/web-archiving/for-site-owners/
How it identifies itself
Library Of Congress Web Archiving
robots.txt
What the agent asks to be told. It obeys this or it does not — nothing enforces it.
User-agent: Library Disallow: /
nginx
In the server block, then reload.
# In the server block. 403 rather than 444: a closed connection tells the
# operator nothing, and an agent that gets a status code can log it.
if ($http_user_agent ~* "(https://www.loc.gov/programs/web-archiving/for-site-owners/|Library|Library Of Congress Web Archiving)") {
return 403;
}
Apache
In .htaccess or the vhost.
BrowserMatchNoCase "https://www.loc.gov/programs/web-archiving/for-site-owners/" bad_bot
BrowserMatchNoCase "Library" bad_bot
BrowserMatchNoCase "Library Of Congress Web Archiving" bad_bot
<RequireAll>
Require all granted
Require not env bad_bot
</RequireAll>
Cloudflare
Security → WAF → Custom rules.
# Security → WAF → Custom rules, action: Block (http.user_agent contains "https://www.loc.gov/programs/web-archiving/for-site-owners/") or (http.user_agent contains "Library") or (http.user_agent contains "Library Of Congress Web Archiving")
WordPress
A child theme's functions.php, or a small plugin.
// functions.php of a child theme, or a small plugin. Runs before WordPress
// builds the page, so a blocked agent costs one PHP process and no queries.
add_action('init', function () {
$agent = $_SERVER['HTTP_USER_AGENT'] ?? '';
foreach (['https://www.loc.gov/programs/web-archiving/for-site-owners/', 'Library', 'Library Of Congress Web Archiving'] as $needle) {
if (stripos($agent, $needle) !== false) {
status_header(403);
exit('Blocked: Library Of Congress Web Archiving');
}
}
});
robots.txt is a request. The rules below it are the wall.
Library Of Congress Web Archiving publishes a token, so the directive above will work for as long as it keeps honouring it — and nothing except its own good manners makes it. The server rules match on a header the requester sets, which stops the agent that says who it is and not the one that copies somebody else's name. Botscope verifies identity against DNS and published address ranges, records which check decided each request, and shows you the ones that lied.