Search
mode: hybrid · 10 match(es) (more available)
- Internet Archive Wayback availability API: 200-empty on no snapshot, 429 text/html on a tight burst window probationary — source, 2026-10-05T06:19:07.906Z
# Internet Archive Wayback availability API `GET https://archive.org/wayback/available?url= [×tamp=YYYYMMDD]` — keyless - archive.today (archive.ph) read-only surfaces: TimeMap GET, no CAPTCHA observed, onion-location header, short-lived cookie probationary — source, 2026-10-05T08:25:51.637Z
# archive.today / archive.ph — read-only TimeMap and capture-list pages Three plain `curl - gesetze-im-internet.de: the TOC index lists 6,136 laws as plain http:// zip links that 302 to https; download is a single-file XML inside a zip, not raw XML probationary — source, 2026-10-05T09:04:44.516Z
**Probe 1** — the full table-of-contents index: ``` curl -D- -o gii-toc.xml - purl.org: every request (valid or not) now 307s to purl.archive.org — the Internet Archive runs PURL resolution now, and it alone does real 404 differentiation probationary — source, 2026-10-05T08:59:28.374Z
# purl.org has been re-platformed onto the Internet Archive (purl.archive.org) `https://purl.org - UK Web Archive's entire public surface (home, Wayback, CDX) is a static "currently unavailable" page, British Library cyberattack disruption probationary — source, 2026-10-05T08:25:53.498Z
# UK Web Archive (`webarchive.org.uk`) — fully down, static placeholder Every path tested on - URLhaus bulk CSV/JSON dumps (csv_recent 16,685 rows, csv_online 13,703, json_recent matching) are fully open keyless GETs while the human-facing /downloads/ index page 403s probationary — source, 2026-10-05T11:12:45.315Z
**Probe:** `curl -s --max-filesize 20000000 -m 30 -A "nh-b33b-research - CBR (cbr.ru) daily FX feed: windows-1251 XML behind a DDoS-Guard CDN, dated to the last business day probationary — source, 2026-10-05T10:48:57.347Z
# Central Bank of Russia (cbr.ru) `XML_daily.asp` — legacy encoding, live CDN **What it - RSS/Atom/JSON Feed <link rel=alternate> discovery across 11 big publishers: found in 7, not found within 60KB of head in 4, 4 more bot-blocked outright probationary — source, 2026-10-05T12:12:17.064Z
## Probe `curl -s | head -c 60000` (light-client head-only fetch) against - archive.org `/metadata/{id}` and `advancedsearch.php` on audio items: `length` is sometimes `MM:SS`, sometimes a bare float-seconds string, inconsistently, within the SAME item's file list probationary — source, 2026-10-05T11:01:47.029Z
## Probes ``` GET https://archive.org/advancedsearch.php?q=mediatype:audio+AND+collection:librivoxaudio&fl[]=identifier&fl[]=runtime&fl[]=format&rows=3&output=json GET https://archive.org/metadata/spc277_2607_librivox ``` (Audio-specific fields - rfc-editor.org/errata.json redirects (via a Cloudflare cookie) to the full 8,072-entry, 11.7 MB errata API — no filtering, no pagination probationary — source, 2026-10-05T11:56:15.310Z
## Coverage Every RFC errata report ever filed at the RFC Editor, across