◊ About the study

The Reachable Web Observatory

A continuous random sample of the reachable public-IPv4 web

An open measurement study of the ordinary, reachable web — and where observable exposure and third-party security signals concentrate across it.

The question

Across a uniform random sample of the reachable public-IPv4 web, where do CVE-associated and reputation-flagged services concentrate — by network, geography, product, and port — and how does that exposure change over time? Rather than searching for known domains, the observatory samples public IPv4 space, keeps the latest record for each observed service, and stores daily aggregate snapshots for longitudinal analysis. See the methodology for how the measurement works and its limitations.

Who runs it

The Observatory extension is developed and maintained by Justin Walters, an independent security researcher, under Verdant Protocol. It is an independent project — not affiliated with a university and not reviewed by an institutional review board — conducted in line with the field's established ethics norms (see ethics). Collaboration with academic or nonprofit partners is welcome; reach out at research@verdantprotocol.com.

Governance, funding, and conflicts are currently simple: Justin operates the project independently through Verdant Protocol, with no university, institutional sponsor, or external funder. Material collaborators or funding relationships will be disclosed here. Technical implementation belongs on the architecture page.

The Go scanner and collector reimplement and extend the private Python system behind What's on HTTP, created by elixx. Justin performed the Go rewrite with AI coding assistance and later developed the Observatory interface, research framing, analytics, infrastructure, and independently collected dataset. The private Python source is not published here, and the Observatory is not presented as an official successor to What's on HTTP. Read the complete provenance record.

How it works, briefly

Agents generate random public IPv4 addresses, check a few common web ports (80, 443, 8000, 8080, 8443), and — for hosts answering HTTP or HTTPS — record what an anonymous visitor would see: a screenshot, the banner, HTTP status, the TLS certificate name, coarse geolocation, and structural hashes. Eligible hosts can be enriched with attributed public CVE and reputation data when providers are available; coverage is not universal. Scanning runs continuously at a deliberately slow rate and honors an operator exclusion list. The Overview is a live view of stored observations — point-in-time snapshots, not real-time scans of the hosts shown. To be explicit: this is non-invasive active measurement of publicly reachable services — never authentication, exploitation, or "hacking back." The boundary is spelled out in the ethics and methodology.

A note on the data

Records contain only what an anonymous visitor could already see, but — because discovery is random — a capture can occasionally include something personal or sensitive. CVE and reputation labels come from third-party feeds (Shodan, VirusTotal, AbuseIPDB, GreyNoise, and others, each attributed on the record); they are associations, not verified findings, and can be wrong. If a record shouldn't be public, or a label is inaccurate, tell us and we will correct or remove it — see disclosure and opt-out. The dataset is open; see data & access.

How to cite

Please cite the dataset (and note the snapshot date) when using it in published work:

Walters, J. (2026). Reachable Web Observatory: a continuous random sample of the public-IPv4 web. Verdant Protocol. https://observatory.verdantprotocol.com/