Research

Residential IP Reputation: 1,000 Published Observations

By Published Updated 10 min read
Residential IP Reputation: 1,000 Published Observations

TL;DR

What 1,000 redacted route observations across 25 ASNs show about DNSBL labels, Tor and DROP checks, network classification, and dataset limits.

On this page

Start With the Downloadable Artifact#

The public residential IP reputation CSV contains 1,000 rows associated with 25 claimed residential ASNs. The rows comprise 624 IPv4 observations and 376 IPv6 observations. IPv4 addresses are represented as /24 network prefixes and IPv6 addresses as /48 prefixes, so the file does not disclose full exit addresses.

The artifact supports recalculating counts for its populated columns. It does not support independently replaying the original capture because it omits full addresses, per-row timestamps, raw DNS or API responses, software versions, resolver details, failed candidate attempts, and the capture harness. The correct unit is therefore a published route observation, not a publicly verifiable unique full IP. The file contains 981 distinct redacted prefixes; multiple full addresses can legitimately share one prefix.

What the CSV Directly Shows#

All 1,000 rows have a claimed ASN, carrier, country, redacted prefix, address family, Cloudflare-observed ASN and country, legacy Cloudflare threat score, and known-bot boolean. The applicable Spamhaus DROP column is populated for each address family, and the ASN-level DROP and Tor fields are populated for every row.

What the CSV Directly Shows: data table 1
Published observationCountInterpretation
Rows1,000Published route observations
Claimed ASNs2540 rows per claimed ASN
IPv4 / IPv6624 / 376Address-family split
Tor exit matches0No published row matched the captured Tor list result
Applicable Spamhaus DROP matches0No published row is marked true
ASN DROP matches0No claimed ASN is marked true
Non-empty DNSBL labels501494 Spamhaus ZEN labels and 7 DroneBL labels
Claimed/observed ASN agreement96896.8% of rows
Claimed/observed country agreement99899.8% of rows

A DNSBL Label Is Not a General Web-Reputation Score#

The dnsbl_listed_in column is non-empty for 501 of the 624 IPv4 rows: 494 contain spamhaus_zen and 7 contain dronebl. That is 80.3% of the IPv4 sample. The CSV does not publish the raw DNS answers or individual Spamhaus subzone return codes, so it cannot show which ZEN component produced each label.

This distinction matters because Spamhaus ZEN aggregates several email-oriented lists, including the Policy Block List. Consumer broadband can be listed for SMTP policy reasons without being classified as malicious web traffic. A combined ZEN label should not be converted into a universal “bad IP” verdict or used to predict whether a website will accept an HTTP request.

The zero DROP results answer a different question: none of the published rows is marked as belonging to an applicable DROP network in this snapshot. They do not prove that an address was safe, authorized, household-operated, or free from destination-specific restrictions.

Four Reserved Enrichment Columns Are Empty#

The CSV includes columns named ip2location_proxy_type, abuseipdb_confidence, ipinfo_hosting, and greynoise_class. Every value in those four columns is blank across all 1,000 rows. They are placeholders, not completed measurements, and they are intentionally excluded from this page's Dataset variableMeasured list.

A blank value does not mean zero, clean, residential, benign, or unknown according to the named provider. It means the public file contains no result. Any comparison involving those services requires a new, authorized capture using their current APIs or licensed datasets, with the response fields and timestamps preserved.

How to Read the Cloudflare Fields#

The CSV records cf_threat_score=0 and cf_client_bot=false for every row. Cloudflare's current documentation states that cf.threat_score is now always zero, so this field has no variance and cannot rank the routes. It should not be presented as proof that Cloudflare considered an address trustworthy.

cf.client.bot indicates whether a request came from a known good bot or crawler. A value of false does not mean human, safe, or accepted; it only means the request was not identified through that known-bot field. Cloudflare Bot Management's granular 1–99 score is a separate Enterprise feature and is not present in this CSV.

The destination-side ASN and country fields are more useful here. They agree with the claimed ASN in 968 rows and claimed country in 998 rows. Those are classification comparisons for this capture, not guarantees about later routes or every geolocation database.

This Dataset Does Not Produce a Best-ASN Leaderboard#

Forty rows are associated with each claimed ASN, but the populated fields do not measure destination success, latency, abuse history, consent, session quality, or long-term stability. DNSBL coverage is dominated by an email-policy aggregate, while the Tor and DROP columns are uniformly false. Ranking ASNs from those fields would create a score without a defensible outcome variable.

The dataset can support narrower questions: address-family mix, whether the applicable published lists marked a row, how often claimed and Cloudflare-observed ASN or country agreed, and which fields are absent. It cannot establish which carrier is “best” for scraping, accounts, advertising, purchasing, or any other third-party workflow.

Recalculate the Published Aggregates#

A reader can reproduce the counts above from the CSV without contacting any proxy or third-party reputation service. Count rows by ip_version; count distinct asn_claimed; filter non-empty dnsbl_listed_in; compare asn_claimed after removing its AS prefix with cf_asn_observed; and compare the two country columns.

Preserve the downloaded file and record its SHA-256 hash before analysis. Treat blank strings as missing, not false. Apply spamhaus_drop_v4 only to IPv4 rows and spamhaus_drop_v6 only to IPv6 rows. Do not infer full-address uniqueness from redacted prefixes.

Reproducing the original network capture would require additional materials that are not distributed: the full addresses under appropriate access controls, capture timestamps, raw lookup responses, resolver and software configuration, candidate-selection records, retry logic, and a versioned harness.

Responsible Use and Practical Meaning#

Reputation sources answer different questions. Tor lists identify published exits. DROP lists identify networks Spamhaus recommends dropping. DNSBLs are largely designed for messaging abuse and policy. ASN and country databases classify network origin. Bot-management products evaluate a request using additional client, account, sequence, and behavior signals.

Use the source designed for the decision you need to make, document its timestamp, and avoid turning a missing or unrelated field into a broad trust score. For systems you operate, combine narrowly relevant signals with rate limits, authentication, monitoring, and a review path. For third-party data access, prefer official APIs, feeds, licenses, and written authorization.

Frequently Asked Questions

Does the CSV publish 1,000 full IP addresses?
No. It publishes 1,000 rows with IPv4 /24 or IPv6 /48 prefixes. The full addresses and per-row timestamps are not distributed.
Were AbuseIPDB, GreyNoise, IPinfo, and IP2Location measured?
No published values are present. All four reserved columns are blank across every row, so they are excluded from the Dataset measurement list.
Why are 501 IPv4 rows labeled by a DNSBL?
The combined field contains 494 Spamhaus ZEN labels and 7 DroneBL labels. Because raw DNS response codes are not published, the CSV cannot decompose ZEN into its individual components. A ZEN label is not a universal web-risk score.
Does cf_threat_score zero mean Cloudflare trusted the route?
No. Cloudflare documents that the legacy cf.threat_score field is now always zero. It cannot distinguish these rows or represent an Enterprise Bot Management score.
Does cf_client_bot false mean the request was human?
No. The field identifies known good bots or crawlers. False means that known-bot designation was not present; it does not prove a human or predict acceptance.
Can the original capture be reproduced from the CSV alone?
No. The aggregates can be recalculated, but replaying the capture requires full addresses, timestamps, raw responses, resolver and software details, candidate records, and the harness.

Related reading

Ready to scale your data collection?

Join 8,000+ customers on Databay: 34M+ residential IPs across 200+ countries, pay as you go.

Pricing, order minimums, and traffic validity vary by product.