← All resources

Web Scraping for Sales: Compliance, Accuracy, and ROI

Quick answer

Scraping public B2B data is not automatically illegal under US federal computer-fraud law, a point the Ninth Circuit settled in hiQ v. LinkedIn. But hiQ still lost the case overall, paid a $500,000 judgment, and was permanently barred from LinkedIn, on breach-of-contract and state tort claims that never depended on the CFAA at all. LinkedIn sued another scraping vendor, Proxycurl, in January 2025 on nearly the same grounds, and that case also ended in settlement. The legal risk sits in your platform's terms of service and in GDPR's notice obligations, not in whether the data was "public." Treat scraped contact data as roughly 70 to 80% accurate before verification, and build your ROI case from your own labor and tooling costs rather than a vendor's headline number.

Why this is a compliance question now, not just a tools one

I'm Hlib Storchak. I build and run outbound systems for B2B founders and sales teams, 2000+ meetings booked for B2B clients, and list building is the part of that system I get the most nervous questions about. Not "which tool is best," but "am I actually allowed to do this." That question used to get waved off with "it's public data, so it's fine." It shouldn't be anymore.

Two things changed the calculus. First, LinkedIn has kept litigating scraping cases well past the point most sales teams assume the matter was settled by hiQ back in 2019. Second, GDPR enforcement against B2B contact data specifically, not just consumer data, has caught up with what most outbound teams have been doing for years. Neither of those facts means you should stop building lists. It means the honest version of "should I scrape this" has more steps in it than a lot of tool marketing pages let on.

What "scraping" actually covers for a sales team

"Scraping" gets used loosely, and the loose usage hides where the real risk sits. For a sales team, it usually means one of three different things, with three different risk profiles:

  • Browser-extension or logged-in-account scraping. A Chrome extension or automation tool that runs through your own LinkedIn or Sales Navigator session, pulling profile and company data as you browse or via automated crawling. This is the highest-risk category, because you're acting under your own account's contractual terms while you do it.
  • Third-party API access to a platform's data. A vendor scrapes on your behalf and resells structured data through its own API, without your account ever touching the target platform directly. The vendor absorbs most of the platform-contract risk, but not the privacy-law risk, since the underlying personal data is still yours to process once you have it.
  • Scraping public, non-platform sources. Company websites, press releases, job postings, public filings. This carries the lowest platform-contract risk, since there's usually no login and no user agreement in the way, but GDPR's notice obligations still apply if the data includes identifiable people.

Most of the actual legal exposure discussed below concentrates in the first category. Most of the accuracy and ROI discussion later applies to all three.

What hiQ v. LinkedIn actually settled, and what it didn't

hiQ Labs v. LinkedIn is the case almost everyone in outbound cites, and almost everyone cites the part that helps their case while skipping the part that doesn't. The Ninth Circuit did rule, most decisively in an April 2022 order, that scraping publicly accessible web pages doesn't violate the federal Computer Fraud and Abuse Act, since the CFAA is an anti-hacking statute aimed at unauthorized access, not a general ban on automated reading of pages anyone can already view. That part of the story is true and it's real precedent, at least within the Ninth Circuit.

What gets left out: hiQ lost anyway. The case ended in a stipulated judgment in December 2022, six years after it started, with a $500,000 judgment entered against hiQ, a finding of liability under California's common-law torts of trespass to chattels and misappropriation, and an injunction permanently barring hiQ from scraping LinkedIn again. The CFAA claim was the one hiQ won. Breach of contract based on LinkedIn's User Agreement, and the state tort claims, were the ones that actually ended the company's scraping business.

The CFAA myth. "The CFAA doesn't cover public data scraping" is true and often repeated as if it means "scraping public data is legal." It means one specific federal criminal and civil statute doesn't apply. Contract law and state tort law are separate exposure, and in hiQ's own case, they're what actually cost the company its business.

The 2025 case that matters more than hiQ now

If hiQ is the case everyone already knows, LinkedIn v. Nubela (the company behind the Proxycurl API) is the one that's more relevant to what a sales team is likely buying today, since Proxycurl sold structured LinkedIn profile data through an API the way a lot of enrichment tools still do. LinkedIn filed suit in the Northern District of California on January 24, 2025 (case 3:25-cv-00828), alleging breach of contract, fraud and deceit, CFAA violations, California Unfair Competition Law violations, trademark dilution, and misappropriation. The case ended in a settlement, with the specific terms undisclosed. Proxycurl's founder has since described building a successor product with what he calls a "zero-LinkedIn-data posture."

The pattern across both cases is the same: LinkedIn's core legal weapon isn't the CFAA, it's the contract you or your vendor already agreed to when creating an account, plus a set of state-law claims that don't require proving unauthorized computer access at all. A vendor that scrapes LinkedIn through logged-in accounts, its own or resold ones, carries that same contract exposure regardless of how the data eventually reaches your CRM.

What LinkedIn's own User Agreement bans, in writing

This is worth reading directly rather than assuming. LinkedIn's own User Agreement, which every account holder agrees to on signup, states that members commit to using one account under their real name, and separately prohibits using "bots or other automated methods" to access the service, add connections, or send messages without LinkedIn's express permission. That clause doesn't distinguish between scraping public profile fields and scraping private ones. It's a contract term tied to the account, not a data-classification rule.

That's the mechanism both hiQ and Proxycurl actually lost on: not a finding that the underlying data was private, but a finding that the account or the vendor's method of access breached a contract LinkedIn is entitled to enforce. Any tool that logs into LinkedIn, yours or a third party's, to pull data inherits that same exposure the moment it runs.

GDPR Article 14: the obligation almost nobody follows

If your list building touches any EU-based contact, and most B2B ICPs eventually do, GDPR's Article 14 applies specifically to scraping, because it governs personal data obtained from a source other than the person themselves. It requires the controller (your company) to give the data subject a defined set of disclosures, including who you are, why you're processing their data, and their rights, within a reasonable period and at the latest within one month of obtaining the data, or before first contacting them if that comes sooner.

There are exemptions, but they're narrower than most outbound teams assume. Article 14(5) allows skipping notice if the person already has the information, if giving notice would take "disproportionate effort" (a standard built mainly for research and archival use, not commercial list building), or if a law requires non-disclosure. Publicly available data is not, on its own, an exemption. A prospect's job title being on their public LinkedIn profile doesn't remove your obligation to tell them you hold their data, if your list includes EU-based contacts.

Source of exposureWhat it actually governsWhat "it's public data" does NOT excuse
US federal law (CFAA)Unauthorized computer accessNothing extra beyond access itself, per the Ninth Circuit's hiQ ruling
Platform contract (User Agreement)What your or a vendor's account is allowed to doAutomated collection, even of visible fields, if the agreement bans it
GDPR Article 14Notice to the data subject when data comes from a source other than themThe one-month disclosure obligation, which public visibility does not remove
US state privacy law (CCPA/CPRA)A California resident's rights over their own dataOpt-out, deletion, and disclosure rights, regardless of how the data was sourced

US state privacy exposure: CCPA and CPRA

For a US-focused outbound program, California residents' data carries its own obligations under the CCPA as amended by the CPRA, independent of GDPR. Any California-resident contact in your list, however it was sourced, is entitled to know what you hold, request deletion or correction, and opt out of your sale or sharing of their data, and limit use of anything classed as sensitive personal information. Other states (Virginia, Colorado, Connecticut, and a growing list of others) have passed broadly similar frameworks with their own thresholds and enforcement mechanisms. None of these statutes ask how the data was collected before granting the rights; a scraped record carries the same obligations as a purchased one.

A compliance checklist before you turn on a scraper

Run through this before adding a new scraping tool or vendor to your stack, not after a prospect or a platform's legal team asks a question you can't answer:

  1. Check the source's own terms of service for automated-access and bot clauses, not just a general "no scraping" line. LinkedIn's language, banning bots and unauthorized automated methods, is common across major platforms.
  2. Ask any vendor directly how they access the data. Logged-in account scraping carries platform-contract risk that a public-page, no-login scrape doesn't. If a vendor won't answer plainly, that's itself an answer.
  3. Confirm whether EU contacts are in scope, and if so, build an Article 14 notice process (a first-touch email footer or a data-source disclosure page is a common approach) rather than assuming public visibility covers you.
  4. Confirm your CCPA/CPRA posture for California residents specifically: a documented way to receive and honor a deletion or opt-out request.
  5. Keep a record of source and collection method per list, not just the contact fields. If a platform or a regulator asks where a record came from, "I don't know, the tool just gave it to us" is the answer that turns a routine question into a real problem.
  6. Re-check this list whenever you add a new source or vendor, since terms of service and enforcement postures change; LinkedIn's own suit against Proxycurl came years after hiQ, not because the law shifted, but because LinkedIn kept enforcing.

How much of what you scrape is actually correct

Even fully compliant scraping has a second, separate problem: a meaningful share of what you collect is wrong the moment you collect it, and more of it goes wrong every month after. Cleanlist's 2026 benchmark, built on 500 real B2B leads tested against identical inputs across providers, found single-source databases landing in a 70 to 80% verified-email range, versus Cleanlist's own 25-plus-provider waterfall claiming 98% on the same test set (Cleanlist, B2B data providers tested). Treat that 98% with the same caution you'd give any vendor grading its own homework, but the 70 to 80% single-source range is a reasonable planning number for a typical scraper or single database, verified before you trust it.

Decay compounds that gap after collection. ZeroBounce's 2026 Email List Decay Report puts annual email list decay at roughly 23%, meaning close to a quarter of a list you scrape today reads as invalid a year from now, separate from whatever inaccuracy existed on day one (cited via Cleanlist's same 2026 report). A scraped list isn't a one-time cost, it's a depreciating asset that needs a refresh cadence built into your list-building budget from the start, not bolted on once reply rates start slipping.

Verify before you trust a vendor's own accuracy number. The Cleanlist report says it plainly: "where a vendor publishes an accuracy figure, it is theirs, not ours." Ask any data or scraping vendor for a benchmark against 100 to 500 of your own real records before accepting a self-reported percentage, the same way I'd ask before trusting any comparison page in this space.

Scraping vs buying data vs an enrichment waterfall

These three approaches solve overlapping problems with different cost and risk shapes. None is universally right; which one fits depends on your list size, refresh frequency, and legal risk tolerance.

DimensionBuild your own scraperBuy a data provider subscriptionEnrichment waterfall (multi-provider)
Upfront costEngineering or ops time; low cash, high laborSubscription fee, check current pricingPer-record or subscription fee across providers, check current pricing
Ongoing costMaintenance every time a target site changes its layout or blocks the methodRenewal, plus decay eating a share of records yearlyRenewal across multiple vendors, plus a waterfall fee if using a managed one
Platform-contract riskHighest if it logs into an account (LinkedIn, Sales Navigator)Shifted to the vendor, though not eliminated for you as the data userShifted to the vendors, spread across several
Typical accuracyDepends entirely on build quality and verification stepRoughly 70-80% single-source per Cleanlist's 2026 test, before your own verificationCan exceed single-source figures by combining providers, verify against your own records regardless
Best fitA narrow, unusual data need no vendor covers, run by a team with engineering capacitySteady, predictable volume from a well-covered segmentTeams that need higher match rates across varied ICPs and can absorb multi-vendor cost and complexity

Building your own ROI case, with the assumptions shown

Here's a way to build your own number rather than trust a vendor's "300-800% ROI" claim, which is a real range some vendors publish but one that depends entirely on inputs you control, not ones they do. Assume: an ops or SDR person costs $60/hour fully loaded (salary plus on-costs, adjust for your market), building and maintaining a basic scraper eats roughly 15 hours a month once you include fixing breakage when a target site changes, a verification tool costs $50 to $150 a month for a mid-size list, and 20 to 30% of scraped records need replacing within 6 months on the decay and accuracy figures above.

Input (your own figure, this is an illustration)Assumed valueMonthly cost driver
Labor: build and maintenance hours~15 hrs/month at $60/hr~$900/month
Verification tooling$50-150/month$50-150/month
Replacement cost for decayed/bad records20-30% of list replaced within 6 monthsScales with your per-record acquisition or verification cost
Illustrative total, mid-size listSum of the aboveRoughly $950-1,050/month before replacement cost, excluding legal/compliance review time

Compare that all-in number, not just the "free" sticker price of a browser extension, against a subscription's actual renewal cost plus its own decay-driven replacement need. The honest comparison usually isn't "scraping is free and buying costs money," it's "scraping shifts the cost from a subscription line to a labor and maintenance line, plus a legal-risk line that's real but harder to price." Swap in your own hourly rate, your own maintenance hours, and your own list size before you decide which side of that trade you're actually on.

The mistake I see most often when a client hands me a scraped list

The mistake I see most often when I take over an account isn't the scraping method itself, it's the missing paper trail. A client hands over a list built from three different tools over eighteen months, and nobody can say which records came from a logged-in LinkedIn scrape, which came from a public-page crawl, and which came from a purchased database. That gap matters twice: once if a platform or a regulator ever asks, and once for basic list hygiene, since I can't tell which segment is decaying fastest or which source is worth renewing without knowing where each record came from in the first place. The fix costs almost nothing: tag source and collection date on every record when it enters your CRM, before it gets mixed into one undifferentiated list.

Key takeaways

  • Scraping public data isn't automatically illegal under the CFAA, the Ninth Circuit settled that in hiQ v. LinkedIn, but hiQ still lost the case overall and paid a $500,000 judgment on contract and tort claims.
  • LinkedIn sued another scraping vendor, Proxycurl, in January 2025 on similar grounds; that case also ended in a settlement, showing the enforcement pattern didn't stop with hiQ.
  • The real exposure sits in platform terms of service (LinkedIn's User Agreement bans bots and automated methods outright) and in GDPR Article 14's one-month notice obligation, not in whether the data was publicly visible.
  • CCPA/CPRA rights apply to California residents' data regardless of how it was collected, scraped or purchased.
  • Plan for roughly 70-80% single-source accuracy before your own verification, per Cleanlist's 2026 benchmark, and around 23% annual list decay per ZeroBounce, before deciding a scraper or database is "accurate enough."
  • Build your own ROI number from your actual labor, tooling, and replacement costs rather than a vendor's headline ROI percentage.

FAQ

Is it legal to scrape LinkedIn for B2B sales leads?

Scraping publicly visible pages isn't automatically a federal computer-fraud crime, per the Ninth Circuit's ruling in hiQ v. LinkedIn. But LinkedIn's own User Agreement bans bots and unauthorized automated access as a contract term, and hiQ still lost its overall case on breach-of-contract and state tort grounds, ending in a $500,000 judgment and a permanent scraping ban. LinkedIn sued another scraping vendor, Proxycurl, on similar claims in January 2025, which also settled. Treat "it's public" as legally insufficient on its own.

Does GDPR apply to data I scraped rather than collected directly from the person?

Yes. GDPR Article 14 specifically governs personal data obtained from a source other than the data subject, which is what scraping is. It requires giving the person a defined set of disclosures within a reasonable period, at the latest within one month of obtaining the data. Public availability of the data doesn't remove this obligation on its own; the exemptions are narrower, mainly covering cases where notice would take disproportionate effort or the person already has the information.

How accurate is scraped B2B contact data before I verify it?

Plan for roughly 70 to 80% verified-email accuracy from a single source before your own verification step, per Cleanlist's 2026 benchmark testing 500 real B2B leads across providers. Multi-provider waterfalls can do better on the same test set, but any vendor's own accuracy claim about itself should be checked against a sample of your own real records before you trust it.

Should I build my own scraper or pay for a data provider?

It depends on your list size, refresh frequency, and risk tolerance. Building your own shifts cost into labor and maintenance hours plus platform-contract risk, especially for anything that logs into an account like LinkedIn. Buying a subscription shifts most of that contract risk to the vendor but still carries GDPR and CCPA obligations on your side as the data user, plus its own decay. Build the full monthly cost, including maintenance hours and replacement of decayed records, before comparing either option to a subscription's sticker price.

What should I actually track to stay compliant when building lists from multiple sources?

Tag the source and collection date on every record when it enters your CRM, not after the fact. Confirm whether each source's platform terms allow automated collection, build an Article 14 notice process if any EU contacts are in scope, and keep a documented way to honor CCPA/CPRA deletion and opt-out requests regardless of how a record was originally sourced.

Want a list-building process that won't blow up on you later?

There are three ways to work with me: done-for-you outbound where I build and source lists with the compliance and verification steps already built in, fractional Head of GTM where I own that process as your GTM lead, or standing up the process inside your own team so it keeps running safely without me.

Book a call