Anti-Scraping Techniques: Which Layer Blocked Your Scraper
You deployed the scraper that worked on your laptop and got back 403 Forbidden. Or worse, a clean 200 OK with an empty results array.
Anti-scraping techniques are the checks a site runs to tell an automated client apart from a real user. They fire in a fixed order: IP reputation, then the TLS handshake, then HTTP/2 framing, then request headers, then the JavaScript environment, then session behaviour. A request that fails one of them never reaches the next. Which layer failed is the entire diagnosis. What you do about it is usually a library swap.
Every source on this page was fetched on 2026-09-09. Where an anti-bot vendor does not publish how its own detection works, this page says so instead of filling the gap.
Which layer blocked you?
Start from the symptom, not from the vendor. Two sites behind the same WAF can block you at different layers depending on how the customer configured it.
| What you observe | Layer that most likely fired | First thing to change |
|---|---|---|
403 on the very first request, every URL, no JavaScript involved | Network: IP reputation and ASN | Get off datacenter IPs |
Works from curl, fails from requests, httpx or aiohttp | TLS handshake | curl-cffi or got-scraping |
Works from curl-cffi, fails from a hand-rolled HTTP/2 client | HTTP/2 framing | Real browser, or an impersonation library |
403 only when you set a Chrome User-Agent | Header contradiction | Stop hand-editing headers |
| Challenge page after a handful of requests | JS environment scoring | Patched browser plus residential IP |
| First 50 pages fine, then everything blocks | Rate limiting or behavioural scoring | Concurrency and pacing |
200 OK with empty, partial or obviously wrong data | Silent block | Diff against a real browser render |
| Timeout or socket hangup, no response at all | Firewall drop | Back to IP and ASN |
The last three rows are the ones people miss. Apify's own academy documents them as first-class block types alongside the 403: a "request timeout/socket hangup" is "the cheapest defense mechanism where the website won't even respond to the request," and a site can return "empty results" while pretending "to not find any results," or serve results that are "totally fake" (docs.apify.com/academy/anti-scraping/techniques, checked 2026-09-09). If your scraper's output quality dropped without an error rate to match, you are being blocked, not debugged.
What do the anti-bot vendors actually publish?
Most articles about anti-bot detection describe the internals of DataDome, Kasada and PerimeterX with a confidence that nobody's documentation supports. Here is what each vendor publishes about its own mechanism, checked on 2026-09-09.
| Vendor | What the vendor publishes about its own detection | Source fetched 2026-09-09 |
|---|---|---|
| Cloudflare Bot Management | A bot score "from 1 to 99." Four engines: heuristics ("pattern matching against a database of known malicious fingerprints"), machine learning ("Produces most scores between 2 and 99"), JavaScript Detections ("Catches headless browsers... and other automation tools"), and anomaly detection, which the page marks deprecated. | developers.cloudflare.com/bots/concepts/bot-score |
| Cloudflare Turnstile | "proof-of-work (computational puzzles), proof-of-space, probing for web APIs, and various other challenges for detecting browser-quirks and human behavior." The Managed widget "automatically decides whether to show a checkbox based on visitor risk level"; the Non-interactive and Invisible widgets never show one. | developers.cloudflare.com/turnstile |
| DataDome | Public product docs exist (Bot Protect, Account Protect, Ad Protect and others), but no page in them describes the signals or the scoring. | docs.datadome.co/docs |
| HUMAN (formerly PerimeterX) | Nothing retrieved. humansecurity.com returned 403 to a non-browser fetch. | — |
| Akamai Bot Manager | Nothing retrieved. techdocs.akamai.com 302-redirects to an Auth0 login; the Bot Manager docs are behind a customer account. | — |
| Imperva | Nothing retrieved. docs.imperva.com now 302-redirects to docs-cybersec.thalesgroup.com, and the bot-protection page returned only the portal shell. | — |
| Kasada | No public detection documentation found. | — |
Cloudflare is the only vendor in that list that documents its own detection in usable detail. Everything you have read about what Kasada "checks" or how DataDome scores a session is reconstructed from observation, including anything specific this page says below. Treat it as an observation with a shelf life, not a spec.
There is a small joke in the table worth pulling out: two of the seven vendors blocked the fetch that was trying to read their marketing pages. That is the actual state of the web in 2026, and it is why the diagnostic table above is more durable than any vendor profile.
The stack, and how to check each layer yourself
Rather than trusting a vendor description, measure your own client. Each row here is something you can run against your scraper's real HTTP stack, from the machine it will actually run on.
| Order | Layer | What it inspects | How to check your own client |
|---|---|---|---|
| 1 | Network | IP reputation, ASN classification, proxy detection | Resolve your egress IP from inside the runner. AWS, GCP, Hetzner and DigitalOcean ASNs are pre-classified as datacenter |
| 2 | TLS | ClientHello cipher and extension order (JA3, JA4) | GET https://tls.browserleaks.com/json from your actual HTTP library. It returns ja3_hash, ja4 and akamai_hash |
| 3 | HTTP/2 | SETTINGS frame values, window update, pseudo-header order | The akamai_hash field from the same endpoint |
| 4 | HTTP headers | Casing, order, Sec-Fetch-*, Sec-Ch-Ua-* client hints | Echo your own headers and diff them against a real Chrome capture |
| 5 | JS environment | navigator.webdriver, CDP leaks, canvas, WebGL, audio, fonts | Point your automated browser at creepjs and read the report it generates |
| 6 | Behaviour | Mouse deltas, scroll cadence, session depth | No public test exists. Track your own block rate against session length |
The consequence that saves the most time: a better proxy cannot fix a TLS fingerprint, and a better browser cannot fix a behavioural score. Find the layer before you buy anything.
1. IP reputation and rate limiting
A single IP walking /products/?page=1…500 in order is trivial to block, and a datacenter ASN is flagged before the first request is even parsed. Slowing down does not rehabilitate a datacenter IP.
Apify Proxy publishes the trade-off plainly: datacenter is "the fastest and cheapest option," but "other users' activity can get these IPs blocked," while residential IPs "are the least likely to be blocked" (docs.apify.com/platform/proxy, checked 2026-09-09).
The number that catches people out is session lifetime, and the two pools differ by two orders of magnitude. A datacenter IP/session pairing "is persisted and expires 26 hours later," with each request resetting the clock (docs.apify.com/platform/proxy/datacenter-proxy). A residential session lasts roughly 30 minutes. If your login-to-checkout flow takes longer than 30 minutes on residential IPs, it will break in the middle and look like a block.
For sizing rotation and choosing between pools, see the proxy strategy learning path. For the Crawlee code, see rotating proxies and sessions.
2. TLS fingerprinting (JA3 and JA4)
This is the reason the "works in curl, fails in Python" scraper exists. JA3 hashes the ClientHello, in order: TLS version, cipher suites, extensions, elliptic curves, EC point formats.
JA4 is its successor, created by John Althouse at FoxIO. Its structural change matters more than the algorithm: instead of one opaque hash, JA4 uses "an a_b_c format, delimiting the different sections," so a defender can hunt on a partial fingerprint rather than an exact match (github.com/FoxIO-LLC/ja4). Randomising your extension order to dodge a JA3 hash does not dodge a JA4 prefix. The JA4 TLS-client method is BSD 3-Clause; the rest of the JA4+ suite is patent-pending under the FoxIO License 1.1, which is why you see JA4 adopted at CDNs faster than the other JA4+ methods.
python-requests, httpx, aiohttp and Node's built-in https each emit a fingerprint no browser produces. Cloudflare and Akamai inspect this at the edge, so a Chrome User-Agent on a Python TLS stack is a contradiction that resolves before your code runs.
What works:
curl-cffi(Python) binds to a fork of curl-impersonate and advertises "JA3/TLS and http2 fingerprints impersonation." It is MIT-licensed and ships 37 preset browser profiles, selectable per request (chrome124,chrome146,safari260,safari260_ios), plus custom JA3 and Akamai strings (github.com/lexiforest/curl_cffi, preset list, both checked 2026-09-09). Additional fingerprints are sold separately through impersonate.pro and pulled in withcurl-cffi update.got-scraping(Node) is what Crawlee uses underneath, with Chrome-like TLS and HTTP/2 fingerprints by default.- Real browser automation uses Chromium's or Firefox's own TLS stack, so the fingerprint matches without configuration.
- Do not try to fix
requestswith configuration. There is no setting that changes the TLS layer.
3. HTTP/2 and header fingerprinting
Past TLS, the HTTP/2 handshake is itself a fingerprint. Akamai's hash covers SETTINGS frame values, the WINDOW_UPDATE increment, priority frame structure and pseudo-header order. Hand-written HTTP/2 clients rarely reproduce Chrome's exact sequence, which is why tls.browserleaks.com/json returns akamai_hash next to the JA3 and JA4 fields.
On HTTP/1.1 the giveaways are cheaper to spot:
- Header casing. Browsers send
User-Agent; Python'srequestslowercases everything. - Header order. Chrome's order is stable and libraries do not match it.
- Missing client hints.
Sec-Ch-Ua,Sec-Ch-Ua-MobileandSec-Ch-Ua-Platformaccompany a real Chromium request. Their absence under a Chrome User-Agent is the contradiction. Sec-Fetch-*semantics. These describe request context.Sec-Fetch-Site: noneon what claims to be a click-through is self-refuting.
What works: browser automation, or a library built for impersonation. Manual header reordering in requests does nothing, because the underlying urllib3 normalises the order back.
4. Browser and environment fingerprinting
Once JavaScript executes, the page reads your runtime. Cloudflare names this layer directly: its JavaScript Detections engine "catches headless browsers (browsers controlled by software, with no visible window or human operator) and other automation tools."
The signals that carry the most weight:
navigator.webdriveristruein unpatched Playwright and Puppeteer.- CDP leaks. Driving Chrome over the DevTools Protocol leaves observable artefacts. Patchright and Nodriver patch these; stock Playwright does not.
- Canvas, WebGL and audio hashes. Headless Chrome on Linux reports a
SwiftShaderrenderer that a real desktop never has, and produces a canvas hash shared across a very large population of scrapers. - Font enumeration. A headless container ships a handful of fonts; a real desktop has hundreds.
- Timezone against geo-IP.
Intl.DateTimeFormat().resolvedOptions().timeZonehas to match the exit node you are using.
What works: Patchright or Nodriver, which patch at the binary level rather than by injecting JavaScript, because the act of injecting is itself detectable. Camoufox is the Firefox-based equivalent. Crawlee's fingerprint injection rotates internally consistent fingerprints across sessions, which matters more than any single value being unusual.
If your target also renders its content client-side, the browser you brought to clear fingerprinting is also your rendering engine. See scraping dynamic websites with Playwright.
5. Behavioural analysis and honeypots
Behavioural scoring watches a session rather than a request: mouse trajectory (real movement has jitter and acceleration, page.mouse.move() is a straight line), scroll velocity, time to first click, focus and blur. Randomising static fingerprints does not help, because the signal is in the dynamics. Realistic pacing does: do not click 50 ms after DOM ready, scroll an element into view before clicking it, and keep one coherent behavioural profile per cookie jar.
Honeypots are the cheap version of the same idea. Hidden links (display:none, zero opacity, off-screen), form fields that should stay empty, and endpoints advertised only as Disallow entries in robots.txt all catch crawlers that follow every <a href>. Crawlee's enqueueLinks honours CSS visibility in a browser crawler. In an HTTP-only crawler you cannot compute styles, so target selectors that describe real navigation (nav a, main a[href^="/product/"]) instead of a blanket a[href].
How do I fix a 403 error on Apify?
Apify's help centre answers this in two short articles, and both are shorter than the question deserves. "How to resolve 403 response status errors" says "Using Apify and Crawlee's default configuration with some best practices generally resolves this blocking problem," and gives one concrete instruction: on a block, "simply throw an error (ideally with session.retire()) to initiate a retry request with a new session" (help.apify.com/en/articles/7029516). "How to bypass Cloudflare" says "Apify's default scrapers are a great choice for Cloudflare protected sites" and warns to "be especially mindful of the Cloudflare browser check and challenge, it might require some additional steps in your code" (help.apify.com/en/articles/7029488). Both then hand you off to a blog post.
Here is the escalation ladder those articles imply, in the order that costs you least:
- Confirm it is a block. Re-run one failing URL through a real browser. If the browser gets data and your crawler gets a
403, an empty array or a redirect to the homepage, it is a block. - Retire the session, do not retry the request.
session.retire()discards the session permanently and takes its IP with it;markBad()only lowers the score for transient failures like a timeout. Crawlee'sSessionPoolbinds cookies to a proxy URL through the session ID, so retiring is what actually gets you a fresh identity (crawlee.dev session management). - Change the proxy tier. Datacenter to residential. This alone clears layer 1 and nothing else.
- Change the HTTP client. If the block survives a residential IP and appears before any JavaScript runs, you are at layer 2 or 3. Swap to
curl-cffiorgot-scraping. - Change the browser. If the block only appears after JavaScript, you are at layer 5. Move to a patched browser with fingerprint injection.
- Stop building and buy the unblock. Below.
When is buying an unblocker cheaper than building one?
Apify Proxy has a fourth group most people never find. Unblocker "automatically bypasses anti-bot and anti-captcha systems using built-in smart routing" (docs.apify.com/platform/proxy, checked 2026-09-09), and you reach it by putting groups-UNBLOCKER in the proxy username on proxy.apify.com:8000 (docs.apify.com/platform/proxy/unblocker, checked 2026-09-09). It bills $1.50/1,000 successful requests subject to change · verified 2026-09-09 on the Free and Starter plans, and only on success.
All four prices below were read from the vendors' own pricing pages on 2026-09-09.
| Option | Price | Billed on | The catch |
|---|---|---|---|
| Apify Unblocker | $1.50/1,000 (Free, Starter), $1.25 (Scale), $1.00 (Business) | Successful requests | No session parameter, so you cannot hold one IP across requests. Multi-step logged-in flows are out |
| Apify residential proxy | $8/GB (Free, Starter), $7.50 (Scale), $7 (Business) | Traffic | ~30-minute sessions, and you still supply the browser and the fingerprint |
| Apify datacenter proxy | 5 IPs included on Free, 30 on Starter, then $1/IP | IPs held | Pre-classified ASNs. Useful only against targets with no protection |
| Bright Data Web Unlocker | $1.50/1,000 pay-as-you-go; $499/month for 383K requests, then $1.30/1,000 | Successful requests | The $1.30 rate requires the $499 monthly commitment. Free tier is 5,000 requests/month |
Sources: apify.com/pricing and brightdata.com/pricing/web-unlocker, both 2026-09-09.
The two pay-per-success products list at the same rate, so that is not the decision. The decision is against rolling your own on residential IPs, and it turns on page weight, because residential is billed on traffic while unblockers are billed on requests. Measure your own before you budget. At $8/GB, a 500 KB page costs $4.00 per 1,000 pages; a 2 MB browser session that loads images and fonts costs $16.00 per 1,000. Blocking subresources in Playwright drops the same run back under the unblocker rate, which is usually the cheapest optimisation available to anyone already running browsers.
If you need a logged-in, multi-step session, Unblocker cannot do it. It has no session parameter, and no amount of budget changes that. Residential with sticky sessions is the only Apify path for that shape of job.
If Bright Data is the better buy, buy Bright Data. If unblocking is the only thing you need and you already run your own scrapers, storage and schedules, Web Unlocker is a standalone product with a free tier and a published commitment tier, and you should not pay for a platform to get at it. Apify's Unblocker is a proxy group inside a platform, which is worth it when you want the run, the dataset and the schedule in the same place as the unblock.
The Free plan is $0 with $5 of monthly usage (apify.com/pricing, 2026-09-09), which covers running steps 3 through 6 above against your real target. That is enough to learn which layer is blocking you, which is the only thing worth knowing before you pick tools. Start at apify.com.
"A pre-built scraper won't fit my target." Often true, and it is worth 10 minutes to find out: Apify publishes "70,000+ ready-to-run" Actors (apify.com/store, checked 2026-09-09), and someone maintaining one against a hard target is absorbing the detection updates you would otherwise chase. Search it for your target before you build: Apify Store search.
Where Apify's own pages disagree
Concurrency is the setting most likely to trip a rate limit, and Apify's two first-party sources do not agree on what the Free plan allows. apify.com/pricing lists 5 concurrent runs on Free; docs.apify.com/platform/limits lists 25. Both were checked on 2026-09-09. They agree on Starter, at 32.
Plan for 5, and confirm the real number on your account's Limits page before you build a schedule that depends on it. If you size a crawl for 25 parallel runs and get 5, the run does not fail, it just takes five times as long and misses its window.
Ethics and compliance
Every technique here can be pointed at a target that does not want to be scraped. That is a legal question, not a technical one, and the answer changes by jurisdiction. Before you deploy: read the target's terms, respect robots.txt where it applies, rate-limit below what the site can comfortably serve, scrape public data only, and take actual legal advice on anything touching PII, copyrighted content or CFAA-relevant jurisdictions.
Conservative pacing is both the honest approach and the most reliable one, because it keeps you under the thresholds that trigger layers 1 and 6 in the first place. For the platform-level view of each obstacle, see 14 web scraping challenges and how to solve them. For a worked Cloudflare walkthrough, see bypassing Cloudflare and CAPTCHAs, and for WAF architecture, engineering bypass architecture for deep-packet WAFs.
TLS fingerprinting. curl and python-requests present different ClientHello messages, and Cloudflare and Akamai inspect that at the edge before any HTTP logic runs. Fetch https://tls.browserleaks.com/json from inside your own script to see the ja3_hash, ja4 and akamai_hash your client actually emits. Switching to curl-cffi (Python) or got-scraping (Node) fixes most of these without touching your parsing code.
Apify's help centre says Apify and Crawlee's default configuration with best practices generally resolves it, and gives one concrete step: call session.retire() on a confirmed block so the retry gets a new session and IP (help.apify.com/en/articles/7029516). If the 403 survives a residential IP, the block is at the TLS or header layer, not the network layer, so change the HTTP client rather than the proxy. If it only appears after JavaScript runs, move to a patched browser.
No. Per Cloudflare's own documentation, Turnstile runs proof-of-work, proof-of-space, web-API probing and browser-quirk checks without any interaction, and the Managed widget decides whether to surface a checkbox based on visitor risk. The Non-interactive and Invisible widget types never show one. A clean browser session on a residential IP usually passes silently.
JA3 (2017) produces one hash of the TLS ClientHello. JA4, created by John Althouse at FoxIO, uses an a_b_c format that delimits the sections, so a defender can match on part of the fingerprint instead of the whole thing. That is why shuffling your extension order defeats a JA3 blocklist but not a JA4 prefix. The JA4 TLS-client method is BSD 3-Clause; the wider JA4+ suite is patent-pending under the FoxIO License 1.1.
Mostly not. As of 2026-09-09, Cloudflare documents its bot score range and its four detection engines publicly. DataDome publishes product documentation but nothing on signals or scoring. Akamai's Bot Manager docs sit behind a customer login, Imperva's have moved to Thales's documentation portal, and Kasada publishes no detection documentation. Any specific claim about those vendors' internals, including on this page, is reconstructed from observation.
Compare on the billing unit, not the sticker price. Apify Unblocker and Bright Data Web Unlocker both list at $1.50 per 1,000 successful requests (checked 2026-09-09). Residential proxies bill on traffic instead, so at $8/GB a 500 KB page costs $4.00 per 1,000 and a 2 MB browser session costs $16.00 per 1,000. Measure your page weight first. Unblocker has no session parameter, so if your job needs a logged-in multi-step flow, residential with sticky sessions is the only option regardless of cost.
Re-run one failing URL in a real browser and diff the HTML. Apify's academy documents empty results, deliberately fake results, redirects to the homepage and socket hangups as block types alongside the 403. If your success rate looks fine but your data quality dropped, assume a block before you assume schema drift.