August 09, 2026 · By YasKad
D4Vinci/Scrapling

Scrapling: adaptive web scraping across HTTP, browser, and crawling

D4Vinci/Scrapling · 83,616★ · 8,546 forks

Everything you need to know about D4Vinci/Scrapling: a Python framework for scraping isolated pages or running concurrent crawls, with adaptive selectors, sessions, and browser options.


What Scrapling is

Scrapling is a web scraping framework for Python. Its README places it on a continuum from a single request to a full crawl: it combines a parser with CSS, XPath, and BeautifulSoup-like selectors; HTTP and browser clients; and a spider API with asynchronous requests and responses.

The distinguishing feature is adaptive mode. By saving a selector with auto_save=True, the project keeps information about the element; in a later run you can request adaptive=True to try to locate it if the page’s structure changed. It is an aid against layout changes, not a guarantee of correct extraction: the retrieved data still needs validation.

The project also offers clients with TLS fingerprint and header spoofing, Chromium or Chrome automation via Playwright, persistent sessions, proxy rotation, and block detection. The README claims that StealthyFetcher can solve Cloudflare Turnstile challenges and interstitials; that is a claim made by the project, not a certification that it will work against any protection.

Split-screen composition: on one side, a fast data stream piercing a digital wall representing a direct HTTP request; on the other, a stealthy browser interface with a holographic shield bypassing a security gate.

The origin: a first public release in October 2024

The GitHub API records the creation of D4Vinci/Scrapling on October 13, 2024, at 20:29 UTC. That same day, the history contains the initial commit and, a few hours later, commit a721073b, whose message describes the release of the first public version. The citation entry included by the repository itself attributes Scrapling to Karim Shoair in 2024.

The D4Vinci GitHub account identifies Shoair as a person from Egypt interested in computer science, information security, and web scraping; he had 45 public repositories and 3,558 followers at the time of the query. The history from that first day shows the initial technical context: compatibility fixes for Python 3.6 and 3.7, publishing the package and CI workflow, documentation, and, the following day, version 0.1.2. No external launch announcement with a broader narrative was retrieved, so no other personal motivation is attributed to the author.

Glowing digital timeline where a node labeled "Oct 2024" sparks to life a sequence of interconnected commits branching toward a wider network.

The tension the design addresses is explicit in the documentation: an HTML-parsing library alone does not solve dynamic pages, structural changes, sessions, rate limits, or anti-automation defenses. Scrapling brings those layers together under a common API while keeping an interface familiar to Scrapy, Parsel, and BeautifulSoup users.

Philosophy and principles

Four verifiable principles emerge from the documentation and contribution rules:

  • One library, multiple levels of work. You can use Selector just to parse HTML, Fetcher for HTTP, DynamicFetcher or StealthyFetcher for the browser, and Spider for crawls.
  • Resilience to change, with traceability. Adaptive selectors and similar-element search try to reduce fragility; the documentation does not present that recovery as infallible.
  • Scaling without giving up operational control. Spiders include concurrency, per-domain limits, pause and resume with checkpoints, automatic wait tuning, retries, and result exporting.
  • User responsibility. The README restricts use to educational and research purposes, requires compliance with privacy laws, site terms, and robots.txt. In spiders, robots_txt_obey is an explicit option that can respect Disallow, Crawl-delay, and Request-rate.

Futuristic control panel with four holographic panels representing an HTML parser, an HTTP client, a stealthy browser mask, and a spider network, all connected to a central core.

How it works

Installation and choosing a client

The parsing core installs with pip install scrapling. For clients, spiders, and browsers, the README requires pip install "scrapling[fetchers]" followed by scrapling install; for MCP, pip install "scrapling[ai]"; and for the full set, pip install "scrapling[all]". It also publishes pyd4vinci/scrapling images on Docker Hub and ghcr.io/d4vinci/scrapling:latest.

The documented flow breaks down as follows:

NeedDocumented interfacePattern example
HTML already availableSelectorSelector("<html>...</html>") → .css() or .xpath()
HTTP requestFetcher or FetcherSessionFetcher.get(url)
Dynamic browserDynamicFetcher or DynamicSessionDynamicFetcher.fetch(url)
Browser with stealth optionsStealthyFetcher or StealthySessionStealthyFetcher.fetch(url, headless=True)
CrawlingSpiderdefine start_urls and async def parse()

Sessions preserve cookies and state; spiders can register several sessions and route a request to one of them via sid. Crawl mode provides CrawlSpider, SitemapSpider, and ShopifySpider templates; it also includes link extraction, resumable checkpoints, progressive streaming with async for item in spider.stream(), and exporting to JSON, JSONL, CSV, or XML.

Adaptive selector, CLI, and AI

The core adaptation pattern is saving a selection, for example page.css('.product', auto_save=True), and retrieving it later with adaptive=True. Besides CSS and XPath, the parser offers text search, regular expressions, filters, and navigation between related elements.

AI scanner with a neon reticle locking onto a glowing element inside an HTML code wall that is reorganizing itself.

The CLI avoids writing code for a simple extraction:

scrapling shell
scrapling extract get 'https://example.com' content.md
scrapling extract fetch 'https://example.com' content.md --no-headless
scrapling extract stealthy-fetch 'https://example.com' captchas.html --solve-cloudflare

The README documents a built-in MCP server, activated with the ai extra, that prepares selected content before handing it to an assistant, keeps browser sessions alive across calls, captures screenshots, and can control a remote browser via CDP. There is also an agent skill inside agent-skill/ so that coding agents can consult the current Scrapling API.

Futuristic terminal with green neon text running scrapling extract, with holographic panels showing integration with an AI assistant and an MCP server.

Performance: limits of the comparison

The repository publishes its own benchmarks.py. In its text-extraction test over 5,000 nested elements, it reports 1.98 ms for Scrapling, 1.99 ms for Parsel/Scrapy, 2.48 ms for lxml, and higher figures for PyQuery, Selectolax, MechanicalSoup, and BeautifulSoup. In the similarity test it reports 2.29 ms for Scrapling and 12.46 ms for AutoScraper. These are averages the project attributes to more than 100 runs; they are not an independent evaluation and do not substitute for a test under the load and policies of the target site.

Two holographic speedometers racing on a dark stage; one bearing a Python icon dominates the race, leaving the others behind, above blocks of nested data.

Official and semi-official status

The scrapling package is officially distributed via PyPI according to the README’s links, and the images the project says it generates on each release are offered from Docker Hub and the GitHub Container Registry. The same README links a skill maintained in the repository itself and the D4Vinci/scrapling-official profile on Clawhub.

This establishes distribution channels and a skill presented by the maintainer; it does not establish that Python, Anthropic, OpenAI, Cloudflare, Playwright, or any other vendor endorses Scrapling’s effectiveness against anti-bot systems. No evidence was retrieved that it has been accepted as an official extension of Claude Code, Cursor, or another vendor’s marketplace. Adoption in third-party MCP projects and skills is, therefore, semi-official or community-driven, not a vendor certification.

The ecosystem

Materials and translations maintained by the project

The README links documentation in Arabic, Spanish, Brazilian Portuguese, French, German, Simplified Chinese, Japanese, Russian, and Korean. These are translations hosted in the docs/ tree of the main repository; they are not separate repositories. It also includes agent-skill/, documentation on Read the Docs, the MCP server, a Docker image, and a project Discord.

As directly related material from the same author, a search of D4Vinci’s repositories returned D4Vinci/Scrapling-Arabic-Crash-Course (6 stars): described as the files for an Arabic-language course on Scrapling. That search did not turn up an evaluation lab or a sibling marketplace dedicated to Scrapling.

Glowing digital globe in a dark space, covered in interconnected nodes representing documentation in multiple languages and community forks.

Ports, forks, and community extensions

The GitHub search and the retrieved list of forks allowed verification of the following projects. The figures are GitHub API star counts as of the measurement date, not a guarantee of maintenance, compatibility, or authorization from the author.

  • dorisoy/Scrapling — a fork translated into Chinese, with 17 stars. Its description presents the adaptive selector, anti-bot clients, and spiders in Chinese; it is the most clearly identifiable non-English community translation among the retrieved results.
  • Cedriccmh/claude-code-skill-scrapling — a skill for Claude Code that claims to automatically choose a Scrapling client and document site patterns; 377 stars.
  • cyberchitta/scrapling-fetch-mcp — a third-party MCP server letting assistants access text from protected sites using Scrapling; 106 stars.
  • sangamsharma/Scrapling-openclaw — a fork or adaptation for OpenClaw; 2 stars. Its low star count and nearly identical description to the original do not support a claim of sustained independent support.
  • Among the highest-starred forks in the API response are smirk-dev/Scrapling (6), ahmed-el-halawani/scrapling (5), and dmore/Scrapling-red-stealth-python (4). The API flags them as derivatives or presents them with a Scrapling description; not enough documentation was retrieved to treat them as maintained alternatives.

The ecosystem also connects with projects the documentation names as API or benchmark references: Scrapy, Parsel, BeautifulSoup, lxml, and Playwright. These names establish compatibility, inspiration, or technical comparison in the retrieved sources; they do not imply corporate affiliation.

Repo numbers

Measured: August 3, 2026, GitHub API.

MetricValue
Stars72,226
Forks7,174
Subscribers263
Commits1,560
Open issues reported by the API6
Primary languagePython
LicenseBSD-3-Clause
CreatedOctober 13, 2024
Latest code pushJuly 30, 2026
Latest metadata updateAugust 3, 2026
Latest releasev0.4.12, July 26, 2026

The total of 1,560 commits comes from the final pagination link of the commits API. The top contributors returned by the API are D4Vinci (Karim Shoair, 1,490 contributions), yetval (14), AbdullahY36 (10), mhillebrand (4), haosenwang1018 (4), and Bortlesboat (4). GitHub duplicates the star total in watchers_count; that is why subscribers_count is reported here as the real subscriber figure. The open_issues_count field may include open pull requests, so it does not necessarily equal issues alone.

How to contribute

The CONTRIBUTING.md guide requires forking the repository, cloning the fork, and working from dev; a pull request against main is rejected. The documented initial flow is:

git clone https://github.com/<username>/Scrapling.git
cd Scrapling
git checkout dev
python -m venv .venv
pip install -e ".[all]"
pip install -r tests/requirements.txt
scrapling install
pre-commit install

Feature contributions must include tests; fixes must include code that reproduces the bug. The guide uses tox and GitHub CI across supported Python versions, plus mypy, pyright, ruff, bandit, vermin, and conventional commit messages. It instructs running pytest tests -n auto; to avoid browser conflicts, it separates tests unrelated to DynamicFetcher or StealthyFetcher from sequential browser tests. It also welcomes translations and spider templates for platforms with a uniform structure across many domains, but not single-site scrapers.

How the community received it

Verifiable external reception is limited on Hacker News, but GitHub threads offer concrete examples of use, problems, and fixes:

  • On Hacker News, 41832425 was posted by d4vinci on October 13, 2024, and reached 4 points and 1 comment. That single comment is from the author himself and presents Scrapling as an adaptive, high-performance library; it is not an independent review. The second submission, 43852100, posted by d4vinci on April 30, 2025, got 1 point and 0 comments. No broad Hacker News discussion was found from which to infer consensus.
  • In issue #50, restlessronin thanked the maintainers for the work and explained a practical use case: retrieving OpenAI documentation he had been unable to fetch otherwise. His concrete objection was that messages sent to standard output interfered with an stdio-based MCP server. He ultimately said he would keep a stdout redirect in place while the integration was resolved; it is praise from a user, but also evidence of a real protocol friction.
  • In #215, nuclei-masta reported that, with ProxyRotator, the browser would not open and the spider ended with zero elements. yetval first identified a method-resolution-order conflict and then a leak in the proxy rotator’s page pool; he opened pull request #223. After the merge, nuclei-masta replied that it worked fine. The thread illustrates a collaborative fix, not that the combination was bug-free before the fix.
  • In #366, chcodex objected that the MCP server did not sanitize control characters before serializing a response and that recommending the local CLI did not solve a remote MCP deployment. yetval proposed a fix; chcodex verified it went from an XML compatibility error to an HTTP 200 response, and the maintainer reported his integration. It is an improvement confirmed by the reporter, though the issue reveals that MCP integrations need to be tested with real content.

Scrapling versus other proposals

ProposalVerifiable overlapVerifiable difference
ScrapyThe Spider API is documented as similar to Scrapy, and Scrapling offers integration via scrapling_response.Scrapling integrates HTTP clients, browser clients, adaptive selectors, and sessions in the same README; the source does not support concluding general superiority.
ParselBoth appear in the parser performance comparison and use CSS/XPath-related selectors.Scrapling states it adapted Parsel code for its translation submodule and adds clients, spiders, and element adaptation.
BeautifulSoupScrapling’s parser offers find_all in a form similar to BeautifulSoup.The retrieved source describes XPath, node navigation, sessions, spiders, and browsers in Scrapling, capabilities outside that similar parsing interface.
AutoScraperBoth appear in the internal element-similarity benchmark.The README compares adaptive-location time; we did not retrieve a source that would let us equate their architectures or results beyond that project-run test.
PlaywrightDynamicFetcher uses Playwright for Chromium and Chrome, and sessions support remote CDP.Playwright appears as the automation engine inside Scrapling; it is not a peer-level competitor in the retrieved documentation.

Use cases and who this repository can help

  • Anyone who needs to scrape pages that change frequently can save selectors and request adaptive recovery, always validating the result to avoid mistaking a similar element for the expected data.
  • Teams alternating between static sites, dynamic applications, and stateful pages can start with Fetcher, move up to DynamicFetcher or StealthyFetcher, and keep cookies and state with sessions, without switching parsing libraries.
  • Long-running crawl operations can use Spider, per-domain limits, checkpoints, pause and resume, progressive streaming, and built-in exporters; development mode also allows reusing on-disk responses while iterating on parse().
  • Developers building assistants or MCP automation can use the MCP server and agent skill to deliver selected content and maintain sessions, but should test the stdio transport, control characters, and site policy, as issues #50 and #366 show.
  • Maintainers already using Scrapy can evaluate the scrapling_response decorator to parse responses with Scrapling’s parser without rewriting their whole spider.

Holographic GitHub repository interface with neon metrics orbiting a central core: a "72.2K" star counter and data streams representing 1,560 commits.

Resources


Note: this article combines the README, contribution guide, and history of D4Vinci/Scrapling, the GitHub API, and the Hacker News API, retrieved on August 3, 2026. Figures and link status change over time.

Comments