Scrapling: adaptive web scraping across HTTP, browser, and crawling
D4Vinci/Scrapling · 83,616★ · 8,546 forks
Everything you need to know about D4Vinci/Scrapling: a Python framework for scraping isolated pages or running concurrent crawls, with adaptive selectors, sessions, and browser options.
What Scrapling is
Scrapling is a web scraping framework for Python. Its README places it on a continuum from a single request to a full crawl: it combines a parser with CSS, XPath, and BeautifulSoup-like selectors; HTTP and browser clients; and a spider API with asynchronous requests and responses.
The distinguishing feature is adaptive mode. By saving a selector with auto_save=True, the project keeps information about the element; in a later run you can request adaptive=True to try to locate it if the page’s structure changed. It is an aid against layout changes, not a guarantee of correct extraction: the retrieved data still needs validation.
The project also offers clients with TLS fingerprint and header spoofing, Chromium or Chrome automation via Playwright, persistent sessions, proxy rotation, and block detection. The README claims that StealthyFetcher can solve Cloudflare Turnstile challenges and interstitials; that is a claim made by the project, not a certification that it will work against any protection.

The origin: a first public release in October 2024
The GitHub API records the creation of D4Vinci/Scrapling on October 13, 2024, at 20:29 UTC. That same day, the history contains the initial commit and, a few hours later, commit a721073b, whose message describes the release of the first public version. The citation entry included by the repository itself attributes Scrapling to Karim Shoair in 2024.
The D4Vinci GitHub account identifies Shoair as a person from Egypt interested in computer science, information security, and web scraping; he had 45 public repositories and 3,558 followers at the time of the query. The history from that first day shows the initial technical context: compatibility fixes for Python 3.6 and 3.7, publishing the package and CI workflow, documentation, and, the following day, version 0.1.2. No external launch announcement with a broader narrative was retrieved, so no other personal motivation is attributed to the author.

The tension the design addresses is explicit in the documentation: an HTML-parsing library alone does not solve dynamic pages, structural changes, sessions, rate limits, or anti-automation defenses. Scrapling brings those layers together under a common API while keeping an interface familiar to Scrapy, Parsel, and BeautifulSoup users.
Philosophy and principles
Four verifiable principles emerge from the documentation and contribution rules:
- One library, multiple levels of work. You can use
Selectorjust to parse HTML,Fetcherfor HTTP,DynamicFetcherorStealthyFetcherfor the browser, andSpiderfor crawls. - Resilience to change, with traceability. Adaptive selectors and similar-element search try to reduce fragility; the documentation does not present that recovery as infallible.
- Scaling without giving up operational control. Spiders include concurrency, per-domain limits, pause and resume with checkpoints, automatic wait tuning, retries, and result exporting.
- User responsibility. The README restricts use to educational and research purposes, requires compliance with privacy laws, site terms, and
robots.txt. In spiders,robots_txt_obeyis an explicit option that can respectDisallow,Crawl-delay, andRequest-rate.

How it works
Installation and choosing a client
The parsing core installs with pip install scrapling. For clients, spiders, and browsers, the README requires pip install "scrapling[fetchers]" followed by scrapling install; for MCP, pip install "scrapling[ai]"; and for the full set, pip install "scrapling[all]". It also publishes pyd4vinci/scrapling images on Docker Hub and ghcr.io/d4vinci/scrapling:latest.
The documented flow breaks down as follows:
| Need | Documented interface | Pattern example |
|---|---|---|
| HTML already available | Selector | Selector("<html>...</html>") → .css() or .xpath() |
| HTTP request | Fetcher or FetcherSession | Fetcher.get(url) |
| Dynamic browser | DynamicFetcher or DynamicSession | DynamicFetcher.fetch(url) |
| Browser with stealth options | StealthyFetcher or StealthySession | StealthyFetcher.fetch(url, headless=True) |
| Crawling | Spider | define start_urls and async def parse() |
Sessions preserve cookies and state; spiders can register several sessions and route a request to one of them via sid. Crawl mode provides CrawlSpider, SitemapSpider, and ShopifySpider templates; it also includes link extraction, resumable checkpoints, progressive streaming with async for item in spider.stream(), and exporting to JSON, JSONL, CSV, or XML.
Adaptive selector, CLI, and AI
The core adaptation pattern is saving a selection, for example page.css('.product', auto_save=True), and retrieving it later with adaptive=True. Besides CSS and XPath, the parser offers text search, regular expressions, filters, and navigation between related elements.

The CLI avoids writing code for a simple extraction:
scrapling shell
scrapling extract get 'https://example.com' content.md
scrapling extract fetch 'https://example.com' content.md --no-headless
scrapling extract stealthy-fetch 'https://example.com' captchas.html --solve-cloudflare
The README documents a built-in MCP server, activated with the ai extra, that prepares selected content before handing it to an assistant, keeps browser sessions alive across calls, captures screenshots, and can control a remote browser via CDP. There is also an agent skill inside agent-skill/ so that coding agents can consult the current Scrapling API.

Performance: limits of the comparison
The repository publishes its own benchmarks.py. In its text-extraction test over 5,000 nested elements, it reports 1.98 ms for Scrapling, 1.99 ms for Parsel/Scrapy, 2.48 ms for lxml, and higher figures for PyQuery, Selectolax, MechanicalSoup, and BeautifulSoup. In the similarity test it reports 2.29 ms for Scrapling and 12.46 ms for AutoScraper. These are averages the project attributes to more than 100 runs; they are not an independent evaluation and do not substitute for a test under the load and policies of the target site.

Official and semi-official status
The scrapling package is officially distributed via PyPI according to the README’s links, and the images the project says it generates on each release are offered from Docker Hub and the GitHub Container Registry. The same README links a skill maintained in the repository itself and the D4Vinci/scrapling-official profile on Clawhub.
This establishes distribution channels and a skill presented by the maintainer; it does not establish that Python, Anthropic, OpenAI, Cloudflare, Playwright, or any other vendor endorses Scrapling’s effectiveness against anti-bot systems. No evidence was retrieved that it has been accepted as an official extension of Claude Code, Cursor, or another vendor’s marketplace. Adoption in third-party MCP projects and skills is, therefore, semi-official or community-driven, not a vendor certification.
The ecosystem
Materials and translations maintained by the project
The README links documentation in Arabic, Spanish, Brazilian Portuguese, French, German, Simplified Chinese, Japanese, Russian, and Korean. These are translations hosted in the docs/ tree of the main repository; they are not separate repositories. It also includes agent-skill/, documentation on Read the Docs, the MCP server, a Docker image, and a project Discord.
As directly related material from the same author, a search of D4Vinci’s repositories returned D4Vinci/Scrapling-Arabic-Crash-Course (6 stars): described as the files for an Arabic-language course on Scrapling. That search did not turn up an evaluation lab or a sibling marketplace dedicated to Scrapling.

Ports, forks, and community extensions
The GitHub search and the retrieved list of forks allowed verification of the following projects. The figures are GitHub API star counts as of the measurement date, not a guarantee of maintenance, compatibility, or authorization from the author.
dorisoy/Scrapling— a fork translated into Chinese, with 17 stars. Its description presents the adaptive selector, anti-bot clients, and spiders in Chinese; it is the most clearly identifiable non-English community translation among the retrieved results.Cedriccmh/claude-code-skill-scrapling— a skill for Claude Code that claims to automatically choose a Scrapling client and document site patterns; 377 stars.cyberchitta/scrapling-fetch-mcp— a third-party MCP server letting assistants access text from protected sites using Scrapling; 106 stars.sangamsharma/Scrapling-openclaw— a fork or adaptation for OpenClaw; 2 stars. Its low star count and nearly identical description to the original do not support a claim of sustained independent support.- Among the highest-starred forks in the API response are
smirk-dev/Scrapling(6),ahmed-el-halawani/scrapling(5), anddmore/Scrapling-red-stealth-python(4). The API flags them as derivatives or presents them with a Scrapling description; not enough documentation was retrieved to treat them as maintained alternatives.
The ecosystem also connects with projects the documentation names as API or benchmark references: Scrapy, Parsel, BeautifulSoup, lxml, and Playwright. These names establish compatibility, inspiration, or technical comparison in the retrieved sources; they do not imply corporate affiliation.
Repo numbers
Measured: August 3, 2026, GitHub API.
| Metric | Value |
|---|---|
| Stars | 72,226 |
| Forks | 7,174 |
| Subscribers | 263 |
| Commits | 1,560 |
| Open issues reported by the API | 6 |
| Primary language | Python |
| License | BSD-3-Clause |
| Created | October 13, 2024 |
| Latest code push | July 30, 2026 |
| Latest metadata update | August 3, 2026 |
| Latest release | v0.4.12, July 26, 2026 |
The total of 1,560 commits comes from the final pagination link of the commits API. The top contributors returned by the API are D4Vinci (Karim Shoair, 1,490 contributions), yetval (14), AbdullahY36 (10), mhillebrand (4), haosenwang1018 (4), and Bortlesboat (4). GitHub duplicates the star total in watchers_count; that is why subscribers_count is reported here as the real subscriber figure. The open_issues_count field may include open pull requests, so it does not necessarily equal issues alone.
How to contribute
The CONTRIBUTING.md guide requires forking the repository, cloning the fork, and working from dev; a pull request against main is rejected. The documented initial flow is:
git clone https://github.com/<username>/Scrapling.git
cd Scrapling
git checkout dev
python -m venv .venv
pip install -e ".[all]"
pip install -r tests/requirements.txt
scrapling install
pre-commit install
Feature contributions must include tests; fixes must include code that reproduces the bug. The guide uses tox and GitHub CI across supported Python versions, plus mypy, pyright, ruff, bandit, vermin, and conventional commit messages. It instructs running pytest tests -n auto; to avoid browser conflicts, it separates tests unrelated to DynamicFetcher or StealthyFetcher from sequential browser tests. It also welcomes translations and spider templates for platforms with a uniform structure across many domains, but not single-site scrapers.
How the community received it
Verifiable external reception is limited on Hacker News, but GitHub threads offer concrete examples of use, problems, and fixes:
- On Hacker News, 41832425 was posted by
d4vincion October 13, 2024, and reached 4 points and 1 comment. That single comment is from the author himself and presents Scrapling as an adaptive, high-performance library; it is not an independent review. The second submission, 43852100, posted byd4vincion April 30, 2025, got 1 point and 0 comments. No broad Hacker News discussion was found from which to infer consensus. - In issue #50,
restlessroninthanked the maintainers for the work and explained a practical use case: retrieving OpenAI documentation he had been unable to fetch otherwise. His concrete objection was that messages sent to standard output interfered with anstdio-based MCP server. He ultimately said he would keep astdoutredirect in place while the integration was resolved; it is praise from a user, but also evidence of a real protocol friction. - In #215,
nuclei-mastareported that, withProxyRotator, the browser would not open and the spider ended with zero elements.yetvalfirst identified a method-resolution-order conflict and then a leak in the proxy rotator’s page pool; he opened pull request #223. After the merge,nuclei-mastareplied that it worked fine. The thread illustrates a collaborative fix, not that the combination was bug-free before the fix. - In #366,
chcodexobjected that the MCP server did not sanitize control characters before serializing a response and that recommending the local CLI did not solve a remote MCP deployment.yetvalproposed a fix;chcodexverified it went from an XML compatibility error to an HTTP 200 response, and the maintainer reported his integration. It is an improvement confirmed by the reporter, though the issue reveals that MCP integrations need to be tested with real content.
Scrapling versus other proposals
| Proposal | Verifiable overlap | Verifiable difference |
|---|---|---|
| Scrapy | The Spider API is documented as similar to Scrapy, and Scrapling offers integration via scrapling_response. | Scrapling integrates HTTP clients, browser clients, adaptive selectors, and sessions in the same README; the source does not support concluding general superiority. |
| Parsel | Both appear in the parser performance comparison and use CSS/XPath-related selectors. | Scrapling states it adapted Parsel code for its translation submodule and adds clients, spiders, and element adaptation. |
| BeautifulSoup | Scrapling’s parser offers find_all in a form similar to BeautifulSoup. | The retrieved source describes XPath, node navigation, sessions, spiders, and browsers in Scrapling, capabilities outside that similar parsing interface. |
| AutoScraper | Both appear in the internal element-similarity benchmark. | The README compares adaptive-location time; we did not retrieve a source that would let us equate their architectures or results beyond that project-run test. |
| Playwright | DynamicFetcher uses Playwright for Chromium and Chrome, and sessions support remote CDP. | Playwright appears as the automation engine inside Scrapling; it is not a peer-level competitor in the retrieved documentation. |
Use cases and who this repository can help
- Anyone who needs to scrape pages that change frequently can save selectors and request adaptive recovery, always validating the result to avoid mistaking a similar element for the expected data.
- Teams alternating between static sites, dynamic applications, and stateful pages can start with
Fetcher, move up toDynamicFetcherorStealthyFetcher, and keep cookies and state with sessions, without switching parsing libraries. - Long-running crawl operations can use
Spider, per-domain limits, checkpoints, pause and resume, progressive streaming, and built-in exporters; development mode also allows reusing on-disk responses while iterating onparse(). - Developers building assistants or MCP automation can use the MCP server and agent skill to deliver selected content and maintain sessions, but should test the
stdiotransport, control characters, and site policy, as issues #50 and #366 show. - Maintainers already using Scrapy can evaluate the
scrapling_responsedecorator to parse responses with Scrapling’s parser without rewriting their whole spider.

Resources
- Repository: https://github.com/D4Vinci/Scrapling
- Documentation and installation: https://scrapling.readthedocs.io/en/latest/
- Contribution and testing: https://github.com/D4Vinci/Scrapling/blob/main/CONTRIBUTING.md
- Official skills: https://github.com/D4Vinci/Scrapling/tree/main/agent-skill
- MCP server: https://scrapling.readthedocs.io/en/latest/ai/mcp-server.html
- Official package: https://pypi.org/project/scrapling/
- Docker image: https://hub.docker.com/r/pyd4vinci/scrapling
- Reviews and conversations: https://news.ycombinator.com/item?id=41832425, https://news.ycombinator.com/item?id=43852100, https://github.com/D4Vinci/Scrapling/issues/50, https://github.com/D4Vinci/Scrapling/issues/215, https://github.com/D4Vinci/Scrapling/issues/366
- Discord community: https://discord.gg/EMgGbDceNQ
Note: this article combines the README, contribution guide, and history of D4Vinci/Scrapling, the GitHub API, and the Hacker News API, retrieved on August 3, 2026. Figures and link status change over time.
Comments