Research & Data · Developer Tools

scrapling

d4vinci/scrapling

Its parser learns from website changes and automatically relocates your elements when pages update. Its fetchers bypass anti-bot systems like Cloudflare Turnstile out of the box.

Install

pip install scrapling
Author
@d4vinci
Category
Research & Data, Developer Tools
License
BSD-3-Clause
Updated
Oct 6, 2026

About scrapling

And its spider framework lets you scale up to concurrent, multi-session crawls with pause/resume, automatic proxy rotation, and a crawl speed that adapts to how fast each website responds and backs off when it starts blocking you - all in a few lines of Python. One library, zero compromises.

Scrapling is an adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl.

Features

  • Scrapy-like Spider API: Define spiders with start_urls, async parse callbacks, and Request/Response objects.

  • Concurrent Crawling: Configurable concurrency limits, per-domain throttling, and download delays.

  • Multi-Session Support: Unified interface for HTTP requests, and stealthy headless browsers in a single spider - route requests to different sessions by ID.

  • Pause & Resume: Checkpoint-based crawl persistence. Press Ctrl+C for a graceful shutdown; restart to resume from where you left off.

  • Streaming Mode: Stream scraped items as they arrive via async for item in spider.stream() with real-time stats - ideal for UI, pipelines, and long-running crawls.

  • Blocked Request Detection: Automatic detection and retry of blocked requests with customizable logic.

  • AutoThrottle: Stop guessing delays. The spider tunes the delay of each domain on its own from how fast the website responds, then doubles it (or waits what Retry-After asks) whenever the website…

  • Robots.txt Compliance: Optional robots_txt_obey flag that respects Disallow, Crawl-delay, and Request-rate directives with per-domain caching.

  • Development Mode: Cache responses to disk on the first run and replay them on subsequent runs - iterate on your parse() logic without re-hitting the target servers.

  • Ready-made Spider Templates: Skip the boilerplate with CrawlSpider for rule-based link following, SitemapSpider for sitemap/robots.txt-driven crawls, XMLFeedSpider/CSVFeedSpider for iterating XML/RSS…

Details on this page are taken from the project's README. Open README

Supported clients

Clients mentioned in this server's README:

View all
Cursor logo

Cursor

Desktop · Freemium · Proprietary

The first agentic IDE. The Cursor editor truly merges how developers and AI work together, delivering a magical coding experience.

WindowsMacOSLinux

Related MCP servers

More servers
Firecrawl MCP Server logo

Firecrawl MCP Server

mendableai/firecrawl-mcp-server

2.8k

A Model Context Protocol (MCP) server that brings Firecrawl to MCP-compatible AI agents — search, scrape, and interact with the live web for clean, agent-ready context.

Research & Data
Exa Web Search logo

Exa Web Search

exa-labs/exa-mcp-server

1.4k

Connect AI agents to Exa for web search, content fetching, and multi-step research.

Research & Data
ArXiv logo

ArXiv

blazickjp/arxiv-mcp-server

1k

A local MCP server for agent literature work. The differentiator is original-LaTeX section reads, BibTeX from arXiv metadata, and topic watches. Papers stay on disk.

Research & Data
sail logo

sail

lakehq/sail

730

Sail is a drop-in Apache Spark replacement written in Rust, unifying batch processing, stream processing, and compute-intensive AI workloads on a distributed, multimodal compute engine.

Research & Data

OpenAPI Proxy

janwilmake/openapi-mcp-server

554

A Model Context Protocol (MCP) server for Claude/Cursor that enables searching and exploring OpenAPI specifications through oapis.org.

Research & Data
Excel MCP Server logo

Excel MCP Server

haris-musa/excel-mcp-server

393

A Model Context Protocol server that lets AI assistants create, read and edit Excel workbooks. It needs no Microsoft Excel installation.

Research & Data