
Cursor
Desktop · Freemium · Proprietary
The first agentic IDE. The Cursor editor truly merges how developers and AI work together, delivering a magical coding experience.
Research & Data · Developer Tools
d4vinci/scrapling
Its parser learns from website changes and automatically relocates your elements when pages update. Its fetchers bypass anti-bot systems like Cloudflare Turnstile out of the box.
pip install scraplingAnd its spider framework lets you scale up to concurrent, multi-session crawls with pause/resume, automatic proxy rotation, and a crawl speed that adapts to how fast each website responds and backs off when it starts blocking you - all in a few lines of Python. One library, zero compromises.
Scrapling is an adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl.
Scrapy-like Spider API: Define spiders with start_urls, async parse callbacks, and Request/Response objects.
Concurrent Crawling: Configurable concurrency limits, per-domain throttling, and download delays.
Multi-Session Support: Unified interface for HTTP requests, and stealthy headless browsers in a single spider - route requests to different sessions by ID.
Pause & Resume: Checkpoint-based crawl persistence. Press Ctrl+C for a graceful shutdown; restart to resume from where you left off.
Streaming Mode: Stream scraped items as they arrive via async for item in spider.stream() with real-time stats - ideal for UI, pipelines, and long-running crawls.
Blocked Request Detection: Automatic detection and retry of blocked requests with customizable logic.
AutoThrottle: Stop guessing delays. The spider tunes the delay of each domain on its own from how fast the website responds, then doubles it (or waits what Retry-After asks) whenever the website…
Robots.txt Compliance: Optional robots_txt_obey flag that respects Disallow, Crawl-delay, and Request-rate directives with per-domain caching.
Development Mode: Cache responses to disk on the first run and replay them on subsequent runs - iterate on your parse() logic without re-hitting the target servers.
Ready-made Spider Templates: Skip the boilerplate with CrawlSpider for rule-based link following, SitemapSpider for sitemap/robots.txt-driven crawls, XMLFeedSpider/CSVFeedSpider for iterating XML/RSS…
Details on this page are taken from the project's README. Open README
Clients mentioned in this server's README:

mendableai/firecrawl-mcp-server
A Model Context Protocol (MCP) server that brings Firecrawl to MCP-compatible AI agents — search, scrape, and interact with the live web for clean, agent-ready context.

exa-labs/exa-mcp-server
Connect AI agents to Exa for web search, content fetching, and multi-step research.

blazickjp/arxiv-mcp-server
A local MCP server for agent literature work. The differentiator is original-LaTeX section reads, BibTeX from arXiv metadata, and topic watches. Papers stay on disk.

lakehq/sail
Sail is a drop-in Apache Spark replacement written in Rust, unifying batch processing, stream processing, and compute-intensive AI workloads on a distributed, multimodal compute engine.
janwilmake/openapi-mcp-server
A Model Context Protocol (MCP) server for Claude/Cursor that enables searching and exploring OpenAPI specifications through oapis.org.

haris-musa/excel-mcp-server
A Model Context Protocol server that lets AI assistants create, read and edit Excel workbooks. It needs no Microsoft Excel installation.