Eight open-source projects collect web pages, each at a different point in the pipeline. Crawlee and Crawlee for Python are frameworks for writing a crawler, over plain HTTP or with a headless browser. Camoufox, a modified Firefox built to get past bot detection, and Rota, which rotates proxies behind a single port, add to a crawler rather than replace it. OpenSERP returns results from six search engines as JSON, changedetection.io and PriceGhost watch the pages you point them at with no code to write, and Agent Reach installs the tools an AI agent needs to read YouTube, Reddit or X. In their open-source form, none of them supplies outgoing IP addresses: getting past blocks means bringing your own proxies.
Crawlee for Node.js or Python: where the two editions part ways
Apify, which sells a cloud scraping platform, publishes Crawlee in two repositories, both under Apache-2.0. The two editions share one model: a request queue, a handler called for each page, a dataset for the results, and the same API for HTTP crawlers, which parse the HTML without running the page's JavaScript, and for browser crawlers. Both ship an AdaptivePlaywrightCrawler that falls back to plain HTTP on pages that don't need the browser. Browser crawlers generate fingerprints by default, and the default HTTP client passes itself off as a browser: got-scraping in Node.js, and impit, a Rust library from Apify, in Python.
Only the Node.js edition drives Puppeteer as well as Playwright. The Python edition adds SQL storage (SQLite, PostgreSQL, MySQL, MariaDB) and Redis, both still experimental, which Crawlee 3.18 for Node.js doesn't have. Its default file storage isn't safe with more than one process, and it is wiped on every run unless purge_on_start is set to False. Python also keeps tiered proxies (tiered_proxy_urls), ranked from cheapest to most reliable, which version 4 of the Node.js edition, a release candidate since 13 August 2026, replaces with per-session rotation.
Crawlee for Python pitches itself against Scrapy: its README makes the case (asyncio, type hints, scripts that run without a dedicated launcher, persisted state), and the docs include a migration guide. The Node.js edition needs Node.js 16 for v3, and Node.js 22.13 with native ESM for v4. Neither has a user interface; both are libraries for developers.
Camoufox and Rota: the pieces you add against blocking
Crawlee's anti-blocking guide concedes that its built-in protections aren't enough against the Cloudflare challenge and points to Camoufox, through camoufox-js, a JavaScript port that Apify maintains and that its repository labels experimental. Crawlee for Python ships a Camoufox example and a project template. Camoufox rewrites the values anti-bot scripts read inside Firefox's C++ code instead of injecting JavaScript, hides Playwright from the page, and draws its identities from a Bayesian network trained on real traffic. Its README sets two limits. Camoufox is still a Firefox: it can't fully pass for Chromium, and some WAFs probe SpiderMonkey, Firefox's JavaScript engine, which the README calls impossible to spoof. And out of the thousands of values that have to agree, Camoufox doesn't always get every one of them right.
Rota is the only network-layer piece in the set: a Go proxy server that takes clients on a single HTTP port, 8000, and sends each request out through an upstream HTTP, HTTPS or SOCKS proxy it has already tested and geolocated. Crawlee already rotates proxies inside its own process; Rota moves that rotation out of the program, so several scrapers can share the same pool and the same health checks. Its pools by country, city or ISP pair with Camoufox's geoip option, which matches timezone, language and location to the exit IP. Camoufox's README expects the browser to run behind rotating proxies, preferably residential ones, which the project doesn't supply, and Rota's database starts out empty.
Ready-made tools: search results, page monitoring, AI agents
OpenSERP starts from a query and gives back addresses. A Go server and CLI open the results page of Google, Bing, Yandex, Baidu, DuckDuckGo or Ecosia in a headless Chromium and return it in a single JSON schema; /mega/search queries several engines at once and records each URL's rank on each of them. The README pitches it as a search tool for LLMs or the backend of an SEO rank tracker.
changedetection.io and PriceGhost work the other way round: you give them URLs, they keep revisiting them and alert you when something changes. changedetection.io compares text, JSON, a price, stock levels or a screenshot, every three hours by default over plain HTTP, and sends alerts through the Apprise library to Discord, email or a webhook. PriceGhost sticks to product pages, one URL per item, with dedicated scrapers aimed mostly at US retailers, several accounts per instance, and alerts to Telegram, Discord, Pushover, ntfy or Gotify. Its README sets it against Keepa, CamelCamelCamel and Honey.
Agent Reach doesn't read anything itself. The Python CLI sets up a coding agent (Claude Code, Cursor, OpenClaw) by picking, installing and testing one tool per platform: yt-dlp for YouTube, gh for GitHub, Jina Reader for web pages. The web channel goes through the hosted r.jina.ai API and stops at a Cloudflare challenge, and search runs through Exa's remote MCP server. Agent Reach has no anti-detection technique of its own and doesn't scrape any search engine.
What each tool needs to run
Read from the repositories and documentation on 7 October 2026.
| Tool | Licence | What you run | Browser | Proxies | Latest release |
|---|---|---|---|---|---|
| Crawlee | Apache-2.0 | Node.js 16 (v3) or 22.13 (v4); local storage, no database | Playwright or Puppeteer, installed separately | bring your own; rotation and sessions built in | 3.18.2, 29 Sep 2026; 4.0.0-rc.1 on 6 Oct |
| Crawlee for Python | Apache-2.0 | Python 3.10; local files, SQL or Redis optional | Playwright, optional | bring your own; plain or tiered rotation | 1.10.4, 6 Oct 2026 |
| Camoufox | MPL-2.0; launchers under MIT | Python 3.10 or Node.js 22.15, Playwright; a download of about 1.3 GB | it is the browser: a modified Firefox 156 | bring your own; built to run behind residential proxies | v156.0.1-beta.36, 6 Oct 2026 |
| Rota | Apache-2.0 | Caddy, a Go service, a Next.js dashboard, TimescaleDB; TimescaleDB even without Docker | none | rotating them is its job; none included | 2.3.0, 20 Sep 2026 |
| OpenSERP | MIT | a Go binary and Chromium; 2 GB of /dev/shm with Compose | Chromium by default, the only mode for Bing and DuckDuckGo | bring your own; "usually required" on a VPS | 0.8.12, 22 Jul 2026 |
| changedetection.io | Apache-2.0 | one Python 3.11 container, no database | Chrome optional, required for screenshots and Browser Steps | one per watch, bring your own | 0.60.8, 28 Sep 2026 |
| PriceGhost | "MIT" in the README, no licence file | PostgreSQL, a Node.js 20 backend, a React frontend | bundled Chromium with the stealth plugin | not covered in the docs | 1.0.6, 26 Jan 2026; no GitHub release |
| Agent Reach | MIT | Python 3.10, Node.js, gh, mcporter, one tool per platform | desktop Chrome for several channels | residential advised on a server | 1.5.0, 11 Jun 2026 |
The headless browser is the heaviest part. Each Camoufox v156.0.1-beta.36 archive weighs about 1.3 GB, against some 650 MB on Linux for v152; the project claims about 200 MB of memory against 800 MB or more for Chrome, a figure it states without a published measurement. OpenSERP runs up to six Chrome processes in parallel by default, and each authenticated proxy gets its own; its Compose file reserves 2 GB of /dev/shm, without which Chrome runs out of memory. changedetection.io does without a browser by default, but its visual selector, Browser Steps and screenshot comparison all need one. Crawlee for Python allows itself 25% of available memory by default when it scales its concurrency.
The IP address is the first cause of blocking that Crawlee's anti-blocking guide names. OpenSERP runs straight into it: on a VPS, its maintainer writes, datacenter IPs may be blacklisted and proxies are "usually required", and an issue opened in July 2026 reports a Google CAPTCHA on the very first request. Google also swaps organic links for encrypted tokens when it judges the client automated; a fix was merged on 22 September 2026 but no release includes it yet. PriceGhost handles no proxies and only spaces out its checks. changedetection.io takes one proxy per watch and recommends Bright Data through a referral link, while Agent Reach advises a residential proxy on a server, at about a dollar a month according to its README.
What to lock down before exposing an instance
The Compose files of three projects leave services open. Rota publishes port 8000 on the host with authentication turned off: any request goes through until a proxy account exists, which makes it an open proxy. OpenSERP's API declares no authentication, accepts any CORS origin and, since v0.8.12, lets a caller impose their own proxy; its Compose file publishes port 7000 on every interface, and the release notes advise setting proxies.allow_request_proxy_url back to false on an exposed instance. PriceGhost publishes PostgreSQL on port 5432 with the password postgres and, without a JWT_SECRET, signs sessions with a value written in the file, which is enough to forge a token according to an issue opened in September 2026. changedetection.io guards the whole instance with a single password, off by default.
What the projects say about the target sites' terms
The changedetection.io README is the most explicit: its Disclaimer section leaves it to the user to respect each site's terms of service, its robots.txt and the law. Both Crawlee editions can skip URLs that robots.txt disallows (respectRobotsTxtFile in Node.js, respect_robots_txt_file in Python), but the option is off by default. In Python, the ThrottlingRequestManager also enforces the file's crawl-delay, and the Node.js docs link from their parallelization guide to Apify's post on ethical web scraping.
Agent Reach talks about accounts rather than sites: its README warns that platforms can detect scripted calls and ban the account, recommends using a secondary one, and says nothing about the platforms' terms. PriceGhost only deals with the risk of getting blocked. Camoufox, OpenSERP and Rota don't raise the subject in their READMEs or their documentation.
Where the projects stand
Crawlee for Python ships close to one release a week: 35 versions in the 1.x line since 1.0.0 on 29 September 2025. Crawlee for Node.js keeps maintaining its 3.x branch (3.18.2 on 29 September 2026) while v4, a release candidate, warns that "there are many" breaking changes. changedetection.io has stayed at 0.x since 2021 but releases often: 67 versions in 2025, 40 since January 2026, and its maintainer wrote 397 of the 514 commits on the main branch in 2026.
Camoufox has never left beta, and its README warns it may not suit stable production use. No release came out between March 2025 and January 2026, and until late August 2026 the README placed active development on a Clover Labs fork; the main repository now publishes a pre-release for each tested merge, five builds between 28 September and 6 October 2026. Rota has shipped five releases since July 2026, after nearly eight months without one, and its 2.1.0 owes its sources, pools, accounts and geolocation to a pull request from an outside contributor; its 3.0 is still a draft. OpenSERP is at 0.8.12 from 22 July, its Google fix is waiting for a release, and its maintainer has 135 commits to the next contributor's 17. Agent Reach labels itself beta, its latest release dates from 11 June 2026, and installation pulls the main branch. Its catalogue keeps shifting: v1.4.2 dropped Douyin, Weibo and WeChat official accounts for lack of a reliable upstream tool.
PriceGhost hasn't had a commit since 3 February 2026. A single developer wrote all 122 commits, the last maintainer reply in the issues dates from 10 February, no outside contribution has been merged, and the README says the app was written entirely with Claude Code.
Licences and paid plans
Crawlee, Crawlee for Python, Rota and changedetection.io are under Apache-2.0, OpenSERP and Agent Reach under MIT. Camoufox mixes three licences: MPL-2.0 for the browser, LGPL-3.0-or-later for the mouse trajectories taken from Cursory, and MIT for the Python and TypeScript launchers. According to its website, releases up to v135.0.1-beta.24 contained a closed-source Canvas patch, and all of the code has been public since January 2026. PriceGhost says "MIT" at the end of its README, but the repository has no licence file and GitHub detects none.
Three vendors sell a service around the open-source code. Apify sells its platform: crawler execution, storage, proxies for subscribers only, scheduling and webhooks, with 5 dollars of usage a month on a free account. The Crawlee site says "Crawlee, by Apify, works anywhere, but Apify offers the best experience", and the docs state that "Crawlee is and will always be open source". changedetection.io's hosted plan costs 8.99 dollars a month for up to 5,000 URLs, with a Chrome instance and European, US and Tor proxies; every feature it lists is in the open-source code, so what you pay for is the infrastructure. OpenSERP Cloud keeps the same API, at 0.009 dollars per credit of up to ten results after 60 free credits, and adds managed API keys and scheduled searches; its page says nothing about proxies or CAPTCHAs. Camoufox, Rota, PriceGhost and Agent Reach have no paid offering, and twelve of Camoufox's nineteen sponsors sell proxies.
