- TypeScript 68.6%
- Python 15.1%
- Rust 4.7%
- Java 3.1%
- PHP 1.8%
- Other 6.5%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| .github | ||
| apps | ||
| examples | ||
| firecrawl-cli | ||
| firecrawl-cli-skills | ||
| firecrawl-skills | ||
| firecrawl-workflows | ||
| img | ||
| upstream-merge/apps | ||
| .gitattributes | ||
| .gitignore | ||
| .gitmodules | ||
| AGENTS.md | ||
| CLAUDE.md | ||
| CONTRIBUTING.md | ||
| docker-compose.yaml | ||
| docker-compose.yaml.bak.2 | ||
| docker-compose.yaml.bak.3 | ||
| LICENSE | ||
| README.md | ||
| SELF_HOST.md | ||
Firecrawl Unleashed
An unofficial Firecrawl fork for self-hosters who want more operator control: CloakBrowser instead of the stock browser path, and fewer upstream guardrails.
Warning
This fork intentionally diverges from upstream Firecrawl behavior.
It is not affiliated with, endorsed by, or supported by Firecrawl or the Firecrawl team.
If you want the official managed product, cloud-only capabilities, or upstream support guarantees, use the upstream project at
firecrawl/firecrawland the hosted service at firecrawl.dev.
What this repo is
Firecrawl Unleashed starts from the upstream Firecrawl codebase and preserves the familiar scrape / crawl / map / search workflow where practical, but is tuned for:
- self-hosting
- operator control
- local experimentation
- harder scraping targets
- more permissive defaults for private or authorized use cases
This fork keeps the Firecrawl API shape where possible, but makes different tradeoffs than upstream.
Why this fork exists
Upstream Firecrawl is built to serve both open-source users and a hosted product. This fork exists for people who want a version that is easier to run locally, easier to patch, and less opinionated about what operators should or should not be allowed to do.
Goals:
- keep the Firecrawl API ergonomics where possible
- reduce friction for private/self-hosted deployments
- improve the browser scraping stack for anti-bot-heavy targets
- make local behavior easier to inspect and modify
- allow more aggressive customization than upstream is likely to accept
Key differences from upstream
This fork currently differs from upstream in several important ways:
- Swapped the stock browser backend for CloakBrowser (patched, anti-detection Chromium)
- Relaxed robots.txt enforcement at every layer — robots checks are bypassed, and robot-blocked crawl URLs log a warning instead of failing the crawl (
IGNORE_ROBOTS_TXT=truealso forces the account-levelignoreRobotsflag) - Removed the default backward-crawl path restriction — crawls are no longer confined to the subtree of the starting URL
- Added a
waitUntilscrape option on the browser engine (domcontentloaded|load|networkidle) — see below - Custom
OPENAI_BASE_URLendpoints are automatically routed to Chat Completions — upstream's Responses API default breaks against vLLM and other OpenAI-compatible providers that don't implement it; this fork detects a non-default base URL and switches - Bundled a local web UI service
- Relaxed some self-host feature gating
This means upstream documentation is still a useful baseline, but behavior in this fork may differ materially from upstream Firecrawl.
Browser stack
The browser microservice in this fork is built around CloakBrowser, rather than the stock upstream browser path.
Why that matters:
- better anti-bot posture for browser-driven scraping
- improved compatibility with harder targets
- no need to rely on Firecrawl’s cloud-only browser infrastructure
- preserves the existing Firecrawl microservice pattern instead of replacing it with a completely different browser API
This repo is therefore best thought of as:
Firecrawl API + Firecrawl worker model + CloakBrowser-backed browser service
rather than a stock upstream self-host.
Scrape option: waitUntil
Upstream Firecrawl only offers waitFor — a fixed delay (ms) applied after page load. This fork adds waitUntil, which controls which load state the initial navigation waits for before scraping:
| Value | Behavior |
|---|---|
domcontentloaded |
Return as soon as the HTML is parsed. Fastest; best for pages that hang on load because of trackers/ads, or when you only need the rendered DOM, not every subresource. |
load (default) |
Wait for the full load event. Matches upstream behavior. |
networkidle |
Wait until the network goes quiet. Slowest, but the most complete for JS-heavy pages. |
curl --request POST \
--url http://localhost:3002/v2/scrape \
--header 'Authorization: Bearer mykey' \
--header 'Content-Type: application/json' \
--data '{
"url": "https://example.com",
"formats": ["markdown"],
"waitUntil": "domcontentloaded"
}'
Supported on both the v1 and v2 scrape APIs. It only takes effect when the request is handled by the CloakBrowser (Playwright) engine. Combine with waitFor if you also need a settle delay after the load state fires — waitUntil picks the event, waitFor adds time.
Included stack
The local compose stack currently includes:
- Firecrawl API
- CloakBrowser-backed browser microservice
- Redis
- RabbitMQ
- NuQ Postgres
- optional FoundationDB components
- a community web UI service for local browsing/inspection
Quick start
1. Clone the repo
git clone <your-fork-url>
cd firecrawl-unleashed
2. Create or update .env
Set the values you care about, for example:
TEST_API_KEY=mykey
USE_DB_AUTHENTICATION=false
LOGGING_LEVEL=info
If your local patch set expects it, you may also use:
IGNORE_ROBOTS_TXT=true
If you use an upstream proxy:
PROXY_SERVER=http://host:port
PROXY_USERNAME=username
PROXY_PASSWORD=password
To use OpenRouter as the LLM provider (e.g. for /extract):
OPENROUTER_BASE_URL=https://openrouter.ai/api/v1
OPENROUTER_API_KEY=sk-or-...
3. Start the stack
docker compose up --build -d
4. Useful local endpoints
- API:
http://localhost:3002 - UI:
http://localhost:8080 - Browser service (direct):
http://localhost:3003— the CloakBrowser microservice is exposed on the host (/health,/scrape), handy for testing the browser path in isolation without going through the API
5. Smoke test
curl --request POST \
--url http://localhost:3002/v2/scrape \
--header 'Authorization: Bearer mykey' \
--header 'Content-Type: application/json' \
--data '{
"url": "https://example.com",
"formats": ["markdown"]
}'
What to expect
This fork is optimized for people who are comfortable running their own scraping infrastructure and making deliberate tradeoffs.
You should expect:
- more local control
- more divergence from upstream defaults
- more willingness to favor successful scraping over conservative policy enforcement
- more need to understand your own deployment
You should not expect:
- cloud parity with the official hosted product
- official Firecrawl support
- upstream compatibility guarantees for every rebase
Upstream relationship
This repository is derived from the upstream Firecrawl project:
- Upstream repo: https://github.com/firecrawl/firecrawl
- Upstream docs: https://docs.firecrawl.dev
- Upstream hosted service: https://firecrawl.dev
This fork should preserve required attribution and licensing notices from upstream.
If you want the official product direction, official support, and cloud-only features, upstream is the right source of truth.
Intended use
Firecrawl Unleashed is intended for:
- self-hosted scraping and crawling
- research and experimentation
- internal tools
- private or operator-authorized automation
It is your responsibility to comply with applicable laws, site terms, policies, and authorization requirements.
Contributing
This is a fork optimized for local control rather than strict upstream compatibility.
Contributions are welcome, especially if they improve:
- self-host ergonomics
- observability
- browser reliability
- anti-breakage during upstream rebases
- documentation of local deviations
Rebase strategy
If you track upstream, expect occasional merge pain in areas such as:
apps/api/src/controllers/auth.tsapps/api/src/controllers/v1/types.tsapps/api/src/controllers/v2/types.tsapps/api/src/lib/generic-ai.tsapps/api/src/lib/robots-txt.tsapps/api/src/scraper/WebScraper/crawler.tsapps/api/src/scraper/scrapeURL/engines/playwright/index.tsapps/api/src/services/worker/scrape-worker.tsapps/playwright-service-ts/*docker-compose.yaml
Document local deltas clearly and keep them small where possible.
License
This fork remains subject to the upstream licensing terms unless explicitly stated otherwise in a given directory or component. See the repository LICENSE and per-directory notices where applicable.