No description
  • TypeScript 68.6%
  • Python 15.1%
  • Rust 4.7%
  • Java 3.1%
  • PHP 1.8%
  • Other 6.5%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-08-21 11:40:58 -07:00
.github fix(ci): restore GitHub-hosted image builds (#4306) 2026-08-13 17:07:44 -07:00
apps my fix 2026-08-16 04:09:08 -07:00
examples docs: align self-hosting guidance with the public quickstart (#4209) 2026-08-07 10:52:08 +05:30
firecrawl-cli Add links to cli, skills, and workflows 2026-05-14 22:35:15 -04:00
firecrawl-cli-skills Add links to cli, skills, and workflows 2026-05-14 22:35:15 -04:00
firecrawl-skills Add links to cli, skills, and workflows 2026-05-14 22:35:15 -04:00
firecrawl-workflows Add links to cli, skills, and workflows 2026-05-14 22:35:15 -04:00
img Update firecrawl_logo.png 2026-03-29 19:42:28 -04:00
upstream-merge/apps my fix 2026-08-15 23:30:04 -07:00
.gitattributes Initial commit 2024-04-15 17:01:47 -04:00
.gitignore remove archive-solve submodule, update gitignore 2026-08-19 06:32:27 -07:00
.gitmodules mendableai -> firecrawl 2025-08-18 20:46:41 +02:00
AGENTS.md fix(api): recover crawl-finish context when FDB sheds member data (#3783) 2026-06-13 16:46:57 -07:00
CLAUDE.md chore(api): fix knip unused-export failures; forbid bypassing knip 2026-06-15 14:52:22 -07:00
CONTRIBUTING.md docs: align self-hosting guidance with the public quickstart (#4209) 2026-08-07 10:52:08 +05:30
docker-compose.yaml Remove archive-solve service and update docker-compose configuration 2026-08-21 11:40:58 -07:00
docker-compose.yaml.bak.2 remove archive-solve submodule, update gitignore 2026-08-19 06:32:27 -07:00
docker-compose.yaml.bak.3 Remove archive-solve service and update docker-compose configuration 2026-08-21 11:40:58 -07:00
LICENSE Update SDKs to MIT license 2024-07-08 13:37:53 -04:00
README.md readme update 2026-08-16 13:17:41 -07:00
SELF_HOST.md docs: align self-hosting guidance with the public quickstart (#4209) 2026-08-07 10:52:08 +05:30

Firecrawl Unleashed

An unofficial Firecrawl fork for self-hosters who want more operator control: CloakBrowser instead of the stock browser path, and fewer upstream guardrails.

Warning

This fork intentionally diverges from upstream Firecrawl behavior.

It is not affiliated with, endorsed by, or supported by Firecrawl or the Firecrawl team.

If you want the official managed product, cloud-only capabilities, or upstream support guarantees, use the upstream project at firecrawl/firecrawl and the hosted service at firecrawl.dev.


What this repo is

Firecrawl Unleashed starts from the upstream Firecrawl codebase and preserves the familiar scrape / crawl / map / search workflow where practical, but is tuned for:

  • self-hosting
  • operator control
  • local experimentation
  • harder scraping targets
  • more permissive defaults for private or authorized use cases

This fork keeps the Firecrawl API shape where possible, but makes different tradeoffs than upstream.


Why this fork exists

Upstream Firecrawl is built to serve both open-source users and a hosted product. This fork exists for people who want a version that is easier to run locally, easier to patch, and less opinionated about what operators should or should not be allowed to do.

Goals:

  • keep the Firecrawl API ergonomics where possible
  • reduce friction for private/self-hosted deployments
  • improve the browser scraping stack for anti-bot-heavy targets
  • make local behavior easier to inspect and modify
  • allow more aggressive customization than upstream is likely to accept

Key differences from upstream

This fork currently differs from upstream in several important ways:

  • Swapped the stock browser backend for CloakBrowser (patched, anti-detection Chromium)
  • Relaxed robots.txt enforcement at every layer — robots checks are bypassed, and robot-blocked crawl URLs log a warning instead of failing the crawl (IGNORE_ROBOTS_TXT=true also forces the account-level ignoreRobots flag)
  • Removed the default backward-crawl path restriction — crawls are no longer confined to the subtree of the starting URL
  • Added a waitUntil scrape option on the browser engine (domcontentloaded | load | networkidle) — see below
  • Custom OPENAI_BASE_URL endpoints are automatically routed to Chat Completions — upstream's Responses API default breaks against vLLM and other OpenAI-compatible providers that don't implement it; this fork detects a non-default base URL and switches
  • Bundled a local web UI service
  • Relaxed some self-host feature gating

This means upstream documentation is still a useful baseline, but behavior in this fork may differ materially from upstream Firecrawl.


Browser stack

The browser microservice in this fork is built around CloakBrowser, rather than the stock upstream browser path.

Why that matters:

  • better anti-bot posture for browser-driven scraping
  • improved compatibility with harder targets
  • no need to rely on Firecrawl’s cloud-only browser infrastructure
  • preserves the existing Firecrawl microservice pattern instead of replacing it with a completely different browser API

This repo is therefore best thought of as:

Firecrawl API + Firecrawl worker model + CloakBrowser-backed browser service

rather than a stock upstream self-host.


Scrape option: waitUntil

Upstream Firecrawl only offers waitFor — a fixed delay (ms) applied after page load. This fork adds waitUntil, which controls which load state the initial navigation waits for before scraping:

Value Behavior
domcontentloaded Return as soon as the HTML is parsed. Fastest; best for pages that hang on load because of trackers/ads, or when you only need the rendered DOM, not every subresource.
load (default) Wait for the full load event. Matches upstream behavior.
networkidle Wait until the network goes quiet. Slowest, but the most complete for JS-heavy pages.
curl --request POST \
  --url http://localhost:3002/v2/scrape \
  --header 'Authorization: Bearer mykey' \
  --header 'Content-Type: application/json' \
  --data '{
    "url": "https://example.com",
    "formats": ["markdown"],
    "waitUntil": "domcontentloaded"
  }'

Supported on both the v1 and v2 scrape APIs. It only takes effect when the request is handled by the CloakBrowser (Playwright) engine. Combine with waitFor if you also need a settle delay after the load state fires — waitUntil picks the event, waitFor adds time.


Included stack

The local compose stack currently includes:

  • Firecrawl API
  • CloakBrowser-backed browser microservice
  • Redis
  • RabbitMQ
  • NuQ Postgres
  • optional FoundationDB components
  • a community web UI service for local browsing/inspection

Quick start

1. Clone the repo

git clone <your-fork-url>
cd firecrawl-unleashed

2. Create or update .env

Set the values you care about, for example:

TEST_API_KEY=mykey
USE_DB_AUTHENTICATION=false
LOGGING_LEVEL=info

If your local patch set expects it, you may also use:

IGNORE_ROBOTS_TXT=true

If you use an upstream proxy:

PROXY_SERVER=http://host:port
PROXY_USERNAME=username
PROXY_PASSWORD=password

To use OpenRouter as the LLM provider (e.g. for /extract):

OPENROUTER_BASE_URL=https://openrouter.ai/api/v1
OPENROUTER_API_KEY=sk-or-...

3. Start the stack

docker compose up --build -d

4. Useful local endpoints

  • API: http://localhost:3002
  • UI: http://localhost:8080
  • Browser service (direct): http://localhost:3003 — the CloakBrowser microservice is exposed on the host (/health, /scrape), handy for testing the browser path in isolation without going through the API

5. Smoke test

curl --request POST \
  --url http://localhost:3002/v2/scrape \
  --header 'Authorization: Bearer mykey' \
  --header 'Content-Type: application/json' \
  --data '{
    "url": "https://example.com",
    "formats": ["markdown"]
  }'

What to expect

This fork is optimized for people who are comfortable running their own scraping infrastructure and making deliberate tradeoffs.

You should expect:

  • more local control
  • more divergence from upstream defaults
  • more willingness to favor successful scraping over conservative policy enforcement
  • more need to understand your own deployment

You should not expect:

  • cloud parity with the official hosted product
  • official Firecrawl support
  • upstream compatibility guarantees for every rebase

Upstream relationship

This repository is derived from the upstream Firecrawl project:

This fork should preserve required attribution and licensing notices from upstream.

If you want the official product direction, official support, and cloud-only features, upstream is the right source of truth.


Intended use

Firecrawl Unleashed is intended for:

  • self-hosted scraping and crawling
  • research and experimentation
  • internal tools
  • private or operator-authorized automation

It is your responsibility to comply with applicable laws, site terms, policies, and authorization requirements.


Contributing

This is a fork optimized for local control rather than strict upstream compatibility.

Contributions are welcome, especially if they improve:

  • self-host ergonomics
  • observability
  • browser reliability
  • anti-breakage during upstream rebases
  • documentation of local deviations

Rebase strategy

If you track upstream, expect occasional merge pain in areas such as:

  • apps/api/src/controllers/auth.ts
  • apps/api/src/controllers/v1/types.ts
  • apps/api/src/controllers/v2/types.ts
  • apps/api/src/lib/generic-ai.ts
  • apps/api/src/lib/robots-txt.ts
  • apps/api/src/scraper/WebScraper/crawler.ts
  • apps/api/src/scraper/scrapeURL/engines/playwright/index.ts
  • apps/api/src/services/worker/scrape-worker.ts
  • apps/playwright-service-ts/*
  • docker-compose.yaml

Document local deltas clearly and keep them small where possible.


License

This fork remains subject to the upstream licensing terms unless explicitly stated otherwise in a given directory or component. See the repository LICENSE and per-directory notices where applicable.