- Python 94.6%
- Shell 2.1%
- CSS 1.3%
- HTML 1%
- JavaScript 0.7%
- Other 0.3%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| tests | ||
| url2md | ||
| .dockerignore | ||
| .gitignore | ||
| Dockerfile | ||
| pyproject.toml | ||
| README.md | ||
| url2md.sh | ||
url2md
Turn an article URL into a nicely formatted Markdown document.
Give it a link and it fetches the page, extracts the article, filters out the site chrome, and writes Markdown to stdout. It handles one article at a time — it's an extractor, not a crawler.
How it works
Live fetches go cheapest-first. A short teaser is not enough to stop, so url2md collects candidates and stops early only when a result looks like a full, ungated article (long and free of paywall signals):
- Direct — a plain HTTP request with a Chrome TLS fingerprint (curl_cffi; falls back to
requests). Because no JavaScript runs, client-side paywalls never get a chance to truncate the page. - Firecrawl — an opt-in rendered fetch via a self-hosted Firecrawl endpoint (
FIRECRAWL_ENDPOINT, enable withURL2MD_FIRECRAWL_ENABLED=1). - Headless browser — an opt-in playwright-service-py render (enable with
URL2MD_PLAYWRIGHT_ENABLED=1) that returns HTML and any cookies the page set.
Archive.today is not tier 4. It costs minutes and third-party transcription credit per call, so it's scheduled, at most once per URL: if a live tier comes back as a paywalled preview, archive jumps the queue — the remaining live tiers would just re-fetch the same truncated page from the same origin — and it's also the fallback when every live tier is blocked (a cookieless NYT fetch is often HTTP 403). With --archive-first it's tried before the live tiers, for publishers that never serve a useful live page at all. ArchiveBox cache hits are vetted by the same paywall and length rules as live candidates: a gated or thin hit remains available as a fallback but no longer preempts fetching.
The sidecar loads the newest snapshot, and url2md extracts the article from that fragment, unwrapping archive-routed links and rebuilding paragraphs the extractor split. An archive fragment has no <head>, so metadata (title, authors, site, date, hero image) is grafted in from whichever live tier saw the publisher's own markup — a paywalled preview is the ideal donor, since only its body was truncated.
After extraction, the body is cleaned whitelist-first: blocks survive only if their text is vouched for by the extracted article baseline, and everything after the article's last paragraph is cut, which removes recirculation rails, "read more" link lists, and newsletter forms. Figures the extractor relocated are restored to source order when they can be placed confidently. The finished document starts with a header — title, authors, date, source link, and an archive link when the body came from a snapshot — followed by the article body.
Every successful fetch's raw HTML is kept in a local capture cache, so re-running conversion (after retuning a filter, say) costs no network — and no second archive call. A fresh, non-paywalled local capture short-circuits all tiers on repeat requests, so paywalled URLs do not re-spend archive-fetch or CAPTCHA-transcription credit within the TTL. --force-refetch bypasses automatic cache hits, including local captures and ArchiveBox; explicit prefer_cache remains an escape hatch that serves any available capture regardless of age or paywall state.
Requirements
- Python 3.11 or newer.
- Optional but recommended: a running playwright-service-py instance for the headless-browser tier and the Archive.today tier. Without it, url2md still works — it just skips both, so paywalled sites have no recovery path.
- Optional: a Firecrawl-compatible scrape endpoint for tier 2.
- The
newspaper4k[nlp]dependency needs NLTK tokenizer data. It downloads automatically on first use; the Docker image pre-bakes it.
Install
git clone <repo-url> && cd url2md
pip install ".[api]" # CLI + HTTP API
pip install -e ".[dev,api]" # development: editable install + test deps
CLI
url2md "https://www.example.com/article" # markdown to stdout
url2md "https://www.example.com/article" > article.md
url2md "https://paywalled.example.com/story" --token my-jar
url2md "https://www.example.com/article" -v # verbose logging
url2md "https://www.nytimes.com/..." --archive-first # snapshot before live tiers
Run url2md --help for the full list of options, including importing cookies from a Netscape-format file (the kind browser extensions like "Get cookies.txt" export), --json for a structured result with per-tier diagnostics, and --rerender, --list-captures URL, or --list-all-captures to work with the capture cache offline.
A total extraction failure exits 1, with the URL and what was tried on stderr.
Cookies and paywalls
Some sites only serve the full article to a browser that's logged in. url2md keeps cookie jars for this: named, per-site cookie stores persisted in a local SQLite database.
- A jar is selected by a token — any string you choose. Pass the same token (CLI
--token, APIAuthorization: Bearer) and you get the same jar. Use different tokens to keep sites or accounts separate. Omit the token for a fully cookieless, ephemeral run. - When the headless-browser tier renders a page, any cookies the site sets during that render are harvested back into the jar, so sessions refresh themselves over time. Expired cookies are dropped automatically. DataDome clearance cookies (
datadomeand_dd_s_v2) are intentionally neither stored nor replayed: they are fingerprint- and network-bound, and reusing them in a later browser context can turn a readable page into a 403. - You can seed a jar by importing a Netscape-format cookies file through the CLI. Match the
USER_AGENTinurl2md/config.pyto the browser you exported from — anti-bot vendors may bind cookies to your UA and IP.
Cookies get you past bot walls and consent gates. They don't unlock a subscriber paywall: a site that truncates the article server-side sends live tiers the preview and nothing more. That's what the archive tier is for — the snapshot is the recovery, which is why a gated preview triggers it immediately. If there's no usable snapshot (or the snapshot is itself a preview), you get the preview.
HTTP API
python -m url2md.api # or: uvicorn url2md.api.main:app --host 0.0.0.0 --port 8000
The API exposes the same fetch pipeline as the CLI. Once the server is running, interactive docs with the exact routes and schemas are at /docs.
Main routes:
| Route | Purpose |
|---|---|
POST /v1/extract |
Live fetch and convert. The archive tier is rejected here (400) — use jobs. ?format=markdown returns raw Markdown. |
POST /v1/convert |
Convert HTML you already have. No network access. |
POST /v1/jobs |
Submit a background fetch (archive allowed, on by default); returns a job_id to poll. |
GET /v1/jobs/{id} |
Poll a job: queued, running, done, or failed. |
GET /v1/captures |
List every stored capture; pass ?url=… to restrict it to one URL. |
POST /v1/captures/{id}/render |
Re-run conversion on a stored capture, no network. |
GET /health |
Liveness; never touches the fetch pools. |
Archive fetches can run for minutes (CAPTCHA solving on the archive mirrors is rate-limited per IP, and solved serially), which is why they go through submit-and-poll instead of a synchronous request. Duplicate in-flight jobs — same URL, options, token, and cookies — are deduplicated, so a second submission returns the existing job_id rather than spending a second archive call.
Extraction failure is a 422 with the URL and per-tier attempts, matching the CLI's exit-1 message. Job results are kept in memory (URL2MD_JOB_TTL_SECONDS, default one hour) and don't survive a restart.
Configuration
All configuration is via environment variables:
| Variable | Purpose | Default |
|---|---|---|
PLAYWRIGHT_SERVICE_ENDPOINT |
Scrape endpoint of the playwright-service sidecar; also used for archive fetches | http://100.64.64.11:3003/scrape |
URL2MD_PLAYWRIGHT_ENABLED |
Enable the Playwright live-fetch and Archive.today tiers (0/false/no disables them) |
0 |
FIRECRAWL_ENDPOINT |
Firecrawl scrape endpoint for tier 2 | http://100.64.64.11:3002/v2/scrape |
URL2MD_FIRECRAWL_ENABLED |
Enable the Firecrawl live-fetch tier (0/false/no disables it) |
0 |
URL2MD_ALLOW_LOCAL_WEBHOOKS |
Allow direct fetches to local/private targets; redirects remain protocol-validated | 0 |
URL2MD_ARCHIVEBOX_URL |
ArchiveBox API base URL; enables the integration | unset |
URL2MD_ARCHIVEBOX_CACHE |
Use ArchiveBox as a cache lookup layer | 1 |
URL2MD_ARCHIVEBOX_FEED |
Add the winning valid source URL to ArchiveBox's scraping feed | 1 |
URL2MD_ARCHIVEBOX_API_KEY |
Bearer token for ArchiveBox API requests | unset |
URL2MD_ARCHIVEBOX_TIMEOUT |
ArchiveBox request timeout (seconds) | 3 |
URL2MD_ARCHIVEBOX_CACHE_TTL_SECONDS |
Maximum age of an ArchiveBox cache hit | 2592000 |
URL2MD_ARCHIVEBOX_HOST_HEADER |
Optional Host header for ArchiveBox requests | unset |
ARCHIVE_MIRRORS |
Comma-separated archive.today mirrors, tried in order | archive.ph,archive.is,archive.today,archive.li,archive.md |
ARCHIVE_TIMEOUT |
Per-call budget for an archive fetch, including CAPTCHA solving (seconds) | 240 |
ARCHIVE_ANNOTATE |
Add an "Archived on …" link to the byline of snapshot bodies | 1 |
URL2MD_CACHE_DIR |
Where the capture cache lives | ~/.cache/url2md/captures |
URL2MD_CACHE_MAX_MB |
Capture cache size cap; oldest captures evicted first | 2048 |
URL2MD_CACHE_TTL_SECONDS |
Maximum age for automatic fresh, non-paywalled local capture hits; 0 disables automatic hits |
86400 |
URL2MD_SITE_COOKIES_DB |
Where token cookie jars live | ~/.cache/url2md/site_cookies.db |
URL2MD_ARCHIVE_COOKIES_DB |
Where archive-mirror cookies live | ~/.cache/url2md/archive_cookies.db |
URL2MD_MAX_CONCURRENT |
Parallel live fetches in the API | 4 |
URL2MD_REQUEST_TIMEOUT_SECONDS |
Synchronous request budget (/v1/extract, /v1/convert); timed-out workers are terminated |
120 |
URL2MD_LOG_LEVEL |
API log verbosity | INFO |
Thresholds (minimum/confident body length, paywall and filter tuning) and the user-agent live in url2md/config.py.
Running with firecrawl-unleashed
url2md's browser and archive tiers depend on playwright-service-py, a Python/FastAPI port of firecrawl's Playwright microservice. That sidecar is itself part of firecrawl-unleashed, a refactored firecrawl deployment — which is also what url2md's tier 2 expects to talk to.
The intended setup is to run all of it together in a single docker-compose.yaml stack. url2md joins the stack's network and talks to the sidecars by service name — no published ports or extra wiring needed:
url2md:
build: ./url2md
environment:
PLAYWRIGHT_SERVICE_ENDPOINT: ${PLAYWRIGHT_SERVICE_ENDPOINT:-http://playwright-service:3000/scrape}
FIRECRAWL_ENDPOINT: ${FIRECRAWL_ENDPOINT:-http://firecrawl-api:3002/v2/scrape}
URL2MD_LOG_LEVEL: ${URL2MD_LOG_LEVEL:-INFO}
volumes:
- url2md-data:/data
ports:
- "${URL2MD_PORT:-8086}:8000" # host access; delete to keep it internal
networks:
- backend
depends_on:
playwright-service:
condition: service_started
restart: unless-stopped
volumes:
url2md-data:
Adjust service names and ports to match your stack. Set URL2MD_CACHE_DIR and the two URL2MD_*_COOKIES_DB variables to paths under /data if you want the capture cache and cookie jars on the named volume.
Then:
docker compose up -d --build url2md
docker compose logs -f url2md
The CLI is also installed in the image, so one-off jobs can run against the same state:
docker compose exec url2md url2md "https://www.example.com/article" --token my-jar
You can also run the image standalone (outside the compose stack) by pointing the endpoint variables at any reachable sidecars:
docker build -t url2md .
docker run -d -p 8086:8000 \
-e PLAYWRIGHT_SERVICE_ENDPOINT=http://host.docker.internal:3003/scrape \
url2md
Limitations
- No live-tier subscriber access. Bot walls and consent gates yield to cookies and the fingerprinted sidecar; server-side paywalls don't. Archive is the recovery — and if a URL was never archived, or the snapshot is a preview, the preview is what you get.
- Conservative by design. The body filter prefers dropping a decorative module over keeping a "More on…" rail. Text that never made the extraction baseline (captions outside
altattributes, some promotional cards) is rejected as unvouched. - Scriptio continua is fuzzier. For Chinese, Japanese, Thai, Lao, and Khmer, the filter compares characters instead of words, so it's more permissive there than for space-delimited languages.
- Jobs are per-process. The in-memory job store doesn't span workers or restarts; a multi-worker deployment needs external job storage.
Development
python -m pytest tests/ -q
The test suite is hermetic: cookie stores and the capture cache are redirected to temp directories, and no test touches the network.
Related projects
- playwright-service-py — the headless-browser sidecar url2md uses for JavaScript-heavy pages, cookie harvests, and Archive.today snapshots.
- firecrawl-unleashed — the refactored firecrawl stack that playwright-service-py ships with, and the compose deployment url2md is designed to join.