No description
  • Python 94.6%
  • Shell 2.1%
  • CSS 1.3%
  • HTML 1%
  • JavaScript 0.7%
  • Other 0.3%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-09-28 14:46:58 -07:00
tests Accept substantial archive extraction 2026-09-28 14:46:58 -07:00
url2md Accept substantial archive extraction 2026-09-28 14:46:58 -07:00
.dockerignore Add Dockerfile and dockerignore for containerized deployment 2026-08-21 19:55:34 -07:00
.gitignore Refactor frontend UI and remove vendored Typora themes 2026-08-22 08:09:12 -07:00
Dockerfile Add Dockerfile and dockerignore for containerized deployment 2026-08-21 19:55:34 -07:00
pyproject.toml Refactor frontend UI and remove vendored Typora themes 2026-08-22 08:09:12 -07:00
README.md Quarantine fingerprint-bound DataDome cookies 2026-09-28 14:36:43 -07:00
url2md.sh Add support for listing all stored captures 2026-08-28 11:40:01 -07:00

url2md

Turn an article URL into a nicely formatted Markdown document.

Give it a link and it fetches the page, extracts the article, filters out the site chrome, and writes Markdown to stdout. It handles one article at a time — it's an extractor, not a crawler.

How it works

Live fetches go cheapest-first. A short teaser is not enough to stop, so url2md collects candidates and stops early only when a result looks like a full, ungated article (long and free of paywall signals):

  1. Direct — a plain HTTP request with a Chrome TLS fingerprint (curl_cffi; falls back to requests). Because no JavaScript runs, client-side paywalls never get a chance to truncate the page.
  2. Firecrawl — an opt-in rendered fetch via a self-hosted Firecrawl endpoint (FIRECRAWL_ENDPOINT, enable with URL2MD_FIRECRAWL_ENABLED=1).
  3. Headless browser — an opt-in playwright-service-py render (enable with URL2MD_PLAYWRIGHT_ENABLED=1) that returns HTML and any cookies the page set.

Archive.today is not tier 4. It costs minutes and third-party transcription credit per call, so it's scheduled, at most once per URL: if a live tier comes back as a paywalled preview, archive jumps the queue — the remaining live tiers would just re-fetch the same truncated page from the same origin — and it's also the fallback when every live tier is blocked (a cookieless NYT fetch is often HTTP 403). With --archive-first it's tried before the live tiers, for publishers that never serve a useful live page at all. ArchiveBox cache hits are vetted by the same paywall and length rules as live candidates: a gated or thin hit remains available as a fallback but no longer preempts fetching.

The sidecar loads the newest snapshot, and url2md extracts the article from that fragment, unwrapping archive-routed links and rebuilding paragraphs the extractor split. An archive fragment has no <head>, so metadata (title, authors, site, date, hero image) is grafted in from whichever live tier saw the publisher's own markup — a paywalled preview is the ideal donor, since only its body was truncated.

After extraction, the body is cleaned whitelist-first: blocks survive only if their text is vouched for by the extracted article baseline, and everything after the article's last paragraph is cut, which removes recirculation rails, "read more" link lists, and newsletter forms. Figures the extractor relocated are restored to source order when they can be placed confidently. The finished document starts with a header — title, authors, date, source link, and an archive link when the body came from a snapshot — followed by the article body.

Every successful fetch's raw HTML is kept in a local capture cache, so re-running conversion (after retuning a filter, say) costs no network — and no second archive call. A fresh, non-paywalled local capture short-circuits all tiers on repeat requests, so paywalled URLs do not re-spend archive-fetch or CAPTCHA-transcription credit within the TTL. --force-refetch bypasses automatic cache hits, including local captures and ArchiveBox; explicit prefer_cache remains an escape hatch that serves any available capture regardless of age or paywall state.

Requirements

  • Python 3.11 or newer.
  • Optional but recommended: a running playwright-service-py instance for the headless-browser tier and the Archive.today tier. Without it, url2md still works — it just skips both, so paywalled sites have no recovery path.
  • Optional: a Firecrawl-compatible scrape endpoint for tier 2.
  • The newspaper4k[nlp] dependency needs NLTK tokenizer data. It downloads automatically on first use; the Docker image pre-bakes it.

Install

git clone <repo-url> && cd url2md
pip install ".[api]"        # CLI + HTTP API
pip install -e ".[dev,api]" # development: editable install + test deps

CLI

url2md "https://www.example.com/article"                 # markdown to stdout
url2md "https://www.example.com/article" > article.md
url2md "https://paywalled.example.com/story" --token my-jar
url2md "https://www.example.com/article" -v              # verbose logging
url2md "https://www.nytimes.com/..." --archive-first     # snapshot before live tiers

Run url2md --help for the full list of options, including importing cookies from a Netscape-format file (the kind browser extensions like "Get cookies.txt" export), --json for a structured result with per-tier diagnostics, and --rerender, --list-captures URL, or --list-all-captures to work with the capture cache offline.

A total extraction failure exits 1, with the URL and what was tried on stderr.

Cookies and paywalls

Some sites only serve the full article to a browser that's logged in. url2md keeps cookie jars for this: named, per-site cookie stores persisted in a local SQLite database.

  • A jar is selected by a token — any string you choose. Pass the same token (CLI --token, API Authorization: Bearer) and you get the same jar. Use different tokens to keep sites or accounts separate. Omit the token for a fully cookieless, ephemeral run.
  • When the headless-browser tier renders a page, any cookies the site sets during that render are harvested back into the jar, so sessions refresh themselves over time. Expired cookies are dropped automatically. DataDome clearance cookies (datadome and _dd_s_v2) are intentionally neither stored nor replayed: they are fingerprint- and network-bound, and reusing them in a later browser context can turn a readable page into a 403.
  • You can seed a jar by importing a Netscape-format cookies file through the CLI. Match the USER_AGENT in url2md/config.py to the browser you exported from — anti-bot vendors may bind cookies to your UA and IP.

Cookies get you past bot walls and consent gates. They don't unlock a subscriber paywall: a site that truncates the article server-side sends live tiers the preview and nothing more. That's what the archive tier is for — the snapshot is the recovery, which is why a gated preview triggers it immediately. If there's no usable snapshot (or the snapshot is itself a preview), you get the preview.

HTTP API

python -m url2md.api    # or: uvicorn url2md.api.main:app --host 0.0.0.0 --port 8000

The API exposes the same fetch pipeline as the CLI. Once the server is running, interactive docs with the exact routes and schemas are at /docs.

Main routes:

Route Purpose
POST /v1/extract Live fetch and convert. The archive tier is rejected here (400) — use jobs. ?format=markdown returns raw Markdown.
POST /v1/convert Convert HTML you already have. No network access.
POST /v1/jobs Submit a background fetch (archive allowed, on by default); returns a job_id to poll.
GET /v1/jobs/{id} Poll a job: queued, running, done, or failed.
GET /v1/captures List every stored capture; pass ?url=… to restrict it to one URL.
POST /v1/captures/{id}/render Re-run conversion on a stored capture, no network.
GET /health Liveness; never touches the fetch pools.

Archive fetches can run for minutes (CAPTCHA solving on the archive mirrors is rate-limited per IP, and solved serially), which is why they go through submit-and-poll instead of a synchronous request. Duplicate in-flight jobs — same URL, options, token, and cookies — are deduplicated, so a second submission returns the existing job_id rather than spending a second archive call.

Extraction failure is a 422 with the URL and per-tier attempts, matching the CLI's exit-1 message. Job results are kept in memory (URL2MD_JOB_TTL_SECONDS, default one hour) and don't survive a restart.

Configuration

All configuration is via environment variables:

Variable Purpose Default
PLAYWRIGHT_SERVICE_ENDPOINT Scrape endpoint of the playwright-service sidecar; also used for archive fetches http://100.64.64.11:3003/scrape
URL2MD_PLAYWRIGHT_ENABLED Enable the Playwright live-fetch and Archive.today tiers (0/false/no disables them) 0
FIRECRAWL_ENDPOINT Firecrawl scrape endpoint for tier 2 http://100.64.64.11:3002/v2/scrape
URL2MD_FIRECRAWL_ENABLED Enable the Firecrawl live-fetch tier (0/false/no disables it) 0
URL2MD_ALLOW_LOCAL_WEBHOOKS Allow direct fetches to local/private targets; redirects remain protocol-validated 0
URL2MD_ARCHIVEBOX_URL ArchiveBox API base URL; enables the integration unset
URL2MD_ARCHIVEBOX_CACHE Use ArchiveBox as a cache lookup layer 1
URL2MD_ARCHIVEBOX_FEED Add the winning valid source URL to ArchiveBox's scraping feed 1
URL2MD_ARCHIVEBOX_API_KEY Bearer token for ArchiveBox API requests unset
URL2MD_ARCHIVEBOX_TIMEOUT ArchiveBox request timeout (seconds) 3
URL2MD_ARCHIVEBOX_CACHE_TTL_SECONDS Maximum age of an ArchiveBox cache hit 2592000
URL2MD_ARCHIVEBOX_HOST_HEADER Optional Host header for ArchiveBox requests unset
ARCHIVE_MIRRORS Comma-separated archive.today mirrors, tried in order archive.ph,archive.is,archive.today,archive.li,archive.md
ARCHIVE_TIMEOUT Per-call budget for an archive fetch, including CAPTCHA solving (seconds) 240
ARCHIVE_ANNOTATE Add an "Archived on …" link to the byline of snapshot bodies 1
URL2MD_CACHE_DIR Where the capture cache lives ~/.cache/url2md/captures
URL2MD_CACHE_MAX_MB Capture cache size cap; oldest captures evicted first 2048
URL2MD_CACHE_TTL_SECONDS Maximum age for automatic fresh, non-paywalled local capture hits; 0 disables automatic hits 86400
URL2MD_SITE_COOKIES_DB Where token cookie jars live ~/.cache/url2md/site_cookies.db
URL2MD_ARCHIVE_COOKIES_DB Where archive-mirror cookies live ~/.cache/url2md/archive_cookies.db
URL2MD_MAX_CONCURRENT Parallel live fetches in the API 4
URL2MD_REQUEST_TIMEOUT_SECONDS Synchronous request budget (/v1/extract, /v1/convert); timed-out workers are terminated 120
URL2MD_LOG_LEVEL API log verbosity INFO

Thresholds (minimum/confident body length, paywall and filter tuning) and the user-agent live in url2md/config.py.

Running with firecrawl-unleashed

url2md's browser and archive tiers depend on playwright-service-py, a Python/FastAPI port of firecrawl's Playwright microservice. That sidecar is itself part of firecrawl-unleashed, a refactored firecrawl deployment — which is also what url2md's tier 2 expects to talk to.

The intended setup is to run all of it together in a single docker-compose.yaml stack. url2md joins the stack's network and talks to the sidecars by service name — no published ports or extra wiring needed:

  url2md:
    build: ./url2md
    environment:
      PLAYWRIGHT_SERVICE_ENDPOINT: ${PLAYWRIGHT_SERVICE_ENDPOINT:-http://playwright-service:3000/scrape}
      FIRECRAWL_ENDPOINT: ${FIRECRAWL_ENDPOINT:-http://firecrawl-api:3002/v2/scrape}
      URL2MD_LOG_LEVEL: ${URL2MD_LOG_LEVEL:-INFO}
    volumes:
      - url2md-data:/data
    ports:
      - "${URL2MD_PORT:-8086}:8000"   # host access; delete to keep it internal
    networks:
      - backend
    depends_on:
      playwright-service:
        condition: service_started
    restart: unless-stopped

volumes:
  url2md-data:

Adjust service names and ports to match your stack. Set URL2MD_CACHE_DIR and the two URL2MD_*_COOKIES_DB variables to paths under /data if you want the capture cache and cookie jars on the named volume.

Then:

docker compose up -d --build url2md
docker compose logs -f url2md

The CLI is also installed in the image, so one-off jobs can run against the same state:

docker compose exec url2md url2md "https://www.example.com/article" --token my-jar

You can also run the image standalone (outside the compose stack) by pointing the endpoint variables at any reachable sidecars:

docker build -t url2md .
docker run -d -p 8086:8000 \
  -e PLAYWRIGHT_SERVICE_ENDPOINT=http://host.docker.internal:3003/scrape \
  url2md

Limitations

  • No live-tier subscriber access. Bot walls and consent gates yield to cookies and the fingerprinted sidecar; server-side paywalls don't. Archive is the recovery — and if a URL was never archived, or the snapshot is a preview, the preview is what you get.
  • Conservative by design. The body filter prefers dropping a decorative module over keeping a "More on…" rail. Text that never made the extraction baseline (captions outside alt attributes, some promotional cards) is rejected as unvouched.
  • Scriptio continua is fuzzier. For Chinese, Japanese, Thai, Lao, and Khmer, the filter compares characters instead of words, so it's more permissive there than for space-delimited languages.
  • Jobs are per-process. The in-memory job store doesn't span workers or restarts; a multi-worker deployment needs external job storage.

Development

python -m pytest tests/ -q

The test suite is hermetic: cookie stores and the capture cache are redirected to temp directories, and no test touches the network.

  • playwright-service-py — the headless-browser sidecar url2md uses for JavaScript-heavy pages, cookie harvests, and Archive.today snapshots.
  • firecrawl-unleashed — the refactored firecrawl stack that playwright-service-py ships with, and the compose deployment url2md is designed to join.