A comprehensive tool for scraping, deduplicating, analyzing, and downloading U.S. Fish and Wildlife Service (FWS) Biological Opinion documents from the ECOSphere reporting system.
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
Sangye Ince-Johannsen 2f45f9aa1d feat: add resume functionality and improve filename handling in FWS scraper
- Add --resume flag to skip already-scraped URLs when interrupted
- Implement unique filename tracking with automatic version suffixes (_v2, _v3)
- Improve filename generation priority: date+title → title → doc_id → UUID
- Enhance analyze command to detect duplicate filenames and show statistics
- Add better console output showing scraped data for each row
- Document JSONL format in help text
- Track used filenames to prevent collisions during scraping

This brings FWS scraper to feature parity with NOAA scraper and ensures
reliable resume capability and unique filenames for all downloaded PDFs.
2025-10-20 17:09:21 -07:00
.gitignore initial commit 2025-09-05 20:51:28 -07:00
fws_scrape.py feat: add resume functionality and improve filename handling in FWS scraper 2025-10-20 17:09:21 -07:00
README.md update README 2025-09-05 20:57:53 -07:00
upload_biops.py Maintenance 2025-10-20 16:39:54 -07:00

FWS Biological Opinion Scraper

A comprehensive tool for scraping, deduplicating, analyzing, and downloading U.S. Fish and Wildlife Service (FWS) Biological Opinion documents from the ECOSphere reporting system.

🎯 Purpose

This tool automates the collection of biological opinion documents from the FWS ECOSphere database, providing researchers, consultants, and environmental professionals with easy access to complete datasets of consultation records with rich metadata.

✨ Features

  • Complete Data Extraction: Scrapes all 9 metadata fields including dates, offices, titles, project codes, species information, and more
  • Intelligent Deduplication: Automatically removes duplicate entries while preserving the best quality records
  • Parallel Downloads: Downloads PDFs with configurable concurrency (4-8 simultaneous downloads recommended)
  • Smart Filename Generation: Creates meaningful filenames using dates and titles (e.g., 2023.05.15 Habitat Restoration Project.pdf)
  • Robust Error Handling: Handles network timeouts, SSL issues, and malformed data gracefully
  • Progress Tracking: Real-time progress updates and comprehensive statistics
  • Multiple Download Methods: Both async (faster) and sync (more reliable) download options
  • Comprehensive Analysis: Built-in tools to analyze data quality and identify issues

📋 Requirements

System Requirements

  • Python 3.7+
  • 2GB+ RAM (for large datasets)
  • 10GB+ free disk space (biological opinions can be large)

Required Python Packages

pip install playwright asyncio aiohttp aiofiles requests urllib3

Install Playwright Browsers

playwright install chromium

🚀 Installation

  1. Clone or download the script:

    curl -O https://raw.githubusercontent.com/yourusername/fws-biop-scraper/main/fws_tool.py
    # or download directly
    
  2. Install dependencies:

    pip install playwright aiohttp aiofiles requests urllib3
    playwright install chromium
    
  3. Make executable (optional):

    chmod +x fws_tool.py
    

📖 Usage

Quick Start

# 1. Scrape all biological opinions (takes 30-60 minutes)
python fws_tool.py scrape

# 2. Remove duplicates (recommended)
python fws_tool.py dedupe

# 3. Download all PDFs (takes several hours)
python fws_tool.py download --async --metadata fws_biop_metadata_clean.json

Command Reference

Scraping Mode

# Basic scraping
python fws_tool.py scrape

# Custom timeout and output files
python fws_tool.py scrape --timeout 120000 --metadata custom_metadata.json --urls custom_urls.txt

Download Mode

# Async downloads (faster, recommended)
python fws_tool.py download --async

# Sync downloads (more reliable for unstable connections)
python fws_tool.py download --sync

# Custom concurrency and output folder
python fws_tool.py download --async --max-concurrent 8 --output-folder MyDownloads

# Download from specific metadata file
python fws_tool.py download --async --metadata fws_biop_metadata_clean.json

Deduplication Mode

# Basic deduplication
python fws_tool.py dedupe

# Custom input/output files
python fws_tool.py dedupe --input-file raw_data.json --output-file clean_data.json

Analysis Mode

# Analyze problematic files
python fws_tool.py analyze

# Analyze specific metadata file
python fws_tool.py analyze --metadata fws_biop_metadata_clean.json

Complete Workflow Example

# Step 1: Scrape all data (30-60 minutes)
python fws_tool.py scrape --max-concurrent 4

# Step 2: Clean up duplicates  
python fws_tool.py dedupe

# Step 3: Analyze data quality
python fws_tool.py analyze --metadata fws_biop_metadata_clean.json

# Step 4: Download PDFs (2-6 hours depending on connection)
python fws_tool.py download --async --metadata fws_biop_metadata_clean.json --max-concurrent 6

📁 Output Files

Generated Files

File Description Size
fws_biop_metadata.json Raw scraped metadata (JSONL format) ~15-20MB
fws_biop_metadata_clean.json Deduplicated metadata ~10-15MB
fws_pdf_urls.txt Simple list of PDF URLs ~1MB
fws_biop_metadata_clean_urls.txt Clean URL list ~500KB
FWS_Opinions/ Downloaded PDF files ~50-100GB

Metadata Structure

Each record in the JSON files contains:

{
  "date": "03/15/2023",
  "formatted_date": "2023.03.15",
  "fws_office": "Sacramento Fish and Wildlife Office",
  "title": "Highway 99 Bridge Replacement Project",
  "project_codes": "08ESMF00-2023-F-0123",
  "project_types": "Transportation - Road / Highway",
  "project_location": "CA: Sacramento",
  "lead_agencies": "California Dept. of Transportation",
  "species": [
    {
      "name": "Oncorhynchus mykiss",
      "url": "https://ecos.fws.gov/ecp/species/4284"
    }
  ],
  "url": "https://ecos.fws.gov/tails/pub/document/12345678",
  "filename": "2023.03.15 Highway 99 Bridge Replacement Project.pdf"
}

⚙️ Configuration Options

Command Line Arguments

Argument Description Default
--metadata Metadata file path fws_biop_metadata.json
--urls URLs file path fws_pdf_urls.txt
--output-folder Download directory FWS_Opinions
--max-concurrent Concurrent downloads 4
--timeout Page timeout (ms) 90000
--async Use async downloads False
--sync Use sync downloads False

Performance Tuning

Scraping Performance:

  • Increase --timeout for slow connections
  • The scraper is single-threaded by design to be respectful to the server

Download Performance:

  • --max-concurrent 4: Conservative, reliable
  • --max-concurrent 6: Good balance
  • --max-concurrent 8: Fast but may trigger rate limiting
  • Use --async for better performance on fast connections
  • Use --sync for better reliability on slow/unstable connections

🔧 Troubleshooting

Common Issues

SSL Certificate Errors:

✗ Error downloading: SSL certificate verify failed

The script automatically disables SSL verification for government sites. If you still see SSL errors, use the --sync download option.

Page Load Timeouts:

Timeout waiting for the next page to load

Increase timeout: python fws_tool.py scrape --timeout 120000

Out of Memory:

MemoryError during scraping

The script processes pages incrementally to minimize memory usage. If you still have issues, try running on a machine with more RAM.

Empty Results:

Found 0 PDF links on this page

This usually indicates the website structure has changed. Check if the site is accessible manually.

Debug Mode

For additional debugging information:

# Enable verbose output (if debugging is needed)
python -u fws_tool.py scrape 2>&1 | tee scraper.log

Recovery from Interruptions

The scraper appends data as it goes, so interruptions won't lose all progress:

# Resume downloads (existing files are automatically skipped)
python fws_tool.py download --async

# Check what was already downloaded
ls -la FWS_Opinions/ | wc -l

📊 Expected Results

Based on recent runs:

  • Total biological opinions: ~4,400-4,600 documents
  • Unique species references: ~8,000-10,000 entries
  • Total PDF size: ~50-100GB
  • Scraping time: 45-90 minutes
  • Download time: 3-8 hours (depending on connection and concurrency)
  • Duplicate records: ~60-100 (mostly header row artifacts)

🤝 Contributing

Reporting Issues

Please include:

  • Command used
  • Error message (full traceback)
  • Python version (python --version)
  • Operating system

Submitting Improvements

  1. Test your changes thoroughly
  2. Update the README if adding features
  3. Follow the existing code style
  4. Add appropriate error handling
  • This tool accesses publicly available government data
  • Please be respectful of server resources (don't increase concurrency excessively)
  • The biological opinion documents are public domain
  • Always check the FWS website's terms of use
  • Consider the server load impact of your scraping frequency

📄 License

This project is released into the public domain. Use freely for research, conservation, and educational purposes.

📈 Changelog

v1.0.0 (Current)

  • Initial release with full scraping, deduplication, and download functionality
  • Support for all 9 metadata fields
  • Async and sync download options
  • Comprehensive error handling and recovery
  • Built-in analysis tools

Questions? Open an issue or check the troubleshooting section above.

Need help with the data? The biological opinion documents contain detailed information about species consultations, habitat impacts, and conservation measures. Each PDF typically includes project descriptions, species analysis, and regulatory determinations.