- Python 100%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
- Add --resume flag to skip already-scraped URLs when interrupted - Implement unique filename tracking with automatic version suffixes (_v2, _v3) - Improve filename generation priority: date+title → title → doc_id → UUID - Enhance analyze command to detect duplicate filenames and show statistics - Add better console output showing scraped data for each row - Document JSONL format in help text - Track used filenames to prevent collisions during scraping This brings FWS scraper to feature parity with NOAA scraper and ensures reliable resume capability and unique filenames for all downloaded PDFs. |
||
| .gitignore | ||
| fws_scrape.py | ||
| README.md | ||
| upload_biops.py | ||
FWS Biological Opinion Scraper
A comprehensive tool for scraping, deduplicating, analyzing, and downloading U.S. Fish and Wildlife Service (FWS) Biological Opinion documents from the ECOSphere reporting system.
🎯 Purpose
This tool automates the collection of biological opinion documents from the FWS ECOSphere database, providing researchers, consultants, and environmental professionals with easy access to complete datasets of consultation records with rich metadata.
✨ Features
- Complete Data Extraction: Scrapes all 9 metadata fields including dates, offices, titles, project codes, species information, and more
- Intelligent Deduplication: Automatically removes duplicate entries while preserving the best quality records
- Parallel Downloads: Downloads PDFs with configurable concurrency (4-8 simultaneous downloads recommended)
- Smart Filename Generation: Creates meaningful filenames using dates and titles (e.g.,
2023.05.15 Habitat Restoration Project.pdf) - Robust Error Handling: Handles network timeouts, SSL issues, and malformed data gracefully
- Progress Tracking: Real-time progress updates and comprehensive statistics
- Multiple Download Methods: Both async (faster) and sync (more reliable) download options
- Comprehensive Analysis: Built-in tools to analyze data quality and identify issues
📋 Requirements
System Requirements
- Python 3.7+
- 2GB+ RAM (for large datasets)
- 10GB+ free disk space (biological opinions can be large)
Required Python Packages
pip install playwright asyncio aiohttp aiofiles requests urllib3
Install Playwright Browsers
playwright install chromium
🚀 Installation
-
Clone or download the script:
curl -O https://raw.githubusercontent.com/yourusername/fws-biop-scraper/main/fws_tool.py # or download directly -
Install dependencies:
pip install playwright aiohttp aiofiles requests urllib3 playwright install chromium -
Make executable (optional):
chmod +x fws_tool.py
📖 Usage
Quick Start
# 1. Scrape all biological opinions (takes 30-60 minutes)
python fws_tool.py scrape
# 2. Remove duplicates (recommended)
python fws_tool.py dedupe
# 3. Download all PDFs (takes several hours)
python fws_tool.py download --async --metadata fws_biop_metadata_clean.json
Command Reference
Scraping Mode
# Basic scraping
python fws_tool.py scrape
# Custom timeout and output files
python fws_tool.py scrape --timeout 120000 --metadata custom_metadata.json --urls custom_urls.txt
Download Mode
# Async downloads (faster, recommended)
python fws_tool.py download --async
# Sync downloads (more reliable for unstable connections)
python fws_tool.py download --sync
# Custom concurrency and output folder
python fws_tool.py download --async --max-concurrent 8 --output-folder MyDownloads
# Download from specific metadata file
python fws_tool.py download --async --metadata fws_biop_metadata_clean.json
Deduplication Mode
# Basic deduplication
python fws_tool.py dedupe
# Custom input/output files
python fws_tool.py dedupe --input-file raw_data.json --output-file clean_data.json
Analysis Mode
# Analyze problematic files
python fws_tool.py analyze
# Analyze specific metadata file
python fws_tool.py analyze --metadata fws_biop_metadata_clean.json
Complete Workflow Example
# Step 1: Scrape all data (30-60 minutes)
python fws_tool.py scrape --max-concurrent 4
# Step 2: Clean up duplicates
python fws_tool.py dedupe
# Step 3: Analyze data quality
python fws_tool.py analyze --metadata fws_biop_metadata_clean.json
# Step 4: Download PDFs (2-6 hours depending on connection)
python fws_tool.py download --async --metadata fws_biop_metadata_clean.json --max-concurrent 6
📁 Output Files
Generated Files
| File | Description | Size |
|---|---|---|
fws_biop_metadata.json |
Raw scraped metadata (JSONL format) | ~15-20MB |
fws_biop_metadata_clean.json |
Deduplicated metadata | ~10-15MB |
fws_pdf_urls.txt |
Simple list of PDF URLs | ~1MB |
fws_biop_metadata_clean_urls.txt |
Clean URL list | ~500KB |
FWS_Opinions/ |
Downloaded PDF files | ~50-100GB |
Metadata Structure
Each record in the JSON files contains:
{
"date": "03/15/2023",
"formatted_date": "2023.03.15",
"fws_office": "Sacramento Fish and Wildlife Office",
"title": "Highway 99 Bridge Replacement Project",
"project_codes": "08ESMF00-2023-F-0123",
"project_types": "Transportation - Road / Highway",
"project_location": "CA: Sacramento",
"lead_agencies": "California Dept. of Transportation",
"species": [
{
"name": "Oncorhynchus mykiss",
"url": "https://ecos.fws.gov/ecp/species/4284"
}
],
"url": "https://ecos.fws.gov/tails/pub/document/12345678",
"filename": "2023.03.15 Highway 99 Bridge Replacement Project.pdf"
}
⚙️ Configuration Options
Command Line Arguments
| Argument | Description | Default |
|---|---|---|
--metadata |
Metadata file path | fws_biop_metadata.json |
--urls |
URLs file path | fws_pdf_urls.txt |
--output-folder |
Download directory | FWS_Opinions |
--max-concurrent |
Concurrent downloads | 4 |
--timeout |
Page timeout (ms) | 90000 |
--async |
Use async downloads | False |
--sync |
Use sync downloads | False |
Performance Tuning
Scraping Performance:
- Increase
--timeoutfor slow connections - The scraper is single-threaded by design to be respectful to the server
Download Performance:
--max-concurrent 4: Conservative, reliable--max-concurrent 6: Good balance--max-concurrent 8: Fast but may trigger rate limiting- Use
--asyncfor better performance on fast connections - Use
--syncfor better reliability on slow/unstable connections
🔧 Troubleshooting
Common Issues
SSL Certificate Errors:
✗ Error downloading: SSL certificate verify failed
The script automatically disables SSL verification for government sites. If you still see SSL errors, use the --sync download option.
Page Load Timeouts:
Timeout waiting for the next page to load
Increase timeout: python fws_tool.py scrape --timeout 120000
Out of Memory:
MemoryError during scraping
The script processes pages incrementally to minimize memory usage. If you still have issues, try running on a machine with more RAM.
Empty Results:
Found 0 PDF links on this page
This usually indicates the website structure has changed. Check if the site is accessible manually.
Debug Mode
For additional debugging information:
# Enable verbose output (if debugging is needed)
python -u fws_tool.py scrape 2>&1 | tee scraper.log
Recovery from Interruptions
The scraper appends data as it goes, so interruptions won't lose all progress:
# Resume downloads (existing files are automatically skipped)
python fws_tool.py download --async
# Check what was already downloaded
ls -la FWS_Opinions/ | wc -l
📊 Expected Results
Based on recent runs:
- Total biological opinions: ~4,400-4,600 documents
- Unique species references: ~8,000-10,000 entries
- Total PDF size: ~50-100GB
- Scraping time: 45-90 minutes
- Download time: 3-8 hours (depending on connection and concurrency)
- Duplicate records: ~60-100 (mostly header row artifacts)
🤝 Contributing
Reporting Issues
Please include:
- Command used
- Error message (full traceback)
- Python version (
python --version) - Operating system
Submitting Improvements
- Test your changes thoroughly
- Update the README if adding features
- Follow the existing code style
- Add appropriate error handling
⚖️ Legal & Ethical Use
- This tool accesses publicly available government data
- Please be respectful of server resources (don't increase concurrency excessively)
- The biological opinion documents are public domain
- Always check the FWS website's terms of use
- Consider the server load impact of your scraping frequency
📄 License
This project is released into the public domain. Use freely for research, conservation, and educational purposes.
🔗 Related Resources
- FWS ECOSphere Database
- Endangered Species Act Consultation Handbook
- FWS Environmental Conservation Online System
📈 Changelog
v1.0.0 (Current)
- Initial release with full scraping, deduplication, and download functionality
- Support for all 9 metadata fields
- Async and sync download options
- Comprehensive error handling and recovery
- Built-in analysis tools
Questions? Open an issue or check the troubleshooting section above.
Need help with the data? The biological opinion documents contain detailed information about species consultations, habitat impacts, and conservation measures. Each PDF typically includes project descriptions, species analysis, and regulatory determinations.