You will get Pull Public Government Data Into a Clean, Filtered CSV

Project details
What's included
| Service Tiers |
Starter
$55
|
Standard
$130
|
Advanced
$300
|
|---|---|---|---|
| Delivery Time | 3 days | 5 days | 7 days |
Number of Pages Mined/Scraped | 5000 | 50000 | 100000 |
Number of Sources Mined/Scraped | 1 | 1 | 3 |
Number of Revisions | 1 | 2 | 3 |
About Fox
Web Scraping & PDF Data Extraction, Clean CSV, Delivered Once
Burbank, United States - 5:09 pm local time
WHAT I DO
I take a one-time job — a site to scrape, a stack of PDFs to extract, a messy spreadsheet to clean, a public dataset to pull and reshape — and deliver a finished file. You get the data, the field definitions, and a short note on how it was collected and what it does and doesn't cover.
WHAT I'M GOOD AT
• PDFs that don't cooperate. Encrypted files, wrapped table rows, page breaks splitting records in half. I wrote my own PDF text extractor for a government dataset because the standard libraries wouldn't open the files.
• Public and municipal data. Socrata and similar open-data APIs, pagination, deduplication across repeat pulls, code-to-plain-English mapping (NAICS, asset-type codes, and so on).
• Bulk site collection. Hundreds to thousands of pages, verified by actually fetching them, rate-limited so nothing gets blocked or blocklisted.
• Saying what the data won't tell you. If 10% of a document set is unreadable, that 10% arrives in a separate file with the reason logged. You will not get silent gaps.
HOW I WORK
• Everything I write is plain Python — standard library, no dependency stack to install or maintain.
• I quote fixed price for a defined deliverable, and I confirm the exact columns with you before I start so there's no surprise at handoff.
• I send progress updates without being asked, and a sample of the first rows early so you can correct the shape before I run the whole job.
WHAT I DON'T TAKE ON
I deliver finished datasets, not systems you run
Steps for completing your project
After purchasing the project, send requirements so Fox can start the project.
Delivery time starts when Fox receives requirements from you.
Fox works on your project following the steps below.
Revisions may occur after the delivery date.
Confirm the source and the filters
I check the source is actually available and machine-readable, confirm the fields you need exist in it, and flag anything that simply is not in the data - before any work starts.
Send you a sample of the first rows
You see the output shape early, while it is still cheap to change. You correct the columns before I run the whole pull, not after.