Python/AWS Dev: Finish Probate Lead-Gen Pipeline
Worldwide
# Python/AWS Developer Needed: Complete Probate Lead-Gen Pipeline (Gemini AI + SQLite + Google Sheets) — 2 Day Turnaround ## About the project I run a single-family real estate investment company in Metro Detroit. I already have a working, in-production lead-generation pipeline that: - Runs 24/7 on an AWS EC2 instance as a systemd service - Monitors PDF issues for mortgage foreclosure notices - Uses Google Gemini AI to extract structured data from the PDFs - Matches extracted leads against a 1.3M-row property database (SQLite, built from ATTOM bulk data: tax assessor, AVM, and recorder files) - Writes enriched leads to Google Sheets I'm extending this same system to also capture **probate leads** — decedent estate and trust notices from the same PDFs — matched against the same property database by **owner name** instead of address (since probate notices give you a name, not an address). A lot of the groundwork is already built. I need someone experienced to test it against real data, fix what breaks, close the remaining gaps, and ship it. ## Current state (being fully transparent about this) The following code has been written but **not yet tested against the live Gemini API or the production database** (developed in an environment without API/network access): 1. `gemini_notice_extractor.py` — sends the PDF to Gemini, extracts probate creditor notices, trust notices, and appointment hearings as structured JSON via a Pydantic schema. Passes a syntax check; has not been run against real Gemini output. 2. `unified_enrichment_builder.py` / `unified_enrichment_service.py` — patched versions of my existing enrichment-database code, adding owner-name search (new columns: owner name, trust/company flag, owner mailing address). Tested locally against synthetic data matching the real schema; not yet run against my actual 1.3M-row production file. 3. `dln_probate_processor.py` — ties extraction and property-matching together into one pipeline. All code, my existing repo's relevant files (`unified_enrichment_service.py`, `unified_enrichment_builder.py`, `pdf_monitor.py`, `DEPLOYMENT.md`), and a written spec of the matching/confidence-scoring logic will be provided to whoever takes this on. ## What I need you to do 1. **Test and debug the Gemini extraction** against real PDF issues. Fix any schema mismatches or extraction accuracy issues (missing fields, misclassified notice types, etc.). 2. **Rebuild my enrichment database** with the new owner-name columns (script is already written, ~15-20 min one-time job) and verify match quality against real decedent names. 3. **Wire up Google Sheets output** — a 4-tab structure (raw notices / unique cases / property matches / qualified leads), using my existing service-account credentials. 4. **Decide and implement the integration point**: either a standalone script/cron job, or — likely better — folded directly into my existing `pdf_monitor.py` polling loop, since it may already be fetching this same PDF for foreclosure notices. Your call once you've seen the code. 5. **Handle deduplication**: notices are sometimes printed twice in the same issue (same case number) and must collapse to one case, not one row per printing. 6. **Deploy as a systemd service** on my existing EC2 instance, matching the pattern my current `pdf-monitor.service` already uses. 7. **Basic QA** across a handful of real PDF issues — confirm accuracy on decedent name, PR/trustee contact info, and property match quality before calling it done. ## Tech stack Python 3, Google Gemini API (`google-genai` SDK), Pydantic, SQLite, Google Sheets API (`gspread` + service account auth), AWS EC2 (Ubuntu, systemd), Google Drive API. ## Required experience - Strong Python — comfortable reading and extending an existing codebase quickly, not just writing from scratch - Experience with LLM structured-output extraction (Gemini, OpenAI function calling, or similar) — you should know what schema mismatches typically look like and how to fix them fast - SQLite / data pipeline experience at the "a few million rows" scale - Google Sheets API / service account authentication - AWS EC2 + Linux systemd (comfortable SSHing in, managing a service, reading logs) - Familiarity with real estate data, ATTOM property data, or probate/foreclosure lead generation specifically ## Deliverables - Working end-to-end pipeline, tested against at least 3-5 real DLN PDF issues - Enrichment database rebuilt and verified with owner-name search working - Google Sheets output live and populating correctly - Deployed and running unattended as a systemd service - Short written summary of what you changed/fixed and any known limitations - Available for a brief handoff call ## Timeline & budget **Need this completed within 2 days of start.** Given the amount of existing groundwork, this should be realistic for someone with the right background — but please review the attached code before committing to the timeline, since the honest state is "built but untested," not "fully working." I'd rather have an accurate estimate than a missed deadline. Please quote a fixed price for the full scope above, and let me know your availability to start immediately. ## To apply Please include: - Relevant experience with LLM-based data extraction and/or property data pipelines - Your proposed fixed price and confirmation you can start right away
$300.00
Fixed-price- IntermediateExperience Level
- Remote Job
- One-time projectProject Type
Skills and Expertise
Activity on this job
- Proposals:50+
- Last viewed by client:4 weeks ago
- Interviewing:0
- Invites sent:0
- Unanswered invites:0
About the client
- United StatesRoyal Oak1:42 PM
- $146K total spent34 hires, 5 active
- 11,521 hours
- Real EstateMid-sized company (10-99 people)
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by