Python/AWS Dev: Finish Probate Lead-Gen Pipeline

Posted 4 weeks ago

Worldwide

Summary

# Python/AWS Developer Needed: Complete Probate Lead-Gen Pipeline (Gemini AI + SQLite + Google Sheets) — 2 Day Turnaround ## About the project I run a single-family real estate investment company in Metro Detroit. I already have a working, in-production lead-generation pipeline that: - Runs 24/7 on an AWS EC2 instance as a systemd service - Monitors PDF issues for mortgage foreclosure notices - Uses Google Gemini AI to extract structured data from the PDFs - Matches extracted leads against a 1.3M-row property database (SQLite, built from ATTOM bulk data: tax assessor, AVM, and recorder files) - Writes enriched leads to Google Sheets I'm extending this same system to also capture **probate leads** — decedent estate and trust notices from the same PDFs — matched against the same property database by **owner name** instead of address (since probate notices give you a name, not an address). A lot of the groundwork is already built. I need someone experienced to test it against real data, fix what breaks, close the remaining gaps, and ship it. ## Current state (being fully transparent about this) The following code has been written but **not yet tested against the live Gemini API or the production database** (developed in an environment without API/network access): 1. `gemini_notice_extractor.py` — sends the PDF to Gemini, extracts probate creditor notices, trust notices, and appointment hearings as structured JSON via a Pydantic schema. Passes a syntax check; has not been run against real Gemini output. 2. `unified_enrichment_builder.py` / `unified_enrichment_service.py` — patched versions of my existing enrichment-database code, adding owner-name search (new columns: owner name, trust/company flag, owner mailing address). Tested locally against synthetic data matching the real schema; not yet run against my actual 1.3M-row production file. 3. `dln_probate_processor.py` — ties extraction and property-matching together into one pipeline. All code, my existing repo's relevant files (`unified_enrichment_service.py`, `unified_enrichment_builder.py`, `pdf_monitor.py`, `DEPLOYMENT.md`), and a written spec of the matching/confidence-scoring logic will be provided to whoever takes this on. ## What I need you to do 1. **Test and debug the Gemini extraction** against real PDF issues. Fix any schema mismatches or extraction accuracy issues (missing fields, misclassified notice types, etc.). 2. **Rebuild my enrichment database** with the new owner-name columns (script is already written, ~15-20 min one-time job) and verify match quality against real decedent names. 3. **Wire up Google Sheets output** — a 4-tab structure (raw notices / unique cases / property matches / qualified leads), using my existing service-account credentials. 4. **Decide and implement the integration point**: either a standalone script/cron job, or — likely better — folded directly into my existing `pdf_monitor.py` polling loop, since it may already be fetching this same PDF for foreclosure notices. Your call once you've seen the code. 5. **Handle deduplication**: notices are sometimes printed twice in the same issue (same case number) and must collapse to one case, not one row per printing. 6. **Deploy as a systemd service** on my existing EC2 instance, matching the pattern my current `pdf-monitor.service` already uses. 7. **Basic QA** across a handful of real PDF issues — confirm accuracy on decedent name, PR/trustee contact info, and property match quality before calling it done. ## Tech stack Python 3, Google Gemini API (`google-genai` SDK), Pydantic, SQLite, Google Sheets API (`gspread` + service account auth), AWS EC2 (Ubuntu, systemd), Google Drive API. ## Required experience - Strong Python — comfortable reading and extending an existing codebase quickly, not just writing from scratch - Experience with LLM structured-output extraction (Gemini, OpenAI function calling, or similar) — you should know what schema mismatches typically look like and how to fix them fast - SQLite / data pipeline experience at the "a few million rows" scale - Google Sheets API / service account authentication - AWS EC2 + Linux systemd (comfortable SSHing in, managing a service, reading logs) - Familiarity with real estate data, ATTOM property data, or probate/foreclosure lead generation specifically ## Deliverables - Working end-to-end pipeline, tested against at least 3-5 real DLN PDF issues - Enrichment database rebuilt and verified with owner-name search working - Google Sheets output live and populating correctly - Deployed and running unattended as a systemd service - Short written summary of what you changed/fixed and any known limitations - Available for a brief handoff call ## Timeline & budget **Need this completed within 2 days of start.** Given the amount of existing groundwork, this should be realistic for someone with the right background — but please review the attached code before committing to the timeline, since the honest state is "built but untested," not "fully working." I'd rather have an accurate estimate than a missed deadline. Please quote a fixed price for the full scope above, and let me know your availability to start immediately. ## To apply Please include: - Relevant experience with LLM-based data extraction and/or property data pipelines - Your proposed fixed price and confirmation you can start right away

  • $300.00

    Fixed-price
  • Intermediate
    Experience Level
  • Remote Job
  • One-time project
    Project Type
Skills and Expertise
Mandatory skills
Python
Amazon Web Services
Nice-to-have skills
Amazon EC2
Amazon S3
Activity on this job
  • Proposals:50+
  • Last viewed by client:4 weeks ago
  • Interviewing:
    0
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since Feb 28, 2016
  • United States
    Royal Oak1:42 PM
  • $146K total spent
    34 hires, 5 active
  • 11,521 hours
  • Real Estate
    Mid-sized company (10-99 people)

Explore similar jobs on Upwork

JavaScript
Node.js
PHP
Web Application
AI App Development
DevOps
API
Git
MySQL
Web Design
WordPress

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo