Senior Python Backend Engineer — PDF/Data Extraction, Supabase & AI
Worldwide
We are building an internal AI assisted web application for a talent management company. The core backend challenge is converting large, complex agency PDFs and other unstructured sources into reliable, structured relational data. We are looking for an experienced Python Backend / Data Extraction Engineer to own the backend and ingestion side of the MVP. This is not a generic chatbot or simple RAG project. The main challenge is extracting accurate structured data, preserving relationships between projects, roles, people and contacts, preventing duplicates and incorrect associations, and storing validated results in Supabase/PostgreSQL. Responsibilities • Extract structured data from complex multi page PDFs and web documents • Build deterministic Python parsers where document structure is reliable • Use LLM extraction only where contextual interpretation is needed • Normalize, validate and deduplicate extracted data before database storage • Resolve ambiguous or duplicate projects, people, companies and roles • Design and maintain Supabase/PostgreSQL schemas • Implement RLS and appropriate backend access controls • Build FastAPI endpoints for project lookup, search, updates and AI assisted queries • Maintain source provenance and extraction/review states • Support natural language queries and ambiguous project references • Build reliable tests for extraction accuracy and regression • Work in parallel with a frontend developer using a clearly defined data/API contract Preferred Experience Strong experience with: • Python • FastAPI • PostgreSQL • Supabase • ETL/data pipelines • Complex PDF/document extraction • Structured LLM outputs • Data validation and normalization • Entity resolution and deduplication • AI coding tools such as Claude Code or Cursor Experience extracting data from financial, healthcare, legal, insurance, regulatory, scientific or similarly complex documents is highly relevant. We especially value engineers who understand that valid JSON does not necessarily mean correct data. You should know how to prevent hallucinated values, silent omissions, incorrect entity relationships and ambiguous data from becoming trusted production records. We expect the initial engagement to be close to full time for approximately 1–2 months, with potential for continued work. When applying, please briefly describe the most difficult document extraction system you have built and what made the extraction challenging.
- More than 30 hrs/weekHourly
- 3-6 monthsDuration
- ExpertExperience Level
- Remote Job
- Ongoing projectProject Type
Skills and Expertise
Activity on this job
- Proposals:50+
- Last viewed by client:yesterday
- Interviewing:8
- Invites sent:10
- Unanswered invites:2
About the client
- United StatesPrattville8:31 AM
- $1.5K total spent20 hires, 8 active
- 153 hours
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by