Senior Python Engineer — production review + live smoke test of an LLM data-collection pipeline
Worldwide
Summary I run a research data-collection pipeline (open-source, Python) that discovers, scrapes, and uses LLMs (OpenAI Batch API) to extract structured records from government websites at scale. The code is mature — ~12k LOC, ~540 tests, 88% coverage, CI, packaged — but it has never had a live production run. Before I commit to one large, expensive, irreversible sweep, I want an experienced engineer to independently review it, run a small live end-to-end test, and harden the specific things that would waste money or corrupt output at full scale. This is a review-and-harden engagement, not a greenfield build. I have a written scope with ranked, discrete deliverables. What you'll do (ranked) Run a small live end-to-end test (~100 institutions) against live search + OpenAI Batch APIs using throwaway, budget-capped keys I provide — and fix every real-world failure it surfaces. Add a cost circuit-breaker so a run aborts if projected/actual spend crosses a ceiling. Close a short list of known correctness gaps at the LLM-output boundaries (foreign-key integrity into the final dataset, date validation, schema-rule enforcement). Validate the concurrency path under real parallelism (or tell me why it should stay single-threaded for the sweep). Deliver a prioritized production-hardening findings memo (provider abstraction, observability for a multi-day run, error-path tests). Must-have skills Expert Python (typing, packaging, pytest); comfortable in a well-tested existing codebase. Hands-on OpenAI Batch API experience (chunking, polling, cost, cached tokens). Web scraping at scale (requests, Playwright/headless, robots.txt, rate limiting). Pydantic / schema validation. Track record hardening pipelines for cost-safe, resumable, idempotent long runs. Nice to have Research-data / reproducibility background. Experience running LLM extraction against messy multilingual web content. How I'll choose Please answer in your proposal (see screening questions). I'll shortlist, share the public repo link + full written scope, and award a paid first milestone (a review memo) before the larger hardening milestones. Screening questions (answer briefly) Describe a data pipeline you made cost-safe or resumable for a long/expensive run. What specifically did you add? You have an LLM emitting a structured record that includes an ID meant to key back to a master table. How do you guarantee the emitted ID is trustworthy before it lands in the final dataset? How do you design a hard cost ceiling for an OpenAI Batch job that may run for days across many chunks? How would you run a small, cheap, live end-to-end test to de-risk a large run without spending much? Important All research/methodology decisions stay with me; you flag, you don't change them. You will not run the full sweep or receive production credentials or the full dataset — only a small sample + capped test keys.
$2,000.00
Fixed-price- ExpertExperience Level
- Remote Job
- Ongoing projectProject Type
Skills and Expertise
Activity on this job
- Proposals:50+
- Interviewing:0
- Invites sent:0
- Unanswered invites:0
About the client
- United StatesMountain View12:38 PM
- $11K total spent40 hires, 7 active
- 671 hours
- Individual client
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by