Senior Python Engineer for LLM Cost Optimization Across 4 Production Features
Worldwide
Our OpenAI bill went from $2,140 in February to $13,890 in July. Usage over the same period grew about 40%. Something is wrong with how we're calling these models and nobody here can tell us what. We're a 14-person B2B SaaS company (contract lifecycle management). We have four features calling the API, all through a shared llm_client.py wrapper that one of our engineers wrote a year ago and has been extended by three different people since: 1. Contract summarizer. User uploads a PDF, we return a structured summary. Uses gpt-4o, roughly 8,000 requests/month. We pass the entire extracted document text in the prompt, up to about 90k tokens on the long ones. 2. Clause classifier. Tags each clause against 22 categories. Also gpt-4o. This is our highest-volume caller by far, around 240,000 requests/month, and it's a fixed-taxonomy classification task. 3. Chat assistant. Questions over the user's contract library. Uses gpt-4o with a 20-message rolling history resent on every turn. Around 31,000 requests/month. 4. Internal search reranking. Uses text-embedding-3-large, re-embeds the full corpus nightly whether or not documents changed. What we've already checked: it isn't a runaway loop or a leaked key. Requests-per-month has grown roughly in line with customers. The spend curve is much steeper than the request curve. What we'd like: An audit of all four call paths with a breakdown of where the money actually goes. We can't currently attribute spend per feature, which is part of the problem. Implementation of the fixes you recommend. Per-feature cost attribution in Datadog so we can see this ourselves going forward. A short written handoff for our team. Target: under $5,000/month at current volume without users noticing a quality drop. If you think that target is unrealistic, say so in your proposal. We'd rather hear it now. Our stack: Python 3.11, FastAPI, openai Python SDK v1.x, Postgres 15, Redis (used for sessions only, not for LLM caching), Celery for the nightly embedding job, Datadog for APM, deployed on AWS ECS Fargate. Monorepo, decent test coverage on the API layer, none on llm_client.py. Who we're looking for: You've cut inference cost on a production system and can tell us by how much. Comfortable reading someone else's LLM integration code and finding the expensive decisions in it. Opinions on model routing, prompt caching, and semantic caching, specifically when each is worth the complexity. You'll tell us if the fix is "stop doing this feature this way" rather than a config change. Nice to have: experience with OpenAI's Batch API, prompt caching, or structured outputs. Familiarity with contract or legal document workflows. How we work: async, Slack, one 30-minute call to kick off and one at handoff. You'd work with our lead backend engineer, who'll review your PRs. To apply: look at the four features above and tell us which one you think is burning the most money and why. We have our own guess and we're curious if it matches.
- Less than 30 hrs/weekHourly
- 1-3 monthsDuration
- ExpertExperience Level
$85.00
-
$120.00
Hourly- Remote Job
- Ongoing projectProject Type
Skills and Expertise
Activity on this job
- Proposals:50+
- Last viewed by client:2 days ago
- Interviewing:0
- Invites sent:0
- Unanswered invites:0
About the client
- USASan Jose4:38 AM
- $95 total spent6 hires, 3 active
- 15 hours
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by