LLM Router Engineer — Multi-Provider Routing Service on GCP (LiteLLM, Python, Postgres)
Worldwide
I'm building a provider-neutral LLM routing service and need a strong backend engineer to implement it to spec on GCP. The architecture and decision logic are already designed. Your job is to build it cleanly, not to design it. What the service does Requests come in, get classified, and route to the cheapest capable model instead of defaulting to a frontier model every time. A quality gate decides whether to accept the cheap answer or escalate. Every request writes a cost and latency trace used for economics reporting. What you'll build Stateless Python routing service (FastAPI or similar) on a single GCP VM Self-hosted LiteLLM proxy configured across multiple providers plus one custom OpenAI-compatible endpoint Postgres schema and writes for request traces, including token and cost accounting Redis response caching, exact match plus optional embedding similarity A custom cost hook so hourly-billed routes are priced correctly alongside per-token APIs An offline evaluation harness that runs a synthetic corpus through the router and through a fixed baseline, then scores and compares No Kubernetes, no training, no frontend. One VM, two processes, three backing services. Required Strong Python and production API service experience You have personally shipped something that calls multiple LLM provider APIs behind one interface, with routing, retry, and fallback Postgres schema design and Redis Comfortable deploying and running a service on a Linux VM in GCP Careful about token and cost accounting, since the output numbers have to hold up to scrutiny Nice to have LiteLLM specifically LLM-as-judge or evaluation harness experience Embedding similarity search Engagement Roughly 6 weeks of build, starting immediately Milestone-based delivery You work to my spec and my architecture decisions. I review and I own client delivery. An NDA is required before any project detail is shared. This is a client requirement, not optional. You'll get the full spec after it's signed. Tell me how you'd prefer to structure it, fixed price or hourly, and what you'd quote for the scope above. Screening Shortlisted candidates do a live video interview where I'll ask technical questions about your actual experience and walk through the same problems below. Please answer in your own words from your own work. A polished LLM-written proposal that you can't back up on a call wastes both our time, and I will be able to tell within five minutes. Two questions: Tell me about a specific service you built that called two or more LLM providers behind one interface. What was the fallback behavior when a provider failed mid-request, and what actually broke in production? You're serving a model on an hourly GPU instance and comparing its cost against a per-token API. How do you make those two numbers comparable, and what assumption moves the result the most? Short and specific beats long and general.
- Less than 30 hrs/weekHourly
- 1-3 monthsDuration
- ExpertExperience Level
$30.00
-
$50.00
Hourly- Remote Job
- Ongoing projectProject Type
Skills and Expertise
Activity on this job
- Proposals:50+
- Last viewed by client:3 weeks ago
- Interviewing:5
- Invites sent:2
- Unanswered invites:0
About the client
- United StatesAtlanta9:31 PM
- $8.5K total spent32 hires, 6 active
- 212 hours
- Tech & ITSmall company (2-9 people)
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by