LLM Engineer to cut our AI costs (FastAPI + LangChain + RAG)

Posted 3 days ago

Worldwide

Summary

We run an AI chat assistant with some actions on a FastAPI backend. It uses LangChain, a RAG pipeline over a Milvus vector DB, and Claude models. It works, but our LLM cost per request is too high, mostly because some of our prompts are very large (one is around 12,000 input tokens per call). We need someone to bring that cost down without changing how the system actually behaves. That last part is the whole point — we don't want "shorter prompts that give worse answers." We want proof the answers stay the same. What we need done: 1. Go through our 4 main prompts and reduce token size where it's genuinely safe. We've already done an analysis that found real duplication (repeated examples, dead formatting instructions, etc.), so there's a starting point. 2. Build an automated test suite BEFORE making changes. Run our real queries through the current system, record what it does, then re-run after changes to prove decisions didn't change. English and Japanese both matter to us. 3. Evaluate 2–3 cheaper/alternative models against that same test suite, so we can see if switching models saves money while keeping quality. Deliver everything as reviewable pull requests with before/after numbers — token counts, cost, and test results. Our stack: Python, FastAPI, LangChain, Milvus, AWS Bedrock (Claude), some Node.js in front. A few things we care about: - You've worked on real production RAG/LLM systems, not just demos - You're comfortable using AI coding tools like Claude Code or Cursor — that's how we work - You can explain your reasoning to a non-technical founder in plain language When you apply, please skip the generic pitch. Just tell me: how would you prove a shorter prompt didn't hurt quality? That one answer tells me if you're the right person. This is a scoped project to start. If it goes well there's more ongoing work .

  • More than 30 hrs/week
    Hourly
  • 6+ months
    Duration
  • Expert
    Experience Level
  • Remote Job
  • Ongoing project
    Project Type

Contract-to-hire opportunity

This lets talent know that this job could become full time.
Learn more
Skills and Expertise
Mandatory skills
Artificial Intelligence
Activity on this job
  • Proposals:20 to 50
  • Last viewed by client:3 days ago
  • Interviewing:
    0
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since Dec 17, 2020
  • India
    Surat8:08 AM
  • $5.4K total spent
    23 hires, 5 active
  • 143 hours

Explore similar jobs on Upwork

iOS
Camera
Android Smartphone
Artificial Intelligence
No-Code Platform for Multimodal ModelsFixed-price‐ Posted 6 days ago
Multimodal Large Language Model
Machine Learning
MLOps

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo