Python/AI Engineer — Review & Harden a Production LLM Agent (Ongoing, Part-Time)

Posted 7 hours ago

Worldwide

Summary

Python/AI Engineer — Review & Harden a Production LLM Agent (Ongoing, Part-Time) Description What this is We run a production LLM agent that answers operational questions for business users — it pulls from a live database, runs calculations, and returns numeric answers people make decisions on. It works. Roughly half of production traffic now goes through fast deterministic paths at sub-500ms with a clean abort rate. The other half is where we need you. The actual problem We have a strong Python developer building this. What we don't have is a second person whose job is to ask "is this actually correct, and can we prove it?" before anything ships. That gap has produced a repeatable set of failures: Deterministic work handed to the model. The agent has hand-summed a 32-row list and returned a total that was off by ~$27. Arithmetic that belongs in code went to a language model. Loop budget with no hard terminal. In one case the correct answer was already computed at step 2. The loop kept going, made an unnecessary call, and reported a number that was ~40x wrong. The system prompt already said "stop once you have a routed answer" — prompt wording didn't hold the line. Guardrails nobody calibrated. A self-check we added to catch wrong numbers now false-blocks a majority of our key metrics because of a 0–1 vs 0–100 scale mismatch nobody tested against real data. A success metric that lies in both directions. Our current "did this turn succeed" flag marks complete, correct answers as failures, and marks real refusals as fine. We have no trustworthy scoreboard. Things built and never proven live. An evaluation harness that has never once been run end-to-end. A change switched on in production while its own notes still said "untested." None of these are exotic. They're the ordinary failure modes of agentic LLM systems that nobody is reviewing at depth. What you'd actually do You are not the primary implementer — our developer stays hands-on-keyboard. You are the person who decides what gets fixed first, reviews the work before it merges, and can be held to a number. Concretely, week to week: Read PRs at real depth — async Python, control flow, tool-call loops, error paths — and block merges that can't be proven correct Decide which work belongs in deterministic code vs. the model, and enforce that boundary in code (hard terminals, not prompt wording) Recalibrate our guardrails against real production data instead of assumed ranges Define and stand up an honest success metric, then run it — a baseline computed on real traffic, not a harness that exists but never executes Own a small number of headline metrics: answer correctness, abort rate, p50/p95 latency, cost per resolved turn Write short, direct weekly notes: what moved, what's still broken, what you recommend next Who this fits 4+ years Python, and you're genuinely comfortable in async code — not "I've used asyncio," but you can spot a bug in a concurrent tool-call loop during review You've shipped and maintained an LLM agent in production, not just a demo. You know what a loop budget costs and why retries get expensive You've built or repaired an evaluation setup and can explain how you knew your numbers were real You can read someone else's code and say "this won't hold" with a specific reason You write clearly and briefly in English. Most of this role is written judgment Who this does not fit If your instinct for fixing a wrong number is to rewrite the system prompt, this isn't the role If you've only worked on greenfield demos, you'll find this frustrating If you need a fully specified ticket to start, this seat is the opposite of that Engagement Ongoing, part-time: 10–20 hours/week Overlap of at least 3 hours/day with 9am–6pm IST Starts with a paid 1-week trial: read the codebase, produce a prioritized remediation plan, and compute an honest baseline on our current traffic. If the plan is sharp, we continue indefinitely. Long-term relationship expected — we're not shopping for a one-off audit To apply Skip the template. Answer the screening questions. Applications that open with "I am a passionate full-stack developer" get archived unread. Skills Python, Asyncio, LLM, OpenAI API, Anthropic Claude, Prompt Engineering, AI Agent Development, Model Evaluation, Code Review, PostgreSQL Scope Ongoing project · More than 6 months · Part time (less than 30 hrs/week) · Intermediate · Hourly · $10.00–$12.00/hr · Worldwide · Freelancers only

  • More than 30 hrs/week
    Hourly
  • 1-3 months
    Duration
  • Expert
    Experience Level
  • $10.00

    -

    $12.00

    Hourly
  • Remote Job
  • Ongoing project
    Project Type

Contract-to-hire opportunity

This lets talent know that this job could become full time.
Learn more
Skills and Expertise
Mandatory skills
Python
Machine Learning
Activity on this job
  • Proposals:10 to 15
  • Interviewing:
    2
  • Invites sent:
    4
  • Unanswered invites:
    2
About the client
Member since Sep 15, 2021
  • India
    Jaipur8:41 PM
  • $5K total spent
    31 hires, 9 active
  • 450 hours
  • Tech & IT
    Mid-sized company (10-99 people)

Explore similar jobs on Upwork

Gen AI Developer (Contract)Fixed-price‐ Posted 2 months ago
AI Agent Development
Python
JavaScript
API
Node.js
Deep Learning
React
PostgreSQL
Football Prediction Tool DeveloperHourly‐ Posted 4 weeks ago
JavaScript
PHP
HTML5
jQuery

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo