Senior AI Engineer, Autonomous Agents & LLM Evaluation

Posted yesterday

Worldwide

Summary

You'll own the intelligence layer of our GenAI platform for e-commerce marketing. We run growth for a portfolio of in-house brands and for client accounts, and AI now sits at the center of how we produce, optimize, and scale that work: product content, ad copy, audience and campaign strategy, creative variation, and performance analysis. This is a senior, architecture-owning role. You'll design the tool-using autonomous agents that execute real marketing work, and you'll build the evaluation infrastructure that tells us, rigorously rather than anecdotally, whether those agents actually move the numbers. In performance marketing, "the copy sounds good" isn't the bar. Output has to be on-brand, conversion-oriented, safe to publish, and measurably tied to results across many brands with different voices and different audiences. You'll be the person who makes that true. You'll sit at the intersection of agent design and evaluation science, and you'll have the standing to say when something isn't ready to touch a live account. What You'll Own Agent architecture. Design and build tool-using autonomous agents (planning, tool selection, retrieval, multi-step reasoning, and graceful failure) that handle real marketing tasks: generating product descriptions and ad variants, analyzing campaign performance, researching audiences, and pulling from catalog, analytics, and ad-platform data. Decide where agents act autonomously and where a human approves before anything ships. Evaluation infrastructure. Architect LLM evaluation dashboards spanning quality, brand-safety, and robustness. Build eval harnesses (OpenAI Evals, TruLens, or custom) that run continuously so quality doesn't silently drift across dozens of brands and campaigns. Outcome correlation. Go beyond offline scoring. Design the instrumentation that connects eval metrics to real marketing outcomes such as CTR, conversion rate, ROAS, and engagement, so we know our quality scores actually predict performance rather than just reading well. Brand safety and robustness. Build testing for the failure modes that matter here: off-brand voice, false or unverifiable product claims, non-compliant advertising language, prompt injection, and behavior drift across model versions. Brand grounding. Ensure agents stay anchored to each brand's voice, guidelines, product catalog, and claims constraints, with traceability from output back to source, so a given brand's agent never sounds like another's. Technical leadership. Set the standard for how the team reasons about agent quality and evaluation. Document methodology so results are reproducible and repeatable across the portfolio. What We're Looking For 5+ years of production ML experience, including 2+ years hands-on with LLMs in production (not just prototypes) Demonstrated experience building tool-using or autonomous agent systems (planning, tool orchestration, multi-step reasoning) Deep experience designing LLM evaluation: quality, safety, and robustness metrics, and the harnesses to measure them (OpenAI Evals, TruLens, or custom) Strong ability to correlate model/eval metrics with downstream business outcomes Deep Python Solid AWS experience for deploying and scaling ML systems Clear written and verbal communication. You'll be defending evaluation methodology to technical and non-technical stakeholders, including marketing leads Strongly Preferred Experience building AI for marketing, advertising, e-commerce, or content-at-scale use cases Familiarity with RAG architectures and grounding LLM outputs in brand guidelines, catalogs, or authoritative sources Experience integrating with e-commerce or ad platforms and their APIs (catalog, analytics, ad managers) Experience with adversarial testing, red-teaming, or robustness evaluation of LLM systems Awareness of advertising-compliance and claims constraints across channels Why This Role This isn't a "wrap an API and generate some copy" job. You'll be building the agent and measurement infrastructure that lets us scale AI-driven marketing across many brands without quality, voice, or compliance falling apart. It's the layer that determines whether we can trust these systems on live client and in-house accounts. If you care about making AI systems provably good rather than plausibly good, this is that role.

  • Less than 30 hrs/week
    Hourly
  • 1-3 months
    Duration
  • Expert
    Experience Level
  • $30.00

    -

    $75.00

    Hourly
  • Remote Job
  • Ongoing project
    Project Type

Contract-to-hire opportunity

This lets talent know that this job could become full time.
Learn more
Skills and Expertise
Mandatory skills
Artificial Intelligence
Activity on this job
  • Proposals:50+
  • Last viewed by client:yesterday
  • Interviewing:
    4
  • Invites sent:
    8
  • Unanswered invites:
    4
About the client
Member since Jul 21, 2026
  • United Kingdom
    Ilford4:39 AM

Explore similar jobs on Upwork

iOS
Camera
Android Smartphone
Artificial Intelligence
No-Code Platform for Multimodal ModelsFixed-price‐ Posted 5 days ago
Multimodal Large Language Model
Machine Learning
MLOps

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo