AI Agent and LLM Evals Course Content Creator

Posted 3 hours ago

Worldwide

Summary

# AI Agent & LLM Evaluation Course Content Creator We are looking for an experienced AI developer to create a short, practical course on evaluating LLM applications and AI agents. The core theme is **evaluation reliability**: how teams determine whether an evaluation result reflects real AI performance or noise from model variability, judge variability, weak datasets, or poor evaluation design. This is a technical content role. You should have hands-on experience building or evaluating LLM applications, RAG systems, or agents. ## Course Topics * Why traditional software testing alone is insufficient for LLM applications * Defining multidimensional quality across correctness, relevance, grounding, safety, task completion, and other dimensions * Designing evaluation criteria, metrics, rubrics, datasets, and representative test cases * Measuring quality using evaluators, scores, and results across the selected dimensions * **Evaluation reliability: understanding when evaluation results are trustworthy and when they may be misleading** * LLM-as-a-judge, deterministic evaluators, human review, calibration, and trade-offs * RAG evaluation: retrieval quality, context relevance, faithfulness, and answer quality * Agent evaluation: outcomes, tool use, decisions, and trajectory quality * Tracing failures to prompts, retrieval, models, tools, or application logic * Turning failures into regression tests and measuring reliable improvement over time ## Deliverables * Course outline and lesson scripts * Practical evaluation examples and demos * Sample datasets, rubrics, evaluators, and exercises * Supporting diagrams or charts ## Requirements * Hands-on LLM or agent development experience * Strong understanding of LLM, RAG, and agent evaluation * Understanding of evaluation reliability and LLM-as-a-judge variability * Python & Experience with evaluation or observability tools * Ability to explain technical concepts clearly **Budget:** $500 fixed price **Project Duration:** 10 calendar days **Delivery:** Milestone-based through Upwork

  • $500.00

    Fixed-price
  • Intermediate
    Experience Level
  • Remote Job
  • One-time project
    Project Type
Skills and Expertise
Mandatory skills
Content Writing
English
Activity on this job
  • Proposals:20 to 50
  • Interviewing:
    0
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since Dec 15, 2020
  • United States
    Fremont5:58 AM
  • $38K total spent
    75 hires, 16 active
  • Tech & IT
    Small company (2-9 people)

Explore similar jobs on Upwork

Clinical Trial Supply Chain Specialists NeededFixed-price‐ Posted 1 month ago
Supply Chain Management
Clinical Trial
Life Science
Pharmaceuticals
Test job - Write a blog about QAFixed-price‐ Posted 1 month ago
Technical Writing
Content Writing

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo