AI Agent and LLM Evals Course Content Creator
Worldwide
# AI Agent & LLM Evaluation Course Content Creator We are looking for an experienced AI developer to create a short, practical course on evaluating LLM applications and AI agents. The core theme is **evaluation reliability**: how teams determine whether an evaluation result reflects real AI performance or noise from model variability, judge variability, weak datasets, or poor evaluation design. This is a technical content role. You should have hands-on experience building or evaluating LLM applications, RAG systems, or agents. ## Course Topics * Why traditional software testing alone is insufficient for LLM applications * Defining multidimensional quality across correctness, relevance, grounding, safety, task completion, and other dimensions * Designing evaluation criteria, metrics, rubrics, datasets, and representative test cases * Measuring quality using evaluators, scores, and results across the selected dimensions * **Evaluation reliability: understanding when evaluation results are trustworthy and when they may be misleading** * LLM-as-a-judge, deterministic evaluators, human review, calibration, and trade-offs * RAG evaluation: retrieval quality, context relevance, faithfulness, and answer quality * Agent evaluation: outcomes, tool use, decisions, and trajectory quality * Tracing failures to prompts, retrieval, models, tools, or application logic * Turning failures into regression tests and measuring reliable improvement over time ## Deliverables * Course outline and lesson scripts * Practical evaluation examples and demos * Sample datasets, rubrics, evaluators, and exercises * Supporting diagrams or charts ## Requirements * Hands-on LLM or agent development experience * Strong understanding of LLM, RAG, and agent evaluation * Understanding of evaluation reliability and LLM-as-a-judge variability * Python & Experience with evaluation or observability tools * Ability to explain technical concepts clearly **Budget:** $500 fixed price **Project Duration:** 10 calendar days **Delivery:** Milestone-based through Upwork
$500.00
Fixed-price- IntermediateExperience Level
- Remote Job
- One-time projectProject Type
Skills and Expertise
Activity on this job
- Proposals:20 to 50
- Interviewing:0
- Invites sent:0
- Unanswered invites:0
About the client
- United StatesFremont6:52 AM
- $38K total spent75 hires, 16 active
- Tech & ITSmall company (2-9 people)
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by