CS-Background Annotator — Rate AI-Predicted System & Tool Outputs (1 week, ~25 hrs)

Posted last week

Worldwide

Needs to hire 6 Freelancers
Summary

We're running a one-week human evaluation of an AI system that predicts what a computer system returns after an action: the output a command produces, the row a query returns, the JSON an API sends back, the error a bad request should have triggered. You read a conversation, look at the predicted result, and rate how right it is. This is a judgment job, not a labeling job. Almost every item hinges on something a non-engineer can't see: the response is well-formatted and confident, but the error code is the wrong one, or a value silently changed from three steps ago, or the call should have been rejected and instead it returned a plausible-looking success. That's why we're hiring CS backgrounds only. SCOPE One week, roughly 25 hours, 100 tasks. Each task takes up to 15 minutes. Before those, you complete 2 evaluation tasks that we review and score — see below. There's no ongoing commitment; if further rounds run you'll be first on the list, but treat this as a defined one-week engagement. WHAT ONE TASK LOOKS LIKE A JSON-formatted conversation ending in an action. Below it, the predicted result. You rate that prediction on five independent 1-5 dimensions — broadly: is the outcome correct, is it consistent with everything established earlier in the conversation, does it look like a genuine response from that system, was a failure handled as a failure rather than faked as a success, and is it sound enough to train an agent on. You get the full rubric, with anchors for every score, before you start. Two things trip up most people, so we'll say them now: - Score the dimensions independently. One bad answer on one axis does not drag the others down. A prediction can be factually wrong and still perfectly realistic in structure. - Rate against reality, not against an expected answer. The same action can legitimately produce different results — sometimes a call succeeds, sometimes it errors. A realistic error is a correct prediction, not a failed one. WHAT WE NEED FROM YOU - CS degree, and hands-on experience using or building AI agents. - Comfortable with tool-calling LLMs, tool schemas, and how an environment responds to an action. - Working fluency in at least two of: Linux/Bash and filesystem semantics; SQL and relational databases; REST APIs and JSON schemas; HTTP status and errno error semantics. The bar: you can say what a command, query, or call returns — and the correct error when it fails — without running it. - Error semantics specifically: you can tell EACCES from ENOENT, 401 from 403, and say why. - English, professional working proficiency. - 18 or older. - Availability for roughly 25 hours during the week of Mon 10 Aug. NO AI ASSISTANCE Every rating must be your own judgment. Do not use ChatGPT, Claude, Copilot or any other model to produce or check your ratings. Models are not reliable at this task, which is the entire reason we're paying humans to do it. We review submissions for it, and it ends the contract. HOW IT WORKS 1. Paid evaluation round — 2 tasks. Everyone does this first. It's the real work at small scale, and it's how we calibrate you and you calibrate us. 2. We review and tell you where your scores diverged from ours. 3. On approval, the 100-task set opens up and you work through it during the week. You submit through a Google Sheet assigned to you, one row per task. We'll ask you not to share it or discuss items with other annotators — independent judgment is the whole point of the exercise, and comparing notes destroys the signal we're paying for. PAY Hourly, capped at 25 hours for the week. The 2 evaluation tasks are paid at a flat rate of $25 and will take approximately 20 min. CONFIDENTIALITY The data and the rubric are confidential and covered by Upwork's standard terms. Don't paste task content into any external tool or model, and don't share it or the rubric outside the engagement. TO APPLY Answer the four questions below. Question 3 is the one we read first — short and correct beats long.

  • Less than 30 hrs/week
    Hourly
  • < 1 month
    Duration
  • Intermediate
    Experience Level
  • $25.00

    -

    $35.00

    Hourly
  • Remote Job
  • One-time project
    Project Type
Skills and Expertise
Mandatory skills
Machine Learning
Data Annotation
Activity on this job
  • Proposals:20 to 50
  • Last viewed by client:3 hours ago
  • Hires:
    15
  • Interviewing:
    5
  • Invites sent:
    30
  • Unanswered invites:
    10
About the client
Member since Jun 27, 2025
  • Canada
    Toronto11:53 PM
  • $105K total spent
    177 hires, 142 active
  • 1,742 hours

Explore similar jobs on Upwork

Cybersecurity & digital screening!Hourly‐ Posted 2 months ago
Python
Machine Learning
Deep Learning
Amazon SageMaker
PyTorch
Amazon Web Services
Cloud Computing
Google Cloud Platform
Retrieval Augmented Generation
AI Agent Development
Vertex AI
LangChain
Large Language Model
Databricks Platform
PySpark
AI and Automation SpecialistHourly‐ Posted 4 weeks ago
Adobe Illustrator
Graphic Design
Adobe Photoshop
HTML

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo