Software / AI Evaluation Engineer (Terminal-Bench, Docker & Agentic Evals)

Posted 4 days ago

Worldwide

Summary

We are looking for an experienced Software Engineer / AI Evaluation Specialist to design, author, and validate terminal-agent evaluation tasks for Terminus 3 project (built on the open Terminal-Bench 3.0 / Harbor format). In this role, you will create reproducible, isolated environments and complex real-world software engineering challenges to evaluate state-of-the-art AI agents. Key Responsibilities: Author end-to-end tasks containing agent Dockerfiles, verifier test suites, prompt specifications (instruction.md), and oracle reference implementations (solution/solve.sh). Configure multi-container verifier isolation to eliminate reward-hacking, ensure strict artifact sharing, and verify zero data leakage of tests or solutions into the agent's environment. Write deterministic, outcome-based grading tests that inspect the final environment state. Calibrate task difficulty across predefined tiers (Base, Core, Advanced, Frontier) by benchmarking pass rates against leading LLM agents. Run local validations using stb harbor CLI to ensure reproducible builds, pinned dependencies, and clean checklist runs before submission. Required Qualifications & Skills: * Linux & Containerization: Advanced expertise in Docker, multi-stage builds, shell scripting (Bash), and isolated container environments. * Testing & Verification: Proven experience building deterministic Python test suites (e.g., PyTest) that evaluate system states rather than trivial outputs. * Domain Breadth: Strong proficiency in at least one key domain:Software & Systems: Systems programming, OS internals, Databases, Data Engineering, or Compilers. Target Languages: Python, C/C++, Rust, Go, TypeScript, Java, or C#. Specialized Domains: Cybersecurity (AppSec, Forensics, Reverse Engineering), ML Engineering, or Science/Operations. * Prompt Engineering & Specification: Ability to write clear, real-world ticket-style requirements stating the goal rather than the step-by-step solution. * Reproducibility Mindset: Deep attention to detail regarding version locking, deterministic execution, and debugging platform vs. test issues. Preferred Qualifications: ~ Prior experience contributing to Terminal-Bench, SWE-bench, or similar LLM agent evaluation frameworks. ~ Direct familiarity with the stb CLI toolchain. Key Candidate Screening Questions“ * Have you developed benchmarks or evaluation tasks using Docker and Python for autonomous coding agents (like Terminal-Bench or SWE-bench)?” * How do you design a deterministic test suite that verifies a task was truly solved without allowing an LLM agent to reward-hack or leak test data?” * Which programming languages and domain specialties (e.g., Systems, Security, ML, Algorithms) are your strongest?”

  • More than 30 hrs/week
    Hourly
  • 6+ months
    Duration
  • Expert
    Experience Level
  • Remote Job
  • Ongoing project
    Project Type

Contract-to-hire opportunity

This lets talent know that this job could become full time.
Learn more
Skills and Expertise
Mandatory skills
Docker Linux / Bash Shell Scripting Python Automated Software Testing * LLM Evaluation / Benchmarking * Systems Programming * DevOps / CI/CD * Prompt Engineering * Application Security (AppSec) * Test Automation (PyTest)
Activity on this job
  • Proposals:20 to 50
  • Interviewing:
    0
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since Apr 20, 2017
  • United States
    12:21 PM
  • Health & Fitness
    Individual client

Explore similar jobs on Upwork

Amazon Web Services
Amazon EC2
Artificial Intelligence
Large Language Model
Amazon SageMaker
Amazon Bedrock
MLOps
Cloud Architecture
Twilio Expert for quick advise & configurationFixed-price‐ Posted 3 weeks ago
Twilio API
VoIP

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo