Senior AI Agent / Software Engineer – Coding Agent Evaluation & Benchmarks
Worldwide
We are looking for an experienced Software Engineer / AI Agent Developer to help us create and evaluate challenging coding tasks for AI coding agents. The work involves taking real open-source codebases and creating realistic software-engineering challenges such as bug fixes, feature implementations, refactoring, performance issues, and concurrency problems. You will prepare reproducible Docker-based environments, write task instructions, build automated tests and verifiers, and create reference solutions where needed. You will also work with AI coding agents by running multiple agents against the same task and evaluating their implementations. This includes reviewing their code changes, patches, trajectories, test results, and engineering decisions, then determining whether the solution is actually correct and which implementation is stronger. Example tasks you'll be working on 1. AI Coding Benchmark Creation Create a difficult software-engineering benchmark from an existing repository. Define the problem, prepare the Docker/Harbor environment, write the requirements, implement automated verification, and validate that the task is challenging but solvable. 2. Agentic Coding Task Evaluation Give the same coding task to multiple AI coding agents, review their solutions, and compare them based on correctness, edge cases, maintainability, architecture, regression risk, and overall software-engineering quality. Create objective rubrics to evaluate the results. Required skills Strong professional software-engineering experience Python and/or JavaScript/TypeScript Git/GitHub and Linux Docker and containerized environments Automated testing and debugging Code review and software architecture Ability to understand unfamiliar codebases Experience with LLMs, AI agents, or agentic coding tools Nice to have LangChain / LangGraph OpenAI, Claude, or Gemini APIs AI coding agents SWE-bench or similar benchmarks Harbor AI evaluation / benchmark development Automated grading/verifier systems CI/CD Python/FastAPI We are looking for a real software engineer with strong AI-agent experience, not someone focused only on basic n8n/Zapier automation. The ability to understand complex codebases, build reliable tests/verifiers, debug software, and critically evaluate AI-generated code is the most important part of this role.
- Less than 30 hrs/weekHourly
- 3-6 monthsDuration
- IntermediateExperience Level
$15.00
-
$25.00
Hourly- Remote Job
- Ongoing projectProject Type
Skills and Expertise
Activity on this job
- Proposals:50+
- Last viewed by client:1 hour ago
- Hires:2
- Interviewing:5
- Invites sent:1
- Unanswered invites:0
About the client
- United StatesDayton5:32 PM
- 1 hire, 1 active
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by