Senior Software Engineer / AI Coding Benchmark Coach

Posted 5 days ago

Worldwide

Summary

I’m looking for an experienced Senior/Staff Software Engineer to help me work through advanced software-engineering benchmark and AI-evaluation projects, primarily involving Mercor Hard Colosseum–style tasks and Snorkel software-engineering benchmark projects. These projects involve realistic software-engineering problems such as understanding existing repositories, debugging complex issues, implementing fixes, reviewing code, working with tests, understanding grading requirements, and identifying edge cases. I am looking for someone who can work with me as a technical coach, pair programmer, debugger, and code reviewer. You should be comfortable jumping into unfamiliar repositories and quickly understanding: Existing architecture Task requirements Failing tests Hidden edge cases Repository structure Build/test environments Docker or containerized environments CI/CD behavior Git/GitHub workflows Code correctness beyond simply passing tests What I Need Help With You may help me with tasks such as: Understanding difficult benchmark/task requirements Reviewing task specifications before implementation Exploring unfamiliar GitHub repositories Identifying the relevant parts of a large codebase Debugging failing builds and tests Explaining why tests are failing Finding edge cases that may not be immediately obvious Reviewing my implementation before submission Suggesting cleaner or more robust solutions Identifying regressions introduced by a change Running and interpreting test suites Understanding Docker-based development environments Troubleshooting environment/setup issues Reviewing patches and Git diffs Evaluating whether a solution actually satisfies the requirements Helping distinguish between a solution that merely passes visible tests and one that is genuinely correct Improving development speed when working through unfamiliar projects The goal is not simply to generate code, but to reason carefully about each engineering problem and make sure the resulting implementation is technically sound. Typical Workflow For each task, we would generally: Review the task requirements together. Inspect the repository and existing architecture. Identify the files/components involved. Discuss the likely implementation approach. Work through difficult implementation or debugging issues. Run the available tests. Investigate failures and edge cases. Review the final diff. Check for regressions, maintainability issues, and hidden failure scenarios. Make sure I fully understand the solution before completing the task. Required Experience I’m looking for someone with strong real-world software-engineering experience. You should have significant experience with several of the following: Python TypeScript / JavaScript Go Rust C# / .NET React / Node.js Backend development REST APIs SQL / databases Linux Git GitHub Docker CI/CD Automated testing Unit and integration testing Debugging unfamiliar codebases You do not need to be an expert in every language. More important is your ability to understand unfamiliar systems quickly and reason through difficult engineering problems. Strongly Preferred I’m especially interested in engineers who have experience with: Senior or Staff-level software engineering Open-source projects GitHub pull-request reviews SWE-bench or similar coding-agent benchmarks AI coding-agent evaluation LLM evaluation Benchmark creation or validation Repository-level coding tasks Debugging complex test failures Codebase modernization Reviewing AI-generated code Finding subtle bugs that automated tests miss Experience with tools such as Claude Code, Cursor, Codex, GitHub Copilot, or other AI coding agents is also useful. What Makes a Good Fit You are probably a good fit if you can open an unfamiliar repository and quickly answer questions like: Where is the bug most likely located? What behavior is the task actually asking us to change? Which tests should cover this behavior? What could break if we modify this component? Is this fix addressing the root cause or only making the test pass? What edge cases are missing? Is the implementation consistent with the existing architecture? Could a hidden test expose a weakness in this solution? I value reasoning and engineering judgment much more than simply producing code quickly. Working Arrangement This will initially be a small engagement, but it may become ongoing work if we work well together. I would prefer someone who can occasionally collaborate through: Upwork chat Screen sharing Pair-programming sessions Code review Git diffs / patches Repository walkthroughs Please let me know your availability and preferred working style. Important All work must follow the applicable project/platform rules. I am specifically looking for technical guidance, debugging assistance, pair programming, and code review. You should not access my account, impersonate me, or submit work under my identity. When Applying Please include: Your years of professional software-engineering experience. Your strongest programming languages. Your experience debugging unfamiliar repositories. Your experience with GitHub pull requests and code reviews. Any experience with SWE-bench, coding benchmarks, AI evaluation, or AI coding agents. Your experience with Docker/Linux development environments. Your hourly rate. Your availability. A short example of a difficult bug or engineering problem you diagnosed. Screening Question Suppose you receive a task in an unfamiliar repository. Your implementation passes all existing tests, but you suspect that hidden tests may still fail. How would you determine whether your implementation is actually correct rather than simply overfitting to the visible tests? Please give a short but technical answer. Engagement Type Hourly / ongoing as needed Potential for longer-term collaboration if the first few tasks go well.

  • More than 30 hrs/week
    Hourly
  • 6+ months
    Duration
  • Expert
    Experience Level
  • $5.00

    -

    $15.00

    Hourly
  • Remote Job
  • Ongoing project
    Project Type

Contract-to-hire opportunity

This lets talent know that this job could become full time.
Learn more
Skills and Expertise
Mandatory skills
Python
Artificial Intelligence
Activity on this job
  • Proposals:5 to 10
  • Last viewed by client:3 days ago
  • Hires:
    1
  • Interviewing:
    1
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since Sep 4, 2025
  • UKR
    Kyiv5:53 AM
  • $47 total spent
    2 hires, 2 active
  • 11 hours

Explore similar jobs on Upwork

Python
Machine Learning
Artificial Intelligence
Artificial Neural Network
Natural Language Processing
Deep Learning
Neural Network
Deep Neural Network
Convolutional Neural Network
Salesperson for Quant Finance AI ModelHourly‐ Posted 4 weeks ago
Financial Modeling
Financial Analysis
Financial Projection
Financial Accounting

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo