AI Robotics / Embodied AI Engineer — GPU, Model Evaluation, Simulation, RL / VLA

Posted 3 days ago

Worldwide

Summary

We are looking for a strong **AI Robotics / Embodied AI engineer or researcher** to help us train, run, evaluate, and improve robotics models in simulation and, ideally, on real robots. This is intentionally a broad technical role. **We do not expect one person to be an expert in every area listed below.** We are interested in people who have exceptional depth in one or more relevant domains, whether that is model architecture, reinforcement learning, robotics simulation, GPU/model infrastructure, statistical evaluation, VLA models, or real-robot deployment. Someone with strong knowledge across several of these areas would be a major plus. ## Areas of Interest We are particularly interested in experience with: * GPU infrastructure, model training, and high-performance inference * PyTorch, CUDA, GPU memory/performance optimization, and profiling * Robotics model evaluation and benchmarking * Robot simulation and large-scale parallel evaluation * Reinforcement Learning (RL) and imitation learning * Transformer architectures and multimodal models * Vision-Language-Action (VLA) models * Diffusion policies and robotics foundation models * Statistical evaluation of robotics experiments * Sim-to-real transfer and real-robot deployment ## Robotics Simulation Experience with one or more robotics simulation platforms is valuable, including: * NVIDIA Isaac Sim / Isaac Lab * MuJoCo * ManiSkill * robosuite * Gazebo * PyBullet * Habitat * Genesis * Other modern robotics or embodied-AI simulation environments You do not need experience with all of them. We are especially interested in someone who understands the strengths and limitations of different simulators, how to construct meaningful evaluation environments, how to parallelize simulation effectively, and how simulation results should be interpreted before moving to a physical robot. ## Model Evaluation & Statistics Rigorous model evaluation is extremely important to us. We want to understand not only whether a model achieves a certain success rate, but **how confident we should be in that result**. For example, you should ideally be comfortable reasoning about: * How many evaluation trials or episodes should be performed * Whether a sample size is large enough to support a meaningful conclusion * Confidence intervals for success rates * Statistical significance when comparing two policies * Variance across random seeds * Paired vs. independent evaluation * Effect size and statistical power * How many environments, objects, positions, and initial conditions should be tested * How to avoid overfitting an evaluation to a small set of scenarios * How to categorize and analyze different failure modes * How to distinguish genuine model improvement from experimental noise You do not need to be a statistician, but you should understand how to design **reproducible and scientifically defensible robotics evaluations**. ## Model / Algorithm Experience Relevant expertise may include: * Transformers * Vision Transformers * Multimodal architectures * Vision-Language Models * Vision-Language-Action models * Reinforcement Learning * PPO, SAC, or related RL algorithms * Offline RL * Behavior cloning / imitation learning * Diffusion policies * Robotics foundation models * Manipulation policies * Robot perception and control Deep expertise in one or two of these areas is more valuable to us than superficial knowledge of everything. ## GPU / ML Systems Strong practical engineering experience is highly valuable, including: * PyTorch * CUDA / GPU workloads * Multi-GPU systems * Model training * Model inference * Training and inference profiling * GPU memory optimization * Checkpoint management * Experiment tracking * Reproducible evaluation pipelines * Parallel simulation * Distributed evaluation * Linux * Docker * Python Experience making robotics experiments significantly faster—for example, scaling an evaluation from a handful of sequential episodes to hundreds or thousands of parallel simulation episodes—is especially valuable. ## Real Robot Experience Experience deploying learned policies onto physical robots is a major plus, particularly experience involving: * Robot arms / manipulation * Cameras and perception systems * Policy inference on robot hardware * Calibration * Control loops and control frequency * Inference latency * Sim-to-real transfer * Safety constraints * Debugging differences between simulation and physical behavior However, excellent candidates with deep expertise in simulation, models, RL, evaluation, or GPU systems are still strongly encouraged to apply. ## Different Types of Candidates Can Be a Great Fit **Model / Architecture Expert** Deep understanding of transformers, VLA models, multimodal architectures, diffusion policies, RL, imitation learning, or modern robot-learning algorithms. **Simulation / Evaluation Expert** Exceptional experience building robot simulation environments, running large-scale evaluations, designing benchmarks, analyzing failures, and understanding sim-to-real issues. **GPU / ML Systems Expert** Very strong in PyTorch, CUDA, training/inference optimization, GPU profiling, multi-GPU systems, and scaling robotics experiments. **RL / Robotics Research Expert** Deep understanding of RL algorithms, experimental methodology, policy evaluation, reward design, generalization, and robotics research. **Robotics Generalist** Someone capable of connecting models, simulation, evaluation, GPU infrastructure, and physical robots into one functioning system. We are interested in all of these profiles. ## When Applying Please answer the following: 1. **What are the 2–3 technical areas where you consider yourself strongest?** 2. **Describe one robotics / embodied-AI project where you personally trained, implemented, deployed, or evaluated a model. What exactly did you do?** 3. **Which robotics simulators have you used extensively?** 4. **What is the largest robotics evaluation you have personally run?** Please describe the approximate number of episodes, environments, random seeds, GPUs, parallel workers, or other relevant scale. 5. **Suppose Policy A achieves a 70% success rate over 100 trials and Policy B achieves 76% over 100 trials. How would you determine whether Policy B is genuinely better?** 6. **Describe a situation where you optimized GPU training or inference performance. What was the bottleneck and what did you change?** 7. **What experience do you have with VLA models, transformers, RL, imitation learning, diffusion policies, or robotics foundation models?** 8. **Have you deployed a learned policy on a physical robot? If yes, briefly describe the system and the main technical difficulties you encountered.** 9. **Please provide any relevant GitHub repositories, papers, project pages, demos, videos, or other technical work you can share.** ## What We Value We are looking for people who can reason deeply about robotics experiments rather than simply run existing repositories. If model performance increases from 65% to 72%, we want someone who asks: **“Is the model actually better, and how do we prove it?”** If a policy performs well in simulation but fails on a physical robot, we want someone who can systematically investigate why. If an evaluation pipeline is too slow, we want someone who understands how to profile it, identify the bottleneck, and scale it efficiently. And when a new VLA, RL, or robotics approach looks promising, we want someone who can understand the underlying architecture and determine how to evaluate it properly. **You do not need to know everything listed above. If you are truly excellent in one or more of these areas, we want to hear from you.**

  • $750.00

    Fixed-price
  • Intermediate
    Experience Level
  • Remote Job
  • Ongoing project
    Project Type

Contract-to-hire opportunity

This lets talent know that this job could become full time.
Learn more
Skills and Expertise
Mandatory skills
Artificial Intelligence
Activity on this job
  • Proposals:5 to 10
  • Last viewed by client:2 hours ago
  • Interviewing:
    6
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since Aug 26, 2025
  • VNM
    Ha Noi9:58 PM
  • $33K total spent
    18 hires, 13 active

Explore similar jobs on Upwork

AI-Driven Graphic Design SpecialistHourly‐ Posted 4 weeks ago
Adobe Illustrator
Graphic Design
Illustration
Adobe Photoshop
Python
Machine Learning
Artificial Intelligence
Artificial Neural Network
Natural Language Processing
Deep Learning
Neural Network
Deep Neural Network
Convolutional Neural Network

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo