Senior Multimodal AI / Video Intelligence Engineer — Google Gemini / Vertex AI

Posted yesterday

Worldwide

Summary

Senior Multimodal AI / Video Intelligence Engineer — Google Gemini / Vertex AI We are looking for a senior AI engineer with deep experience in multimodal/video LLMs, Google Gemini, and Vertex AI to help improve a production video-analysis platform. MUST WRITE “BOOTS” IN COVER LETTER This is not a basic prompt-engineering project. We currently analyze long-form security/body-camera footage using AI. We need an expert who can significantly improve the accuracy, consistency, and depth of the analysis, reduce false positives/false negatives, and build an architecture that improves over time from human feedback. What We’re Trying to Accomplish Our system analyzes hours of first-person body-camera footage from security officers. We want the AI to accurately identify and timestamp things such as: * Officer activities and locations * Security/compliance violations * Cell-phone usage * Extended conversations * Customer and employee interactions * Potential safety hazards * Blocked emergency exits * Patrol/activity patterns * Post-order compliance * Incidents and suspicious activity * Positive/exemplary employee behavior * Other configurable behaviors and events Accuracy is extremely important. We do not want a system that simply generates convincing summaries. We want an AI system whose findings can be measured, validated, and continuously improved. Responsibilities You will help us: * Audit our existing Gemini video-analysis architecture * Improve system prompts and task-specific prompts * Determine the best way to analyze multi-hour videos * Design video chunking and frame-sampling strategies * Improve temporal reasoning and timestamp accuracy * Break complex video-analysis tasks into specialized AI passes * Create second-pass verification for potential violations * Reduce false positives and false negatives * Improve structured JSON outputs * Develop confidence/evidence mechanisms for AI findings * Build automated LLM/video evaluation pipelines * Create ground-truth test datasets * Benchmark prompts and model versions against historical footage * Implement human-in-the-loop review * Store reviewer corrections as training/evaluation data * Develop a feedback loop so the system becomes more accurate over time * Evaluate RAG, embeddings, fine-tuning, supervised fine-tuning, and other approaches * Determine whether a specialized/fine-tuned model should eventually be used * Optimize performance, latency, and inference cost Continuous Learning / Feedback Loop A major goal of this project is creating a system where human corrections become useful AI training data. For example: AI Finding: Officer was using a cell phone at 01:32:15 Human Review: False positive Correction: Device was a handheld scanner Or: AI Finding: No violation detected Human Review: Missed blocked emergency exit at 02:14:32 We want these corrections captured systematically and used for: 1. Automated model evaluation 2. Prompt optimization 3. Retrieval of similar examples 4. Few-shot learning 5. Dataset development 6. Fine-tuning when appropriate 7. Potentially more advanced reinforcement/feedback-based approaches We are open to building additional specialized models if technically justified, but we do not want to train a new LLM from scratch unless there is a compelling reason. Required Experience You should have strong experience with several of the following: * Google Gemini API * Gemini multimodal/video understanding * Google Vertex AI * Python * Long-context LLM applications * Video LLM architectures * Multimodal prompt engineering * Structured outputs / JSON schemas * LLM evaluation frameworks * Ground-truth datasets * Human-in-the-loop AI systems * RAG * Embeddings / vector databases * Fine-tuning / supervised fine-tuning * FFmpeg * OpenCV * Video preprocessing * Temporal event detection Strong Bonus Experience Experience with any of the following is highly desirable: * Security/surveillance video analytics * Police/security body-camera footage * Retail loss prevention * Computer vision * Activity recognition * Object detection/tracking * Manufacturing video inspection * Sports-video analytics * Large-scale video processing pipelines * Active learning * MLflow or model/prompt versioning * Google Cloud Storage / Cloud Run / Cloud Functions / BigQuery What We Do NOT Want Please do not apply if your primary experience is: * ChatGPT prompt writing * Basic chatbot development * No-code AI workflows * Simple OpenAI API integrations * Generic “AI automation” We need someone who understands multimodal AI architecture and video analysis at a deep technical level. Initial Project The initial engagement will likely involve: 1. Reviewing our current architecture and prompts 2. Reviewing sample body-camera videos and existing AI output 3. Creating a ground-truth evaluation set 4. Establishing our current accuracy baseline 5. Identifying the largest sources of false positives/negatives 6. Designing an improved video-analysis architecture 7. Implementing improved prompting and analysis pipelines 8. Building automated evaluations 9. Creating the human-feedback/correction loop 10. Providing a roadmap for fine-tuning and continuous improvement If successful, this can become an ongoing engagement. Application Questions Please answer all of the following: 1. Describe a multimodal/video LLM system you have personally built. What model did you use and what did the system analyze? 2. What specific experience do you have with Google Gemini video understanding and Vertex AI? 3. Imagine you have an 8-hour chest-mounted security body-camera video and need to detect 20 different behaviors, some lasting 10 seconds and others lasting 30+ minutes. How would you architect the system? 4. Would you send the entire 8-hour video through Gemini with one large prompt? Why or why not? 5. How would you measure false positives, false negatives, precision, recall, and overall performance of a video LLM system? 6. How would you design a system that learns from corrections made by human reviewers? 7. When would you use prompting vs. RAG vs. fine-tuning for this problem? 8. Please provide examples of code, GitHub repositories, architecture diagrams, or previous projects relevant to this work. Important We are looking for someone who can challenge our current architecture rather than simply implement exactly what we ask. If you see a better way to solve the problem, we want you to tell us. The end goal is a production-grade video intelligence system with measurable accuracy that becomes increasingly effective as more footage and human-reviewed data are collected.

  • Less than 30 hrs/week
    Hourly
  • 3-6 months
    Duration
  • Expert
    Experience Level
  • $30.00

    -

    $60.00

    Hourly
  • Remote Job
  • Complex project
    Project Type

Contract-to-hire opportunity

This lets talent know that this job could become full time.
Learn more
Skills and Expertise
Mandatory skills
AI Development
Artificial Intelligence
Activity on this job
  • Proposals:50+
  • Last viewed by client:yesterday
  • Interviewing:
    1
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since Dec 22, 2020
  • United States
    Rancho Palos Verdes7:03 PM
  • $306K total spent
    76 hires, 36 active
  • 6,930 hours
  • Tech & IT
    Individual client

Explore similar jobs on Upwork

AI-Driven Graphic Design SpecialistHourly‐ Posted 4 weeks ago
Adobe Illustrator
Graphic Design
Illustration
Adobe Photoshop
Python
Machine Learning
Artificial Intelligence
Artificial Neural Network
Natural Language Processing
Deep Learning
Neural Network
Deep Neural Network
Convolutional Neural Network

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo