I build AI agents, MCP servers, and automation pipelines that run unattended and do real work.
I work in Claude Code and the Model Context Protocol every day: I run 40+ MCP servers in my own stack and have shipped open-source ones, including a KiCad MCP server that lets an agent design PCBs. If you need an LLM wired into your actual tools (a CRM, a database, an internal API, a file system, a piece of hardware) and you want it reliable instead of flaky, that's the work I do best.
What I deliver:
• Custom MCP servers (Python or TypeScript) that expose your API or data as agent tools, with auth, error handling, and tests
• Agent workflows on Claude Code, the Claude Agent SDK, or OpenAI: multi-step, tool-using, with guardrails and logging
• Data extraction and scraping at scale: Playwright, undocumented APIs, scanned PDFs (OCR + LLM cleanup), delivered as clean CSV, JSON, or Postgres
• Speech and media pipelines: Whisper / faster-whisper transcription, caption alignment, ffmpeg and yt-dlp automation, real-time audio translation
• Vision pipelines: VLM-based image scoring, classification, and culling (I built one that scores and culls macro-photography shoots automatically)
• Local, self-hosted AI: Ollama and open-weight models on your own GPUs when data can't leave the building
Also fluent in Rust (high-performance async services), embedded systems (ESP32 firmware, KiCad PCB design, Raspberry Pi), and Linux / Docker / PostgreSQL.
How I work:
Async-first. Send me a written brief and I'll come back with a scoped fixed-price quote, then ship in focused sprints with written progress notes and working demos you can run yourself. Every deliverable is documented, reproducible (one command to run), and handed over with full source. No meetings required unless you want one.
Recent builds: bulk collection and keyword indexing of public comments from the Federal Register; an automated photo-culling pipeline built on a vision-language model; a real-time Italian-to-English speech translation pipeline for live gameplay; a distributed artificial-life simulation running across an 11-node ESP32 cluster.
If it's weird, undocumented, or "not possible," I'm interested. Send me the brief.
FFmpeg
Python
Artificial Intelligence
Large Language Model
Generative AI
AI Agent Development
Prompt Engineering
Automation
Automated Workflow
API Integration
Bot Development
Data Scraping
Web Crawling
Browser Automation
Data Extraction
Computer Vision
Machine Learning
Rust
ESP32
KiCad
Andrew A.
Downe, New Jersey
$75/hr
5.0
11 jobs
Quick and innovative solutions for the technologies of today.
For over a decade I have designed, managed and built custom software solutions for customers.
I've work in various industries including:
• Banking and Finance
• Traditional Radio & TV Broadcasting
• Advertising & Digital Signage
• Marketing and Sales
• Private Education
• Accounting & Human Resources
• Mobile Entertainment
• News Journalism
• Non Profits and Charitable Organizations
I see software as a solution to real world problems.
My industry skills and experience are:
• Project Management and Planning
• Dynamic / Interactive Website Design
• Custom Business Database Software
• Online Marketing Campaigns & Ad Distribution
• Mobile App Development
• Industry Specific Desktop Software
• GUI User Interface Design & Planning
• Low Level C/C++ Server Architecture
• Full Stack Development (Kernel all the way to User Interface)
In the past I have worked with clients ranging from single owner startups to large international firms with over 7,500+ employees.
No project is too small or too large.
Linux
JavaScript
C++
MySQL
PHP
Firmware
Web Application
Digital Signal Processing
Audio Engineering
Video Processing
Sound Synthesis
Music & Sound Design
Broadcast Engineering
Audio Mastering
AI Consulting
Ilya P.
Leander, Texas
$80/hr
5.0
15 jobs
Senior full-stack engineer and team lead with 20+ years of experience. I specialize in real-time video / WebRTC streaming and in AI/ML systems that ship to production.
I currently lead engineering at a live-video startup, where I architect streaming, scalable real-time chat, and AI features across the entire stack. Across my career I've built high-load platforms, custom HTML5 video players, and developer tooling serving thousands of concurrent users.
I can help you with two things in particular:
Real-time video & streaming: WebRTC, live streaming, HLS, HTML5 and custom video players, low-latency media pipelines, and scalable chat.
AI/ML engineering: LLM-powered apps, MCP servers & plugins, autonomous agents and workflows, RAG with vector databases, model fine-tuning (Hugging Face, PyTorch, LoRA/PEFT), and reinforcement learning (training models and building custom RL environments).
Core stack: TypeScript/JavaScript, Go, Python, Rust, Node.js, React, Next.js; PostgreSQL, MongoDB, Redis; Google Cloud, AWS, Docker, Kubernetes.
Open-source author: I created and maintain StateFlow (stateflow.dev) — a type-safe, predictable state-management library for TypeScript, running in production at the company I lead.
What you get working with me: clean, testable, production-grade code; clear communication; and a senior engineer who can both lead the architecture and do the hands-on work to ship it.
Building a video product, an AI/LLM feature, or an autonomous agent or need a senior lead to take something from idea to launch? Send me a message with a few details about your project and I'll tell you exactly how I'd approach it.
TypeScript
Golang
Python
Rust
PyTorch
WebRTC
Video Stream
Node.js
LLM Prompt Engineering
AI Model Training
AI Agent Development
Reinforcement Learning
Full-Stack Development
SDK
Open Source
Brady H.
Honeyville, Utah
$75/hr
5.0
16 jobs
I build the automations that kill your team's busywork — leads that route themselves, data that stops getting copied between apps, follow-ups that just happen.
I'm an automation engineer. I wire your CRM, forms, spreadsheets, and AI into the tools you already run — using n8n, Make, GoHighLevel, Zapier, and the Claude and OpenAI APIs — so the manual steps disappear.
A few things I've built:
- An AI agent that texts a client about messy QuickBooks transactions and updates the books straight from their reply.
- A voice AI agent for an HVAC company across 6 markets that calls new leads in seconds, gives an estimate, and sends a PDF bid by text and email.
- Stage-based GoHighLevel pipelines with the SMS and email follow-ups firing on their own.
- A full custom CRM with AI lead routing and Twilio messaging for a sales team.
Here's what makes me different: my day job is rocket engineering. So your automation gets built the way flight hardware does — documented, tested, and hard to break — not a brittle Zap that falls over the first time something weird comes through.
I'm a good fit if you're drowning in manual steps, duct-taping tools that don't talk to each other, or you want AI actually wired into your operations instead of just talked about.
Send me a message with whatever's eating your team's time, and I'll tell you straight whether it's worth automating.
Automation
CRM Automation
Task Automation
AI Model Integration
Make.com
API Integration
Business Process Automation
Python
FastAPI
Claude
RESTful API
Zapier
ChatGPT API
AI Agent Development
AI App Development
Twilio
CFD Analysis
ANSYS
COMSOL Multiphysics
OpenFOAM
Vad M.
Hallandale Beach, Florida
$40/hr
5.0
105 jobs
Live Video App Development, AI App Development, MVP Development, Video Surveillance, WebRTC & Live Streaming.
I help founders and product teams plan, scope, and deliver production-grade video and AI applications with Fora Soft’s senior engineering team.
My focus is discovery, product requirements, technical scoping, architecture planning, delivery coordination, and client communication. The engineering work is handled by Fora Soft’s specialized developers, architects, QA engineers, and project managers.
Backed by Fora Soft: $10M+ earned on Upwork, 914 jobs, 401K+ hours, and 250+ projects since 2005.
𝗪𝗵𝗮𝘁 𝗜 𝗖𝗮𝗻 𝗛𝗲𝗹𝗽 𝗬𝗼𝘂 𝗕𝘂𝗶𝗹𝗱
– Live video apps, WebRTC platforms, streaming products, webinars, broadcasting, IPTV/OTT
– AI video apps: object detection, facial analysis, emotion AI, speech-to-speech translation, captions, LLM assistants
– Video surveillance and IP camera platforms: multi-camera ingestion, mobile-as-camera, smart analytics, Smart TV apps
– VoIP, UCaaS, SIP, LiveKit, Twilio, Telnyx, FreeSWITCH, real-time communication systems
– MVPs, rescue projects, architecture reviews, and modernization of existing systems
𝗥𝗲𝗰𝗲𝗻𝘁 𝗣𝗿𝗼𝗷𝗲𝗰𝘁 𝗘𝘅𝗮𝗺𝗽𝗹es
→ Nucleus – secure UCaaS platform by Fibernetics, a Canadian telecom serving 300K+ customers and 5,000+ businesses. Fora Soft worked on iOS and platform delivery for WebRTC + SIP audio/video, real-time voice workflows, CRM/ERP integrations, and enterprise-grade communication features. Project volume: $285K+ combined.
→ TransLinguist – real-time AI interpretation platform for video calls. The platform supports interpretation workflows across 75+ languages, certified interpreter operations, AI speech-to-speech features, and live captions.
→ Doorbell App – LiveKit SIP-to-WebRTC bridge for smart intercoms. Video and audio from SIP intercoms to a mobile app, low-latency media through LiveKit, and DTMF-based remote door opening.
→ EmoProof – on-device emotion AI for iOS using Core ML. Real-time facial emotion and voice sentiment analysis on the phone, with privacy-focused processing and journal-style trend stats.
𝗪𝗵𝗲𝗿𝗲 𝗙𝗼𝗿𝗮 𝗦𝗼𝗳𝘁 𝗜𝘀 𝗦𝘁𝗿𝗼𝗻𝗴𝗲𝘀𝘁
– Live streaming & WebRTC: low-latency broadcasting, webinars, multi-camera apps, IPTV/OTT
– Video surveillance: IP cameras, mobile-as-camera, smart analytics, Smart TV apps
– Real-time AI/ML on video: OpenCV, YOLO, facial analysis, object detection, speech AI, captions
– Mobile live video: iOS, Android, React Native, AVFoundation, Camera2
– VoIP & UCaaS: SIP, FreeSWITCH, LiveKit, Telnyx, Twilio, high-load communication systems
𝗧𝗲𝗰𝗵𝗻𝗼𝗹𝗼𝗴𝗶𝗲𝘀 𝗢𝘂𝗿 𝗧𝗲𝗮𝗺 𝗪𝗼𝗿𝗸𝘀 𝗪𝗶𝘁𝗵
◆ Frontend: React, Next.js, Vue, Angular, TypeScript
◆ Backend: Node.js, NestJS, Python / FastAPI, Go, PHP / Laravel
◆ Video: WebRTC, LiveKit, Wowza, Kurento, MediaSoup, Janus, Jitsi, Agora, FFmpeg, GStreamer, HLS/DASH, RTMP/RTSP, SIP, FreeSWITCH, Telnyx, Twilio
◆ AI/ML: OpenAI API, Claude, Whisper, TensorFlow, PyTorch, OpenCV, YOLO, Azure AI Face, Deepgram
◆ Mobile: Swift, Kotlin, React Native, ARKit, ARCore, AVFoundation, Camera2 API
◆ Databases & Infrastructure: PostgreSQL, MongoDB, Redis, Elasticsearch, AWS, GCP, Docker, Kubernetes, NGINX, CI/CD
𝗪𝗵𝘆 𝗙𝗼𝗿𝗮𝗦𝗼𝗳𝘁
Fora Soft is a specialized video, WebRTC, and telecom software team with:
– $10M+ earned on Upwork
– 914 Upwork jobs
– 401K+ hours
– 250+ projects since 2005
– 400+ clients across 17+ countries
– Deep experience in live video, streaming, VoIP, AI video, and real-time communication
Send me a message or invite me to your job. I’ll review your idea, ask the right scoping questions, and suggest the best next step: discovery, audit, MVP planning, or full development.
WebRTC
JavaScript
AI Development
Mobile App Development
Web Application Development
iOS Development
Node.js
Chat & Messaging Software
Twilio API
Telemedicine
VoIP Software
Streaming Platform
Video Stream
Video Management Software
IPTV
Jarad P.
Port Orange, Florida
$90/hr
4.7
47 jobs
I design and build systems that pair generative AI with production-grade cloud infrastructure. Over 25 years in software engineering and cloud architecture, I've spent the last three years focused intensively on what happens when LLMs move from chat interfaces into autonomous, multi-agent pipelines.
My work bridges three layers: model integration (selecting, deploying, and chaining the right models for a task), agentic orchestration (building state machines that coordinate multiple AI agents through plan/write/review/gate cycles), and edge deployment (compiling models to run on-device via WASM, LiteRT, and browser runtimes).
I architected a state machine framework (sm3) that forks a single model context into three specialized roles — architect, engineer, and hypervisor — each operating under distinct permission boundaries, producing convergent artifacts without shared read access. This system uses DeepSeek V4 Flash, GPT-5, and Gemini as interchangeable engines within the same orchestration pipeline, with FastAPI dashboards and SQLite-backed persistence tracking every phase run, dispatch log, and sprint. A companion daemon (fw-sm) generalizes this into a DB-driven pipeline where states, transitions, guards, agent assignments, and file contracts are fully data-driven — no hardcoded orchestration logic.
On the multimodal side, I have architected video processing workflows: Whisper for speech-to-text, InsightFace for face detection and clustering (DBSCAN), and heuristic-ML hybrid scoring to extract the best subclips. These services ran on GCP with Celery workers, FastAPI/Flask backends, PostgreSQL persistence, and GCS media storage — coordinating through message queues and signed upload URLs.
I led the migration of a multilingual social simulation from a Python-server-dependent architecture (llama-cpp GGUF + Argos Translate + Piper TTS) to a fully browser-based stack where Gemma 4 runs via LiteRT-LM WASM for on-device chat and translation, with wllama, WebLLM, and ONNX Runtime Web as fallback runtimes — completely decoupling inference from any backend. The same project involved Piper TTS, Kitten NPZ-based browser TTS, and on-device phonemization.
Before going deep on AI, I spent years as a cloud architect specializing in GCP, AWS, and Azure — designing high-traffic, multi-region distributed systems for medical, nonprofit, and digital media clients. I have held certifications across all three platforms and have built microservice backends, media processing pipelines, and full-stack web applications.
What sets me apart is the range: contract AI model training and red-teaming, TTS systems from scratch (TensorFlow waveform modeling to browser NPZ engines), Apertium compiled to WASM for browser machine translation, self-hosted LLM inference with llama.cpp, and multi-agent frameworks that decompose engineering work across specialized AI roles. I can work at any level — from designing a cloud network topology to debugging a model's tokenization edge case.
I bring this full-stack AI engineering capability — cloud to edge, model to agent, prototype to production — to every engagement.
AI / LLM Models
GPT-4, GPT-5 / 5.5
Gemini 2.5 Pro
Gemini 3.5
Gemini (Gemini CLI) - bootstrap framework
Gemma 4 (LiteRT-LM)
o3 - Reasoning tasks
DeepSeek R1 - Reasoning tasks
DeepSeek V4 Flash - Primary experiment model
Llama 4 Scout
Claude (Claude Code) - CLI agent workflow
SmolLM2 (135M, 1.7B)
Qwen2.5 0.5B
TinyLlama 1.1B
Whisper (faster-whisper)
Ollama (local)
# CLI & Dev Tooling
OpenCode CLI - Agent dispatch
Google Gemini Agent CLI - CLI bootstrap
Codex - CLI agent workflow
Bun - Custom tool runtime
Docker / Docker Compose
Gitea - Self-hosted Git
llama.cpp - Local GGUF serving
ffmpeg / ffprobe
nginx + certbot
# Python Libraries & Frameworks
Web & API:
FastAPI
Flask
Streamlit
uvicorn
Jinja2
Mako
Bootstrap 5
Authlib
Databases & ORM:
SQLAlchemy - ORM
Alembic - Migrations
psycopg (binary) - PostgreSQL driver
Prisma (JS) - ORM
Task Queue:
Celery
Redis
AI / ML:
llama-cpp-python
InsightFace
tensorflow / Keras
argostranslate
huggingface-hub
transformers
torch
scikit-learn (DBSCAN)
librosa
soundfile
numpy
Game Dev:
pygame-ce
Kivy / KivyMD
Other:
python-avatars
cairosvg / pillow
g2pw / pypinyin / jieba
pydantic
pytest
pycountry
# Frontend & Browser Technologies
Next.js 16
React + Vite - Frontend
TypeScript - Tools & frontends
Tailwind CSS
shadcn/ui
TanStack React Query
NextAuth.js
ESLint / Prettier
Browser AI Runtimes:
phonemizer
# Edge & On-Device ML
Gemma 4 LiteRT
LiteRT / ai_edge_litert
WASM Inference - LiteRT-LM, wllama, WebLLM, ONNX Runtime Web
Apertium WASM
Pygbag WASM
Pyodide
Piper TTS (ONNX)
Kitten TTS
Amazon Web Services
Google Cloud Platform
DevOps
Artificial Intelligence
Machine Learning Framework
Machine Learning
Generative AI
Multimodal Large Language Model
Natural Language Processing
Computer Vision
Cloud Architecture
Python
FastAPI
Docker
PostgreSQL
API Development
Edge AI
WebAssembly
Automatic Speech Recognition
MLOps
Cloudflare
LAMP Administration
Web Analytics
Network Monitoring
Network Security
Service Cloud Development
Database Administration
How it works
Post a job for freePost a job
Tell us what you need. Create your own job post or generate one with AI then filter talent matches.
Hire top talent fast
Consult, interview, and hire quickly, so you can meet the freelancers you're excited about.
Collaborate easily
Use Upwork to chat or video call, share files, and track project progress right from the app.
Payment simplified
Manage payments in one place with flexible billing options. Only pay for approved work, hourly or by milestone.
Don't just take our word for it
“Upwork provides an umbrella-level of security. I can see a talent’s work history and ratings. I can hold payments in escrow. I can communicate through Upwork Messages instead of working through my email address.”
KD
Kim Darling
Emerald Tiger
“Upwork is the best platform to hire skilled professionals when we're not looking for a full-time employee. All the companies in our portfolio use Upwork to find talent across a wide range of fields.”
DM
David Merry
Kinetic Investments
“Our very specific requirements can be a challenge—With Upwork, we’re able to access a bigger community to ensure the success of our projects.”
KK
Katja Krohn
Summa Linguae
How do I hire a FFmpeg Specialist in the United States on Upwork?
You can hire a FFmpeg Specialist in the United States on Upwork in four simple steps:
Create a job post tailored to your FFmpeg Specialist project scope. We'll walk you through the process step by step.
Browse top FFmpeg Specialist talent on Upwork and invite them to your project.
Once the proposals start flowing in, create a shortlist of top FFmpeg Specialist profiles and interview.
Hire the right FFmpeg Specialist for your project from Upwork, the world's largest work marketplace.
At Upwork, we believe talent staffing should be easy.
How much does it cost to hire a FFmpeg Specialist?
Rates charged by FFmpeg Specialists on Upwork can vary with a number of factors including experience, location, and market conditions. See hourly rates for in-demand skills on Upwork.
Why hire a FFmpeg Specialist in the United States on Upwork?
As the world's work marketplace, we connect highly-skilled freelance FFmpeg Specialists and businesses and help them build trusted, long-term relationships so they can achieve more together. Let us help you build the dream FFmpeg Specialist team you need to succeed.
Can I hire a FFmpeg Specialist in the United States within 24 hours on Upwork?
Depending on availability and the quality of your job post, it's entirely possible to sign up for Upwork and receive FFmpeg Specialist proposals within 24 hours of posting a job description.
Find more freelancers
Top states for FFmpeg Specialists in the United States