Talent badge filter
Select talent location
Select talent time zones
$80/hr
100%
Job Success
$700K+ earned
Available now
Start of list.
End of list.
I build production AI agents and LLM systems. 50+ projects across fintech, healthcare, media, legal tech, climate tech, and research: multi-agent systems, agent harnesses, evaluation systems for agentic workflows, Claude Code, Claude Agent SDK, OpenAI Codex, Hermes, Openclaw, MCP servers, RAG at scale, fine-tuning, and document AI.
7+ years shipping ML and AI in production. $700K+ delivered. Expert-Vetted (top 1%). 100% Job Success Score.
I’m not new to ML. Before the LLM wave I spent years building classical ML and NLP systems in Python: search engines, recommendation engines, entity extraction, knowledge graphs, and predictive models. That foundation means I know when a problem needs an LLM and when it doesn’t. I’ve also cofounded two startups, so I think about business outcomes, not just model accuracy.
Recent work:
- N1 Healthcare: Led the Report Generation team on a medical AI platform. Built the agent harness that dispatches 13+ clinical report workflows across Claude Agent SDK, OpenAI Agents SDK, and Agno through a LiteLLM gateway, with a planner coordinating 20+ specialty agents. Report cost went from roughly $80 to $15-25 per run. Also worked on the upstream medical record parsing pipeline.
- Newsweek (4+ years): 15+ AI applications and 6 Azure Function Apps used daily by non-technical editors. Cosmos DB vector search with Reciprocal Rank Fusion hybrid ranking on 3072-dim embeddings, MCP servers exposing agent tools to Copilot Studio, automated article-to-video pipeline at 95%+ success. A live 30-day window measured 497 editorial hours saved at 15.8% AI adoption. Trained editorial staff to run the systems without developer support.
- Stratifi: Built the AI layer on a Django + Postgres backend. Dual-LLM PDF extractor (GPT-4o + Gemini) at 95%+ accuracy on 50+ page brokerage statements, 10-30 positions/second, with a 24-document eval harness gating prompt changes. Multi-agent financial chatbot routing across 4 market data sources, plus a 4-layer portfolio optimization agent and a Lambda support agent with tool calling. Elasticsearch hybrid search over millions of securities.
- CGIAR: Research impact assessment processing 9,166+ records with 6 fine-tuned GPT-4o-mini variants, GPU-accelerated PDF layout detection with Detectron2, 3-tier evidence extraction at 50 concurrent operations.
- Adoro: 7-stage AI automation pipeline for marketing, turning a week-long manual process into under an hour, processing 100K+ items per run with 3-pass LLM extraction (initial, reflection, self-reflection).
- ClimateX: Web crawling and location intelligence pipeline with Flair NER for entity extraction, dual geocoding APIs, and proximity-based deduplication.
- Accelchain: Smart contract vulnerability detection covering 37 SWC vulnerability types with vector-DB retrieval over 33,488 known patterns, RAG plus few-shot plus Chain of Thought for explainable remediation.
- SynMax: NLP pipeline processing 400,000+ historical records with LongT5 fine-tuned at 16K context, 16-category document classification, cross-document entity linking across court records, deeds, and financial documents.
- Epitome (cofounded): HR Tech B2B2C platform with custom Word2Vec on millions of LinkedIn profiles, ArangoDB knowledge graph modeling O*NET across 20+ edge collections, entity resolution cascade (exact, fuzzy, embedding, LLM), Elasticsearch candidate search, 40+ REST endpoints.
Core capabilities: AI Agents and Orchestration, Agent Harnesses, Agent Evals, Multi-Agent Systems, Claude Agent SDK, MCP, LLM Fine-Tuning and Post-Training, RAG and Hybrid Retrieval, Document AI and OCR, Production ML and Forecasting, NLP and Knowledge Graphs, Data Pipelines and ETL, Cloud Infrastructure, and API Development.
Nikhil B.
has worked
.
$35/hr
100%
Job Success
$10K+ earned
Available now
Start of list.
End of list.
Senior Backend & AI Engineer with 5+ years owning high-throughput, event-driven systems across fintech, research, and AI infrastructure. I design and ship production AI systems: multi-tenant LLM gateways (LiteLLM) routing across Bedrock/OpenAI/Azure, agent orchestration with LangGraph and Postgres-backed checkpointing, and MCP (Model Context Protocol) integrations connecting AI assistants to backend tools securely.
Beyond AI, I bring deep backend fundamentals: real-money fintech systems (banking integrations, payment gateways, auditable ledgers), high-scale messaging (500K+ daily SMS/events), and RAG/document pipelines (ingestion, OCR, vector search) built for reliability and auditability at production scale.
Technical Stack: Python, TypeScript, FastAPI, NestJS, Java/Spring Boot, LangChain/LangGraph, PostgreSQL/pgvector, AWS (ECS, Fargate, RDS), Docker, Kubernetes, Terraform, Kafka/SQS/RabbitMQ, Prometheus/Grafana.
If you need someone who can own both the AI layer and the backend infrastructure it runs on end-to-end, secure, and production-grade, let's connect.
$30/hr
$8K+ earned
Available now
Offers consultations
Start of list.
End of list.
I build production-ready AI agents that automate real business workflows, not demos that break after deployment.
My recent project is a LangGraph multi-agent platform that monitors millions of social media posts in real time, performs semantic search with Pinecone, and routes monitoring, reporting, and live Q&A to specialized AI agents. That's the kind of work I deliver: AI systems that continue creating value long after handoff.
I'm a Senior AI Engineer with 7+ years of experience in NLP, LLMs, and production AI systems, currently working as a Lead Data Scientist for a San Francisco-based company.
What I Build
AI Agents & LLM Systems:
• Production AI Agents using LangGraph, OpenAI Agent SDK, CrewAI
• Multi-agent orchestration with approval workflows and tool calling
• RAG systems with hybrid search, reranking, and evaluation pipelines
• Enterprise knowledge assistants
• SQL-to-AI analytics systems
• Browser automation agents using Playwright & Chrome DevTools Protocol
• Self-hosted LLM deployments (Llama, Qwen, Mistral)
Intelligent Document AI
I build high-accuracy document processing systems for:
• Invoices
• Bank statements
• Receipts
• Financial documents
• Government forms
Specialized support for English and Arabic, including:
• OCR & Vision Language Models
• RTL text processing
• Arabic-Indic digits
• Field validation
• Human-in-the-loop review
• Confidence scoring
These systems reduce manual processing while improving extraction accuracy.
Recent Projects:
Real-Time Social Intelligence Platform:
Built an end-to-end AI platform that scrapes, indexes, and semantically searches millions of social media posts. A LangGraph multi-agent layer handles monitoring, automated reporting, and live Q&A.
Enterprise SQL AI Assistant:
Developed a multi-agent system that converts natural language into secure SQL queries across multiple Databricks databases, allowing business users to retrieve insights without writing SQL.
Private Enterprise LLM Infrastructure:
Deployed self-hosted Llama models on private infrastructure, eliminating third-party API dependency while reducing inference costs and meeting data residency requirements.
Trend-Aware AI Assistant:
Built a live RAG assistant combining vector search, Langchain, and real-time market/social data for intelligent answering and automated summarization.
Streaming NLP Platform:
Designed a production pipeline for real-time text classification and content categorization across live data streams.
How I Work:
I believe successful AI projects start with solving the right problem—not choosing the fanciest model.
When you work with me, you get:
• Clear project scoping before development
• Production-ready, maintainable code
• Evaluation and monitoring built into the solution
• API-based or self-hosted deployments
• Secure, scalable architecture designed for long-term use
I focus on systems that remain reliable after launching nonproof-of-concept notebooks.
Technology Stack:
LLMs & AI: OpenAI, Claude, Gemini, LangGraph, Langchain, CrewAI, Hugging Face
Backend: Python, FastAPI, Django, Vector Databases: Pinecone, Qdrant, ChromaDB, pgvector
Infrastructure: AWS, GCP, Kubernetes, Docker, Terraform, Databases: PostgreSQL, Redis, Elasticsearch, Automation: Playwright, Chrome DevTools Protocol, n8n.
Why Clients Hire Me:
• 22+ successful Upwork contracts
• 7+ years building NLP and AI systems
• Production experience with enterprise-scale LLM applications
• Strong background in English and Arabic NLP
• Experience delivering systems used in real business environments
If you need an AI system that saves time, automates complex workflows, or turns your data into actionable insights, I'd be happy to discuss your project.
I'll give you an honest assessment of what's feasible, the expected timeline, and the most practical approach—even if that means recommending a simpler solution.
$120/hr
100%
Job Success
$300K+ earned
Offers consultations
Start of list.
End of list.
Most "AI chatbots" deflect tickets. I build the ones that resolve them.
I'm a Zendesk + AI engineer with 13+ years at one intersection: support platforms on one side, integrations and applied AI on the other. Seven of those years were at Upwork itself — leading engineering for the Help Center, Support Chatbot, and Zendesk integration. You can see that history right here on my profile.
Results from that work, featured in Zendesk's own customer marketing:
- Cut Zendesk API usage 60% by redesigning the integration to call Upwork's GraphQL API directly
- Lifted chatbot self-service resolution 58% by building the GraphQL ↔ Zendesk Guide context layer
- Pioneered the three-way messaging model for dispute cases
What I do for clients now:
- Zendesk engineering: integrations, apps, Guide/Help Center architecture, API cost reduction
- AI in support, done right: RAG over your help-center and ticket data, LLM integration (OpenAI, Claude, Gemini, local models), model routing for cost control, human handoff that works
- Implementation of vendor AI tools (I've integrated Forethought and Ada in production)
- Support-stack audits: find why the chatbot plateaued or the API bill climbed, with a prioritized fix plan
Recent: a paid AI-automation build integrating ChatGPT/Claude APIs (5 stars — see my work history), and my own AI support assistant currently in beta: LiteLLM gateway, hybrid RAG on PostgreSQL + pgvector, Zendesk ticketing handoff.
Stack: Zendesk API · GraphQL · REST · OAuth2 · TypeScript/Node.js · Python · OpenAI/Claude/Gemini/Ollama · LiteLLM · Langfuse · MCP · n8n · PostgreSQL + pgvector · Docker · AWS/GCP
If your support stack has hit its limits, send me a message with what's breaking — I'll tell you honestly whether I'm the right person to fix it.
Associated with
adommo
$500K+
earned
$20/hr
$0 earned
Start of list.
End of list.
Most AI agents die as demos. Yours won't.
I build AI agents and automations for small businesses — the kind that keep working on Monday morning: with cost caps, automatic model fallbacks, monitoring, and testing built in from day one.
What I build:
🔹 WhatsApp/Website AI assistants — answer customer questions, qualify leads, book appointments, chase pending documents (24/7, in your customer's language)
🔹 Back-office automation — connect your CRM, Google Sheets, email and WhatsApp so data moves without anyone copy-pasting
🔹 Custom AI workflows — invoice processing, lead follow-up, price monitoring, report generation
Why me and not the next "AI expert":
Everything I build runs on an your end to end stack or even open-source agent stack published on GitHub — LLM gateway with automatic fallback (your bot doesn't go down when one AI provider does), daily budget caps (no surprise bills), full tracing of every AI call, and automated tests before anything ships.
I come from 15+ years in product and technology leadership, so I start with your business problem, not the tech. You'll get plain-English scoping, a fixed price and a working system — not a science project.
Recent builds:
✅ RBI Circular Intelligence Agent - Production RAG (Retrieval-Augmented Generation) agent that ingests and indexes RBI circulars in real-time, enabling compliance officers, bankers, and financial professionals to query regulatory guidelines in plain English and get accurate, cited answers instantly. Uses custom web-scraper, Hybrid search, Structural chunking with circular identity anchoring, Circular lineage tracking etc. Stack: Python · Playwright · Supabase · pgvector · n8n · LiteLLM · Groq · Gemini · Redis · Cloudflare · Oracle Cloud
✅ Document Chaser — reads a CRM for pending documents, follows up with customers on WhatsApp, validates what they send, updates the CRM. Zero manual chasing.
✅ Grocery Price Agent — send a shopping list on WhatsApp, get the cheapest option per item across three delivery apps in seconds.
✅ Error Handler — any workflow failure triggers an instant WhatsApp alert, so problems get fixed before customers notice.
How we'll work:
Free 20-minute call — you describe the manual work eating your team's time
I send a one-page scope: what gets automated, what it costs, when it ships
Build, test, demo. You pay for outcomes, not hours of "exploration."
If your team is short on people and drowning in repetitive work, that's exactly the problem I solve. Send me a message describing the task you'd most like to never do again.
$90/hr
$0 earned
Start of list.
End of list.
I am a Full-Stack Engineer and AI Systems Architect with over 8 years of experience building high-performance web applications, scalable backend microservices, and enterprise-grade LLM infrastructure. Having worked across fast-moving Y-Combinator startups and complex enterprise platforms, I specialize in shipping production-ready systems from architecture to deployment.
Whether you need to integrate multi-provider AI routing into your existing SaaS, build a secure high-concurrency web application from scratch, or modernize your frontend architecture, I bring an AI-native development workflow (leveraging Cursor, MCP tool orchestration, and agentic pipelines) that accelerates feature delivery by 3–5x without sacrificing code quality or security.
🚀 What I Can Build For You:
Production AI & LLM Infrastructure: Custom OpenAI-compatible gateways, multi-provider model routing (OpenAI, Gemini, Bedrock, Ollama) via LiteLLM, cost/latency fallbacks, and org-scoped quota tracking.
AI Security & Guardrails: Enterprise 3-tier guardrail systems (regex → NLP semantic analysis → LLM judge) protecting against prompt injection, jailbreaks, and PII/secrets exfiltration.
Modern Full-Stack Web Applications: High-complexity frontends using React 19, TypeScript, Vite, Tailwind, and ReactFlow, backed by robust Node.js, Express, Python, Django, or FastAPI REST APIs.
Developer Tooling & Code Analyzers: AI-native SAST engines, custom AST/tree-sitter parsers, and LSP-aware code remediation copilots supporting 15+ languages.
E-Commerce & Live-Streaming Platforms: Custom payment workflows (Stripe integration), real-time live-streaming fitness portals, and scalable telehealth booking systems.
🏆 Key Career Highlights & Impact:
Zero False-Positive AI Security Engine: Developed an AI-native SAST engine (Venus) across 15+ languages that achieved 18/18 true positives with 0 false positives on industry benchmarks.
Massive User Scale: Created a viral ranking predictor application that successfully handled 80,000+ users within 14 days, serving as the company's top lead-generation engine.
Massive Frontend Modernization: Architected and migrated an enterprise API security platform (~86k lines of code, 50+ routed views) from Create React App to Vite 7, cutting new feature rollout times by ~70%.
High-Conversion Campaigns: Engineered gamified promotional engines that acquired 5,000+ new unique users and boosted Play Store app reviews by 504%.
🛠 Technical Stack:
Frontend: React.js (v19), TypeScript, Vite 7, Tailwind CSS, Material UI (MUI), Zustand, ReactFlow, Flutter
Backend & APIs: Node.js, Express.js, Python, FastAPI, Django, REST APIs, MCP (Model Context Protocol), JSON-RPC
AI & Machine Learning: OpenAI, Vertex AI/Gemini, Amazon Bedrock, Ollama, LiteLLM, LangChain, LangGraph, Vector Embeddings
Database & Infrastructure: PostgreSQL, MySQL, Redis, Docker, Kubernetes, AWS, GCP, Socket. io, Stripe
Let’s connect to discuss how I can help architect, build, or scale your next engineering product.
$75/hr
$0 earned
Start of list.
End of list.
I deploy open-weight LLMs (Llama, Qwen, DeepSeek, Mistral) on GPUs you
control — on-prem or rented — and take them to production. You get an
OpenAI-compatible API inside your own perimeter: tuned for your latency
and throughput targets, load-balanced, monitored, and with no data ever
leaving your network.
WHAT I DO
- Production deployment of vLLM / SGLang / llama.cpp on NVIDIA GPUs
(H100, H200, B200) — single-node and multi-GPU setups with tensor and
pipeline parallelism
- Inference tuning to hit your SLA: time-to-first-token, tokens/sec,
concurrency, KV-cache sizing, quantization trade-offs (FP8, AWQ, GPTQ)
- LLM gateway layer: LiteLLM or custom proxy — routing between local
and cloud models, load balancing across heterogeneous GPUs, failover,
per-team budgets and rate limits
- Fine-tuning (LoRA / QLoRA) and taking the resulting model into
production serving
- Observability: Prometheus + Grafana dashboards for the metrics that
matter (TTFT, throughput, queue depth, GPU utilization)
Backend foundation: Python (FastAPI, gRPC), microservices, Redis,
Celery, relational and vector databases, Docker, Linux — so the
inference stack I hand over integrates cleanly with your application,
not just runs in isolation.
WHY SELF-HOSTING
Cloud LLM API costs scale linearly with usage. At steady volume,
serving open-weight models on your own or rented GPUs cuts cost per
million tokens several times over — and keeps sensitive data inside
your infrastructure (GDPR, HIPAA, EU AI Act compliance).
HOW WE CAN START
1. Sizing & PoC (fixed price) — GPU sizing math, engine selection,
cost model: cloud API vs self-hosted for your workload
2. Production deployment — full stack from bare GPUs to a monitored
OpenAI-compatible endpoint
3. Support retainer — model upgrades, capacity planning, incident
response
Send me your model, expected load, and hardware (or budget) — I'll
reply with a concrete architecture proposal and numbers, not a
generic pitch.
$100/hr
$0 earned
Offers consultations
Start of list.
End of list.
Applied AI engineer who ships production LLM systems end-to-end — RAG, agentic assistants, and real-time voice — for a regulated fintech vertical, not demos. Over the past two years, solo across 20+ repositories, I designed, built, and operate an AI-native Arizona tax-compliance platform (CactusComply) and a reinforcement-learning options-trading system (CactusQuant): a pgvector RAG copilot with server-side authorization guards, a 10-tool GPT-4o agent with hallucination and excessive-agency suppression, an OpenAI Realtime + Twilio voice agent at sub-1s latency, and an Arizona Department of Revenue SOAP/XSD e-file gateway built from scratch with an immutable audit trail. I pair hands-on applied-AI engineering — evals, observability (OpenTelemetry / Phoenix), and cost-aware model routing (Claude Opus→Sonnet→Haiku; o3→o4-mini→gpt-4.1-mini via LiteLLM) — with 15+ years translating business requirements into shipped systems at Choice Hotels, City of Phoenix, and American Express, making me a rare builder who can both architect the system and close the last mile with the customer.
$40/hr
$800+ earned
Start of list.
End of list.
I deploy AI systems on your own infrastructure — RAG chatbots, n8n and Dify
agent workflows, LiteLLM gateways, Qdrant vector search — so you own the
data and control the API bill instead of renting someone else's.
I've built and run this stack in production for my own product (Unlkr, a
content monetization platform I operate on Oracle Cloud), not just for
clients. That means I've already hit the failure modes: rate limits,
context bloat, vector drift, cold-start costs, and the moment your n8n
workflow silently stops firing at 3am.
WHERE I'M USEFUL
- You want a chatbot that answers from YOUR documents, accurately
- You're paying too much for GPT-4 calls and want a routing/fallback layer
- You have a manual process that should be an automated workflow
- You need the app around the AI, not just the AI — auth, billing,
dashboard, mobile client
STACK
AI/Automation: n8n, Dify, LiteLLM, Open WebUI, Qdrant, LangChain, AWS
Bedrock, OpenAI, Claude
Backend: Node.js, Express, Python, MongoDB, PostgreSQL, Redis
Frontend: React, Angular, Vue, TypeScript
Mobile: Kotlin (native Android), Swift (iOS), React Native
Infra: AWS, Azure, Oracle Cloud, Docker, Nginx, Linux, WireGuard
I'd rather turn down a project than take one I can't do well. If it's
outside my lane, I'll say so in the first reply.
United Arab Emirates
$35/hr
$0 earned
Start of list.
End of list.
I set up and maintain self-hosted infrastructure and AI automation.
What I do:
- Self-hosted services in Docker — n8n, Nextcloud, Vaultwarden, Immich, media
servers. Reverse proxy, SSL, Cloudflare Tunnel, backups.
- Proxmox — cluster setup, VM and LXC provisioning, GPU passthrough, storage
and networking, migration from bare metal.
- AI automation — n8n workflows with LLM steps, RAG over your documents, Telegram and Slack bots. I run a two-agent system on my own hardware — an advisor and a sandboxed executor — built on the open-source Hermes Agent framework with my own plugin layer: semantic memory, safety gates, and hard cost caps ($20/month for 7.8k messages).
- LLM cost optimization — routing between cheap and expensive models by task
complexity. Most integrations overpay because everything goes to the top model.
- Linux and networking — TCP/IP, NAT, iptables, DNS, Nginx. Debugging things
that are broken for no obvious reason.
I run a Proxmox cluster with GPU passthrough at home. Everything I offer here,
I have already built and broken on my own hardware.
I am new to Upwork, not new to this work, and my rate reflects that.
Dubai, GMT+4. English and Russian.