I build production multimodal AI systems that turn images, documents, and business data into reliable outputs. My recent work includes reference-based visual reasoning and image-generation pipelines with deterministic answer and uniqueness validation, SVG/PNG rendering, bilingual content, and secure FastAPI/Docker APIs.
I also build production AI agents and RAG systems that connect LLMs to private documents, databases, APIs, and enterprise workflows. I own the lifecycle from requirements and architecture through structured model outputs, programmatic validation, deployment, and handover.
SELECTED RESULTS
• Built and deployed Visual Reasoning AI Studio, a reference-based multimodal AI question-generation and deterministic verification platform with a live synthetic showcase in my Upwork portfolio.
• Delivered non-verbal image reconstruction and .NET/API integration for an education workflow.
• Built a real-time multi-agent cybersecurity platform that improved incident-response speed by 40%.
• Delivered an image-to-data pipeline with 95%+ accuracy across complex document types.
• Built natural-language-to-SQL assistants for non-technical teams.
WHAT I BUILD
• Multimodal AI, computer vision, image processing, and document intelligence
• Reference-based image and diagram generation with structured visual outputs
• AI agents and multi-agent workflows using LangGraph, LangChain, AutoGen, and LlamaIndex
• Production RAG, extraction, classification, and structured-output pipelines
• FastAPI, Docker, MCP, API, ERP, and business-workflow integrations
• Azure/AWS deployment, evaluation, security, and human-approval controls
CORE STACK
Python, FastAPI, OpenAI, Azure OpenAI, AWS Bedrock, LangGraph, LangChain, LlamaIndex, AutoGen, MCP, computer vision, OCR, SVG/PNG rendering, PostgreSQL, Docker, Kubernetes, Databricks, and Kafka.
WHY CLIENTS HIRE ME
• Production systems, not isolated prompt demos
• Full-stack AI ownership from architecture through deployment
• Deterministic validation for correctness, uniqueness, and output consistency
• Clear communication, documented decisions, and realistic estimates
If you are building a multimodal AI, computer-vision, visual-reasoning, RAG, or enterprise LLM application, send me the current system and desired outcome. I will propose the smallest reliable path to production.
Ajesh M.
Senior Generative AI Engineer | LLM, RAG & Agentic AI | Trading System
Bengaluru, India
$20/hr$20 per hour5.0 (10) 14 jobs $1K+ total earnings
I build production-grade Generative AI systems — RAG pipelines, LLM-powered agents, and document-intelligence tools that turn unstructured data into measurable business results.
Over the past 4+ years, I've designed and shipped GenAI solutions for a Fortune 500 semiconductor client, including:
A RAG-based compliance system (Azure OpenAI + FAISS, 5,000+ indexed rules) that cut manual document review time by 70% and hit 92% accuracy in flagging policy violations.
A Graph RAG + Vector RAG hybrid multi-agent workflow (Neo4j + LangGraph) enabling multi-hop reasoning across technical documents — well beyond what flat semantic search can do.
LLM-driven ticket summarization & classification pipelines that reduced manual triage effort by 80%.
Vision-Language Model (VLM) extraction pipelines combining OCR, Azure Document Intelligence, and GPT-4o to pull structured data out of engineering drawings and patents.
I work across the full GenAI stack — from prompt engineering and function calling/tool use to vector databases, hybrid search & reranking, and scaled deployment on Azure/Databricks with FastAPI + PostgreSQL backends. I hold an MTech in Data Science from IIT Palakkad and have a Computer Vision engineering background (YOLOv6/7/8, OCR, video analytics), so I'm equally comfortable when a project blends GenAI with vision or classic ML.
What I can help you build:
RAG pipelines & knowledge assistants (FAISS, Pinecone, Neo4j/GraphRAG)
LLM agents & multi-step agentic workflows (LangChain, LangGraph, function calling)
Document AI: extraction, classification, and compliance automation from PDFs/DOCX/scanned docs
LLM-powered automation: summarization, classification, structured data extraction
Integrations with Azure OpenAI / OpenAI API, FastAPI backends, and vector stores
Vision-language and OCR-based extraction pipelines
I care about shipping things that actually reduce manual work and hold up in production — not just demos. Happy to start with a short paid trial task so you can see the quality of my work before committing to a larger engagement.
Let's talk about what you're trying to build.
Skills (tags to add on Upwork)
Generative AI, LLM, RAG (Retrieval-Augmented Generation), LangChain, LangGraph, Agentic AI, Prompt Engineering, Azure OpenAI, GPT-4o, Function Calling, Vector Databases, FAISS, Neo4j, GraphRAG, Semantic Search, FastAPI, Python, Document AI, OCR, Computer Vision, Hugging Face Transformers, PyTorch, Machine Learning, PostgreSQL, Databricks, Docker
Portfolio / Project entries to add
1. RAG-Based Technical Document Compliance Checker FastAPI · Azure OpenAI · FAISS · Azure Blob · Databricks Built a RAG system that evaluates technical documents against a 5,000+ rule policy knowledge base, cutting manual review time by 70% with 92% violation-detection accuracy. Extended into a Graph RAG + Vector RAG hybrid multi-agent architecture (Neo4j + LangGraph) for multi-hop reasoning over relationship-heavy compliance queries.
2. NC Ticket Summarization & Classification Azure OpenAI · FastAPI · PostgreSQL · Databricks Jobs Automated non-conformance ticket summarization and multi-class classification via prompt engineering, deployed as a daily Databricks job, reducing manual triage effort by 80%.
3. LCO Patent Assist — Patent Analysis Tool Azure OpenAI · Azure Document Intelligence · OpenCV · DOCX Processing Hybrid patent analysis tool combining OCR, rule-based parsing, and GPT-4o for structured extraction of claims, descriptions, and figures from patents, with intra-document similarity search for claim-to-description mapping.
4. Warehouse Path Optimization & Analytics TSP · OpenCV · Python · Heatmap Analytics TSP-based path optimization for warehouse picking, reducing average path length by 25%, with heatmap tracking of high-traffic zones.
Atharva P.
Expert Computer Vision Engineer | Behavioural Video Analytic
Gandhinagar, India
$35/hr$35 per hour4.8 (4) 10 jobs $4K+ total earnings
Experienced in building scalable real-time video intelligence systems — from traffic analytics to sports performance tracking, OCR, Speech Processing, and LLMs. With 5 years of expertise in real-time video analytics, document automation, and AI-driven insights, I leverage top technologies like Pytorch, TensorFlow, Deepstream, OpenCV, DeepSpeech, and Hugging Face to deliver scalable, intelligent systems. Let’s transform your vision into reality! 🚀
🔧 Tech Stack :
💻 Computer Vision and Video Analytics
🔹 OpenCV
🔹 TensorFlow
🔹 PyTorch
🔹 YOLO (You Only Look Once)
🔹 MMPose (Pose Estimation)
🔹 GStreamer (Video Processing)
📄 OCR and Document Processing
🔹 Tesseract OCR
🔹 EasyOCR
🔹 Pytesseract
🔹 Google Cloud Vision API
🔹 AWS Textract
🤖 RAG and LLMs (Large Language Models)
🔹 OpenAI GPT-3/4
🔹 Hugging Face Transformers
🔹 LangChain
🔹 Pinecone (Vector Database)
🔹 LLamaIndex
🔹 Ray (Distributed Computing for LLMs)
🎤 Speech Processing and Speech-to-Text
🔹 Google Speech-to-Text
🔹 Amazon Transcribe
🔹 Microsoft Azure Speech Service
🔹 DeepSpeech (Mozilla)
🔹 Coqui
🔹 Whisper (OpenAI)
⚙️ Deep Learning & Hardware Acceleration
🔹 CUDA (GPU Acceleration)
🔹 TensorRT
🔹 NVIDIA DeepStream SDK
🔹 OpenVINO
🔹 Jetson Nano (Edge Computing)
🔹 EdgeTPU
🌐 Cloud, Deployment, and DevOps
🔹 Docker
🔹 Kubernetes
🔹 AWS S3
🔹 GCP AI Platform
🔹 Jenkins
🛠️ Previous Projects :
1. AI-Powered Smart City Traffic Management System 🚦🚘
2. Smart Factory Monitoring and Safety System 🏭🚧
3. Voice-Activated Medical Transcription System 🩺📝
4. Multimodal Document Understanding and Automation 📄🤖
5. Personalized Virtual Assistant for Customer Support 🎙️💬
6. Real-Time Sports Analytics and Player Performance Tracker ⚽📊
7. Smart Home Surveillance and Intruder Detection 🏡🎥
8. AI-Powered Healthcare Data Extraction and Analysis 🏥📈
9. AI-Driven Financial Document Automation 💼📑
10. Real-Time Video Surveillance for Compliance in Factories 🎥🏭
🚀 Let’s Build the Future Together!
With my extensive experience in deploying AI-driven solutions across multiple domains, I'm ready to take your projects to the next level. From Computer Vision to LLMs and OCR, I bring a passion for innovative technology and a track record of delivering high-impact solutions. Let’s transform your ideas into reality and drive success together!
Disha J.
Computer Vision Engineer | AI/ML & Deep Learning Expert
Ahmedabad, India
$35/hr$35 per hour4.8 (146) 257 jobs $200K+ total earnings
With 7+ years of experience in AI/ML, I focus on building solutions that are not just technically sound but also practical and valuable in real-world scenarios. My work is driven by research, experimentation, and selecting the right technologies for each unique problem.
🔹 Where I Add Value
I have strong hands-on experience in Computer Vision / Machine Vision, working on problems such as:
* Object detection and recognition
* Multi-camera tracking and digital twin systems
* Face recognition and human pose estimation
* Edge AI and real-time inference
* Inspection systems and anomaly detection
* OCR and text extraction
* Data preparation including segmentation, augmentation, and synthesis
🔹 Machine Learning Expertise
Depending on the problem, I apply a wide range of ML techniques:
* Supervised Learning: Classification, Regression
* Unsupervised Learning: Clustering, Dimensionality Reduction
* Semi-supervised learning
* Reinforcement learning
🔹 NLP & AI Systems
Alongside vision, I’ve worked on intelligent systems involving:
* GPT-based models (GPT-3, GPT-4) and LangChain workflows
* Text summarization and content generation
* Recommendation systems (feature-based, user-based, hybrid)
* Sentence similarity and semantic analysis
* Conversational AI and dialog systems
* Speech processing (ASR, speech synthesis)
* Machine translation
🔹 Deep Learning Applications
I’ve applied deep learning across multiple real-world use cases:
* Multi-camera tracking systems
* Autonomous drones
* Chatbots and service bots
* Predictive analytics solutions
I’ve also worked closely with startups and businesses to build MVPs and POCs, helping them validate ideas and move faster toward production.
✔️ Open to signing NDAs for confidential projects
✔️ Additional demos and project details available on request
If you’re looking to build something in AI or want to explore a new idea, feel free to reach out happy to discuss and contribute.
Best regards
Disha
How it works
Post a job for freePost a job
Tell us what you need. Create your own job post or generate one with AI then filter talent matches.
Hire top talent fast
Consult, interview, and hire quickly, so you can meet the freelancers you're excited about.
Collaborate easily
Use Upwork to chat or video call, share files, and track project progress right from the app.
Payment simplified
Manage payments in one place with flexible billing options. Only pay for approved work, hourly or by milestone.
Don't just take our word for it
“Upwork provides an umbrella-level of security. I can see a talent’s work history and ratings. I can hold payments in escrow. I can communicate through Upwork Messages instead of working through my email address.”
KD
Kim Darling
Verified
Emerald Tiger
“Upwork is the best platform to hire skilled professionals when we're not looking for a full-time employee. All the companies in our portfolio use Upwork to find talent across a wide range of fields.”
DM
David Merry
Verified
Kinetic Investments
“Our very specific requirements can be a challenge—With Upwork, we’re able to access a bigger community to ensure the success of our projects.”
Top interview questions to help you hire the right Computer Vision Engineers, faster.
How do I hire a Computer Vision Engineer in India on Upwork?
You can hire a Computer Vision Engineer in India on Upwork in four simple steps:
Create a job post tailored to your Computer Vision Engineer project scope. We'll walk you through the process step by step.
Browse top Computer Vision Engineer talent on Upwork and invite them to your project.
Once the proposals start flowing in, create a shortlist of top Computer Vision Engineer profiles and interview.
Hire the right Computer Vision Engineer for your project from Upwork, the world's largest work marketplace.
At Upwork, we believe talent staffing should be easy.
How much does it cost to hire a Computer Vision Engineer?
Rates charged by Computer Vision Engineers on Upwork can vary with a number of factors including experience, location, and market conditions. See hourly rates for in-demand skills on Upwork.
Why hire a Computer Vision Engineer in India on Upwork?
As the world's work marketplace, we connect highly-skilled freelance Computer Vision Engineers and businesses and help them build trusted, long-term relationships so they can achieve more together. Let us help you build the dream Computer Vision Engineer team you need to succeed.
Can I hire a Computer Vision Engineer in India within 24 hours on Upwork?
Depending on availability and the quality of your job post, it's entirely possible to sign up for Upwork and receive Computer Vision Engineer proposals within 24 hours of posting a job description.