I build production AI pipelines that turn messy documents into structured data.
8 years in Python and data engineering. I specialize in document processing systems — the kind where you throw in a stack of scanned PDFs and get clean, validated JSON out the other side.
## What I actually build
### LLM-Powered Extraction Pipelines
Multi-agent systems using Claude, GPT-4, and GPT-4o with structured output enforcement (Instructor, Pydantic). I design domain-specific agents that classify documents, extract fields, validate data, and reconcile conflicts across sources. Not prompt-and-pray — schema-enforced, retry-aware, production-grade.
### OCR & PDF Processing
Amazon Textract (async API), PyMuPDF, PDFPlumber, Camelot, Azure Document Intelligence. I handle the real-world mess — scanned documents with poor quality, mixed orientations, concatenated multi-document PDFs that need splitting and boundary detection. Table extraction, form parsing, layout-aware chunking.
### Serverless & Event-Driven Architecture
AWS Lambda, SNS, SQS, Step Functions, EventBridge, ECS/Fargate. I build pipelines where S3 uploads trigger processing chains — fan-out to parallel extractors, dead-letter queues for failures, CloudWatch monitoring for quality metrics.
### Healthcare Document Systems
Currently running a production pipeline that processes medical records daily — licenses, clinical notes, immunization records, provider credentialing documents. Full pipeline: PDF ingestion → OCR → document classification → LLM extraction → deduplication → normalization → structured output. Built with HIPAA considerations.
### Legal Document Systems
Production pipeline for UK cost-assessment workflows — ingests multi-document bill bundles (PDFs, scans, Excel ledgers, Outlook .msg), classifies them, extracts disbursements, receipts, invoices, and party data, reconciles figures across sources with a 5-case matching algorithm, and generates court-compliant inter-partes narratives. Every extracted figure traces back to its source page for audit defensibility.
### Speech-to-Text Pipeline
- AWS Transcribe Medical (streaming WebSocket + batch S3 jobs) with HIPAA-compliant configuration
- Deepgram integration as an alternate ASR backend
- Speaker diarization (up to 30 speakers), stereo channel separation, custom medical vocabularies, specialty-specific vocab lists (e.g. orthopedics)
- Multi-threshold confidence analysis for dot-phrase and quick-text detection
### LLM Clinical Note Generation
- Instructor + Pydantic schema-enforced structured outputs (Claude / GPT-4 / Gemini)
- Prompt Studio: a template-authoring system where clinicians design reusable note templates with variables, segment bindings, and inline citations back to the raw transcript
- Two output modes — Ambient (multi-category SOAP-style synthesis) and Dictation (verbatim with cleanup)
- DOCX template engine with inline placeholder tokens, dual-archive export, and anchor-based fill for exact-fidelity Word output
### AI Agent Orchestration
LangChain, LangGraph, MCP servers, RAG systems, function calling. I build agents that do real work — not chatbots that summarize, but extraction systems that produce validated, structured data from unstructured sources.
## Anthropic API Cost Optimization
I help teams cut Claude/LLM API spend in half on production deployments without sacrificing product quality.
I work across the levers that actually move the bill:
- **Prompt caching** — designed around real traffic patterns, not just toggled on
- **Model tiering** — Sonnet where reasoning matters, Haiku everywhere else
- **Context discipline** — lazy-loading tools, skills, and instructions instead of dumping everything upfront
- **Tool-call efficiency** — fewer round-trips, smaller payloads
I've done this on a medical-document extraction pipeline and a LangGraph trading agent — both running in production with real users and real budgets. If you've inherited an over-budget Claude deployment, I know where to look first.
## Tech I use daily
Python, FastAPI, Anthropic/OpenAI/Gemini, Instructor, LangChain, LangGraph, Textract, PyMuPDF, PostgreSQL, Redis, S3, Lambda, SNS/SQS, Docker, Pydantic.
20+ projects delivered on Upwork. I'm strongest when the problem involves turning unstructured documents into clean, structured data at scale.
Python
LLM Prompt
OpenAI API
LangChain
OCR Algorithm
AWS Lambda
Azure OpenAI Service
PDF
Data Extraction
OCR Software
OpenCV
Azure Cognitive Services
Retrieval Augmented Generation
Information Retrieval
Amazon Bedrock
Legal
Medical Report
Tuan T.
Quang Ngai, Vietnam
$40/hr
5.0
113 jobs
Tired of unreliable weather guesses costing you money? I deliver precise, actionable forecasts that help businesses mitigate risks and capitalize on weather patterns. With 5 years in operational meteorology, I bridge the gap between complex atmospheric science and your real-world needs.
1. Why My Forecasts Win?
• 90% accuracy rate on high-impact weather events.
• Develop optimized weather forecasting systems for HPC/cloud (EC2), built for 24/7 reliability and easy scaling.
• Turn NOAA data, WRF models, weather radar and satellite imagery into clear business recommendations.
• AI-enhanced modeling that spots risks standard forecasts miss.
• Beyond traditional forecasts: using atmospheric science to predict nature's light shows.
• Making the unpredictable predictable: advancing clear air turbulence forecasts in aviation.
• Assess extreme weather trends with NEX-GDDP across RCP4.5 and RCP8.5.
2. Skills:
• Numerical weather prediction models (NWP): WRF, GFS, ECMWF, ICON-D2, AROME,…
• Python-driven data analysis
• Visualization: weather map, vertical cross section, Skew-T,… by using matplotlib, d3.js,…
• Converting and processing meteorological data format file (NetCDF4, GRIB, HDF5,…)
Data Analysis
Python
Data Visualization
Data Processing
JavaScript
Microsoft Power BI
Modeling
Mathematics
Web Scraping
Geospatial Data
Image Processing
MATLAB
Bash Programming
Climate Science
Web Development
Long T.
Ho Chi Minh City, Vietnam
$25/hr
5.0
6 jobs
I'm a bachelor of Computer Science focused on Machine Learn/Data Science and Web Development. Open for work that's data related (ML, AI, Data Processing) or web related (application, design, scaping).
Recently I've been working on customizing Image Processing DL model such as Pix2pix, Pix2pixHD, and Image exposure fusion model TransMEF, MEF-NET, EMEF for custom dataset of up to 500GBs.
Right now I'm building a Data Gathering platform for Air Writing Recognition to support a Research Paper.
- Web related: JS, TS, HTML, CSS, PHP, React, Ruby On Rails
- Database: mainly MySQL, MongoDB
- Data: EDA, Large Data Processing, Image Fusion, Image Processing, Adversarial DL Model, Image Fusion.
- Communication: I'm confident in my management/leadership skill. Personally I value progress update as a mean to be transparent. So rest assured that I will be in touch.
Python
Web Application
Node.js
React
JavaScript
TypeScript
MySQL
SQL
MongoDB
Ruby on Rails
pandas
NumPy
TensorFlow
Exploratory Data Analysis
Version Control
Cuog N.
Hanoi, Vietnam
$15/hr
5.0
2 jobs
AI Engineer with 5+ years of expertise in computer vision and deep learning. Experienced in leading technical teams, building production ML systems, and implementing MLOps best practices. Passionate about leveraging cutting-edge AI technologies to solve complex business challenges.
Machine Learning
Computer Vision
Git
LLM Prompt Engineering
Retrieval Augmented Generation
Toan Viet D.
Bien Hoa, Vietnam
$60/hr
5.0
41 jobs
ABOUT ME
I'm a Kaggle competition master and a senior AI/ML engineer with over 8 years of experience. My expertise spans Agentic AI Systems, Multi-Agent Architectures, Large Language Models (LLMs), complex Agentic Retrieval-Augmented Generation (RAG) chatbots, Stable Diffusion, Computer Vision (CV), Natural Language Processing (NLP), Machine Learning, and Deep Learning.
HOW I WORK
I run multi-agent Claude Code and Cursor workflows: a task master coordinating 4-5 focused agents handling implementation, testing, and refactoring in parallel. High throughput without losing control.
I pair this with production-ready engineering: unit tests, clear module boundaries, scalable architecture. I focus on shipping reliable systems that actually get used.
EXPERTISE
Agentic AI & Multi-Agent Systems
I build production-grade AI agents across domains (education, tutoring, purchase advisors, real estate, and more), powered by state-of-the-art frameworks and techniques including LangChain, Agno, MCP, advanced agent memory/skills architectures, Langfuse, and LangSmith.
Agentic RAG & LLM Expertise
Built 50+ production RAG systems with advanced architectures such as Agentic RAG and GraphRAG (LightRAG).
Core strengths:
Open-source & self-hosted LLMs (vLLM, Hugging Face, llama.cpp) and closed-source providers (OpenAI, OpenRouter, Gemini, Claude).
Multi-step reasoning, query expansion, dynamic chunking, hybrid search, re-ranking.
Vector DBs: Qdrant, Pinecone, Chroma, PgVector, AWS OpenSearch.
Data sources: CSV, Excel, MongoDB, websites, PDFs/Office files, Google Drive, Notion, Confluence, Email, Microsoft SharePoint.
CloudOps Architecture
Design, deploy, and self-host AI systems across AWS, GCP, Azure, SageMaker, and RunPod.
Focused on scalable, stable, and cost-efficient AI infrastructure.
LLMs Fine-tuning & Deployment
Fine-tuned LLMs from 2B–72B parameters across domains and languages. Applied optimization techniques achieving 3–6× inference speedups.
Main author of the 2k+ stars efficient LLM fine-tuning project: githubcom/stochasticai/xTuring
Don't hesitate to reach out—I am committed to delivering the highest quality of work with cutting-edge AI technologies.
Data Science
Natural Language Processing
AI Agent Development
Large Language Model
Chatbot
Stable Diffusion
Model Tuning
Multimodal Large Language Model
ChatGPT
LLM Prompt
Computer Vision
PyTorch
Artificial Intelligence
LangChain
OpenAI API
Chanh Phuong Trang V.
Ho Chi Minh City, Vietnam
$10/hr
5.0
1 jobs
I specialize in transforming complex data into actionable insights as a Senior Data & BI developer with over 4 years of experience. My expertise lies in designing data warehouses, building robust ETL pipelines, and creating executive dashboards that drive strategic decision-making across departments. With a solid foundation in Natural Language Processing, I leverage my research skills to enhance data interpretation, focusing on areas such as biomedical text mining. I thrive on delivering tailored solutions that support informed business choices. If you're seeking a data expert who can bridge the gap between data analysis and business intelligence, let's connect and discuss how I can help elevate your projects.
ETL Pipeline
Data Warehousing & ETL Software
Microsoft Power BI
Data Analytics
Analytics Dashboard
Microsoft Power BI Data Visualization
Business Intelligence
Data Analytics Expertise
Dashboard
Advanced Analytics
How it works
Post a job for freePost a job
Tell us what you need. Create your own job post or generate one with AI then filter talent matches.
Hire top talent fast
Consult, interview, and hire quickly, so you can meet the freelancers you're excited about.
Collaborate easily
Use Upwork to chat or video call, share files, and track project progress right from the app.
Payment simplified
Manage payments in one place with flexible billing options. Only pay for approved work, hourly or by milestone.
Don't just take our word for it
“Upwork provides an umbrella-level of security. I can see a talent’s work history and ratings. I can hold payments in escrow. I can communicate through Upwork Messages instead of working through my email address.”
KD
Kim Darling
Emerald Tiger
“Upwork is the best platform to hire skilled professionals when we're not looking for a full-time employee. All the companies in our portfolio use Upwork to find talent across a wide range of fields.”
DM
David Merry
Kinetic Investments
“Our very specific requirements can be a challenge—With Upwork, we’re able to access a bigger community to ensure the success of our projects.”
Top interview questions to help you hire the right Data Scientists, faster.
How do I hire a Data Scientist in Vietnam on Upwork?
You can hire a Data Scientist in Vietnam on Upwork in four simple steps:
Create a job post tailored to your Data Scientist project scope. We'll walk you through the process step by step.
Browse top Data Scientist talent on Upwork and invite them to your project.
Once the proposals start flowing in, create a shortlist of top Data Scientist profiles and interview.
Hire the right Data Scientist for your project from Upwork, the world's largest work marketplace.
At Upwork, we believe talent staffing should be easy.
How much does it cost to hire a Data Scientist?
Rates charged by Data Scientists on Upwork can vary with a number of factors including experience, location, and market conditions. See hourly rates for in-demand skills on Upwork.
Why hire a Data Scientist in Vietnam on Upwork?
As the world's work marketplace, we connect highly-skilled freelance Data Scientists and businesses and help them build trusted, long-term relationships so they can achieve more together. Let us help you build the dream Data Scientist team you need to succeed.
Can I hire a Data Scientist in Vietnam within 24 hours on Upwork?
Depending on availability and the quality of your job post, it's entirely possible to sign up for Upwork and receive Data Scientist proposals within 24 hours of posting a job description.