Hire the Best Apache Spark Engineers
in Vietnam

More than 3,000 reviews on G2
Rating is 4.5 out of 5.
4.5/5
of Upwork by G2 peer reviewers
Nghi L.

Ho Chi Minh City, Vietnam

$25/hr
5.0
53 jobs

⏰ Available 24/7 – Long-term & High-impact Projects Hi, I’m Nghi, a Senior Data Engineer and Data Architect with a strong backend foundation, now focused on building high-performance analytics platforms, explainable data pipelines, and production-grade cloud architectures. I help companies transform unreliable, slow, or opaque data systems into scalable, well-documented, and business-trustworthy platforms. 🧠 WHAT I SPECIALIZE IN 🏗️ Data Architecture & Platform Design - Designing modern lakehouse & warehouse architectures - dbt-first analytics engineering with testing, freshness & lineage - Event-driven and batch hybrid pipelines - Data quality frameworks & SLA monitoring - Customer-facing data explainability systems Tools: dbt, Dagster, Airflow, Spark, Kafka, Snowflake, BigQuery, Redshift, PostgreSQL, DuckDB, ClickHouse ⚡ Database Performance Engineering - If your queries are slow, costs are high, or dashboards lag, This is my zone - Query plan analysis & index strategies - Warehouse cost optimization (Snowflake, BigQuery, Redshift) - OLTP & OLAP performance tuning - High-concurrency workload design 🔄 Reverse ETL & Operational Analytics - Syncing analytics back to CRMs & internal tools - Building real-time metrics pipelines - Feature-store style transformations 🕷️ Enterprise-grade Web Data Extraction - I don’t just scrape pages, I build durable data acquisition systems: - Complex ASP.NET, JS-heavy, authenticated & paginated systems - Anti-bot bypassing & failure-recovery pipelines - Headless browser automation + async scraping - Real-estate, finance, campaign-finance & marketplace platforms ☁️ Cloud Infrastructure - AWS | Azure | GCP - EMR / Dataproc / Glue / Dataflow / Synapse / BigQuery / Redshift - Terraform-based deployments - Cost-aware architectures - Kubernetes + Dockerized data services 🧪 What You Get Working With Me ✔️ Production-ready pipelines ✔️ Clean, testable dbt models ✔️ Well-documented architecture diagrams ✔️ Transparent data logic for non-technical stakeholders ✔️ Systems that scale beyond MVP ✔️ Honest advice and not over-engineering 🏆 Ideal Projects 👍 Data warehouse migrations 👍 Broken pipelines that need debugging & stabilization 👍 Analytics platforms that lack trust or explainability 👍 Performance bottlenecks costing thousands per month 👍 Long-term data platform ownership ❣️ Why Clients Stay Long-Term 🍀Clear communication 🍀 Business-first thinking 🍀 No black-box systems 🍀 I build systems others can maintain 🇻🇳🇻🇳🇻🇳🇻🇳 If your data platform feels fragile, slow, or impossible to explain to customers, I can fix that. Let’s make your data system something you can confidently stand behind.

  • Python
  • Data Scraping
  • ETL
  • Data Visualization
  • SQL Programming
  • Microsoft Azure
  • Amazon Web Services
  • Web Development
  • Database Administration
  • NoSQL Database
  • Google Cloud Platform
  • Apache Airflow
  • dbt
  • Analytics
Tran T.

Ho Chi Minh City, Vietnam

$50/hr
5.0
164 jobs

Top 1% on Upwork: Top Rated Plus, 100% Job Success, 5.0★ (150+ reviews). I build production AI agents, RAG chatbots, LLM apps, AI automation, ETL data pipelines and web scraping. Founder of Van Data Team: 20+ AI products shipped, 9 of them live public platforms, for clients across the US, EU, and APAC. 🚀 WHAT I'VE SHIPPED (live products, real numbers) - Enterprise AI assistant: LangGraph Supervisor routing sub-agents across 70+ MCP tool integrations; web, PWA, iOS & Android from one React codebase - AI cybersecurity platform: fine-tuned Mistral 7B on 424k training samples; Neo4j knowledge graph + RAG - EdTech RAG platform: raised answer accuracy from ~80% to 95%+ (pgvector, LangChain, Claude) - Multi-tenant enterprise AI suite: 5-agent architecture with persistent memory, running at 94.7% platform health - Sales-KPI AI dashboard: cut LLM cost by 92% with Claude Haiku - My own SaaS: multi-tenant AI content agents, personalized per tenant (live product, see Portfolio section) 💬 AI AGENTS, AI CHATBOTS & LLM / GENERATIVE AI APPS (Full-Stack) - AI Agent Development: autonomous agents with tool calling, multi-step reasoning, and decision-making (LangGraph, LangChain, MCP) - Multi-Agent Systems: supervisor patterns, agent-to-agent communication, task delegation, workflow orchestration - AI Chatbot Development: customer-service bots, document Q&A, and conversational AI grounded in your business data (RAG / Retrieval-Augmented Generation, no hallucinations) - Custom MCP servers: turn your internal tools into capabilities for Claude, Cursor, and other LLM clients - Models & APIs: Claude (Anthropic), OpenAI API (GPT-4/4o), Gemini, DeepSeek, Llama 3, Hugging Face; AWS Bedrock, Azure OpenAI, Vertex AI - Vector databases: Pinecone, Milvus, ChromaDB, pgvector, FAISS + Neo4j graphs - Observability, evals & fine-tuning: LangSmith, Langfuse; LLM fine-tuning and machine learning for domain-specific tasks - Frontend: Next.js, React, TypeScript, Tailwind. Backend: Python (FastAPI, Flask, Django), NestJS, Node.js/Express, GraphQL Perfect for: AI automation, support automation, sales chatbots, document Q&A, intelligent search, workflow and business-process automation. ⚡ DATA ENGINEERING & ETL PIPELINES Production data infrastructure processing millions of records: - Pipelines: Apache Spark, Airflow, Dagster, dbt • Streaming: Kafka - Batch: AWS Glue, Google Dataflow, Azure Data Factory - Warehouses: BigQuery, Snowflake, Redshift, Athena, with optimized schemas plus HIPAA-style audit & lineage - Zero-downtime migrations to cloud warehouses 🕷️ WEB SCRAPING & DATA EXTRACTION - Platforms: Facebook, TikTok, LinkedIn, Shopify, HubSpot, QuickBooks, SaaS portals - Tools: Scrapy, Playwright, Selenium, Puppeteer, BeautifulSoup, Crawl4AI - Anti-bot: Cloudflare, CAPTCHA, rate limiting, solved with proxy rotation & monitoring - Output: clean CSV/JSON, SQL databases, or a ready-to-use API ☁️ CLOUD & DEVOPS AWS (Lambda, ECS Fargate, Glue, Athena, Redshift, SQS, EventBridge, EMR) • GCP (BigQuery, Cloud Run, Dataflow, Dataproc) • Azure (Data Factory, Databricks) • Docker, Kubernetes, Terraform, GitHub Actions & GitLab CI/CD • Monitoring: CloudWatch, Datadog, Prometheus, Grafana ✅ WHY CLIENTS PICK ME - Founder-led, senior-only execution: you work with me directly, no agency layer - AI-native workflow (Claude Code, Cursor, Copilot): 2-3x faster delivery, production quality - 6+ years across AI engineering, data engineering, and full-stack development - End-to-end ownership: architecture → code → deployment → monitoring - Clear communication: regular updates, transparent workflow, overlap with US/EU hours Let's build your AI solution. Send an invite or message me, and I'll reply with a concrete plan within hours.

  • Apache Spark
  • Generative AI
  • Machine Learning
  • Full-Stack Development
  • Data Engineering
  • Web Scraping
  • LangChain
  • Next.js
  • FastAPI
  • AWS Lambda
  • Retrieval Augmented Generation
  • Vector Database
  • AI Chatbot
  • Python
  • Docker
Tram N.

Hanoi, Vietnam

$25/hr
5.0
12 jobs

- Build & maintain ETL pipelines for small to medium sized companies - Visualize data using Power BI - Preprocess data using regex or other techniques - Custom Python scripts

  • Apache Spark
  • AWS Glue
  • ETL Pipeline
  • Python
  • SQL
  • ETL
  • Data Extraction
  • Microsoft Power BI Data Visualization
  • Data Visualization
Lam T.

Hanoi, Vietnam

$40/hr
5.0
1 jobs

I am a highly motivated and passionate data engineer. I often work with Python, Scala, and Java and use the latest big data technologies to solve problems, making tools to improve my and others' work productivity. I am currently certified with AWS Solution Architect and SnowPro. Check out my blog for my work and newsletter lam-tran.dev

  • Apache Spark
  • Python
  • SQL
  • Scala
  • Snowflake
  • Databricks Platform
  • AWS Glue
  • Apache Kafka
  • Apache Hadoop
  • Polars
  • MySQL
  • Docker
  • Apache Airflow
  • Kubernetes
  • Data Engineering
Huy T.

Hanoi, Vietnam

$20/hr
5.0
8 jobs

I have experience with Data Scraping, RDBMS, NoSQL DB, building ETL pipeline, Big Data - Data Scraping with Python with Scrapy, Selenium, Python Request. - Expert SQL with Oracle, MySql. - Expert AWS in Data Engineering field - Experience with NoSQL: ElasticSearch, HBase, Cassandra. - Experience with Minio, AWS, Kafka, Terraform, Redis - Expert with Spark, HDFS Looking for long term contract

  • Apache Hadoop
  • Scala
  • Python
  • SQL
  • Java
  • PySpark
  • Snowflake
  • Oracle Database
  • Big Data
  • PostgreSQL
Ngu N.

Ho Chi Minh City, Vietnam

$30/hr
5.0
20 jobs

✅ 𝙏𝙊𝙋 𝙍𝙖𝙩𝙚𝙙 𝘿𝙚𝙫𝙊𝙥𝙨 𝙀𝙣𝙜𝙞𝙣𝙚𝙚𝙧 | ✅ 𝟱-𝙎𝙩𝙖𝙧 𝙍𝙚𝙫𝙞𝙚𝙬𝙨 | ✅ 𝟳+ 𝙔𝙚𝙖𝙧𝙨 𝙤𝙛 𝙀𝙭𝙥𝙚𝙧𝙞𝙚𝙣𝙘𝙚 I’m a passionate AWS DevOps & AI-Powered Data Engineer with hands-on experience architecting, automating, and optimizing mission-critical cloud solutions. I specialize in serverless AI integration using AWS Bedrock and designing intelligent workflows that scale. 𝙄 𝙝𝙖𝙫𝙚 𝙚𝙭𝙥𝙚𝙧𝙞𝙚𝙣𝙘𝙚 𝙞𝙣 𝙩𝙝𝙚 𝙛𝙤𝙡𝙡𝙤𝙬𝙞𝙣𝙜 𝙖𝙧𝙚𝙖𝙨, 𝙩𝙤𝙤𝙡𝙨 𝙖𝙣𝙙 𝙩𝙚𝙘𝙝𝙣𝙤𝙡𝙤𝙜𝙞𝙚𝙨: 🧠 𝘼𝙄 & 𝙂𝙚𝙣𝙚𝙧𝙖𝙩𝙞𝙫𝙚 𝘼𝙄 𝙬𝙞𝙩𝙝 𝘼𝙒𝙎 𝘽𝙚𝙙𝙧𝙤𝙘𝙠 ► Integrated LLMs (Claude, Mistral, Titan) into serverless applications using AWS Bedrock + Lambda + Step Functions ► Designed prompt orchestration pipelines for data enrichment, customer support, and automated insights ► Built secure, scalable AI workflows for multi-tenant environments ☁️ 𝘾𝙇𝙊𝙐𝘿 𝙄𝙉𝙁𝙍𝘼𝙎𝙏𝙍𝙐𝘾𝙏𝙐𝙍𝙀 ► AWS: EC2, S3, RDS, ECS, EKS, ECR, Lambda, IAM, VPC, Route53, CloudTrail, KMS, CloudWatch, Redshift, DynamoDB, Glue, Athena, Step Functions, Bedrock ► Azure: Data Factory, Blob Storage, Cosmos DB, Azure DevOps 🗄️ 𝘿𝘼𝙏𝘼𝘽𝘼𝙎𝙀𝙎 ► SQL | NoSQL | MySQL | PostgreSQL | SQL Server | MongoDB | Redis | DynamoDB | ElasticCache 🚀 𝘾𝙄/𝘾𝘿 & 𝘼𝙐𝙏𝙊𝙈𝘼𝙏𝙄𝙊𝙉 ► Jenkins | AWS CodePipeline, CodeBuild, CodeDeploy | GitHub Actions | Terraform | CloudFormation 🔐 𝙄𝘿𝙀𝙉𝙏𝙄𝙏𝙔 & 𝙎𝙀𝘾𝙐𝙍𝙄𝙏𝙔 ► AWS IAM Identity Center (SSO) | AWS Cognito | Okta | AWS WAF | Network Firewall 🧩 𝘽𝙄𝙂 𝘿𝘼𝙏𝘼 & 𝘿𝘼𝙏𝘼 𝙀𝙉𝙂𝙄𝙉𝙀𝙀𝙍𝙄𝙉𝙂 ► Apache Spark | Hadoop | EMR | Kinesis | Data Lakes | Glue ETL | Airflow | Athena 📊 𝘽𝙄 & 𝙈𝙊𝙉𝙄𝙏𝙊𝙍𝙄𝙉𝙂 ► Power BI | Tableau | Superset | Grafana | SSMS 🛠️ 𝙏𝙊𝙊𝙇𝙎 & 𝙇𝘼𝙉𝙂𝙐𝘼𝙂𝙀𝙎 ► Python | Docker | Kubernetes | Bash | Ansible | REST APIs | JSON | YAML 🧪 𝙎𝙤𝙢𝙚 𝙤𝙛 𝙢𝙮 𝙢𝙖𝙟𝙤𝙧 𝙥𝙧𝙤𝙟𝙚𝙘𝙩𝙨 𝙞𝙣𝙘𝙡𝙪𝙙𝙚𝙙 ► Built AI-powered data enrichment workflows using AWS Bedrock + Lambda. ► Designed, implemented the entire software deployment system, system security and data assurance of company on both Cloud and Premise. ► Architected fully automated CI/CD pipelines for containerize applications. ► Administered infrastructure and back-end developer for a VPN system with 300+ servers. ► Designed, implemented and operated a Data Platform to centralize all company data ► Designed scalable Data Lake solutions with EMR + Glue + Athena ► Implemented centralized identity access via AWS SSO (IAM Identity Center) ► Automated AWS Route53, DNS, and Glue deployment for 50+ microservices 💬 𝙄 𝙗𝙧𝙞𝙣𝙜 𝙣𝙤𝙩 𝙟𝙪𝙨𝙩 𝙩𝙚𝙘𝙝𝙣𝙞𝙘𝙖𝙡 𝙨𝙠𝙞𝙡𝙡𝙨 𝙗𝙪𝙩 𝙖 𝙘𝙤𝙢𝙢𝙞𝙩𝙢𝙚𝙣𝙩 𝙩𝙤 𝙥𝙧𝙚𝙘𝙞𝙨𝙞𝙤𝙣, 𝙧𝙚𝙨𝙥𝙤𝙣𝙨𝙞𝙫𝙚𝙣𝙚𝙨𝙨, 𝙖𝙣𝙙 𝙘𝙡𝙖𝙧𝙞𝙩𝙮. Whether you're looking to launch a new AI feature, optimize your DevOps pipelines, or need a reliable Linux system engineer to manage your infrastructure—I’m ready to help. 𝙇𝙚𝙩’𝙨 𝙢𝙖𝙠𝙚 𝙮𝙤𝙪𝙧 𝙞𝙣𝙛𝙧𝙖𝙨𝙩𝙧𝙪𝙘𝙩𝙪𝙧𝙚 𝙨𝙢𝙖𝙧𝙩𝙚𝙧, 𝙛𝙖𝙨𝙩𝙚𝙧, 𝙖𝙣𝙙 𝙛𝙪𝙩𝙪𝙧𝙚-𝙧𝙚𝙖𝙙𝙮. Available for long-term or short-term contracts!

  • Amazon Web Services
  • AWS CodeBuild
  • AWS CodePipeline
  • Amazon Elastic Beanstalk
  • AWS CodeDeploy
  • Data Lake
  • Data Warehousing & ETL Software
  • Data Visualization
  • DevOps Engineering
  • Kubernetes
  • CI/CD
  • Data Engineering
  • Database

How it works

Post a job for freePost a job

Tell us what you need. Create your own job post or generate one with AI then filter talent matches.

Hire top talent fast

Consult, interview, and hire quickly, so you can meet the freelancers you're excited about.

Collaborate easily

Use Upwork to chat or video call, share files, and track project progress right from the app.

Payment simplified

Manage payments in one place with flexible billing options. Only pay for approved work, hourly or by milestone.

Don't just take our word for it

How do I hire a Apache Spark Engineer in Vietnam on Upwork?

You can hire a Apache Spark Engineer in Vietnam on Upwork in four simple steps:

  • Create a job post tailored to your Apache Spark Engineer project scope. We'll walk you through the process step by step.
  • Browse top Apache Spark Engineer talent on Upwork and invite them to your project.
  • Once the proposals start flowing in, create a shortlist of top Apache Spark Engineer profiles and interview.
  • Hire the right Apache Spark Engineer for your project from Upwork, the world's largest work marketplace.

At Upwork, we believe talent staffing should be easy.

How much does it cost to hire a Apache Spark Engineer?

Rates charged by Apache Spark Engineers on Upwork can vary with a number of factors including experience, location, and market conditions. See hourly rates for in-demand skills on Upwork.

Why hire a Apache Spark Engineer in Vietnam on Upwork?

As the world's work marketplace, we connect highly-skilled freelance Apache Spark Engineers and businesses and help them build trusted, long-term relationships so they can achieve more together. Let us help you build the dream Apache Spark Engineer team you need to succeed.

Can I hire a Apache Spark Engineer in Vietnam within 24 hours on Upwork?

Depending on availability and the quality of your job post, it's entirely possible to sign up for Upwork and receive Apache Spark Engineer proposals within 24 hours of posting a job description.