Hire the Best Pyspark Developers
in China

More than 3,000 reviews on G2
Rating is 4.5 out of 5.
4.5/5
of Upwork by G2 peer reviewers
Lian X.

Wuhan, China

$25/hr
5.0
15 jobs

I’m a developer with experience in NLP、LLM 、 CV and recommendation system and bigdata. 1. I’m experienced in RAG、langchain、deepspeed、llm、 object dection、video generation 2.I’m experienced in hadoop/hdfs/yarn/hbase/redis/kafka/hive, 2. I’m experienced in tensorflow/kubeflow/tf serving/trtion serving 3.I’m experienced in k8s/docker/html/jqurey/spring boot/mybatis 4.I’m experienced in aws componet, such as emr/ec2/s3/code deploy

  • Apache Spark
  • PySpark
  • TensorFlow
  • Java
  • Apache Flink
  • Kubernetes
  • Spring Boot
  • Apache Hadoop
  • Artificial Intelligence
  • Big Data
  • AWS Application
  • LLM Prompt Engineering
  • LangChain
  • jQuery
Yazhen L.

Beijing, China

$30/hr
5.0
32 jobs

🔝 𝐓𝐨𝐩-𝐑𝐚𝐭𝐞𝐝 𝐅𝐫𝐞𝐞𝐥𝐚𝐧𝐜𝐞𝐫 𝐨𝐧 𝐔𝐩𝐰𝐨𝐫𝐤 🚀 𝗦𝘁𝗿𝘂𝗴𝗴𝗹𝗶𝗻𝗴 𝘄𝗶𝘁𝗵 𝘀𝗹𝗼𝘄 𝗱𝗮𝘁𝗮 𝗽𝗿𝗼𝗰𝗲𝘀𝘀𝗶𝗻𝗴 𝗼𝗿 𝘂𝗻𝗿𝗲𝗹𝗶𝗮𝗯𝗹𝗲 𝗘𝗧𝗟 𝗽𝗶𝗽𝗲𝗹𝗶𝗻𝗲𝘀? 𝗜 𝗯𝘂𝗶𝗹𝗱 𝗮𝗻𝗱 𝗼𝗽𝘁𝗶𝗺𝗶𝘇𝗲 𝗵𝗶𝗴𝗵-𝗽𝗲𝗿𝗳𝗼𝗿𝗺𝗮𝗻𝗰𝗲 𝗱𝗮𝘁𝗮 𝘀𝘆𝘀𝘁𝗲𝗺𝘀 𝘁𝗵𝗮𝘁 𝘁𝘂𝗿𝗻 𝗺𝗮𝘀𝘀𝗶𝘃𝗲 𝗱𝗮𝘁𝗮𝘀𝗲𝘁𝘀 𝗶𝗻𝘁𝗼 𝗮𝗰𝘁𝗶𝗼𝗻𝗮𝗯𝗹𝗲 𝗶𝗻𝘀𝗶𝗴𝗵𝘁𝘀. With 4+ years of professional experience, I have worked as a Data Engineer at world-class tech giants Xiaomi (Fortune Global 500) and Shopee ($80B Market Cap). My expertise lies in transforming complex data challenges into seamless, efficient, and scalable data workflows. ⭐ How I Can Elevate Your Business: ✔ End-to-End ETL/ELT Pipeline Development: I architect and build fully automated data pipelines using Airflow and Apache Spark (Scala/Python/Java), ensuring timely and accurate data for your analytics and machine learning models. ✔ Spark Performance Tuning & Optimization: Is your Spark job running slow or failing? I specialize in deep-diving into Spark applications to diagnose bottlenecks, optimize resource utilization (memory/CPU), and significantly cut down processing time and cost. ✔ Big Data Architecture & Solutions: Leveraging modern data stack tools like Kafka, Flink, Hadoop (HDFS, Hive), and Druid, I design and implement scalable Data Lakes and Data Warehouses tailored to your specific business needs. ✔ Data Quality & Integrity: I implement rigorous data cleaning and validation processes, transforming raw, messy data into a pristine, reliable asset for your decision-making. ⭐ Core Technical Skills: ✔ Big Data Ecosystem: Apache Spark, Airflow, Kafka, Flink, Hadoop, Hive, HDFS, Druid ✔ Programming: Scala, Python, Java, SQL ✔ Databases: HBase, Redis (NoSQL), Relational SQL Databases ✔ Platforms: Data Lake, Data Warehouse, AWS, GCP ⭐ What Sets Me Apart: ✅ Problem-Solver, Not Just a Coder: I focus on understanding your business goals first, then architect a solution that delivers real value. My work at Xiaomi was praised for not just meeting, but exceeding expectations by foreseeing future needs. ✅ Proactive & Clear Communication: You will always be in the loop. I believe in transparent, frequent updates to ensure the project aligns perfectly with your vision. ✅ Partnership & Kindness: I see my clients as partners. I am committed to working collaboratively and kindly to make your project a success and your life easier. Ready to build a data infrastructure that drives your business forward? 𝗝𝗨𝗦𝗧 𝗖𝗟𝗜𝗖𝗞 𝗢𝗡 𝗜𝗡𝗩𝗜𝗧𝗘 𝗕𝗨𝗧𝗧𝗢𝗡 and let's have a quick chat about your project goals.

  • Apache Spark
  • Python
  • Scala
  • Big Data
  • Apache Flink
  • Hive
  • Java
  • Apache Hadoop
  • Apache Kafka
  • ETL Pipeline
  • SQL
  • Data Engineering
  • Data Extraction
  • Data Cleaning
Irbaz S.

Beijing, China

$40/hr
5.0
23 jobs

AWS Certified Data Engineer and Data Analyst with 5+ years of experience designing scalable data pipelines, building robust ETL workflows, and transforming raw data into actionable insights. Proven expertise in AWS cloud services, big data processing, and advanced analytics to drive data-driven decision-making. Dedicated to delivering high-quality data solutions that optimize performance, reduce costs, and unlock business value. Core Skills & Technologies 🔷 Data Engineering: AWS Stack: Redshift, Glue, EMR, Athena, Kinesis, Lambda, S3, CloudFormation ETL/ELT: Apache Spark, PySpark, AWS Glue, Airflow, Kafka Data Warehousing: Snowflake, Redshift optimization Big Data: Hadoop, Hive, Databricks, real-time streaming 🔷 Data Analysis & Visualization: SQL, Python (Pandas, NumPy), R BI Tools: Tableau, Power BI, AWS QuickSight Advanced Analytics: Predictive modeling, regression, clustering 🔷 Infrastructure: IaC (Terraform), Docker, CI/CD pipelines, AWS Security (IAM, KMS) Services Offered 🔷End-to-End Data Pipeline Development → Build automated, fault-tolerant data pipelines (batch/streaming) using AWS services → Migrate on-premise data systems to AWS cloud (e.g., S3 → Redshift) 🔷Data Warehousing & Modeling → Design star/snowflake schemas, optimize Redshift clusters, manage data lakes 🔷ETL Optimization → Refactor legacy ETL jobs to Spark/Glue for 50%+ faster processing 🔷Analytics & Reporting → Create interactive Tableau/Power BI dashboards for KPIs and forecasting 🔷Ad-Hoc Analysis → Clean, analyze, and visualize datasets to uncover growth opportunities Why Hire Me? ✅ AWS Certified Data Engineer- Associate ✅ Performance-Driven: 40% lower query costs | 60% pipeline efficiency gains ✅ Full-Cycle Delivery: Architecture → Deployment → Monitoring → Documentation ✅ Client Focus: Responsive communication | Agile workflow | Transparent timelines AWS Certified Data Engineer, Data Analyst, Big Data Engineer, ETL Developer, Data Pipeline, AWS Glue, Redshift, PySpark, SQL Data Analyst, Tableau Specialist, Data Warehousing, Kinesis, Lambda, S3, Data Lake, Power BI, Data Modeling, Business Intelligence, Machine Learning, Data Visualization, Cloud Migration

  • PySpark
  • SQL
  • Python
  • Data Science
  • ETL
  • Microsoft Power BI
  • Data Analysis
  • Business Analysis
  • Business Intelligence
  • AWS Glue
  • Data Engineering
  • Data Lake
  • Data Warehousing
  • Amazon Redshift
  • Data Ingestion
Xiao P.

Shanghai, China

$80/hr
5.0
9 jobs

Work as tech lead & chief architect for a web company serving 4m+ users. Collect and store user behavior data & transaction data, perform daily and ad hoc ETL on the data, find useful pattern and provide BI visualization; design & develop server backend and providing apis to end users. Able to help a middle sized tech firm to setup OLAP(Online Analytical Processing) & OLTP(Online Transactional Processing) platform from zero. Has rich working experience in Spark, MongoDB, Fastapi, Clickhouse, Pandas, Numpy, Terraform, Agent development, AWS environment(Lambda, Athena, DynamoDB, Glue, EMR, Bedrock, Cognito, QuickSight, Sagemaker, etc). Rich coding experience in Python/Scala/Java. Rich experience in Machine Learning especially NLP. Can work as part-time job, 10~20 hours per week.

  • Apache Spark
  • Scala
  • pandas
  • AWS Lambda
  • AWS Glue
  • MongoDB
  • Terraform
  • Apache Airflow
  • ClickHouse
  • FastAPI
  • Amazon DynamoDB
  • Metabase
  • AI Agent Development
  • SQL
  • Amazon Bedrock

How it works

Post a job for freePost a job

Tell us what you need. Create your own job post or generate one with AI then filter talent matches.

Hire top talent fast

Consult, interview, and hire quickly, so you can meet the freelancers you're excited about.

Collaborate easily

Use Upwork to chat or video call, share files, and track project progress right from the app.

Payment simplified

Manage payments in one place with flexible billing options. Only pay for approved work, hourly or by milestone.

Don't just take our word for it

How do I hire a Pyspark Developer in China on Upwork?

You can hire a Pyspark Developer in China on Upwork in four simple steps:

  • Create a job post tailored to your Pyspark Developer project scope. We'll walk you through the process step by step.
  • Browse top Pyspark Developer talent on Upwork and invite them to your project.
  • Once the proposals start flowing in, create a shortlist of top Pyspark Developer profiles and interview.
  • Hire the right Pyspark Developer for your project from Upwork, the world's largest work marketplace.

At Upwork, we believe talent staffing should be easy.

How much does it cost to hire a Pyspark Developer?

Rates charged by Pyspark Developers on Upwork can vary with a number of factors including experience, location, and market conditions. See hourly rates for in-demand skills on Upwork.

Why hire a Pyspark Developer in China on Upwork?

As the world's work marketplace, we connect highly-skilled freelance Pyspark Developers and businesses and help them build trusted, long-term relationships so they can achieve more together. Let us help you build the dream Pyspark Developer team you need to succeed.

Can I hire a Pyspark Developer in China within 24 hours on Upwork?

Depending on availability and the quality of your job post, it's entirely possible to sign up for Upwork and receive Pyspark Developer proposals within 24 hours of posting a job description.