Hire the Best Big Data Engineers

Clients rate our Big Data Engineers
Rating is 4.8 out of 5.
4.8/5
Based on 386 client reviews
Kashif S.

Gudja, Malta

$25/hr
5.0
8 jobs

I build data platforms that work at scale and keep working as your business grows. Over the past 10 years I've served as the lead or founding data engineer across fintech, e-commerce, ride-hail, legal tech, and cybersecurity companies. That means I've designed systems from scratch, made architecture decisions with no one to fall back on, and delivered platforms that product teams actually use. Here's what I typically get hired to do: → Build greenfield data platforms on AWS or GCP from the ground up → Design and ship production ETL/ELT pipelines (Airflow, Dagster, dbt) → Set up scalable warehouses and governance (Snowflake, BigQuery, Redshift) → Implement real-time streaming pipelines (Kafka, Spark Streaming, CDC) → Build AI-powered data applications (RAG, LLMs, LangChain, vector DBs) → Fix broken or unreliable pipelines and make them production-grade → Architect cloud infrastructure on AWS, GCP, Azure (Terraform, Kubernetes) Recent work includes: - Led data platform engineering for a US e-commerce company processing billions of events daily. I re-architected ingestion pipelines, built Snowflake governance from scratch, introduced Prometheus monitoring and CI/CD standards across the platform. - Built a full data platform on GCP (BigQuery, Dataproc, Airflow) for a music streaming company. Firebase, AppsFlyer, and app store data all flowing into one warehouse within weeks. - Designed an AWS data platform for a ride-hail company managing 500+ streaming and 700+ batch jobs — including a self-serve portal that replaced multi-step CLI workflows for engineers. - Built a legal AI search engine using LangChain, Pinecone, and RAG — full pipeline from document ingestion to LLM-generated answers, deployed on AWS with auto-scaling. - Built an AI inventory insights agent for a US automotive company — multi-source data pipelines, real-time APIs, conversational interface. I work in English daily, communicate proactively, and deliver production- ready code — not prototypes. I'm used to working directly with CTOs and technical leads in US and European time zones. Tools I work with regularly: Python · SQL · Airflow · Dagster · dbt · Snowflake · BigQuery · Spark · Meltano · Kafka · AWS (S3, EMR, Glue, ECS, Lambda, EC2, EKS) · Databricks · GCP · Azure Terraform · Docker · Kubernetes · LangChain · FastAPI · MLflow · Weaviate, Celery If you're building a data platform, fixing one, or adding AI/ML capabilities to your stack, let's talk.

  • Python
  • Google Cloud Platform
  • Microsoft Azure
  • Amazon Web Services
  • Data Engineering
  • Docker
  • DevOps
  • GitHub
  • BigQuery
  • Snowflake
  • Apache Airflow
  • Apache Spark
  • Terraform
  • ETL
  • Apache Kafka
Pranjal P.

Dartford, United Kingdom

$25/hr
5.0
3 jobs

I am a versatile technologist with 4+ years of experience designing and implementing high-performance backend systems and scalable data solutions. As both a Backend Architect (Node.js, NestJS, Firebase, Python) and Data Engineer (Spark, AWS/Azure, ETL/ELT), I bridge the gap between raw data and production-ready applications—building robust APIs, real-time data pipelines, and cloud-native architectures for web, mobile, and analytics platforms. Core Expertise: 🔹 Backend Development & Architecture Design and optimize REST/GraphQL APIs (Node.js, NestJS, Python/Django/FastAPI) and serverless backends (Firebase, AWS Lambda). Build event-driven systems (Kafka, Kinesis) and real-time applications (WebSockets, Firebase Realtime DB). Implement secure, scalable cloud solutions (AWS, Azure, GCP) with containerization (Docker/Kubernetes). 🔹 Data Engineering & Analytics Develop batch/streaming pipelines (Spark, Glue, Databricks) and data lakehouses (Iceberg, Delta Lake). Automate ETL/ELT workflows and migrate legacy systems to modern cloud warehouses (Snowflake, BigQuery). Integrate ML models into data pipelines and enable AI-driven analytics. 🔹 Cross-Functional Leadership Collaborate with product and analytics teams to unify data and application architecture. Ensure data quality, security, and performance across both backend services and data infrastructure. Why Work With Me? ✅ Full-spectrum expertise—From API design to data pipeline orchestration. ✅ Cloud-native mindset—AWS/Azure/GCP, serverless, and IaC (Terraform, CDK). ✅ Impact-driven—Solutions that enhance scalability, speed, and business value. Let’s build systems that don’t just process data—they power decisions.

  • Big Data
  • Data Cleaning
  • Data Extraction
  • Data Analytics
  • ETL Pipeline
  • Data Warehousing
  • Data Lake
  • Software Development
  • Web & Mobile Design Consultation
  • Mobile App Development
  • Apache Spark
  • Data Modeling
  • Big Data File Format
Youness M.

Casablanca, Morocco

$35/hr
4.8
11 jobs

Your data pipeline is either driving decisions, or quietly slowing your business down. Most teams don’t struggle with data volume, they struggle with reliability. Pipelines fail without alerts, dashboards lag behind reality, and engineers spend more time fixing than building. The result is slower decisions, growing technical debt, and missed opportunities. I design and build robust, scalable data systems that simply work, from ingestion to analytics-ready data. The focus is always on clarity, performance, and reliability, so your team can trust the data and move faster without constant firefighting. My tech stack: Python, SQL, Spark, Apache NiFi, Airflow, Kafka, Flink, Snowflake, BigQuery, AWS, GCP, Docker, Terraform, and FastAPI. If you share your current setup or challenge, I’ll break down exactly how to fix or scale it. I usually respond within a few hours.

  • Apache Spark
  • Apache Kafka
  • Apache NiFi
  • Apache Airflow
  • Data Warehousing
  • Apache Flink
  • Amazon Web Services
  • Looker Studio
  • dbt
  • Snowflake
  • BigQuery
  • Google Cloud Platform
  • Kubernetes
  • Apache Superset
  • CI/CD
Mandadi A.

Hyderabad, India

$3/hr
4.3
14 jobs

With a background in GIS services and consultancy, I offer end-to-end solutions for geospatial analysis, satellite and drone image processing, agricultural data extraction, and advanced AI/ML applications. My expertise includes: 1. GIS mapping, data management, and spatial analysis 2. Satellite imagery processing and interpretation 3. Drone image processing for high-precision mapping 4. Agricultural data processing and analytics 5. Statistical analysis for research and decision-making 6. AutoCAD and Creo for technical design and modeling 7. Remote sensing for diverse industry applications 8. AI and ML integration for geospatial workflows I collaborate closely with clients to deliver actionable insights, custom solutions, and top-quality results-on time and within budget. Whether you have a one-time mapping project or ongoing data analysis needs, I am ready to help you achieve your goals.

  • Big Data
  • GIS
  • QGIS
  • Python
  • Machine Learning
  • Spatial Analysis
  • Geospatial Data
  • Agricultural Engineering
  • GIS Software
  • ArcGIS
  • Image Processing
  • Statistical Analysis
  • GeoTIFF
  • Drone
  • Remote Sensing
Mujtaba S.

Karachi, Pakistan

$15/hr
5.0
3 jobs

Updated on 14/08/2026 Most dashboard problems are not dashboard problems. A number that does not match what someone counted by hand usually broke three steps earlier, in ingestion or a transformation nobody tested. That is where I actually spend my time as a Data and AI Engineer, and the chart at the end is the easy part. My pipelines typically run on Airflow or Mage AI, with Kafka and PyFlink handling anything that needs to move in real time. On AWS I work with Lambda, S3, EventBridge, and SNS, and bad records get pulled into a quarantine bucket instead of quietly sitting in a table someone trusts. For transformation, I build dbt models on Snowflake and PostgreSQL, structured bronze through gold, with schema tests and business rule checks written into the models themselves, so a broken assumption gets caught in the pipeline instead of by whoever opens the report next. Reporting comes after the data is solid. I build in Power BI or Tableau around the one question the business actually needs answered, not a stack of generic rollups nobody reads. When the need is document search or research rather than dashboards, I build RAG systems that score their own retrieval accuracy, so a weak answer gets flagged instead of handed over as confident nonsense. I am currently applying this same thinking at Genix Pharma, building AI-assisted workflows on local LLMs through Ollama for model experimentation, evaluation, and automated reporting. A few things I have shipped recently: a district-level KPI dashboard on Snowflake and Power BI built from layered dbt models, incremental dbt pipelines feeding logistics and lending risk reporting, a real-time Kafka and PyFlink pipeline with event-time processing sinking to PostgreSQL, and a serverless AWS pipeline where a quarantine bucket keeps bad files from ever reaching the tables people query. If a tool is not something I have genuinely used, I will say so instead of guessing my way through your job. Tell me what the reporting needs to answer and where your data lives right now, whether that is Excel, PDFs, or a handful of systems that do not talk to each other, and I will give you a straight read on whether it is a small fix or a bigger rebuild. Machine Learning, Database Design, Delta Lake Expert, Databricks Engineer, Big Data Consultant, AWS Data Specialist, Database Architecture, Amazon Web Services, Artificial Intelligence, Deep Learning Modeling, Machine Learning Engineer, Data Analytics & Visualization Software, Data Warehousing & ETL Software Data Processing, Cloud Engineering, GCP Analytics, Data Analytics, Data Visualization, Spark Developer, ETL, SQL, Python, DBT, Snowflake, Apache Airflow, Apache Kafka, AWS, Data Pipeline, Power BI, Python, Snowflake, ETL, Big Data, ETL Pipeline, Data Engineer, ETL Developer, Data Science, Data Analysis, Deep Learning, Data Engineering, Azure Databricks, MLOps Engineer

  • Microsoft Power BI
  • Data Engineering
  • Data Extraction
  • dbt
  • Data Analysis
  • ETL
  • ETL Pipeline
  • API
  • Apache Airflow
  • AWS Lambda
  • Data Modeling
  • Machine Learning
  • Data Quality Assessment
  • ClickUp
  • Snowflake
  • Artificial Intelligence
Muhammad Umer S.

Arlington, Texas

$85/hr
5.0
22 jobs

𝐈 𝐛𝐮𝐢𝐥𝐝 𝐩𝐫𝐨𝐝𝐮𝐜𝐭𝐢𝐨𝐧-𝐠𝐫𝐚𝐝𝐞 𝐝𝐚𝐭𝐚 𝐩𝐥𝐚𝐭𝐟𝐨𝐫𝐦𝐬 𝐟𝐨𝐫 𝐜𝐨𝐦𝐩𝐚𝐧𝐢𝐞𝐬 𝐝𝐞𝐚𝐥𝐢𝐧𝐠 𝐰𝐢𝐭𝐡 𝐛𝐫𝐨𝐤𝐞𝐧 𝐩𝐢𝐩𝐞𝐥𝐢𝐧𝐞𝐬, 𝐬𝐜𝐚𝐭𝐭𝐞𝐫𝐞𝐝 𝐬𝐲𝐬𝐭𝐞𝐦𝐬, 𝐬𝐥𝐨𝐰 𝐫𝐞𝐩𝐨𝐫𝐭𝐢𝐧𝐠, 𝐚𝐧𝐝 𝐮𝐧𝐫𝐞𝐥𝐢𝐚𝐛𝐥𝐞 𝐦𝐞𝐭𝐫𝐢𝐜𝐬. I’m a Senior Data Engineer with 10+ years of experience building cloud data platforms, ETL/ELT pipelines, lakehouses, warehouses, and analytics-ready data layers using Microsoft Fabric, Snowflake, AWS, BigQuery, dbt, Python, SQL, Airflow, Fivetran, Airbyte, and Databricks. My focus is not just moving data from point A to point B. I design reliable data systems that are automated, scalable, well-modeled, and trusted by business teams. 𝐖𝐡𝐚𝐭 𝐈 𝐡𝐞𝐥𝐩 𝐰𝐢𝐭𝐡 ✅ 𝐌𝐢𝐜𝐫𝐨𝐬𝐨𝐟𝐭 𝐅𝐚𝐛𝐫𝐢𝐜 𝐄𝐧𝐠𝐢𝐧𝐞𝐞𝐫𝐢𝐧𝐠 Lakehouse, Warehouse, Dataflows Gen2, pipelines, notebooks, semantic models, Medallion architecture, and Power BI-ready data layers. ✅ 𝐂𝐥𝐨𝐮𝐝 𝐃𝐚𝐭𝐚 𝐖𝐚𝐫𝐞𝐡𝐨𝐮𝐬𝐞 𝐃𝐞𝐯𝐞𝐥𝐨𝐩𝐦𝐞𝐧𝐭 Snowflake, BigQuery, Redshift, Databricks, PostgreSQL, SQL Server, and Azure Synapse architecture. ✅ 𝐄𝐓𝐋/𝐄𝐋𝐓 𝐏𝐢𝐩𝐞𝐥𝐢𝐧𝐞 𝐃𝐞𝐯𝐞𝐥𝐨𝐩𝐦𝐞𝐧𝐭 API ingestion, database replication, SaaS integrations, file ingestion, batch jobs, incremental loads, and scheduled workflows. ✅ 𝐝𝐛𝐭 & 𝐀𝐧𝐚𝐥𝐲𝐭𝐢𝐜𝐬 𝐄𝐧𝐠𝐢𝐧𝐞𝐞𝐫𝐢𝐧𝐠 Staging, intermediate, marts, incremental models, tests, documentation, metric definitions, and business logic standardization. ✅ 𝐏𝐢𝐩𝐞𝐥𝐢𝐧𝐞 𝐀𝐮𝐭𝐨𝐦𝐚𝐭𝐢𝐨𝐧 & 𝐎𝐫𝐜𝐡𝐞𝐬𝐭𝐫𝐚𝐭𝐢𝐨𝐧 Airflow, Dagster, AWS Lambda, Glue, Step Functions, ADF, CI/CD, monitoring, retries, alerts, and production workflow automation. ✅ 𝐃𝐚𝐭𝐚 𝐐𝐮𝐚𝐥𝐢𝐭𝐲 & 𝐆𝐨𝐯𝐞𝐫𝐧𝐚𝐧𝐜𝐞 Deduplication, reconciliation, schema drift handling, validation rules, MDM, Golden Record logic, RBAC, access control, and audit-ready reporting layers. 𝐑𝐞𝐜𝐞𝐧𝐭 𝐩𝐫𝐨𝐣𝐞𝐜𝐭 𝐞𝐱𝐩𝐞𝐫𝐢𝐞𝐧𝐜𝐞: 𝐌𝐢𝐜𝐫𝐨𝐬𝐨𝐟𝐭 𝐅𝐚𝐛𝐫𝐢𝐜 𝐞𝐧𝐭𝐞𝐫𝐩𝐫𝐢𝐬𝐞 𝐚𝐧𝐚𝐥𝐲𝐭𝐢𝐜𝐬 𝐩𝐥𝐚𝐭𝐟𝐨𝐫𝐦 Built a centralized Fabric platform with Lakehouse, Warehouse, Dataflows, pipelines, semantic models, and Power BI reporting layers for a global organization with fragmented reporting sources. 𝐍𝐞𝐚𝐫 𝐫𝐞𝐚𝐥-𝐭𝐢𝐦𝐞 𝐒𝐧𝐨𝐰𝐟𝐥𝐚𝐤𝐞 𝐝𝐚𝐭𝐚 𝐩𝐥𝐚𝐭𝐟𝐨𝐫𝐦 Designed operational pipelines into Snowflake with incremental ingestion, schema change handling, deduplication, retries, and reliable dbt-based reporting models. 𝐌𝐨𝐝𝐞𝐫𝐧 𝐒𝐚𝐚𝐒 𝐝𝐚𝐭𝐚 𝐬𝐭𝐚𝐜𝐤 Centralized HubSpot, Stripe, GA4, Google Ads, Salesforce, MongoDB, and product data into BigQuery/Snowflake using Fivetran, Airbyte, dbt, Dagster, and Metabase. 𝐋𝐞𝐠𝐚𝐜𝐲 𝐦𝐢𝐠𝐫𝐚𝐭𝐢𝐨𝐧 𝐭𝐨 𝐜𝐥𝐨𝐮𝐝 𝐰𝐚𝐫𝐞𝐡𝐨𝐮𝐬𝐞 Led migrations from SQL Server, SAP BW, Oracle, Hadoop, and on-prem systems into modern cloud warehouses with optimized performance and automated workflows. 𝐀𝐖𝐒 𝐝𝐚𝐭𝐚 𝐚𝐮𝐭𝐨𝐦𝐚𝐭𝐢𝐨𝐧 Built automated workflows using S3, Lambda, Glue, Step Functions, IAM, SNS, CloudWatch, and Python to reduce manual reporting and improve pipeline reliability. 𝐓𝐨𝐨𝐥𝐬 𝐈 𝐰𝐨𝐫𝐤 𝐰𝐢𝐭𝐡 Microsoft Fabric, Snowflake, BigQuery, Redshift, Databricks, Azure Synapse, PostgreSQL, SQL Server, Oracle, SAP BW, Hadoop, dbt, Python, SQL, Airflow, Dagster, Fivetran, Airbyte, ADF, SSIS, Talend, AWS S3, Lambda, Glue, Step Functions, IAM, SNS, CloudWatch, Power BI, Tableau, Metabase, and Looker. You should reach out if you need a senior data engineer to: ✅ Build a cloud data warehouse or lakehouse ✅ Migrate legacy systems to Snowflake, Fabric, BigQuery, or AWS ✅ Fix unreliable ETL/ELT pipelines ✅ Design dbt models and trusted reporting layers ✅ Automate manual reporting workflows ✅ Integrate APIs, CRMs, ERPs, databases, and SaaS platforms ✅ Build production-ready data infrastructure for analytics and BI If your data stack is messy, slow, or hard to trust, send me a message. I’ll help you map the cleanest path from scattered systems to a reliable data platform.

  • Big Data
  • Data Engineering
  • Azure DevOps
  • Snowflake
  • Data Analytics
  • ETL
  • Amazon Web Services
  • Data Migration
  • Python
  • SQL
  • Artificial Intelligence
  • Microsoft Power BI
  • Databricks Platform
  • Data Modeling
  • dbt
  • Apache Airflow
  • API Integration
  • Cloud Engineering
  • Apache Kafka
  • Azure Service Fabric

How it works

Post a job for freePost a job

Tell us what you need. Create your own job post or generate one with AI then filter talent matches.

Hire top talent fast

Consult, interview, and hire quickly, so you can meet the freelancers you're excited about.

Collaborate easily

Use Upwork to chat or video call, share files, and track project progress right from the app.

Payment simplified

Manage payments in one place with flexible billing options. Only pay for approved work, hourly or by milestone.

Don't just take our word for it

What does a Big data engineer do?

A big data engineer builds the infrastructure that moves and transforms massive volumes of information for analysis. This role focuses on constructing reliable pipelines that ingest raw data from diverse sources and prepare it for downstream use. You design systems that handle both historical batch records and real-time streaming events without losing fidelity. Your work enables organizations to query large datasets quickly and make decisions based on accurate, up-to-date information.

  • Design and implement scalable data processing pipelines using Apache Spark to handle complex transformations across distributed clusters. You write code that processes terabytes of data in parallel, ensuring jobs complete within acceptable time windows while managing resource consumption. This involves configuring Spark applications to optimize shuffle operations and memory usage for specific workload patterns.
  • Develop extract, transform, and load workflows that clean and structure raw inputs for analytics platforms like BigQuery. You build logic that validates data quality, handles missing values, and standardizes formats before loading results into queryable tables. These pipelines run on managed services such as Google Cloud Dataflow, which automatically scales compute resources to match incoming data volume.
  • Integrate real-time streaming sources such as Google Cloud Pub/Sub with analytics destinations to support live dashboards and alerts. You configure connectors that read event streams continuously and write processed records to storage systems with low latency. This setup requires careful management of windowing strategies and stateful processing to ensure accurate aggregation of time-series data.

How to hire a Big data engineer on Upwork

Step 1: Post a job

Define your data pipeline needs clearly to attract qualified engineers. Use the Job Post Generator powered by Uma™, Upwork's Mindful AI to draft a precise description in seconds. Describe your requirements in a few sentences, and Uma constructs a tailored post for this role. You can write a new post, update a saved draft, or reuse an existing post.

  • Specify whether you need batch processing with Apache Spark or real-time streaming using Google Cloud Dataflow.
  • List required tools such as BigQuery for analytics storage or Pub/Sub for event ingestion.
  • Include expected deliverables like runnable ETL workflows or configured data routing pipelines.

Step 2: Evaluate candidates

Look for proof of experience building scalable data systems. Uma runs instant video interviews and builds shortlists with side-by-side comparisons to help you assess technical fit quickly.

  • Review portfolios for examples of Spark transformations or Dataflow jobs that handle large datasets.
  • Check work history for successful loads into BigQuery or integration with Hadoop clusters.
  • Verify experience with both batch and streaming architectures to ensure they match your volume needs.

Step 3: Interview your top choices

Discuss specific technical challenges related to your data infrastructure. Schedule and conduct interviews within Upwork Messages, which generates an immediate transcript and summary after each session.

  • Ask how they optimize Spark jobs for performance and cost efficiency during peak loads.
  • Request details on handling schema changes in streaming sources like Pub/Sub without downtime.
  • Discuss their approach to testing data quality before loading results into analytics destinations.

Step 4: Agree on scope and begin work

Set clear milestones for pipeline development and data integration tasks. Use Upwork Messages and the contract workroom for communication and project management, plus identity verification, payment protection, hourly tracking, and project funds for security.

  • Define milestones for building initial ETL logic and connecting source systems to sinks.
  • Agree on acceptance criteria for data accuracy and pipeline latency metrics.
  • Establish a schedule for code reviews and deployment to production environments.

Upwork is not affiliated with and does not sponsor or endorse any of the tools or services discussed in this article. These tools and services are provided only as potential options, and each reader and company should take the time needed to adequately analyze and determine the tools or services that would best fit their specific needs and situation.

The rates and information provided in this article are based on current data and industry sources available at the time of publication. Freelance rates can vary depending on factors such as experience, location, project scope, and market conditions. Readers are encouraged to conduct their own research to confirm current rates and trends, as this information may change over time.

How much does hiring a Big data engineer cost?

$500-$1,500 per project is a typical range for focused Big data engineer work. Final pricing depends on scope, technical complexity, required integrations, source-material quality, revision needs, and the freelancer's experience level.

Pipeline architecture design

$500-$1,200/project

Entry-level to mid-level
  • Visual map of data flow and system components
  • Recommended stack for batch or streaming needs
  • Step-by-step guide for building the pipeline

ETL workflow development

$1,200-$2,500/project

Mid-level
  • Code for cleaning and structuring raw data
  • Configured Spark job for scheduled data loads
  • Record of data quality checks and error rates

Streaming integration

$2,500-$4,500/project

Mid-level to senior-level
  • Configured ingestion from real-time data sources
  • Live processing logic for continuous data streams
  • Automated write process to BigQuery or similar store

Large-scale data migration

$4,500-$7,000/project

Senior-level
  • Automated tool for moving historical datasets
  • Detailed translation of old fields to new structure
  • Verification that all records transferred correctly

Custom analytics infrastructure

$7,000-$12,000/project

Expert-level
  • Optimized Hadoop or cloud environment setup
  • Combined batch and streaming processing engine
  • Documentation for maintaining speed and cost efficiency

Frequently asked questions

Is hiring a Big data engineer worth it?

For most businesses, yes: hiring a Big data engineer is worthwhile. These specialists build the pipelines that turn raw data into usable information for analytics teams. You avoid the technical debt of patching together incompatible tools when an expert architects scalable ingestion and transformation from the start.

How do I evaluate Big data engineer candidates?

Evaluate Big data engineer candidates by reviewing their code for specific pipeline implementations rather than general programming skills. Ask them to explain how they handled backpressure in a streaming job using Apache Spark or Google Cloud Dataflow. A strong candidate describes concrete steps they took to optimize a slow query in BigQuery or reduce latency in a Pub/Sub integration.

What tools does a Big data engineer use?

A Big data engineer uses distributed processing frameworks like Apache Spark and managed services such as Google Cloud Dataflow. They also configure storage and querying systems like BigQuery and Hadoop to support large-scale analytics workloads.

How does a Big data engineer handle streaming data?

A Big data engineer builds real-time ingestion pipelines that read from sources like Google Cloud Pub/Sub. They route this live data through transformation logic in Dataflow before loading it into analytics destinations for immediate querying.