Thank you for visiting my profile.
Hi, I'm Rohit Sharma — a Data Automation Engineer with 4+ years of experience designing and building scalable data pipelines, ETL systems, and cloud-based data platforms.
I specialize in transforming raw, messy data into reliable, analytics-ready datasets that power business decisions.
🚀 What I Do
✔ Build end-to-end ETL/ELT pipelines using PySpark, SQL, and Python
✔ Migrate and modernize legacy data pipelines (Azure → GCP, on-prem → cloud)
✔ Design scalable data lake and warehouse architectures
✔ Implement CI/CD pipelines for data workflows (GitLab, dbt)
✔ Automate data ingestion from APIs, SFTP, email, and third-party systems
✔ Ensure data quality, validation, and monitoring in production pipelines
💼 Recent Work Highlights
• Migrated legacy pipelines from Azure & Pentaho to GCP using Apache Spark, improving scalability and reducing costs
• Built end-to-end batch pipelines to process multi-source data with high reliability and performance
• Developed data synchronization between Unity Catalog and Hive Metastore for cross-platform data visibility
• Implemented CI/CD workflows using GitLab and dbt for faster and safer deployments
• Automated repetitive processes using Selenium-based RPA solutions
🛠 Tech Stack
Languages: Python, SQL
Big Data: Apache Spark, PySpark, Spark SQL
Cloud: Azure (ADF, ADLS), Google Cloud (GCS, Dataproc, Composer), Databricks
Tools: Airflow, dbt, Git, Azure DevOps
Other: Web Scraping (Scrapy, Selenium, BeautifulSoup), REST APIs
💡 Why Work With Me?
✔ Strong experience in real-world production data systems
✔ Focus on performance, scalability, and clean architecture
✔ Reliable communication and quick turnaround
✔ Ability to understand business needs and translate them into data solutions
Let’s connect and discuss how I can help you build efficient and scalable data solutions.
“Let your dreams be bigger than your fears and your actions louder than your words.”
Python
Data Engineering
SQL
Data Extraction
API Integration
Data Scraping
Data Mining
RESTful API
NoSQL Database
Email Automation
Robotic Process Automation
Microsoft Azure
Google Cloud Platform
Data Analytics
Data Modeling
GitLab
dbt
Data Lake
Data Migration
Databricks Platform
Abha K.
Mumbai, India
$56/hr
5.0
9 jobs
🚀 Data Engineer & Solution Architect | Scaling Data Platforms 10× Without Breaking Them
I design data systems that don’t just run, they scale, perform, and stay reliable under real-world pressure.
With 7+ years building enterprise-grade platforms, I’ve seen the same story repeat:
A pipeline works at 10M records… then collapses at 100M.
Costs spiral. Latency explodes. Nobody wants to touch the legacy system.
That’s where I come in.
🧠 What I Actually Deliver
I architect cloud-native data platforms built for tomorrow not quick fixes for today.
✔ Migrate fragile legacy systems to modern, resilient architectures
✔ Design scalable data lakes and lakehouses
✔ Optimize pipelines bleeding money and compute
✔ Build real-time analytics for mission-critical decisions
✔ Create foundations ready for AI/ML workloads
Result: Systems that grow with your business instead of holding it back.
⚙️ Deep Technical Expertise Across the Stack
☁️ Cloud Platforms
AWS: Glue, EMR, Redshift, Kinesis, S3, Lambda, Lake Formation, DMS, MSK, RDS
Azure: Data Factory, Synapse, Databricks, DevOps
GCP: Dataflow, Cloud Functions, Cloud Storage
🔥 Big Data & Streaming
Apache Spark (Scala & PySpark) • Kafka • Kinesis • NiFi • Hadoop Ecosystem • Airflow • Delta Lake
💻 Programming
Python • Scala • SQL • Shell • Java
🗄️ Databases & Storage
PostgreSQL • MySQL • Oracle • SQL Server • MongoDB • Cassandra • DynamoDB • Elasticsearch
🛠️ DevOps & Infrastructure
Docker • Kubernetes • OpenShift • Terraform • Jenkins • Ansible • Git
📊 Observability & Governance
CloudWatch • ELK • Grafana • Athena
IAM • Lake Formation • Encryption • Audit Logging • Okta • Cognito
🏢 Enterprise Experience That Matters
I’ve delivered production systems for Fortune 500 organizations across finance, energy, hospitality, and SaaS handling hundreds of millions of records daily.
From ingestion → transformation → real-time analytics → security → DevOps automation — I design the full lifecycle.
🏆 Proven Impact
✔ Re-architected legacy pipelines → 5× performance boost & 60% cost reduction
✔ Built event-driven systems processing 500M+ records/day
✔ Delivered secure data lakes with row-level governance
✔ Reduced MTTR by 70% with end-to-end observability
✔ Led zero-downtime cloud migrations
✔ Secured $2B+ transaction data with encryption platforms
🤝 Best Fit For Organizations That Need
🔹 Cloud migration with strong architectural guidance
🔹 Performance or scalability bottlenecks
🔹 Data platforms for AI/ML initiatives
🔹 Multi-cloud or hybrid strategies
🔹 Long-term reliability over quick hacks
⚠️ Not a Fit For
❌ One-off scripts or basic SQL tasks
❌ Temporary data cleanup work
❌ Short-term patch solutions
I focus where architecture decisions create lasting business value.
💬 What Clients Value Most
Clear thinking on complex problems
Communication executives understand
Engineering teams trust
Systems built to last
👉 If your data platform needs to scale, stabilize, or modernize then let’s talk.
Amazon Web Services
Google Cloud Platform
Elasticsearch
Python
Scala
MongoDB
Microsoft Azure
PostgreSQL
Apache Spark
Apache Kafka
Kibana
Grafana
Big Data
ETL Pipeline
Databricks Platform
PySpark
Apache NiFi
Anup S.
Patna, India
$45/hr
5.0
47 jobs
Slow, unreliable, or expensive data pipelines cost you time and money — and most teams don't realize how much until reporting breaks or cloud bills spike. I help companies fix that by building fast, well-architected data platforms on Snowflake.
I'm a Snowflake Data Engineer with 8+ years in cloud data engineering, including 4+ years hands-on building production data platforms on Snowflake. SnowPro Certified, with deep expertise in Python, dbt, Airflow, and AWS — I design ELT pipelines and analytics platforms that are reliable, cost-optimized, and built to scale.
My work spans the full lifecycle: designing warehouses and pipelines from scratch, migrating legacy systems onto Snowflake, and optimizing performance and cost on platforms that were already live — on more than one project, that optimization work meaningfully cut query times and warehouse spend for teams stuck on inefficient setups. I build the transformation and orchestration layers (dbt, Airflow) that make data reliably usable, backed by data quality checks so issues get caught before they reach a dashboard.
Core Expertise
✅ Snowflake Data Warehouse Design & Implementation
✅ Python-Based Data Engineering & Automation
✅ ELT/ETL Pipeline Development
✅ dbt Development & Analytics Engineering
✅ Snowflake Performance & Cost Optimization
✅ Data Modeling (Star Schema & Dimensional Modeling)
✅ Apache Airflow Workflow Orchestration
✅ AWS Data Engineering Solutions
✅ Data Migration & Modernization
✅ Data Quality, Validation & Monitoring
Whether you're implementing Snowflake for the first time, migrating from a legacy platform, optimizing an existing data warehouse, or building modern ELT pipelines, I can help deliver a scalable data platform that supports both current needs and future growth.
Snowflake
Data Engineering
dbt
SQL
Python
Amazon Web Services
PySpark
Apache Airflow
Data Modeling
Apache Kafka
Data Warehousing
Performance Optimization
ETL Pipeline
Fivetran
Docker
Databricks Platform
ETL
FastAPI
Amazon S3
AWS Glue
Jayant C.
Gandhinagar, India
$20/hr
4.9
31 jobs
✅ Top Rated Plus | 100% JSS | 4x Certified (AWS SA Pro, GCP Pro Architect, Snowflake) | BITS Pilani MTech Data Science | Full Stack Developer & Data Engineer | React, Python, Node.js, Spark | $20K+ earned | 1,845+ hours
I build full-stack web applications and data engineering systems that go to production, not to demo day. SaaS MVPs, Spark-based ETL pipelines, cloud architecture on AWS and GCP, I handle both the application layer and the data infrastructure behind it.
🔹 Full-Stack SaaS & Web Application Development
React, Next.js, Node.js, and Python backends for SaaS platforms, dashboards, internal tools, and customer-facing apps. MVP to production on AWS/GCP with CI/CD, automated testing, and monitoring from day one. 19 Upwork contracts delivered with structured milestones.
🔹 Data Engineering & ETL Pipeline Architecture
End-to-end data pipeline design with Apache Spark, PySpark, Scala, Snowflake, and Airflow. Batch and streaming ETL processing millions of records per run. Data lake architecture, warehouse modeling, analytics-ready output layers. 8+ years building production Spark + Cassandra systems at enterprise scale.
🔹 Cloud Architecture & Infrastructure (AWS + GCP)
4 cloud architecture projects on Upwork, all rated 5.0. $2,300 CloudStack design. AWS architecture advisory. EC2, Lambda, S3, RDS, EMR, Redshift on AWS. BigQuery, Dataflow, Cloud Functions on GCP. Terraform for IaC, Docker and Kubernetes for orchestration, zero-downtime deployments.
🔹 API Development & Backend Systems
REST API and GraphQL backends with Node.js, NestJS, FastAPI, and Django. Microservices, Redis caching, WebSocket integrations, Stripe payment APIs, OAuth/JWT authentication. Backend services handling concurrent users at production scale.
🔹 Database Design & Data Modeling
PostgreSQL, MongoDB, MySQL, Cassandra, DynamoDB, Redis. Schema design, query tuning, indexing, partitioning. Star and snowflake schemas, slowly changing dimensions, SQL optimization for analytics. Architecture decisions balancing performance, throughput, and cost.
🔹 AI Integration & Intelligent Applications
OpenAI API, Hugging Face, NLP pipelines, chatbot systems, text extraction and summarization. Delivered NLP processing on Upwork. AI-powered features built into SaaS products as production features, not standalone experiments.
🔹 Real-Time Processing & Event-Driven Systems
Kafka for event-driven architectures, change data capture, WebSocket dashboards, streaming pipelines for near-real-time analytics. Application events connected to data warehouse layers.
🔹 Frontend Performance & TypeScript Engineering
React and Next.js with SSR/SSG for SEO-friendly rendering. TypeScript full stack. Core Web Vitals optimization, Tailwind CSS, responsive design. Fast-loading frontends that rank and convert.
🔹 DevOps, CI/CD & Production Systems
Docker, Kubernetes, Terraform, GitHub Actions, GitLab CI. Serverless with AWS Lambda and GCP Cloud Functions. Monitoring, logging, alerting for production. Zero-downtime deployment strategies.
🔹 Technical Consulting & Architecture Advisory
TypeScript and AWS Lambda tutor on Upwork, rated 5.0 over 13 hours. Cloud migration advisory, system design review, code audits, performance optimization, engineering mentorship.
📊 AWS Solutions Architect Professional + Associate (Dec 2026) | GCP Pro Cloud Architect (Jul 2026) | Snowflake Core (Jan 2026)
📊 MTech Data Science, BITS Pilani, ranked top 5 engineering institutions in India
📊 19 contracts, 100% JSS, Top Rated Plus, 1,845+ hours tracked, $20K+ earned
📊 "Jay's expertise brought the architecture design to life in ways I hadn't imagined" (5.0 rated)
📊 8+ years: React, Node.js, Python, Java, Scala across SaaS, healthcare, fintech, enterprise
→ Day 1: Requirements call + architecture proposal with tech stack rationale
→ Week 1: Sprint development, daily Loom/Slack updates, working code shipped
→ Ongoing: Weekly demos, priority reviews, transparent tracking, full documentation
→ Delivery: Documented code, CI/CD configured, deployment guide, 2-week post-launch support
Full Stack: React, Next.js, Node.js, NestJS, Express, TypeScript, JavaScript, Python, FastAPI, Django
Data: Apache Spark, PySpark, Scala, Snowflake, Airflow, Kafka, ETL, dbt, SQL, BigQuery
Cloud: AWS (Lambda, EC2, S3, RDS, EMR, Redshift), GCP (BigQuery, Dataflow), Docker, Kubernetes, Terraform
DB: PostgreSQL, MongoDB, MySQL, Redis, Cassandra, DynamoDB, Supabase
AI: OpenAI API, Hugging Face, NLP, LLM Integration, TensorFlow, PyTorch
💬 Message me with your project scope or data challenge. I respond within 4 hours with a free assessment and can start within 48 hours.
Java
Python
Apache Spark
Scala
SQL
React
Node.js
Full-Stack Development
Data Engineering
TypeScript
API Integration
PostgreSQL
Next.js
AWS Lambda
NestJS Development
Generative AI
Snowflake
DevOps
Google Cloud Platform
ETL
Shivam W.
Shahdara, India
$20/hr
5.0
8 jobs
I'm a Senior Data Engineer with 4.5+ years of experience building scalable, cloud-native data platforms that turn raw data into reliable, business-ready insights. I've delivered enterprise solutions across banking (NAB), healthcare (Molina), and CPG (PepsiCo), specializing in end-to-end pipeline architecture, data modeling, and cloud migrations.
What I bring to your project:
🔹 Cloud Data Engineering – Deep expertise in Azure (Databricks, Data Factory, Synapse) and AWS (EMR, Glue, S3, RedShift), with hands-on migration experience from on-prem and Teradata to cloud.
🔹 Pipeline Architecture & ETL – I design and build robust ingestion frameworks handling batch, incremental, and real-time data (Event Hub, Kafka) across formats like JSON, CSV, Parquet, and fixed-width files.
🔹 Data Modeling & Warehousing – Skilled in dimensional modeling, Data Vault, star/snowflake schemas, and silver/gold layer design. I've modeled 50+ tables across Oracle Fusion, SAP S/4, and healthcare domains.
🔹 Transformation & Orchestration – I translate complex business rules into DBT models, orchestrate workflows with Apache Airflow or AutoSys, and automate CI/CD via Jenkins and Azure DevOps.
🔹 Performance & Governance – I tune PostgreSQL and Spark jobs, implement data quality checks, reconciliation frameworks, and ensure compliance with data governance standards.
🔹 Generative AI & MLOps – Databricks-certified in Generative AI, with experience integrating MLflow for experiment tracking and building LLM-based automation using OpenAI and LangChain.
Tech Stack: Python | SQL | Scala | Apache Spark | DBT | PostgreSQL | Snowflake | Airflow | Databricks | Azure | AWS | Git | Jenkins | MLflow | Power BI
Certifications: Databricks Certified Data Engineer Professional | Azure Data Engineer (DP-203) | Snowflake SnowPro Core | Fabric Analytics Engineer (DP-600) | Generative AI Engineer Associate
Whether you need a production-grade pipeline, a cloud migration, or a well-modeled data warehouse, I deliver clean, documented, and scalable solutions — on time and with clear communication. Let's discuss your project!
Data Extraction
Data Mining
Artificial Intelligence
ETL Pipeline
Machine Learning
Database Design
Database Modeling
PySpark
Databricks Platform
Snowflake
Data Warehousing
Apache Airflow
Python
Web Scraping
Data Engineering
Generative AI
Exploratory Data Analysis
Scala
Data Integration
Piyush M.
Bangalore, India
$14/hr
4.6
8 jobs
Helping companies build scalable, reliable, and cost-efficient data platforms.
I'm a Principal Data Engineer with 11+ years of experience designing and implementing modern data engineering solutions for startups, fintech companies, healthcare organizations, and enterprise businesses. I've helped organizations migrate legacy systems, build cloud-native data platforms, optimize processing costs, and deliver production-ready analytics pipelines.
My expertise includes designing end-to-end data architectures, building batch and streaming pipelines, implementing Data Lakes and Lakehouses, and automating infrastructure using Infrastructure as Code.
What I can help you with
✔ Databricks Development & Optimization
✔ Apache Spark (PySpark & Scala)
✔ Azure Data Factory (ADF)
✔ Azure Data Lake Storage (ADLS)
✔ Delta Lake & Delta Live Tables
✔ AWS (EMR, Glue, Athena, Lambda, S3)
✔ Data Warehouse Design
✔ ETL / ELT Pipelines
✔ Data Migration
✔ Data Modeling
✔ Terraform & Infrastructure Automation
✔ SQL Performance Optimization
✔ Python Development
✔ CI/CD for Data Platforms
✔ Airflow Workflow Automation
✔ AI-powered Workflow Automation (Cursor, Claude, MCP, n8n)
Recent accomplishments
• Reduced operational costs by 90% by redesigning SCD implementation using Delta Live Tables.
• Led the architecture and delivery of financial products including Loans and Credit Cards.
• Migrated enterprise data warehouses to cloud-native lakehouse architecture.
• Built scalable reconciliation frameworks using Databricks and Airflow.
• Implemented Terraform-managed Databricks infrastructure for improved governance and scalability.
• Designed enterprise-grade data platforms for healthcare, fintech, and retail organizations.
My Skills Sets are:
SQL, Apache Spark, Hive, Hadoop, Excel, Shell Scripting, AWS EMR, Ec2, S3, cloud formation, Clojure, MongoDB MySQL, Airflow.
Apache Hadoop
Python
SQL
Apache Spark
Clojure
Amazon S3
AWS Lambda
Apache Hive
Amazon EC2
Bash Programming
Databricks Platform
How it works
Post a job for freePost a job
Tell us what you need. Create your own job post or generate one with AI then filter talent matches.
Hire top talent fast
Consult, interview, and hire quickly, so you can meet the freelancers you're excited about.
Collaborate easily
Use Upwork to chat or video call, share files, and track project progress right from the app.
Payment simplified
Manage payments in one place with flexible billing options. Only pay for approved work, hourly or by milestone.
Don't just take our word for it
“Upwork provides an umbrella-level of security. I can see a talent’s work history and ratings. I can hold payments in escrow. I can communicate through Upwork Messages instead of working through my email address.”
KD
Kim Darling
Emerald Tiger
“Upwork is the best platform to hire skilled professionals when we're not looking for a full-time employee. All the companies in our portfolio use Upwork to find talent across a wide range of fields.”
DM
David Merry
Kinetic Investments
“Our very specific requirements can be a challenge—With Upwork, we’re able to access a bigger community to ensure the success of our projects.”
KK
Katja Krohn
Summa Linguae
How do I hire a Hadoop Developer & Programmer in India on Upwork?
You can hire a Hadoop Developer & Programmer in India on Upwork in four simple steps:
Create a job post tailored to your Hadoop Developer & Programmer project scope. We'll walk you through the process step by step.
Browse top Hadoop Developer & Programmer talent on Upwork and invite them to your project.
Once the proposals start flowing in, create a shortlist of top Hadoop Developer & Programmer profiles and interview.
Hire the right Hadoop Developer & Programmer for your project from Upwork, the world's largest work marketplace.
At Upwork, we believe talent staffing should be easy.
How much does it cost to hire a Hadoop Developer & Programmer?
Rates charged by Hadoop Developers & Programmers on Upwork can vary with a number of factors including experience, location, and market conditions. See hourly rates for in-demand skills on Upwork.
Why hire a Hadoop Developer & Programmer in India on Upwork?
As the world's work marketplace, we connect highly-skilled freelance Hadoop Developers & Programmers and businesses and help them build trusted, long-term relationships so they can achieve more together. Let us help you build the dream Hadoop Developer & Programmer team you need to succeed.
Can I hire a Hadoop Developer & Programmer in India within 24 hours on Upwork?
Depending on availability and the quality of your job post, it's entirely possible to sign up for Upwork and receive Hadoop Developer & Programmer proposals within 24 hours of posting a job description.
Find more freelancers
Top cities for Hadoop Developers & Programmers in India