Data Engineer & AI Developer | 3+ Years Financial Industry Experience
I build data pipelines, AI-powered applications, and automation systems that run reliably at scale. My background spans web scraping, LLM integration, computer vision, betting automation, and full-stack data dashboards — delivered to clients across the US, UK, Europe, and Japan.
💼 Background
— 3+ years at a leading Indian bank building risk models, credit scorecards, and AutoML pipelines
— PG Diploma in Big Data Analysis
⚡ What I Deliver
— Web scrapers handling 1.2M+ URLs and 120K daily pipelines
— LLM/AI apps using GPT-4, Gemini, LangChain, RAG, Text-to-SQL
— Full Betting automation for horse racing, golf, and football signals
— Computer vision pipelines with YOLOv8 and PaddleOCR
— Streamlit dashboards, risk scorecards, and AutoML tools
🏆 Notable Work
— PitchBook scraper — 1.2M URLs
— Njuskalo — 120K daily real estate listings
— Text-to-SQL architecture
— BetFare — full Betfair automation
— LLM Notebook — $1,420 solo delivery
— Anti-bot bypass systems
🛠️ Stack
Python · Playwright · Selenium · GPT-4 · Gemini · LangChain · Streamlit · PySpark · SQL · YOLOv8 · PaddleOCR · FastAPI · Betfair API · n8n
Clean code. Clear communication. Delivered on time.
PySpark
Data Analysis
Python
SQL
Java
Front-End Development
Streamlit
Data Science
AI Chatbot
API
Web Scraping
Selenium
PyQt
YOLO
Gowthaman N.
Salem, India
$20/hr
5.0
4 jobs
📊 Data drives every successful business decision. I help organizations design, build, and optimize modern data platforms that transform raw data into reliable business intelligence.
🚀 Whether you need to build a cloud-native data platform, migrate legacy systems, automate ETL pipelines, or deliver executive dashboards, I provide scalable, production-ready solutions that generate measurable business value.
⭐ With 5+ years of experience and a strong track record, I have partnered with startups, SMBs, and enterprise organizations to deliver high-quality data engineering and analytics solutions.
🔧 Core Expertise:
✅ Microsoft Fabric (OneLake, Data Factory, Lakehouse, Warehouse, Notebooks)
✅ Azure Data Engineering, Azure Databricks, Apache Spark, and Delta Lake
✅ Azure Data Factory (ADF) and Azure Synapse Analytics
✅ ETL and ELT Pipeline Development
✅ Data Warehousing and Data Modeling
✅ Data Integration and Data Migration
✅ Power BI Dashboards, Semantic Models, and DAX
✅ Python, SQL, PySpark, and Data Processing
💡 Services:
✅ Design and implement end-to-end data engineering solutions
✅ Build scalable ETL and ELT pipelines
✅ Develop enterprise data warehouses and Lakehouse architectures
✅ Integrate APIs, ERP, CRM, SQL databases, and cloud platforms
✅ Migrate legacy data platforms to Microsoft Fabric, Azure, and Databricks
✅ Implement data quality, validation, governance, and monitoring
✅ Create interactive Power BI dashboards and executive reports
✅ Optimize data platforms for performance, scalability, and cost efficiency
🛠️ Technologies: Microsoft Fabric, Microsoft Azure, Azure Databricks, Azure Synapse Analytics, Azure Data Factory, Snowflake, BigQuery
💻 Programming: Python, SQL, PySpark, DAX
🗄️ Databases: SQL Server, PostgreSQL, MySQL
📈 Analytics: Power BI, Tableau, Looker Studio
🤝 Why Clients Choose Me:
✅ Strong focus on scalable, maintainable, and production-ready data solutions
✅ Extensive experience modernizing legacy data platforms
✅ Expertise across Data Engineering, Business Intelligence, and Cloud Analytics
✅ Clear communication and collaborative project delivery
✅ Proven track record of delivering high-quality solutions on time
📩 If you are looking for an experienced Azure Data Engineer, Microsoft Fabric Consultant, Databricks Engineer, Power BI Developer, or Data Engineering Specialist, I would be pleased to discuss your project.
Apache Spark
PySpark
ETL Pipeline
SQL
Fabric
Microsoft Azure
Databricks Platform
Amazon Web Services
Data Warehousing
API Integration
Python
n8n
Business Process Automation
AI Agent Development
Make.com
LangChain
Automation
CRM Automation
Artificial Intelligence
Abha K.
Mumbai, India
$56/hr
5.0
9 jobs
🚀 Data Engineer & Solution Architect | Scaling Data Platforms 10× Without Breaking Them
I design data systems that don’t just run, they scale, perform, and stay reliable under real-world pressure.
With 7+ years building enterprise-grade platforms, I’ve seen the same story repeat:
A pipeline works at 10M records… then collapses at 100M.
Costs spiral. Latency explodes. Nobody wants to touch the legacy system.
That’s where I come in.
🧠 What I Actually Deliver
I architect cloud-native data platforms built for tomorrow not quick fixes for today.
✔ Migrate fragile legacy systems to modern, resilient architectures
✔ Design scalable data lakes and lakehouses
✔ Optimize pipelines bleeding money and compute
✔ Build real-time analytics for mission-critical decisions
✔ Create foundations ready for AI/ML workloads
Result: Systems that grow with your business instead of holding it back.
⚙️ Deep Technical Expertise Across the Stack
☁️ Cloud Platforms
AWS: Glue, EMR, Redshift, Kinesis, S3, Lambda, Lake Formation, DMS, MSK, RDS
Azure: Data Factory, Synapse, Databricks, DevOps
GCP: Dataflow, Cloud Functions, Cloud Storage
🔥 Big Data & Streaming
Apache Spark (Scala & PySpark) • Kafka • Kinesis • NiFi • Hadoop Ecosystem • Airflow • Delta Lake
💻 Programming
Python • Scala • SQL • Shell • Java
🗄️ Databases & Storage
PostgreSQL • MySQL • Oracle • SQL Server • MongoDB • Cassandra • DynamoDB • Elasticsearch
🛠️ DevOps & Infrastructure
Docker • Kubernetes • OpenShift • Terraform • Jenkins • Ansible • Git
📊 Observability & Governance
CloudWatch • ELK • Grafana • Athena
IAM • Lake Formation • Encryption • Audit Logging • Okta • Cognito
🏢 Enterprise Experience That Matters
I’ve delivered production systems for Fortune 500 organizations across finance, energy, hospitality, and SaaS handling hundreds of millions of records daily.
From ingestion → transformation → real-time analytics → security → DevOps automation — I design the full lifecycle.
🏆 Proven Impact
✔ Re-architected legacy pipelines → 5× performance boost & 60% cost reduction
✔ Built event-driven systems processing 500M+ records/day
✔ Delivered secure data lakes with row-level governance
✔ Reduced MTTR by 70% with end-to-end observability
✔ Led zero-downtime cloud migrations
✔ Secured $2B+ transaction data with encryption platforms
🤝 Best Fit For Organizations That Need
🔹 Cloud migration with strong architectural guidance
🔹 Performance or scalability bottlenecks
🔹 Data platforms for AI/ML initiatives
🔹 Multi-cloud or hybrid strategies
🔹 Long-term reliability over quick hacks
⚠️ Not a Fit For
❌ One-off scripts or basic SQL tasks
❌ Temporary data cleanup work
❌ Short-term patch solutions
I focus where architecture decisions create lasting business value.
💬 What Clients Value Most
Clear thinking on complex problems
Communication executives understand
Engineering teams trust
Systems built to last
👉 If your data platform needs to scale, stabilize, or modernize then let’s talk.
Apache Spark
PySpark
Amazon Web Services
Google Cloud Platform
Elasticsearch
Python
Scala
MongoDB
Microsoft Azure
PostgreSQL
Apache Kafka
Kibana
Grafana
Big Data
ETL Pipeline
Databricks Platform
Apache NiFi
Anup S.
Patna, India
$45/hr
5.0
47 jobs
Slow, unreliable, or expensive data pipelines cost you time and money — and most teams don't realize how much until reporting breaks or cloud bills spike. I help companies fix that by building fast, well-architected data platforms on Snowflake.
I'm a Snowflake Data Engineer with 8+ years in cloud data engineering, including 4+ years hands-on building production data platforms on Snowflake. SnowPro Certified, with deep expertise in Python, dbt, Airflow, and AWS — I design ELT pipelines and analytics platforms that are reliable, cost-optimized, and built to scale.
My work spans the full lifecycle: designing warehouses and pipelines from scratch, migrating legacy systems onto Snowflake, and optimizing performance and cost on platforms that were already live — on more than one project, that optimization work meaningfully cut query times and warehouse spend for teams stuck on inefficient setups. I build the transformation and orchestration layers (dbt, Airflow) that make data reliably usable, backed by data quality checks so issues get caught before they reach a dashboard.
Core Expertise
✅ Snowflake Data Warehouse Design & Implementation
✅ Python-Based Data Engineering & Automation
✅ ELT/ETL Pipeline Development
✅ dbt Development & Analytics Engineering
✅ Snowflake Performance & Cost Optimization
✅ Data Modeling (Star Schema & Dimensional Modeling)
✅ Apache Airflow Workflow Orchestration
✅ AWS Data Engineering Solutions
✅ Data Migration & Modernization
✅ Data Quality, Validation & Monitoring
Whether you're implementing Snowflake for the first time, migrating from a legacy platform, optimizing an existing data warehouse, or building modern ELT pipelines, I can help deliver a scalable data platform that supports both current needs and future growth.
PySpark
Snowflake
Data Engineering
dbt
SQL
Python
Amazon Web Services
Apache Airflow
Data Modeling
Apache Kafka
Data Warehousing
Performance Optimization
ETL Pipeline
Fivetran
Docker
Databricks Platform
ETL
FastAPI
Amazon S3
AWS Glue
Vivek M.
Surat, India
$30/hr
5.0
114 jobs
With 7+ years of experience, I'm Expert in Web Scraping, Data Engineer, AI/ML and Full-Stack Developer specializing in large-scale data extraction, automation, and pipeline engineering. I build robust, scalable systems that transform raw data into actionable insights.
💡 Core Expertise
Web Scraping & Automation: Expert in bypassing anti-bot systems (CAPTCHA, rate limits, IP rotation) using Scrapy, BeautifulSoup, Selenium, Playwright, and rotating proxies.
Automation & Workflow Engineering: Airflow, Prefect, Dagster, n8n, Zapier, Make, Power Automate, UiPath, Step Functions, Logic Apps, GCP Workflows, Business Process Automation, RPA, CI/CD, Jenkins, GitHub Actions, GitLab CI/CD, Monitoring & Alerting.
Data Engineering: Designing and building scalable ETL/ELT pipelines for structured, semi-structured, and unstructured data using Apache Airflow, Apache Spark (PySpark), Pandas, Dask, Databricks, Snowflake, Apache Kafka, Apache Hive, Apache Hadoop, Delta Lake, Apache Iceberg, dbt, AWS Glue, Azure Data Factory, Google Cloud Dataflow, Apache NiFi, Trino, Presto, and Apache Beam.
Experienced in data warehousing, data lakes, lakehouse architectures, data modeling, data transformation, data quality, data governance, batch and real-time processing, streaming data pipelines, orchestration, workflow automation, schema design, partitioning, optimization, and performance tuning. Proficient with cloud platforms including AWS, Azure, and GCP, S3, Redshift, EMR, Athena, Lambda, Azure Synapse Analytics, Azure Data Lake Storage, BigQuery, Cloud Storage, and Pub/Sub. Skilled in SQL, Python, data integration, data migration, CDC, metadata management, monitoring, CI/CD, Docker, Kubernetes, and modern data stack technologies.
Backend Development: High-performance APIs and microservices with FastAPI, Django, Flask, and Celery for async task handling.
AI/ML Integration: Leveraging NLP and LLMs (LangChain, Llama, NLTK) for data enrichment, classification, and intelligent automation.
Cloud & DevOps: Deploying scalable scrapers and data workflows on AWS (Lambda, ECS, S3), GCP, Docker, and Kubernetes.
🛠️ Tech Stack
Data & Scraping:
▸ Scrapy | Selenium | Playwright | Proxies (BrightData, ScraperAPI, etc)
▸ Pandas | PySpark | Apache Airflow | PostgreSQL | MongoDB | Redis
Backend & Cloud:
▸ Python (FastAPI, Django, Flask) | Celery | RabbitMQ
▸ AWS (Lambda, ECS, RDS, S3) | GCP | Docker | Kubernetes
AI/ML:
▸ NLP (NLTK, spaCy) | LLMs (LangChain, OpenAI, Llama) | Data Annotation
Let's turn your data challenges into reliable, scalable solutions. Send me a message to discuss your project!
Python
Data Scraping
Data Mining
Scrapy
Selenium
Scripting
Web Crawling
Data Extraction
JavaScript
AWS Lambda
Node.js
Web Scraping
Data Engineering
Flask
Django
Arvind H.
Bengaluru, India
$50/hr
5.0
2 jobs
Data warehouse developer with 11 years of experience in pyspark and SQL spark.
Currently working as a Palantir foundry implementation specialist.
Have worked on MSBI technology for building data warehouse and datamarts
PySpark
SQL
Data Warehousing & ETL Software
SQL Server Integration Services
SQL Server Reporting Services
How it works
Post a job for freePost a job
Tell us what you need. Create your own job post or generate one with AI then filter talent matches.
Hire top talent fast
Consult, interview, and hire quickly, so you can meet the freelancers you're excited about.
Collaborate easily
Use Upwork to chat or video call, share files, and track project progress right from the app.
Payment simplified
Manage payments in one place with flexible billing options. Only pay for approved work, hourly or by milestone.
Don't just take our word for it
“Upwork provides an umbrella-level of security. I can see a talent’s work history and ratings. I can hold payments in escrow. I can communicate through Upwork Messages instead of working through my email address.”
KD
Kim Darling
Emerald Tiger
“Upwork is the best platform to hire skilled professionals when we're not looking for a full-time employee. All the companies in our portfolio use Upwork to find talent across a wide range of fields.”
DM
David Merry
Kinetic Investments
“Our very specific requirements can be a challenge—With Upwork, we’re able to access a bigger community to ensure the success of our projects.”
KK
Katja Krohn
Summa Linguae
How do I hire a Pyspark Developer in India on Upwork?
You can hire a Pyspark Developer in India on Upwork in four simple steps:
Create a job post tailored to your Pyspark Developer project scope. We'll walk you through the process step by step.
Browse top Pyspark Developer talent on Upwork and invite them to your project.
Once the proposals start flowing in, create a shortlist of top Pyspark Developer profiles and interview.
Hire the right Pyspark Developer for your project from Upwork, the world's largest work marketplace.
At Upwork, we believe talent staffing should be easy.
How much does it cost to hire a Pyspark Developer?
Rates charged by Pyspark Developers on Upwork can vary with a number of factors including experience, location, and market conditions. See hourly rates for in-demand skills on Upwork.
Why hire a Pyspark Developer in India on Upwork?
As the world's work marketplace, we connect highly-skilled freelance Pyspark Developers and businesses and help them build trusted, long-term relationships so they can achieve more together. Let us help you build the dream Pyspark Developer team you need to succeed.
Can I hire a Pyspark Developer in India within 24 hours on Upwork?
Depending on availability and the quality of your job post, it's entirely possible to sign up for Upwork and receive Pyspark Developer proposals within 24 hours of posting a job description.