Hire the Best Apache Spark Engineers
in the United States

More than 3,000 reviews on G2
Rating is 4.5 out of 5.
4.5/5
of Upwork by G2 peer reviewers
Chris K.

Chattanooga, Tennessee

$100/hr
4.9
19 jobs

My name is Chris and I am currently working as a Data Engineering contractor. As a certified Databricks Data Engineer with 10 years of Python experience, I specialize in building scalable data pipelines and leveraging Databricks to process large-scale streaming datasets. My background in Data Engineering includes deploying and managing Databricks instances in a variety of Cloud Environments (including Microsoft Azure), implementing CI/CD pipelines with Databricks Asset Bundles (DAB), and utilizing tools like dbt for analytics engineering. During this time, I have performed data engineering tasks and built data pipelines for a wide variety of organizations, including e-Commerce, Insurance, GRC, and a multitude of federal customers, both in the US and the German Bundeswehr.

  • Databricks Platform
  • Data Engineering
  • Data Science
  • Python
  • CI/CD
  • Analytics Dashboard
  • SQL
  • Database Management
  • ETL Pipeline
  • Machine Learning
  • Tableau
  • dbt
  • Git
  • GitHub
  • Interactive Data Visualization
  • Looker
  • Data Analytics & Visualization Software
  • Adobe Spark
  • Apache Kafka
  • Ecommerce
Matthew D.

Kansas City, Missouri

$100/hr
5.0
2 jobs

Principal AI Engineer | GenAI, Edge AI, RAG & Agentic Workflows I build production-ready AI solutions, not just prototypes and demos. I am Matthew - a Principal AI Engineer and Data Scientist with over 20 years of experience solving complex enterprise technology and data problems. I specialize in Generative AI, Agentic workflows, machine learning, data engineering, and anticipating the next frontier of intelligent automation. Currently serving as a Principal AI Engineer at a Fortune 50 enterprise and holding an M.S. in Data Science from Northwestern University, I bring enterprise-grade architecture and rigor to businesses of all sizes. I don't just connect applications to an API; I understand the entire AI lifecycle. Furthermore, I architect future-proof systems—leveraging emerging paradigms like Edge AI, Small Language Models (SLMs), and Multi-Agent Swarms to ensure your tech stack is ready for the demands of 2027, 2030, and beyond. Here is how I can help you build something that actually works: 🔹 EDGE AI & NEXT-GEN ARCHITECTURE (2027+ Readiness) Edge AI & TinyML: Deploying lightweight, high-performance ML and AI models directly to IoT and edge devices for zero-latency, offline, and privacy-first capabilities. Small Language Models (SLMs) & Local AI: Fine-tuning and deploying highly efficient, domain-specific models that drastically cut cloud compute costs and keep enterprise data secure on-premise. Federated Learning: Architecting decentralized model training across distributed networks to maximize data privacy. Multi-Modal AI: Seamlessly integrating real-time vision, audio, and spatial data streams for advanced physical and ambient AI applications. 🔹 GENERATIVE AI & LLM APPLICATIONS Custom GenAI applications & Enterprise LLM/SLM solution architecture Prompt engineering, optimization, and structured outputs (tool calling) AI-powered document and intelligent workflow automation LLM evaluation, testing, guardrails, and optimization 🔹 AI AGENTS & AGENTIC WORKFLOWS Multi-agent swarms and complex autonomous AI orchestration LangGraph workflows & Tool-enabled agents Human-in-the-loop workflows & Model Context Protocol (MCP) Agent evaluation and enterprise production readiness 🔹 RAG & ENTERPRISE SEARCH Retrieval-Augmented Generation (RAG) architecture Embeddings, vector databases, and semantic search Knowledge-base assistants and document ingestion pipelines Retrieval accuracy and groundedness evaluation 🔹 MACHINE LEARNING & DATA ENGINEERING Predictive modeling, classification, segmentation, and anomaly detection Forecasting, time-series analysis, and recommendation systems Large-scale data processing (PySpark, Apache Spark, Databricks, Snowflake) MLOps architecture, continuous learning, and model validation 🏆 MY BACKGROUND & CREDENTIALS: Experience: 20+ years in enterprise tech, data, analytics, ML, and AI. Current Role: Principal AI Engineer / Data Scientist at a Fortune 50 enterprise. Education: M.S. in Data Science, Northwestern University. Innovation: U.S. Patent Inventor. Tech Stack: Azure, Databricks, Snowflake, Spark, Python, SQL, LangChain/LangGraph, Edge AI Frameworks, and modern AI/ML platforms. I am equally comfortable designing forward-looking AI architecture, building complex agentic workflows hands-on with Python, or translating deep technical concepts into clear business value for executives and stakeholders. Whether you need a cutting-edge Edge AI deployment, a robust RAG solution, or help taking an AI concept from a fragile idea to a secure, scalable production deployment, I am here to help. Have an AI, ML, or data challenge? Hit the "Invite" or "Hire" button, send me a message, and let's discuss what you’re building.

  • Apache Spark
  • Generative AI
  • LLM Prompt Engineering
  • Artificial Intelligence
  • Machine Learning
  • Data Science
  • MLOps
  • Data Engineering
  • Natural Language Processing
  • Edge Computing
  • Prompt Engineering
  • Predictive Analytics
  • Deep Learning
  • Microsoft Azure
  • Databricks Platform
  • Snowflake
  • Vector Database
  • Big Data
  • Python
  • Data Science Consultation
Shaista R.

Lake Grove, New York

$65/hr
5.0
7 jobs

Most data pipelines don’t fail because of code. They fail because they weren't built for scale. With 8+ years of experience engineering data systems at companies like Microsoft and Coreweave, I help businesses move away from "brittle prototypes" to production-grade, scalable infrastructure. I don’t just move data; I build the "Source of Truth" that leadership and AI systems actually trust. 💬What I Solve for You: Productionizing AI Pipelines: Hardening Python prototypes into scalable RAG and LLM infrastructures (AWS/Azure). ➔Infrastructure-as-Code: Building automated, modular ETL/ELT pipelines that don't require daily manual fixes. ➔The "One-Source" Dashboard: Integrating messy data from APIs, SaaS (Shopify, HubSpot), and DBs into clean Snowflake/BigQuery layers. ➔Performance Recovery: Optimizing slow SQL queries and high-cost cloud warehouses to save you thousands in monthly spend. 🛠 Tech Stack: Languages: Python (FastAPI, Pandas, PySpark), SQL Cloud & Warehousing: AWS (Glue, Lambda, S3), Snowflake, BigQuery, Azure Orchestration: Airflow, dbt, GitHub Actions Data Ops: API Integrations, Vector DBs, Data Validation ✅ Why Me? 8+ Years Experience: I’ve seen what breaks at the enterprise level and how to prevent it in your startup. Speed over Perfection: I focus on shipping high-impact systems that drive revenue, not just technical documentation. Transparent Communication: You get regular updates and a partner who challenges requirements to find better solutions. Ready to clean up your data debt? 📩 Message me for a FREE 15-minute technical consultation. Let’s discuss your architecture and see if I’m the right fit for your system.

  • Apache Spark
  • Data Engineering
  • Python
  • ETL Pipeline
  • SQL
  • Apache Airflow
  • Snowflake
  • Amazon Web Services
  • BigQuery
  • Data Warehousing
  • Data Modeling
  • Apache Kafka
  • PostgreSQL
  • Data Integration
  • Tableau
  • Docker
SERGEI K.

Denver, Colorado

$50/hr
5.0
1 jobs

I'm a Data Engineer with 3 years of experience, specializing in distributed databases and SQL optimization. In my recent role, I built a 60TB data warehouse using Greenplum that processes 2 million rides, generating 10-20 million records daily. I have extensive experience with MPP databases like Greenplum, ClickHouse, and Redshift. I'm proficient in SQL and Python, which I use for building ETL pipelines. In my previous role, I completely rebuilt a data processing system, improving performance from 17 hours to 2 hours, helping clients access reports much faster. I also automate workflows with Airflow and work with AWS services like Redshift and S3 for data storage and processing.

  • SQL
  • Python
  • Data Engineering
  • Greenplum
  • ETL Pipeline
  • Data Warehousing
  • pandas
  • PySpark
  • Apache Airflow
  • Snowflake
Matthew D.

New York City, New York

$70/hr
4.7
14 jobs

I’m Matt, a U.S.-based Data Scientist and AI Consultant with an M.S. in Data Science from Columbia University’s Fu Foundation School of Engineering and Applied Science. I help clients understand, visualize, and act on their data—translating advanced machine learning and AI concepts into clear business insights. With experience spanning finance, healthcare, and analytics consulting, I specialize in designing solutions that balance technical depth with practical clarity. Clients hire me to communicate complex models simply, advise on strategy, and deliver production-ready systems that executives can trust. My core services include: AI & ML Consulting: Business problem scoping, model design Machine Learning Engineering: Predictive modeling, feature pipelines, optimization, and deployment Natural Language Processing: Text classification, sentiment analysis, topic modeling, summarization, and retrieval Data Visualization & Storytelling: Dashboards and reports for stakeholders (Plotly, Dash, Streamlit, Power BI, ggplot2) Client Communication: Presenting findings, running client meetings, and translating technical work for non-technical teams My technical skills include: Languages: Python, R, SQL, NoSQL (MongoDB) Frameworks: scikit-learn, PyTorch, TensorFlow, spaCy, Hugging Face, BERTopic Visualization: Plotly, Dash, Streamlit, ggplot2, Power BI MLOps & Cloud: AWS (SageMaker, S3, Lambda), MLflow, Prefect, Docker, Git Databases: PostgreSQL, Hive, MS SQL, MongoDB Selected Experience: Deutsche Bank – Anti-Financial Crime Modeling Developed anomaly-detection models that improved fraud detection precision while maintaining interpretability. Epic Systems – Healthcare Analytics Built readmission risk and quality-metric models using claims and registry data. Political Data Dashboards Created interactive demographic and voter-trend dashboards used by advocacy and policy groups. Financial Forecasting Modeled stock-market and economic indicator trends with advanced time-series and sentiment features. NLP Summarization Deployed transformer-based summarizers for long-form financial reports and research analysis. Communication & Delivery Clients value my ability to bridge the technical and strategic. I routinely: Lead and participate in client meetings to align business goals with technical design Present data findings in clear, jargon-free language to executives and stakeholders Provide written reports, annotated notebooks, and reproducible deliverables Manage timelines, expectations, and transparency from start to finish Approach Every engagement starts with one question: “What decision needs to be made?” I design data workflows and AI systems that make those decisions faster, more accurate, and more explainable. Each project ends with clean, interpretable, and documented outputs—ready for production or presentation.

  • Apache Spark
  • R
  • Deep Learning
  • Python
  • Machine Learning
  • Tableau
  • SQL
  • Apache Hadoop
  • R Shiny
  • Apache Hive
  • Microsoft Power BI
  • PySpark
  • Data Visualization
  • ggplot2
Phani K.

Kansas City, Missouri

$20/hr
5.0
2 jobs

Results-driven Data Engineer and Software Developer with 3+ years of experience in Big Data, ETL, SQL, React and Python. Proven expertise in Azure and AWS cloud platforms, along with a strong background in developing and optimizing ETL pipelines, data migration, and implementing CI/CD processes. Skilled in Spark, React, Python, Docker and various data management tools. One of my notable projects includes "Hire Human," developed during a Microsoft Hackathon, This project involved real-time interview capture and analysis using advanced NLP algorithms and Azure OpenAI services. Experienced in Agile methodologies and am eager to leverage my technical skills and professional experience to drive innovative solutions and contribute to the success of a forward-thinking team that values innovation and excels in niche technologies.

  • Python
  • JavaScript
  • React
  • Microsoft Azure
  • Docker
  • Apache Hadoop
  • PySpark
  • Apache Kafka
  • PostgreSQL
  • HTML
  • CSS
  • API
  • GitHub
  • Databricks Platform
  • LLM Prompt

How it works

Post a job for freePost a job

Tell us what you need. Create your own job post or generate one with AI then filter talent matches.

Hire top talent fast

Consult, interview, and hire quickly, so you can meet the freelancers you're excited about.

Collaborate easily

Use Upwork to chat or video call, share files, and track project progress right from the app.

Payment simplified

Manage payments in one place with flexible billing options. Only pay for approved work, hourly or by milestone.

Don't just take our word for it

How do I hire a Apache Spark Engineer in the United States on Upwork?

You can hire a Apache Spark Engineer in the United States on Upwork in four simple steps:

  • Create a job post tailored to your Apache Spark Engineer project scope. We'll walk you through the process step by step.
  • Browse top Apache Spark Engineer talent on Upwork and invite them to your project.
  • Once the proposals start flowing in, create a shortlist of top Apache Spark Engineer profiles and interview.
  • Hire the right Apache Spark Engineer for your project from Upwork, the world's largest work marketplace.

At Upwork, we believe talent staffing should be easy.

How much does it cost to hire a Apache Spark Engineer?

Rates charged by Apache Spark Engineers on Upwork can vary with a number of factors including experience, location, and market conditions. See hourly rates for in-demand skills on Upwork.

Why hire a Apache Spark Engineer in the United States on Upwork?

As the world's work marketplace, we connect highly-skilled freelance Apache Spark Engineers and businesses and help them build trusted, long-term relationships so they can achieve more together. Let us help you build the dream Apache Spark Engineer team you need to succeed.

Can I hire a Apache Spark Engineer in the United States within 24 hours on Upwork?

Depending on availability and the quality of your job post, it's entirely possible to sign up for Upwork and receive Apache Spark Engineer proposals within 24 hours of posting a job description.