Full-Stack AI Lead with 15+ years of experience in AI/ML/Data Engineering. Specialist in productionizing Document AI, Agentic LLM Workflows, and Scalable Big Data systems. Master’s in Applied Mathematics.
𝐀𝐠𝐞𝐧𝐭𝐢𝐜 𝐀𝐈 & 𝐋𝐋𝐌
- Architecting production-grade agents and multi-intent workflows using LangGraph and LangChain.
- Implementing NeMo Guardrails, DeepEval, and Langfuse to ensure security and near-zero hallucination rates.
- Developing high-performance FastAPI backends for real-time operational data integration.
𝐃𝐨𝐜𝐮𝐦𝐞𝐧𝐭 𝐈𝐧𝐭𝐞𝐥𝐥𝐢𝐠𝐞𝐧𝐜𝐞 & 𝐎𝐂𝐑
- Specialist in extracting structure from "dirty" or non-digital sources using Visual NLP and OCR.
- Lead Developer of Spark OCR.
- Author and main contributor of open-source projects: Spark PDF and ScaleDP.
- Expertise in document layout analysis, data extraction, and semantic search.
𝐏𝐫𝐨𝐝𝐮𝐜𝐭𝐢𝐨𝐧 𝐌𝐋𝐎𝐩𝐬 & 𝐁𝐢𝐠 𝐃𝐚𝐭𝐚
- Architecting scalable infrastructures using Vector DBs (Milvus, Pinecone, Qdrant).
- Optimizing Apache Spark (PySpark, Streaming) and Databricks pipelines; reduced processing time from days to hours.
- Implementing robust CI/CD via Databricks Asset Bundles, Jenkins, and GitHub Actions.
𝐏𝐫𝐢𝐯𝐚𝐜𝐲 & 𝐂𝐨𝐦𝐩𝐥𝐢𝐚𝐧𝐜𝐞
- Automated de-identification and redaction for GDPR/HIPAA.
- Specialist in masking sensitive information in DICOM, PDF, and image formats.
𝐓𝐞𝐜𝐡𝐧𝐨𝐥𝐨𝐠𝐢𝐞𝐬
AI & LLM: LangGraph, LangChain, RAG, Llama 3, Gemini, GPT-4, Vertex AI, Hugging Face, GliNER.
Backend: Python, Scala, FastAPI, Node.js, PostgreSQL, MongoDB.
Vector DBs: Pinecone, Milvus, Qdrant.
Infrastructure: AWS, GCP, Azure, Docker, Kubernetes, Jenkins, AirFlow.
Big Data: Apache Spark (PySpark, Streaming, MLlib), Kafka, Kinesis.
Apache Spark
PySpark
Natural Language Processing
Computer Vision
PyTorch
Scala
Python
Tesseract OCR
Machine Learning
Large Language Model
Hugging Face
Software Architecture & Design
Databricks Platform
LangChain
AI Development
AI Agent Development
Generative AI
Vector Database
Recommendation System
Chatbot
Kamil K.
Bytom, Poland
$35/hr
5.0
2 jobs
I am a Senior Data Engineer with 9+ years of experience architecting high‑performance data pipelines, cloud data warehouses, and enterprise BI solutions. I specialise in DBT‑driven transformation layers, BigQuery optimisation, and large‑scale ETL/ELT orchestration – turning raw, fragmented data into trusted, analytics‑ready assets that power strategic decisions.
My approach combines deep data modelling expertise with modern engineering practices: version‑controlled transformations, automated testing, CI/CD for data, and cost‑aware cloud design. I deliver systems that are not only scalable and reliable but also maintainable and well‑documented for your entire team.
✔️𝐒𝐩𝐞𝐜𝐢𝐚𝐥𝐢𝐳𝐞𝐝 𝐄𝐱𝐩𝐞𝐫𝐭𝐢𝐬𝐞.
- DBT (Data Build Tool): Advanced modelling, incremental strategies, macro development, and rigorous data testing for analytical warehouses.
- BigQuery Optimisation: Partitioning, clustering, slot management, and query cost control – delivering sub‑second responses on petabyte‑scale data.
- Big Data Processing: Spark (PySpark) and Kafka for stream/batch processing, with efficient handling of high‑volume, high‑velocity datasets.
✔️𝐒𝐭𝐫𝐨𝐧𝐠 𝐄𝐱𝐩𝐞𝐫𝐢𝐞𝐧𝐜𝐞 𝐢𝐧
- E‑commerce & retail analytics: building 360° customer views, inventory forecasting, and real‑time sales dashboards.
- Financial data platforms: regulatory reporting, fraud detection pipelines, and GDPR‑compliant data masking.
- Real‑time streaming architectures: Kafka + BigQuery + dbt for near‑real‑time operational intelligence.
✔️𝐏𝐫𝐢𝐦𝐚𝐫𝐲 𝐒𝐭𝐚𝐜𝐤.
- Cloud Platforms: GCP (BigQuery, Composer, Pub/Sub), Azure (Data Factory, Synapse, Databricks), AWS (Glue, Redshift, S3).
- Warehousing & Lakes: BigQuery, Snowflake, Azure Synapse, MSSQL, PostgreSQL, MySQL.
- Transformation & Orchestration: dbt (core/cloud), Apache Airflow, Prefect, Azure Data Factory.
- Languages: Advanced SQL, Python (Pandas, PySpark, Polars), DAX, M (Power Query).
- BI & Visualisation: Power BI (Service, Premium, Paginated Reports), Tableau, Looker Studio.
- Big Data Tech: Apache Spark, Kafka, Flink, Databricks (Delta Lake).
- DevOps for Data: Git, CI/CD (GitHub Actions, Azure DevOps), Docker, Terraform (basic).
- Data Governance: RLS, column‑level security, data lineage, and quality monitoring (Great Expectations).
✔️𝐖𝐡𝐞𝐭𝐡𝐞𝐫 𝐲𝐨𝐮 𝐧𝐞𝐞𝐝 𝐭𝐨
- Migrate legacy ETL to a modern dbt + BigQuery stack with full CI/CD,
- Build a high‑throughput streaming pipeline that ingests millions of events per second,
- Design a self‑service semantic layer that empowers business users with reliable KPIs,
- Or optimise your existing warehouse for cost and performance without sacrificing freshness,
I deliver a clear roadmap, measurable improvements, and a system your team can confidently own.
✔️𝐖𝐡𝐚𝐭 𝐲𝐨𝐮 𝐠𝐞𝐭
✅ A data platform built for scale – not just today, but for 10x growth
✅ Cloud cost optimisation that reduces your monthly bill by 20‑40%
✅ Daily updates (Slack/email) with progress, blockers, and next steps
✅ Clean, modular dbt code with full documentation and automated tests
✅ Production‑ready pipelines with monitoring, alerting, and rollback strategies
Message me or invite me to a quick 15‑min call – I will walk you through a tailored architecture sketch for your project.
Keywords:
Data Engineer, DBT, BigQuery, Big Data, ETL, ELT, Cloud Analytics, Power BI, MSSQL, Python, SQL, GCP, Azure, AWS, Data Warehousing, Snowflake, Databricks, Apache Airflow, Kafka, Spark, PySpark, Data Pipeline, Business Intelligence, Dashboard, Data Modelling, DAX, M, Data Lake, Data Integration, E‑commerce Analytics, Financial Data, Real‑time Streaming, CI/CD, DevOps, Data Governance, Performance Optimisation, Cloud Migration, Senior Data Engineer, Remote, Full‑time.
ETL Pipeline
dbt
SQL
Python
ETL
Data Warehousing & ETL Software
Microsoft Power BI
Snowflake
Cloud Architecture
BigQuery
Performance Optimization
Amazon Web Services
Data Modeling
Microsoft Power BI Data Visualization
Airtable
Data Analysis
Data Migration
Data Science
Business Analysis
Financial Analysis
Artsiom S.
Warsaw, Poland
$20/hr
5.0
4 jobs
- 10+ years of professional experience in JVM-based software development.
- Expertise in Scala’s modern FP stack (Typelevel, ZIO, Akka), data engineering (Hadoop,
AWS, Spark), and Java (Spring) for high-performance solutions.
- A product-oriented mindset, focused on understanding the business domain deeply to
deliver solutions that align with and exceed business expectations.
Apache Spark
PySpark
Scala
Python
Apache Kafka
Apache Flink
Java
Microsoft Azure
Amazon Web Services
Google Cloud Platform
ClickHouse
PostgreSQL
Microsoft SQL Server
Distributed Computing
Distributed Database
Sylwester N.
Warsaw, Poland
$110/hr
5.0
25 jobs
🚀 Build scalable data infrastructure 🛠 Automate data workflows 📊 Deliver actionable analytics
I’m a certified data consultant specializing in Microsoft Fabric, Azure Databricks, SQL, Python, and Power BI. I help companies, from SaaS startups to global enterprises, turn complex, fragmented data into reliable, analytics-ready datasets that drive faster decisions and product growth.
My work spans the full data lifecycle: designing architecture, building ETL pipelines, managing data lakes, and delivering secure reporting layers. I’ve delivered 20+ successful remote projects across logistics, maritime, energy, and real estate - from setting up 5+ Microsoft Fabric environments from scratch to managing infrastructure for Verizon AI’s 80M+ company dataset.
Core Skills
📶 Data Engineering & Analytics: Microsoft Fabric, Azure Databricks, SQL Server, T-SQL, Python, PySpark, DAX, M
☁️ Cloud & Orchestration: Azure Data Factory, ADLS Gen2, Medallion Architecture, CI/CD, Git
🔌 Integrations: REST APIs, CRMs (Salesforce, HubSpot), ERPs (SAP, QuickBooks), IoT data
📊 Data Visualization: Power BI (Embedded), Apache Superset
Example Use Cases
End-to-end SaaS data infrastructure for lead-generation platforms
ETL pipelines consolidating CRM, ERP, and IoT data into analytics-ready datasets
Embedded analytics portals for client KPI tracking
Role-based reporting for operations, finance, and product teams
Microsoft Certified
PL-300 – Power BI Data Analyst Associate
DP-700 – Fabric Data Engineer Associate
DP-500 – Azure Enterprise Data Analyst Associate
Clients choose me because I:
🔸 Deliver on time and to spec
🔸 Take ownership of the whole solution, from ingestion to analytics
🔸 Communicate clearly with both technical and non-technical teams
🔸 Bring both enterprise-scale engineering and startup agility to projects
If you’re looking for a technical data consultant who can architect, build, and optimize your data infrastructure - let’s connect.
Microsoft Excel
SQL
Data Engineering
Microsoft Power BI
Fabric
Data Visualization
Dashboard
Business Intelligence
Power Query
Microsoft Power Automate
Microsoft Power BI Data Visualization
Microsoft Power BI Development
Data Analysis
ETL
Database
Mariusz S.
Brzozowka, Poland
$100/hr
5.0
50 jobs
I have over 9 years of experience in Data Engineering (especially using Spark and pySpark to gain value from massive amounts of data). I worked with analysts and data scientists by conducting workshops on working in Hadoop/Spark and resolving their issues with big data ecosystem. I also have experience on Hadoop maintenance and building ETL, especially between Hadoop and Kafka.
You can find my profile on stackoverflow (link in Portfolio section) - I help mostly in spark and pyspark tagged questions.
Apache Spark
PySpark
Apache Hadoop
Apache Kafka
Apache Airflow
Data Migration
Python
Data Visualization
ETL
Data Scraping
Data Warehousing
MongoDB
Oleh S.
Warsaw, Poland
$40/hr
5.0
5 jobs
🔥 Looking for a Data Engineer who knows how to scale? I help enterprises transform raw data into real-time insights using Databricks, Azure, and streaming architectures that deliver results.
👉 Press “𝐈𝐍𝐕𝐈𝐓𝐄 𝐓𝐎 𝐓𝐇𝐄 𝐉𝐎𝐁” or “𝐒𝐄𝐍𝐃 𝐀 𝐌𝐄𝐒𝐒𝐀𝐆𝐄” and let’s talk about how I can help you build a high-performance data platform.
🚀 Case Study:
I partnered with Databricks Professional Services (via Azure) to deliver a production-grade, structured streaming pipeline for a client that ingests over 4 billion events per day.
The result? A scalable, monitored system that powers real-time analytics and intelligent automation—proving that AI Chatbots and Agents can thrive even in high-volume, enterprise-grade environments.
📜 Certifications:
🧠 Generative AI – Databricks
🏗️ Azure Databricks Platform Architect – Databricks
🧱 Databricks Lakehouse – Databricks
🔧 Databricks Platform – Databricks
🛡️ Platform Administrator – Databricks
I’m a certified Data Engineer with deep expertise in Databricks, Azure, and distributed data systems. I’ve worked with enterprise clients (1000+ employees) to deploy scalable, multi-region data platforms that power real-time analytics and business growth.
✅ As a Data Engineer, I deploy enterprise-grade data infrastructure by rolling out Databricks across 13 regions to support global operations with high availability and performance.
✅ As a Data Engineer, I build streaming pipelines that process billions of records daily with near real-time SLAs (<1 minute), enabling fast, actionable insights.
✅ As a Data Engineer, I drive impact-driven analytics that enabled a client to onboard 500+ new customers by unlocking insights from their data lake through Databricks-powered analysis.
✅ As a Data Engineer, I collaborate with Microsoft Professional Services to optimize Azure and Databricks performance, ensuring cost efficiency and speed at scale.
✅ As a Data Engineer, I design secure and reliable architectures—implementing Medallion Architecture, Delta Lake optimizations, and CI/CD workflows for robust data operations.
🤔 What’s stopping us from building something powerful together? Send me a message, and let’s kick off your next data engineering project!
Apache Spark
SQL
ETL Pipeline
Big Data
Python
Data Analysis
Machine Learning
Data Engineering
Amazon Web Services
Data Science
Data Warehousing
ETL
BigQuery
Data Scraping
Data Extraction
How it works
Post a job for freePost a job
Tell us what you need. Create your own job post or generate one with AI then filter talent matches.
Hire top talent fast
Consult, interview, and hire quickly, so you can meet the freelancers you're excited about.
Collaborate easily
Use Upwork to chat or video call, share files, and track project progress right from the app.
Payment simplified
Manage payments in one place with flexible billing options. Only pay for approved work, hourly or by milestone.
Don't just take our word for it
“Upwork provides an umbrella-level of security. I can see a talent’s work history and ratings. I can hold payments in escrow. I can communicate through Upwork Messages instead of working through my email address.”
KD
Kim Darling
Emerald Tiger
“Upwork is the best platform to hire skilled professionals when we're not looking for a full-time employee. All the companies in our portfolio use Upwork to find talent across a wide range of fields.”
DM
David Merry
Kinetic Investments
“Our very specific requirements can be a challenge—With Upwork, we’re able to access a bigger community to ensure the success of our projects.”
KK
Katja Krohn
Summa Linguae
How do I hire a Pyspark Developer in Poland on Upwork?
You can hire a Pyspark Developer in Poland on Upwork in four simple steps:
Create a job post tailored to your Pyspark Developer project scope. We'll walk you through the process step by step.
Browse top Pyspark Developer talent on Upwork and invite them to your project.
Once the proposals start flowing in, create a shortlist of top Pyspark Developer profiles and interview.
Hire the right Pyspark Developer for your project from Upwork, the world's largest work marketplace.
At Upwork, we believe talent staffing should be easy.
How much does it cost to hire a Pyspark Developer?
Rates charged by Pyspark Developers on Upwork can vary with a number of factors including experience, location, and market conditions. See hourly rates for in-demand skills on Upwork.
Why hire a Pyspark Developer in Poland on Upwork?
As the world's work marketplace, we connect highly-skilled freelance Pyspark Developers and businesses and help them build trusted, long-term relationships so they can achieve more together. Let us help you build the dream Pyspark Developer team you need to succeed.
Can I hire a Pyspark Developer in Poland within 24 hours on Upwork?
Depending on availability and the quality of your job post, it's entirely possible to sign up for Upwork and receive Pyspark Developer proposals within 24 hours of posting a job description.