I am a seasoned Data Engineer and Big Data Specialist with over 7 years of experience in building and optimizing scalable data pipelines, crafting real-time analytics solutions, and leveraging cloud technologies to deliver business value. My expertise spans across Big Data Ecosystems, Cloud Platforms, and Data Warehousing, ensuring end-to-end solutions for data-driven decision-making.
๐๐ผ๐บ๐ฎ๐ถ๐ป ๐๐ ๐ฝ๐ฒ๐ฟ๐๐ถ๐๐ฒ :
๐ฃ๐ต๐ฎ๐ฟ๐บ๐ฎ๐ฐ๐ฒ๐๐๐ถ๐ฐ๐ฎ๐น ๐๐ป๐ฑ๐๐๐๐ฟ๐ โ Extensive experience in clinical data management, working with IQVIA, Sanofi, and optimizing study data pipelines for pharmaceutical research and trials.
๐๐ถ๐ป๐ฎ๐ป๐ฐ๐ถ๐ฎ๐น ๐ฆ๐ฒ๐ฐ๐๐ผ๐ฟ โ Expertise in the insurance domain, handling claims processing, policy data management, and ensuring compliance with industry regulations.
๐ฅ๐ฒ๐ฐ๐ฟ๐๐ถ๐๐บ๐ฒ๐ป๐ ๐๐ป๐ฑ๐๐๐๐ฟ๐ โ Worked with job advertisement data, company insights, geographical talent search, and salary analytics to enhance recruitment strategies.
๐๐๐ฒ ๐๐ข๐ ๐ก๐ฅ๐ข๐ ๐ก๐ญ๐ฌ:
Extensive experience in Big Data Analytics, specializing in Apache Spark, Kafka, NIFI, and Databricks, with a strong focus on Delta Lake, Streaming, and Batch Processing.
Proficient in building robust ETL/ELT pipelines and designing OLAP data models using technologies Redshift, Snowflake, and BigQuery with dbt transformation.
Hands-on expertise in AWS (Glue, S3, EMR, Lambda, Step Functions) and Azure (Databricks, ADF, CosmosDB) to build scalable, cloud-native architectures.
Experienced in Docker containerization and deploying distributed applications across various environments.
Adept at orchestrating data pipelines with Apache Airflow, Jenkins, and automated CI/CD workflows to streamline production deployments.
Skilled in programming with Python and Pandas, with exposure to Scala for efficient data transformation and processing.
๐๐๐๐๐๐ซ๐ฌ๐ก๐ข๐ฉ & ๐๐จ๐ฅ๐ฅ๐๐๐จ๐ซ๐๐ญ๐ข๐จ๐ง:
Proven track record in mentoring and leading technical teams, ensuring high-quality deliverables and fostering innovation.
Experienced in collaborating with cross-functional teams, translating complex requirements into scalable solutions, and ensuring alignment with business goals.
Passionate about adopting Agile (Scrum) and DevOps methodologies to enhance development efficiency and team collaboration.
๐๐ก๐ฒ ๐๐จ๐ง๐ง๐๐๐ญ ๐๐ข๐ญ๐ก ๐๐?
I thrive on solving complex data challenges, driving impactful insights. Whether itโs optimizing data architectures, building real-time analytics solutions, or mentoring teams,
Letโs connect and discuss how we can shape the future of data together!
Big Data
Data Warehousing & ETL Software
Data Engineering
Apache Spark
Python
SQL
Apache Airflow
Databricks Platform
Amazon Web Services
Data Analytics
Snowflake
GitHub
CI/CD Platform
Azure Cosmos DB
Data Modeling
Faizan K.
Berlin, Germany
$60/hr
5.0
63 jobs
Top Rated Plus | 100% Job Success | $400K+ earned | 11,890 hours | 5.0 from 29 reviews. Your dashboards are slow, your sources disagree, and leadership no longer trusts the numbers. I rebuild the analytics layer on dbt, Snowflake and BigQuery.
Three things usually break at once. Numbers that disagree between tools, because the same metric is defined three times in three places. Reports that arrive too late to act on, because the pipeline is manual or fragile. A warehouse nobody can extend, because the modelling was never written down or tested.
I fix them in that order: agree the definitions, model them once in dbt with tests and documentation attached, then put the reporting layer on top. After that, adding a source is a day's work, not a project.
DATA MODELLING AND TRANSFORMATION
dbt models on Snowflake and BigQuery that turn raw multi-source data into analytics-ready datasets. Tests on every source and key, documentation generated from the models themselves, modular design so the layer stays maintainable. dbt Core, dimensional and star-schema modelling, incremental models, custom macros. A public dbt case study sits in the portfolio below if you want to read the models themselves.
PIPELINES AND ORCHESTRATION
ETL and ELT built with SQL, Python, Apache Airflow and Dagster across GCP, AWS and Azure. Ingestion from APIs, SaaS tools and ad platforms using Airbyte, Fivetran and Supermetrics, GA4 BigQuery exports and serverless pulls on Cloud Functions. Monitoring, retries and validation, so failures surface before your stakeholders find them.
BI AND DASHBOARDS
Power BI, Tableau and Looker Studio, plus Metabase, Apache Superset, Sigma Computing, Klipfolio. Semantic models and row-level security rather than charts bolted onto raw tables. Executive reporting, KPI tracking and self-service analytics, designed for the people who actually use them.
WAREHOUSE, CLOUD AND DATA PROTECTION
Snowflake, BigQuery, PostgreSQL, Redshift, ClickHouse and Microsoft Fabric. Migrations off legacy systems, query and cost tuning, access separation per client or entity. I work from Germany under EU rules, so GDPR handling, PII-aware modelling and EU data residency are part of the design, not a later patch.
RECENT WORK
1. CGIAR - centralized enterprise analytics platform on Microsoft Fabric and Power BI, replacing fragmented reporting across a global research consortium.
2. Markets4U - migrated a fintech off Metabase, Oracle BI and Vertica onto an analytics stack built for real-time decisions.
3. Wellness Studio - Fortnox API into BigQuery on a scheduled Cloud Functions pipeline, with income, expenses and contribution margins per site across 29 locations in Power BI.
4. Pierce - GA4 into BigQuery on Airbyte, modelled in dbt with a custom macro for channel grouping, so every store metric reconciles to GA4 itself.
5. Flipstream - three fragmented warehouses consolidated into one Cube semantic model with Metabase on top, a REST API serving the frontend, and role-based access control.
6. AML Foods - on-prem Sybase migrated to a PostgreSQL warehouse with Airflow syncs and monitoring, delivered pilot first, making RFM and market basket analysis possible.
Also Sojo Industries (continuous intelligence, retail), ECAP (financial modelling, private real estate) and Westwise (marketing and sales data, legal services).
HOW I WORK
End to end ownership, from ingestion through dbt modelling to the dashboard. I document what I build and hand it over, so your team can extend it without me. One engagement was maternity cover for a client's own senior data engineer - the clearest test of working inside an existing team and codebase. I ran Power BI, Azure and SQL workshops as a Microsoft Student Ambassador, which is why handover and enablement are part of what I deliver.
Send me your current setup - the warehouse, the tools, the report nobody trusts - and I will tell you what I would change first, and why. Available 30+ hrs/week, replies inside 0-4 hours, open to projects, retainers and contract-to-hire.
KEYWORDS
analytics engineer, data engineer, data architect, dbt, Snowflake, BigQuery, PostgreSQL, Redshift, ClickHouse, Microsoft Fabric, Apache Airflow, Dagster, Airbyte, Fivetran, Supermetrics, ETL, ELT, data pipeline, data warehouse, data modelling, star schema, semantic layer, Cube, Power BI, DAX, Tableau, Looker Studio, Metabase, Apache Superset, Sigma, Klipfolio, SQL, Python, GA4, Google Analytics 4, Google Ads, Meta Ads, LinkedIn Ads, TikTok Ads, HubSpot, Salesforce, Stripe, marketing analytics, attribution, ROAS, cohorts, LTV, CAC, churn, RFM, GCP, AWS, Azure, GDPR
Data Engineering
BigQuery
Data Warehousing
ETL Pipeline
Python
SQL
Tableau
Data Visualization
Marketing Analytics
Snowflake
dbt
Looker Studio
Business Intelligence
Apache Airflow
Data Modeling
Microsoft Power BI
Data Analysis
Google Analytics
GDPR
AI Data Analytics
Ekaterina S.
Berlin, Germany
$20/hr
5.0
1 jobs
Iโm Ekaterina, a Junior Big Data Engineer based in Berlin with a Masterโs degree in Engineering Science. My focus is on data processing, analysis, and building reliable data workflows. Skilled in Python, SQL, and modern data engineering technologies, I design and automate efficient data solutions with a structured, detail-oriented mindset. Fluent in English, German, and Russian, and motivated by solving technical challenges with clarity and precision.
Big Data
Python
SQL
German
Russian
English
Madiha K.
Hamburg, Germany
$185/hr
4.9
8 jobs
I architect data platforms for companies where bad data isn't a reporting problem โ it's a regulatory risk, a compliance failure, or a million-dollar pipeline outage. My clients include Novo Nordisk, Danske Bank, 10XCRM and enterprise companies across maritime, e-commerce, and SaaS. I've worked alongside AWS Professional Services on mission-critical cloud migrations and built AI systems that are running in production today โ not proofs of concept.
Why Me:
You're not getting a pipeline developer. You're getting the person who decides what gets built, how it scales, and why it won't break at 3 AM. I've designed systems processing 60M+ events/day, migrated 8 billion records with zero data loss, and replaced 6-8 hour manual workflows with 30-minute automated pipelines.
Recent engagements
โ Novo Nordisk โ Event-driven scientific intelligence platform on AWS. Ingesting research documents from 6+ sources (PubMed, Springer, BioArxiv) through a Bronze/Silver/Gold lakehouse architecture. MSK Kafka with 10+ topics, Iceberg tables, OpenSearch, DynamoDB caching layer. Processing 30M+ events/day and 4TB+ monthly.
โ Danske Bank โ Enterprise cloud migration alongside AWS Professional Services. Architected ETL pipelines for metadata integration covering VMs, storage, databases, and wave planning. Built Python automation that replaced Excel-driven migration processes โ 6-8 hours down to 30 minutes. Also developed an EC2 recommendation engine for optimal instance sizing.
โ Maritime AI โ AI-powered document processing for the maritime industry. Extracts maintenance plans, spare parts, and BOM data from complex S1000D technical manuals. RAG system with PDF citation traceability and ERP integration. Deployed on-prem for data sovereignty requirements.
โ E-commerce brand (Hamburg) โ Full Redshift to Snowflake migration. Converted hundreds of legacy SQL Jinja pipelines to Snowflake dialect. Implemented Time Travel, Zero-Copy Cloning, and automated warehouse scaling. Zero downtime during cutover.
โ 10xCRM (B2B SaaS) โ Multi-tenant CRM analytics on GCP. Airbyte, dbt, BigQuery, Looker Studio. Replaced an 8-hour Airbyte sync with a custom Python pipeline running in 12 minutes on Cloud Run. Scales from 1 to N customers automatically.
โ Jimdo (4.5 years) โ Led the Redshift to Snowflake migration handling billions of daily events. Transformed 500+ legacy ETL jobs into dbt models. Integrated data streams from HubSpot (7M+ records), Facebook, Google AdWords, Zendesk, and Stripe. Achieved 60-70% ETL performance improvement through a 3-phase optimization strategy.
What every engagement includes:
โ Architecture-first approach. I design the platform before writing code โ lakehouse layers, governance model, cost projections, scaling strategy. You get a blueprint, not just scripts.
โ Production-grade deliverables. Proper error handling, monitoring, alerting, and documentation. Not notebook code pushed to production.
โ Knowledge transfer. Your team can maintain and extend everything I build. I don't create dependency.
โ One week post-deployment support included.
Stack: Python, PySpark, dbt, Airflow, AWS (S3, Glue, MSK Kafka, Bedrock, Lake Formation, Lambda, Redshift, EMR, ECS), Snowflake, Databricks, GCP BigQuery, Docker, Kubernetes, Terraform, RAG/LLM integration.
Background: Erasmus Mundus MSc in Big Data (UPC Barcelona, ULB Brussels, UFRT France). AWS Certified Data Engineer Associate and AI Practitioner. 10+ years across pharma, banking, energy, maritime, telecom, and SaaS. Based in Hamburg, Germany.
Big Data
Data Engineering
Apache Spark
Data Integration
ETL Pipeline
Python
Amazon Web Services
Apache Kafka
Machine Learning
Data Analysis
Data Annotation
Data Lake
Snowflake
Databricks Platform
Fivetran
Krishnasis M.
Berlin, Germany
$50/hr
5.0
8 jobs
Iโm a Senior Backend and Platform Engineer with 7+ years of experience building production software, cloud infrastructure, microservices, and data platforms.
My main expertise is Python backend development with FastAPI and Flask, combined with Kubernetes, Azure, Terraform, Docker, CI/CD, and distributed systems.
I can help you build a backend service from scratch, modernize an existing system, deploy applications to Kubernetes, automate your infrastructure, or build internal platforms and developer tooling.
My core areas of expertise include:
* Python backend development with FastAPI and Flask
* REST APIs and microservices
* Kubernetes and AKS
* Azure cloud infrastructure
* Terraform and infrastructure as code
* Docker, Helm and ArgoCD
* CI/CD with GitHub Actions and Azure DevOps
* PostgreSQL, Redis, Kafka and other distributed systems
* Databricks and PySpark
* Data pipelines and data platforms
* AI-powered applications and internal tools
* Ollama, LLM applications, RAG and vector databases
* Performance and cloud cost optimization
Some examples of what I have built:
* Bootstrapped an AKS-based data platform from 0 to 1, including private networking, Azure Container Registry, Key Vault, DNS, ingress, monitoring and GitOps deployment.
* Built multiple production Python microservices using FastAPI and Flask and deployed them to Kubernetes.
* Built an AI-enabled Internal Developer Portal for onboarding, access requests, compute requests and internal tooling.
* Built a self-hosted AI agent system using Ollama and vector search for troubleshooting and internal documentation.
* Developed a Python data-loading framework on top of dlt, enabling YAML-driven data ingestion between different sources and destinations.
* Built CI/CD infrastructure using Azure DevOps, GitHub Actions, ArgoCD and Databricks Asset Bundles.
* Built private PyPI infrastructure integrated with Databricks network isolation, removing a major manual package installation bottleneck.
* Optimized Databricks infrastructure and workloads, saving over โฌ100,000 per year and reducing development costs by 97% in one project.
* Reduced Azure Log Analytics costs by approximately โฌ4,000 per month through ingestion optimization.
* Optimized a Python document-processing system using multiprocessing, reducing processing time by 10x.
* Optimized a financial document-processing pipeline and reduced total processing time by approximately two thirds.
Iโm comfortable working across the entire engineering lifecycle: understanding requirements, designing the architecture, writing the application, provisioning infrastructure, setting up CI/CD, deploying to production, adding monitoring and documentation, and improving the system afterwards.
If you need a Python engineer who can also understand the infrastructure and deployment side of the system, I can help.
Big Data
Python
SQL
Scala
Java
Amazon Web Services
Game Development
Apache Hadoop
Android
Image Processing
Syed Hasan A.
Berlin, Germany
$70/hr
5.0
36 jobs
Iโm a senior data engineer/solutions architect with a rigorous background in mathematical physics and over $70,000 earned on Upwork. I build production-grade lakehouse platforms, scalable ELT/ETL, and ML/MLOpsโwith deep, hands-on experience in Google BigQuery and Azure Databricks. Iโve delivered in both enterprise (EY-Parthenon, PwC, Amazon) and startup environments, combining Fortune-500 reliability with startup speed and pragmatism.
Why clients hire me:
- BigQuery at scale (fast & cost-efficient). Designed serverless ELT on BigQuery + dbt with CI/CD that cut processing costs by ~5ร and powered near-real-time marketing analytics (GA4, Google Ads, Facebook Ads, Bing Ads, Shopware) in Looker Studio.
- Healthcare & interoperability on Azure Databricks. Built an openEHR-driven research data platform on Azure Databricks, standardized multi-team deployments with Databricks Asset Bundles (DABs), and enforced Unity Catalog governance (RBAC + row filters). End-to-end MLOps with MLflow/Model Registry for risk models.
- Enterprise-grade warehousing & CI/CD. Parameterized Synapse deployments; incremental ELT from SAP/CRM APIs with pagination/parallelismโruntime reductions ~5ร.
- Analytics acceleration. Improved Power BI refresh/report performance ~6ร via Azure SQL tuning plus robust ADLS + ADF pipelines with CI/CD.
- Proven Upwork delivery. Long-term API integration & data engineering: $44,246.67 on a single engagement (consistent 5โ track record across projects).
Core services
- BigQuery & GCP: Serverless ELT/ETL with dbt, Cloud Functions, Dataflow; marketing/CRM ingestion to BigQuery; cost/perf modeling; near-real-time dashboards (Looker Studio).
- Azure & Databricks: Lakehouse (medallion), Unity Catalog governance, Synapse warehousing, ADF (schema drift/SCD), MLflow MLOps, DABs for IaC.
- Interoperability & APIs: HL7 โ FHIR โ openEHR, REST integrations (SAP/CRM), Ads/GA4 pipelines, Google Apps Script/Node.js.
Selected project highlights (Upwork & enterprise)
- $44K+ long-term API integration & data engineering (Apps Script + Node.js).
- BigQuery marketing & analytics stack: multi-channel ingestion, dbt modeling, CI/CD; costs down ~5ร, near-real-time dashboards.
- Azure lakehouse & MLOps: openEHR data hub on Databricks with DABs + Unity Catalog; MLflow-driven pipelines.
- E-commerce data warehouse: end-to-end build with ML features (k-means, ARIMA/LSTM) for decision support.
Certifications
- Databricks Certified Data Engineer Professional
- Microsoft Certified: Azure Data Engineer Associate (DP-203)
- GCP coursework: Big Data & ML fundamentals / modern data warehousing
Toolbox
- Languages: SQL (advanced SPs/windowing/dynamic SQL), Python (pandas/PySpark/FastAPI), Spark
- GCP: BigQuery (advanced), Cloud SQL, Dataflow, Cloud Functions
- Azure: Synapse, ADF, ADLS Gen2, Unity Catalog, Purview, Fabric
- Databricks: DABs, Workflows, MLflow, Lakehouse governance
- Orchestration: dbt (advanced macros), ADF (schema drift/SCD), Airflow (basic)
- Visualization: Power BI, Looker Studio, Superset, Tableau
Python
SQL
Machine Learning
Academic Writing
Financial Writing
Financial Analysis
Statistical Analysis
Mathematics Tutoring
Technical Writing
Visual Basic for Applications
How it works
Post a job for freePost a job
Tell us what you need. Create your own job post or generate one with AI then filter talent matches.
Hire top talent fast
Consult, interview, and hire quickly, so you can meet the freelancers you're excited about.
Collaborate easily
Use Upwork to chat or video call, share files, and track project progress right from the app.
Payment simplified
Manage payments in one place with flexible billing options. Only pay for approved work, hourly or by milestone.
Don't just take our word for it
โUpwork provides an umbrella-level of security. I can see a talentโs work history and ratings. I can hold payments in escrow. I can communicate through Upwork Messages instead of working through my email address.โ
KD
Kim Darling
Emerald Tiger
โUpwork is the best platform to hire skilled professionals when we're not looking for a full-time employee. All the companies in our portfolio use Upwork to find talent across a wide range of fields.โ
DM
David Merry
Kinetic Investments
โOur very specific requirements can be a challengeโWith Upwork, weโre able to access a bigger community to ensure the success of our projects.โ
KK
Katja Krohn
Summa Linguae
How do I hire a Big Data Engineer in Germany on Upwork?
You can hire a Big Data Engineer in Germany on Upwork in four simple steps:
Create a job post tailored to your Big Data Engineer project scope. We'll walk you through the process step by step.
Browse top Big Data Engineer talent on Upwork and invite them to your project.
Once the proposals start flowing in, create a shortlist of top Big Data Engineer profiles and interview.
Hire the right Big Data Engineer for your project from Upwork, the world's largest work marketplace.
At Upwork, we believe talent staffing should be easy.
How much does it cost to hire a Big Data Engineer?
Rates charged by Big Data Engineers on Upwork can vary with a number of factors including experience, location, and market conditions. See hourly rates for in-demand skills on Upwork.
Why hire a Big Data Engineer in Germany on Upwork?
As the world's work marketplace, we connect highly-skilled freelance Big Data Engineers and businesses and help them build trusted, long-term relationships so they can achieve more together. Let us help you build the dream Big Data Engineer team you need to succeed.
Can I hire a Big Data Engineer in Germany within 24 hours on Upwork?
Depending on availability and the quality of your job post, it's entirely possible to sign up for Upwork and receive Big Data Engineer proposals within 24 hours of posting a job description.