Hire the Best DataLife Engine Specialists

More than 3,000 reviews on G2
Rating is 4.5 out of 5.
4.5/5
of Upwork by G2 peer reviewers
Muhammad H.

Karachi, Pakistan

$10/hr
5.0
1 jobs

Slow pipelines, unreliable data, or a warehouse that breaks every time the source changes? I build data systems that don't. I'm a ๐— ๐—ถ๐—ฐ๐—ฟ๐—ผ๐˜€๐—ผ๐—ณ๐˜ ๐—–๐—ฒ๐—ฟ๐˜๐—ถ๐—ณ๐—ถ๐—ฒ๐—ฑ ๐—™๐—ฎ๐—ฏ๐—ฟ๐—ถ๐—ฐ ๐——๐—ฎ๐˜๐—ฎ ๐—˜๐—ป๐—ด๐—ถ๐—ป๐—ฒ๐—ฒ๐—ฟ with 5 years of experience delivering end-to-end data engineering and BI solutions. I currently work at Pakistan's largest payment gateway, a high volume fintech environment where ๐—ง๐—•-๐˜€๐—ฐ๐—ฎ๐—น๐—ฒ ๐˜๐—ฟ๐—ฎ๐—ป๐˜€๐—ฎ๐—ฐ๐˜๐—ถ๐—ผ๐—ป๐—ฎ๐—น ๐—ฑ๐—ฎ๐˜๐—ฎ, strict governance, and zero tolerance for pipeline failures are the daily reality. My specialty is building systems that are ๐—ฎ๐—ฟ๐—ฐ๐—ต๐—ถ๐˜๐—ฒ๐—ฐ๐˜๐—ฒ๐—ฑ ๐—ฝ๐—ฟ๐—ผ๐—ฝ๐—ฒ๐—ฟ๐—น๐˜† ๐—ณ๐—ฟ๐—ผ๐—บ ๐˜๐—ต๐—ฒ ๐˜€๐˜๐—ฎ๐—ฟ๐˜ metadata-driven, layered, monitored, and built to scale. ๐—ช๐—›๐—”๐—ง ๐—œ ๐—•๐—จ๐—œ๐—Ÿ๐—— โœฆ ๐— ๐—ฒ๐˜๐—ฎ๐—ฑ๐—ฎ๐˜๐—ฎ-๐——๐—ฟ๐—ถ๐˜ƒ๐—ฒ๐—ป ๐—˜๐—ง๐—Ÿ/๐—˜๐—Ÿ๐—ง ๐—ฃ๐—ถ๐—ฝ๐—ฒ๐—น๐—ถ๐—ป๐—ฒ๐˜€ Control logic lives in configuration, not hardcoded. One framework handles dozens of sources with built-in logging, error handling, and restartability. Proven: reduced ETL runtime by ๐Ÿฏ๐Ÿด% on a production enterprise warehouse by eliminating redundant mapping layers. โœฆ ๐—˜๐—ป๐˜๐—ฒ๐—ฟ๐—ฝ๐—ฟ๐—ถ๐˜€๐—ฒ ๐——๐—ฎ๐˜๐—ฎ ๐—ช๐—ฎ๐—ฟ๐—ฒ๐—ต๐—ผ๐˜‚๐˜€๐—ฒ ๐——๐—ฒ๐˜€๐—ถ๐—ด๐—ป End-to-end warehouse design across ๐—ฆ๐˜๐—ฎ๐—ด๐—ถ๐—ป๐—ด โ†’ ๐—–๐—ผ๐—ฟ๐—ฒ โ†’ ๐—š๐—ผ๐—น๐—ฑ (Medallion Architecture), with star/snowflake schema modeling, incremental loading, duplicate handling, and structured audit logging baked in. โœฆ ๐— ๐—ถ๐—ฐ๐—ฟ๐—ผ๐˜€๐—ผ๐—ณ๐˜ ๐—™๐—ฎ๐—ฏ๐—ฟ๐—ถ๐—ฐ ๐—ฆ๐—ผ๐—น๐˜‚๐˜๐—ถ๐—ผ๐—ป๐˜€ Lakehouse and Warehouse design on OneLake, Fabric Data Factory pipelines, semantic models with ๐—ฅ๐—ผ๐˜„-๐—Ÿ๐—ฒ๐˜ƒ๐—ฒ๐—น ๐—ฆ๐—ฒ๐—ฐ๐˜‚๐—ฟ๐—ถ๐˜๐˜† (๐—ฅ๐—Ÿ๐—ฆ), and report publishing as Fabric Apps for internal teams and external stakeholders. โœฆ ๐—”๐˜‡๐˜‚๐—ฟ๐—ฒ & ๐——๐—ฎ๐˜๐—ฎ๐—ฏ๐—ฟ๐—ถ๐—ฐ๐—ธ๐˜€ ๐—ฃ๐—ถ๐—ฝ๐—ฒ๐—น๐—ถ๐—ป๐—ฒ๐˜€ ADF orchestrated cloud pipelines and PySpark based distributed data processing on Databricks for large-scale, partitioned datasets. โœฆ ๐—ฃ๐—ผ๐˜„๐—ฒ๐—ฟ ๐—•๐—œ & ๐—ฆ๐—ฆ๐—ฅ๐—ฆ ๐—ฅ๐—ฒ๐—ฝ๐—ผ๐—ฟ๐˜๐—ถ๐—ป๐—ด Semantic model design, DAX measures, drill-through dashboards, RLS enforcement, SSRS and Report Builder reports, and Fabric App deployment for enterprise stakeholders. โœฆ ๐—Ÿ๐—ฒ๐—ด๐—ฎ๐—ฐ๐˜† ๐— ๐—œ๐—ฆ ๐— ๐—ถ๐—ด๐—ฟ๐—ฎ๐˜๐—ถ๐—ผ๐—ป Migrated 20+ reports from legacy systems into a centralized, modern BI architecture without disrupting ongoing operations. ๐—ฅ๐—˜๐—–๐—˜๐—ก๐—ง ๐—ฅ๐—˜๐—ฆ๐—จ๐—Ÿ๐—ง๐—ฆ โ€ข Reduced ETL runtime by ๐Ÿฏ๐Ÿด% (4 hrs โ†’ 2.5 hrs) by optimizing metadata-driven SSIS pipelines โ€ข Built automated SFTP ingestion pipelines with archive logic to ensure ๐—ถ๐—ป๐—ฐ๐—ฟ๐—ฒ๐—บ๐—ฒ๐—ป๐˜๐—ฎ๐—น, ๐—ฑ๐˜‚๐—ฝ๐—น๐—ถ๐—ฐ๐—ฎ๐˜๐—ฒ-๐—ณ๐—ฟ๐—ฒ๐—ฒ data loading โ€ข Delivered ๐—บ๐˜‚๐—น๐˜๐—ถ๐—ฝ๐—น๐—ฒ ๐—˜๐—ป๐˜๐—ฒ๐—ฟ๐—ฝ๐—ฟ๐—ถ๐˜€๐—ฒ ๐——๐—ฎ๐˜๐—ฎ ๐—ช๐—ฎ๐—ฟ๐—ฒ๐—ต๐—ผ๐˜‚๐˜€๐—ฒ๐˜€ supporting different business products across fintech, billing, and payments โ€ข Published ๐Ÿญ๐Ÿฑ+ ๐—ฃ๐—ผ๐˜„๐—ฒ๐—ฟ ๐—•๐—œ ๐—ฟ๐—ฒ๐—ฝ๐—ผ๐—ฟ๐˜๐˜€ as Fabric Apps with Row Level Security for external stakeholders โ€ข Onboarded 10+ new source tables into a redesigned data warehouse while improving ETL performance and storage efficiency โ€ข Worked extensively with ๐—ง๐—•-๐˜€๐—ฐ๐—ฎ๐—น๐—ฒ ๐˜๐—ฟ๐—ฎ๐—ป๐˜€๐—ฎ๐—ฐ๐˜๐—ถ๐—ผ๐—ป๐—ฎ๐—น ๐—ฑ๐—ฎ๐˜๐—ฎ in a high-volume payment processing environment. ๐—–๐—ข๐—ฅ๐—˜ ๐—ฆ๐—ง๐—”๐—–๐—ž ๐— ๐—ถ๐—ฐ๐—ฟ๐—ผ๐˜€๐—ผ๐—ณ๐˜ ๐—™๐—ฎ๐—ฏ๐—ฟ๐—ถ๐—ฐ | ๐—”๐˜‡๐˜‚๐—ฟ๐—ฒ ๐——๐—ฎ๐˜๐—ฎ ๐—™๐—ฎ๐—ฐ๐˜๐—ผ๐—ฟ๐˜† | ๐—”๐˜‡๐˜‚๐—ฟ๐—ฒ ๐——๐—ฎ๐˜๐—ฎ๐—ฏ๐—ฟ๐—ถ๐—ฐ๐—ธ๐˜€ | ๐—ฃ๐˜†๐—ฆ๐—ฝ๐—ฎ๐—ฟ๐—ธ | ๐—”๐—ฝ๐—ฎ๐—ฐ๐—ต๐—ฒ ๐—ฆ๐—ฝ๐—ฎ๐—ฟ๐—ธ | ๐—ฆ๐—ฆ๐—œ๐—ฆ | ๐—ฆ๐—ค๐—Ÿ ๐—ฆ๐—ฒ๐—ฟ๐˜ƒ๐—ฒ๐—ฟ | ๐—ข๐—ฟ๐—ฎ๐—ฐ๐—น๐—ฒ | ๐—ฃ๐—ผ๐˜€๐˜๐—ด๐—ฟ๐—ฒ๐—ฆ๐—ค๐—Ÿ | ๐—ฃ๐—ผ๐˜„๐—ฒ๐—ฟ ๐—•๐—œ | ๐—ฆ๐—ฆ๐—ฅ๐—ฆ | ๐—ง-๐—ฆ๐—ค๐—Ÿ | ๐—ฃ๐—Ÿ/๐—ฆ๐—ค๐—Ÿ | ๐——๐—ฎ๐˜๐—ฎ ๐—ช๐—ฎ๐—ฟ๐—ฒ๐—ต๐—ผ๐˜‚๐˜€๐—ถ๐—ป๐—ด | ๐— ๐—ฒ๐—ฑ๐—ฎ๐—น๐—น๐—ถ๐—ผ๐—ป ๐—”๐—ฟ๐—ฐ๐—ต๐—ถ๐˜๐—ฒ๐—ฐ๐˜๐˜‚๐—ฟ๐—ฒ | ๐—ฆ๐˜๐—ฎ๐—ฟ ๐—ฆ๐—ฐ๐—ต๐—ฒ๐—บ๐—ฎ | ๐—ฆ๐—ป๐—ผ๐˜„๐—ณ๐—น๐—ฎ๐—ธ๐—ฒ ๐—ฆ๐—ฐ๐—ต๐—ฒ๐—บ๐—ฎ | ๐—˜๐—ง๐—Ÿ/๐—˜๐—Ÿ๐—ง | ๐—Ÿ๐—ฎ๐—ธ๐—ฒ๐—ต๐—ผ๐˜‚๐˜€๐—ฒ ๐—•๐—˜๐—ฆ๐—ง-๐—™๐—œ๐—ง ๐—ฃ๐—ฅ๐—ข๐—๐—˜๐—–๐—ง๐—ฆ โ€ข Data warehouse or lakehouse design from scratch โ€ข ETL/ELT pipeline build, optimization, or troubleshooting โ€ข Microsoft Fabric or Azure migration from legacy on-prem systems โ€ข Power BI, SSRS, or Fabric App reporting solutions โ€ข SQL performance tuning, stored procedures, and indexing โ€ข Production pipeline monitoring, job scheduling, and failure resolution ๐—›๐—ข๐—ช ๐—œ ๐—ช๐—ข๐—ฅ๐—ž I understand your business process, data sources, and reporting needs first. Then I design a practical architecture, build clean and observable pipelines, validate the data, and deliver reporting ready models your team can actually trust with ๐—น๐—ผ๐—ด๐—ด๐—ถ๐—ป๐—ด, ๐—ฒ๐—ฟ๐—ฟ๐—ผ๐—ฟ ๐—ต๐—ฎ๐—ป๐—ฑ๐—น๐—ถ๐—ป๐—ด, and ๐—ท๐—ผ๐—ฏ ๐˜€๐—ฐ๐—ต๐—ฒ๐—ฑ๐˜‚๐—น๐—ถ๐—ป๐—ด built in from day one, not added as an afterthought. ๐Ÿ“ฉ ๐—ฆ๐—ฒ๐—ป๐—ฑ ๐—บ๐—ฒ ๐—ฎ ๐—บ๐—ฒ๐˜€๐˜€๐—ฎ๐—ด๐—ฒ ๐˜„๐—ถ๐˜๐—ต ๐˜†๐—ผ๐˜‚๐—ฟ ๐—ฝ๐—ฟ๐—ผ๐—ท๐—ฒ๐—ฐ๐˜ ๐—ฑ๐—ฒ๐˜๐—ฎ๐—ถ๐—น๐˜€. ๐—œ ๐—ฟ๐—ฒ๐˜€๐—ฝ๐—ผ๐—ป๐—ฑ ๐—พ๐˜‚๐—ถ๐—ฐ๐—ธ๐—น๐˜† ๐—ฎ๐—ป๐—ฑ ๐˜„๐—ถ๐—น๐—น ๐—ผ๐˜‚๐˜๐—น๐—ถ๐—ป๐—ฒ ๐—ฎ ๐—ฐ๐—น๐—ฒ๐—ฎ๐—ฟ ๐—ฎ๐—ฝ๐—ฝ๐—ฟ๐—ผ๐—ฎ๐—ฐ๐—ต ๐—ณ๐—ผ๐—ฟ ๐˜†๐—ผ๐˜‚๐—ฟ ๐—ฝ๐—ฟ๐—ผ๐—ท๐—ฒ๐—ฐ๐˜.

  • Data Engineering
  • ETL Pipeline
  • Microsoft Azure
  • Microsoft Power BI
  • Databricks Platform
  • Data Warehousing
  • Data Lake
  • SQL
  • Data Modeling
  • SQL Server Integration Services
  • SQL Server Reporting Services
  • Microsoft SQL Server
  • Oracle
  • Fabric
  • Database Development
  • PySpark
  • Business Intelligence
  • PostgreSQL
  • Microsoft Power BI Data Visualization
  • Big Data
Daniel Fabrico S.

Chaco Pora, Argentina

$30/hr
5.0
1 jobs

ยธยธโ™ฌยทยฏยทโ™ชยทยฏยทโ™ซยธยธ ๐—ช๐—ฒ๐—น๐—ฐ๐—ผ๐—บ๐—ฒ ๐˜๐—ผ ๐—บ๐˜† ๐—ฝ๐—ฟ๐—ผ๐—ณ๐—ถ๐—น๐—ฒ! ยธยธโ™ซยทยฏยทโ™ชยธโ™ฉยทยฏยทโ™ฌยธยธ I'm a Senior Data Engineer & Cloud Data Architect. I bridge the gap between fragmented raw data and high-performance, analytics-ready infrastructure. Whether you need to build a scalable data warehouse from scratch, transition from legacy ETL to modern dbt/Databricks stack, optimize costly cloud queries, or power real-time AI/ML applications, I specialize in architecting reliable, zero-downtime data pipelines across AWS, GCP, Azure, Snowflake, and BigQuery. โšก ๐‚๐จ๐ซ๐ž ๐’๐ž๐ซ๐ฏ๐ข๐œ๐ž๐ฌ 1. End-to-End Modern Data Stack (MDS) & ETL/ELT Pipelines Designing and deploying automated, resilient pipelines that extract, clean, transform, and load petabyte-scale data into centralized analytics hubs. โ—พ Batch & Stream Ingestion: Building automated ingestion jobs from SaaS applications, REST APIs, webhooks, and legacy DBs using Fivetran, Airbyte, Kafka, and Debezium (Change Data Capture - CDC). โ—พ Analytics Engineering: Modular, version-controlled transformations with dbt (Data Build Tool), custom SQL, and PySpark-complete with automated documentation and lineage tracking. โ—พ Workflow Orchestration: Designing DAGs, automated retries, and monitoring alerts using Apache Airflow, Prefect, Dagster, and AWS Step Functions. 2. Cloud Data Warehousing & Lakehouse Architecture (Snowflake, Databricks, BigQuery) Structuring high-efficiency, cost-optimized databases designed for instant analytical querying and BI dashboard performance. โ—พ Warehouse Optimization: Clustering keys, partitioning, materialization, micro-partitioning, and query tuning to cut monthly cloud compute/storage costs by 30%โ€“60%. โ—พ Lakehouse & Open Table Formats: Architecting Delta Lake, Apache Iceberg, and Hudi layers on AWS S3/GCP Cloud Storage using Medallion Architecture (Bronze -> Silver -> Gold). โ—พ Data Modeling: Dimensional modeling (Kimball methodology), Star/Snowflake Schemas, Data Vault 2.0, and Wide Flat Tables (OBT) optimized for Looker, Tableau, and PowerBI. 3. Real-Time Data Streaming & Event-Driven Systems Enabling millisecond-latency processing for live dashboards, fraud detection, dynamic pricing, and real-time operational metrics. โ—พ Event Streaming: Setting up Apache Kafka clusters, AWS Kinesis, GCP Pub/Sub, and RabbitMQ with event serialization (Avro, Protobuf). โ—พ Real-Time Analytics: Developing continuous stream-processing engines using Apache Flink, Spark Streaming, and ClickHouse/RisingWave for immediate insight delivery. 4. Data Quality, Governance, MLOps & AI Infrastructure Ensuring every byte of data entering your reporting systems is accurate, secure, compliant, and ready for advanced analytics or LLM applications. โ—พ Data Quality & Observability: Automated schema validation, anomaly detection, and data testing using Great Expectations, Soda, and dbt test suites. โ—พ AI/ML Infrastructure: Vector database setup (Pinecone, Weaviate, Qdrant, Milvus), RAG pipeline data ingestion, and feature store integration (Feast) for AI model training. โ—พ Governance & Compliance: Role-Based Access Control (RBAC), Column/Row-level masking, PII obfuscation, and automated lineage mapping for GDPR/HIPAA compliance. โšก ๐“๐ž๐œ๐ก๐ง๐จ๐ฅ๐จ๐ ๐ข๐ž๐ฌ & ๐…๐ซ๐š๐ฆ๐ž๐ฐ๐—ผ๐—ฟ๐ค๐ฌ ๐ˆ ๐‡๐—ฎ๐˜ƒ๐ž ๐Œ๐š๐ฌ๐ญ๐ž๐ซ๐ž๐ - ๐—ช๐—ผ๐—ฟ๐—ธ๐—ณ๐—น๐—ผ๐˜„ ๐—ข๐—ฟ๐—ฐ๐—ต๐—ฒ๐˜€๐˜๐—ฟ๐—ฎ๐˜๐—ถ๐—ผ๐—ป: Apache Airflow, Prefect, Dagster, Mage, AWS Step Functions, MWAA - ๐——๐—ฎ๐˜๐—ฎ ๐—ช๐—ฎ๐—ฟ๐—ฒ๐—ต๐—ผ๐˜‚๐˜€๐—ฒ๐˜€ & ๐—˜๐—ป๐—ด๐—ถ๐—ป๐—ฒ๐˜€: Snowflake, Google BigQuery, AWS Redshift, ClickHouse, Trino/Presto, DuckDB - ๐——๐—ฎ๐˜๐—ฎ ๐—Ÿ๐—ฎ๐—ธ๐—ฒ / ๐—Ÿ๐—ฎ๐—ธ๐—ฒ๐—ต๐—ผ๐˜‚๐˜€๐—ฒ: Databricks, Apache Iceberg, Delta Lake, Apache Hudi, AWS Glue, PySpark, Apache Spark - ๐—˜๐—ง๐—Ÿ / ๐—˜๐—Ÿ๐—ง & ๐—ง๐—ฟ๐—ฎ๐—ป๐˜€๐—ณ๐—ผ๐—ฟ๐—บ๐—ฎ๐˜๐—ถ๐—ผ๐—ป: dbt (Core & Cloud), Airbyte, Fivetran, Kafka Connect, Debezium, Meltano - ๐—ฆ๐˜๐—ฟ๐—ฒ๐—ฎ๐—บ๐—ถ๐—ป๐—ด & ๐— ๐—ฒ๐˜€๐˜€๐—ฎ๐—ด๐—ถ๐—ป๐—ด: Apache Kafka, AWS Kinesis, GCP Pub/Sub, Apache Flink, Spark Streaming, RabbitMQ - ๐——๐—ฎ๐˜๐—ฎ๐—ฏ๐—ฎ๐˜€๐—ฒ๐˜€ (๐—ก๐—ผ๐—ฆ๐—ค๐—Ÿ & ๐—ฅ๐——๐—•๐— ๐—ฆ): PostgreSQL, MySQL, MongoDB, Redis, Cassandra, DynamoDB, Pinecone, Qdrant - ๐—Ÿ๐—ฎ๐—ป๐—ด๐˜‚๐—ฎ๐—ด๐—ฒ๐˜€ & ๐—ฆ๐—พ๐—น: Python (Pandas, Polars, PySpark, SQLAchemy), SQL (Advanced Dialects), Scala, Bash, Go - ๐—œ๐—ป๐—ณ๐—ฟ๐—ฎ๐˜€๐˜๐—ฟ๐˜‚๐—ฐ๐˜๐˜‚๐—ฟ๐—ฒ & ๐——๐—ฒ๐˜ƒ๐—ข๐—ฝ๐˜€: Terraform, Docker, Kubernetes, AWS (S3, EC2, ECS, Lambda), GCP, Azure, GitHub Actions, CI/CD - ๐—ค๐˜‚๐—ฎ๐—น๐—ถ๐˜๐˜† & ๐—ข๐—ฏ๐˜€๐—ฒ๐—ฟ๐˜ƒ๐—ฎ๐—ฏ๐—ถ๐—น๐—ถ๐˜๐˜†: Great Expectations, Soda, Monte Carlo, OpenLineage, Datahub โšก ๐—›๐—ผ๐˜„ ๐—œ ๐—ช๐—ผ๐—ฟ๐—ธ: I am hired to design production-grade data pipelines, modernize legacy data stacks, fix slow analytics queries, and bring software engineering best practices (Git, CI/CD, unit testing, modular code) into data infrastructure. I prioritize clean lineage, cost efficiency, ironclad data security, and zero-downtime migrations. โšก ๐—ช๐—ต๐˜† ๐—–๐—น๐—ถ๐—ฒ๐—ป๐˜๐˜€ ๐—–๐—ต๐—ผ๐—ผ๐˜€๐—ฒ ๐— ๐—ฒ: โœ”๏ธ Pipelines built to scale โœ”๏ธ Massive cloud bill reduction โœ”๏ธ Production-grade reliability โœ”๏ธ Software engineering rigor โœ”๏ธ Clear communication ๐Ÿ‘‰ ๐—–๐—น๐—ถ๐—ฐ๐—ธ ๐— ๐—ฒ๐˜€๐˜€๐—ฎ๐—ด๐—ฒ - ๐—น๐—ฒ๐˜'๐˜€ ๐˜๐—ฎ๐—น๐—ธ.

  • SQL
  • Python
  • ETL Pipeline
  • Data Mining
  • Data Integration
  • Data Analysis
  • ETL
  • Big Data
  • Data Engineering
  • Data Warehousing & ETL Software
  • Database Architecture
  • Database Design
  • Machine Learning
  • BigQuery
  • Apache Spark
  • Data Warehousing
  • Amazon Web Services
  • Data Scraping
  • Data Migration
  • dbt
Mochammad Arie N.

Jakarta, Indonesia

$15/hr
5.0
7 jobs

Most data pipelines donโ€™t fail because of code. They fail because they weren't built for scale. With 5+ years of experience engineering data systems at companies like Danone and Zurich, I help businesses transform fragile prototypes into resilient, production-grade infrastructure. I donโ€™t just move data; I build the "Source of Truth" that leadership and AI systems actually trust. โž” Productionizing AI Pipelines: Hardening Python prototypes into scalable RAG and LLM infrastructures (Azure). โž” Infrastructure-as-Code: Building automated, modular ETL/ELT pipelines that don't require daily manual fixes. โž” The "One-Source" Dashboard: Integrating messy data from APIs, SaaS (Shopify, HubSpot), and databases into clean Snowflake/BigQuery layers. โž” Performance Recovery: Optimizing slow SQL queries and high-cost cloud warehouses to save you thousands in monthly spend. โž” Technical Writing for Data & AI Teams: Creating product documentation, implementation guides, architecture documentation, data dictionaries, knowledge bases, and thought leadership content that makes complex systems easier to understand and adopt. ๐Ÿ›  Tech Stack Languages: Python (FastAPI, Pandas, PySpark), SQL Data Engineering: ETL/ELT Pipelines, Data Warehousing, Data Modeling, Data Quality, Data Governance Cloud & Warehousing: Snowflake, BigQuery, Databricks, Azure Data Factory, Azure Data Lake, AWS (S3, Athena, Glue) Orchestration & Transformation: Apache Airflow, dbt Analytics & BI: Tableau, Power BI Development & Collaboration: Git, GitHub, VS Code Data Ops: API Integrations, Data Validation, Workflow Automation Technical Writing: Product Documentation, API Documentation, User Guides, Knowledge Bases, Data Dictionaries, Technical Blog Content โœ… Why Me? 5+ Years Experience: I've seen what breaks at the enterprise level and how to prevent it in your startup. Hands-On Builder & Technical Writer: I can both build the system and explain it clearly to engineers, stakeholders, and customers. Speed over Perfection: I focus on shipping high-impact systems that drive revenue, not just technical documentation. Transparent Communication: You get regular updates and a partner who challenges requirements to find better solutions. Ready to clean up your data debt?

  • Data Engineering
  • Python
  • SQL
  • ETL Pipeline
  • Databricks Platform
  • Snowflake
  • dbt
  • Apache Airflow
  • BigQuery
  • Data Migration
  • LLM Prompt
  • AI Content Writing
  • Microsoft Power BI
  • Machine Learning
  • Microsoft Azure
  • Data Warehousing & ETL Software
  • Technical Writing
  • Microsoft Power Automate
  • Data Warehousing
  • Azure Service Fabric
Vebri S.

Jakarta, Indonesia

$25/hr
5.0
1 jobs

Are your data pipelines failing silently, or is your cloud data warehouse bill spiraling out of control? I help data teams and startups design, build, and optimize reliable modern data platforms across AWS and GCP - ensuring zero data loss, predictable pipeline runs, and cost-efficient query performance. Core Focus Areas & Solutions: - Pipeline Orchestration & Modeling: Production-grade Apache Airflow DAGs, dbt transformations, modular ELT architectures. - Data Lake & Warehouse Design: Modern storage and modeling on Amazon Redshift, BigQuer, and S3. - Cost & Performance Optimization: Partitioning/clustering tuning, cluster resizing, query debugging, and infrastructure refactoring to reduce monthly cloud spend. - Automated & Resilient Ingestion: CDC ingestion, API connectors, error handling, automated alerts, and schema evolution handling. Tech Stack: Cloud: AWS (S3, Redshift, ECS, Lambda, IAM), GCP (BigQuery, Cloud Storage, GCF) Data Tools: Apache Airflow, dbt, Docker, Airbyte, Kafka Languages: Python (Pandas, Polars, PySpark), Advanced SQL, Bash Whether you need to migrate legacy pipelines, fix brittle ETL workflows, or optimize your data warehouse infrastructure from the ground up, let's connect and discuss your architecture.

  • Data Engineering
  • Data Lake
  • Data Ingestion
  • ETL
  • Python
  • SQL
  • Apache Airflow
  • Google Cloud Platform
  • Amazon Web Services
  • Amazon Redshift
  • BigQuery
  • Data Warehousing
  • Data Integration
  • dbt
  • Kubernetes
Andrew F.

Chernivtsi, Ukraine

$35/hr
5.0
2 jobs

I'm a senior software engineer specializing in backend development, cloud infrastructure, and data-intensive applications. I help companies build reliable systems, optimize existing platforms, and solve challenging technical problems that require more than just writing code. My expertise includes Python, PHP (Laravel), MySQL, Elasticsearch, Google Cloud Platform (Cloud SQL, BigQuery, Compute Engine, App Engine, IAM), and API design. I also work extensively with AI and data processing pipelines, including named entity recognition (NER), large-scale data cleansing, search optimization, and knowledge graph technologies. I enjoy tackling complex engineering challenges such as: * Designing scalable backend services and REST APIs * Optimizing databases and SQL queries for large datasets * Building ETL and data processing pipelines * Cloud architecture, migrations, and production troubleshooting on GCP * Elasticsearch search tuning and indexing * AI-powered applications and data enrichment workflows * Performance optimization, debugging, and infrastructure automation Clients appreciate that I quickly understand unfamiliar systems, identify root causes instead of treating symptoms, and deliver practical, maintainable solutions. Whether it's fixing a production issue, designing a new service, optimizing cloud costs, or building a data platform from scratch, I focus on writing clean, reliable code that can be maintained long after the project is finished. I'm comfortable working independently, collaborating with distributed teams, and communicating technical concepts clearly throughout a project.

  • MySQL
  • Laravel
  • Python
  • JavaScript
  • PHP
  • MongoDB
  • RESTful API
  • Google Cloud Platform
  • Elasticsearch
  • Google Dataflow
  • Google APIs
  • Zend
  • BigQuery
  • Google App Engine
  • Apache Airflow
Danish V.

Mithi, Pakistan

$25/hr
4.8
84 jobs

Most data problems show up the same way, whether it's a spreadsheet someone updates by hand every week or a pipeline that's quietly started giving numbers nobody trusts. I build and fix ETL pipelines, warehouses, and API integrations on GCP and AWS, from replacing manual copy-paste with automated pulls to cutting a client's warehouse costs by ~50%. Tell me what's manual or what feels off, and I'll give you a straight read on what it'll take to fix it. Over the last 3 years I've delivered 55+ data projects across finance, healthcare, energy, and e-commerce, from one-off ETL jobs to platforms processing billions of records a day. I work GCP-first (BigQuery, Airflow, dbt, Dataflow), and I'm comfortable across AWS, Postgres, and the messy real-world stack most teams actually have. What I build: - End-to-end ETL/ELT pipelines in Python, SQL, Airflow, and dbt - BigQuery / Snowflake / Redshift warehouses and data models that stay clean as they grow - Migrations off legacy jobs and on-prem databases โ€” without losing data in the move - Metabase, Looker Studio, and Power BI dashboards your team will actually open - Query and cost optimization when your warehouse bill stops making sense - API integrations with proper logging, retries, and checkpoints, so failures are visible instead of silent You probably need me if: - Your pipelines break and you hear it from a stakeholder, not an alert - Reports run slow, cost too much, or quietly disagree with each other - A previous developer left and nobody fully understands the setup anymore - You're scaling fast and the current data stack is starting to crack A few real results: - Architected pipelines processing 5B+ records daily at 99% reliability - Cut a client's warehouse costs ~50% by migrating legacy jobs to BigQuery - 4ร— throughput and 70% faster ingestion on an API pipeline pulling 2K+ domains a day - 40% faster pipeline runs through Airflow optimization How I work: a clear yes/no on feasibility before you commit, regular updates, and no disappearing mid-project. Most clients come back โ€” usually because fixing one thing surfaces the next. If that sounds like your situation, send a short note on what's breaking or what you're trying to build, and I'll tell you straight what it'll take.

  • Python
  • Data Engineering
  • SQL
  • Google Cloud Platform
  • Amazon Web Services
  • BigQuery
  • Apache Airflow
  • dbt
  • MySQL
  • Amazon Redshift
  • PySpark
  • Big Data
  • ETL Pipeline
  • Data Warehousing
  • API
  • Spreadsheet Skills
  • Automation
  • Cloud Database
  • MySQL Programming
  • PostgreSQL

How it works

Post a job for freePost a job

Tell us what you need. Create your own job post or generate one with AI then filter talent matches.

Hire top talent fast

Consult, interview, and hire quickly, so you can meet the freelancers you're excited about.

Collaborate easily

Use Upwork to chat or video call, share files, and track project progress right from the app.

Payment simplified

Manage payments in one place with flexible billing options. Only pay for approved work, hourly or by milestone.

Don't just take our word for it

What does a DataLife engine specialist do?

A DataLife engine specialist builds and maintains news portals and content-heavy websites using the DataLife Engine content management system. This role focuses on configuring the PHP and MySQL environment to support high-traffic publishing workflows. The specialist customizes templates and develops modules to match specific editorial requirements. They resolve technical conflicts between the core engine code and third-party integrations.

  • Install the DataLife Engine script on a web server and set correct writable permissions for required folders such as engine/data and engine/cache. Configure the MySQL database connection during the initial setup process to align the database structure with the versioned application files. Verify the installation by loading the site front end and accessing the Administration Panel through the admin.php entry point.
  • Edit DLE template files like mainpage.tpl to implement custom site layouts and integrate external style sheets or scripts. Use template tags such as {include file=...} to pull in dynamic content blocks and maintain a consistent visual theme across all pages. Adjust theme assets to ensure the design renders correctly on various devices while preserving the structural integrity of the underlying PHP code.
  • Develop or modify DLE modules and plugins to add functionality that the default installation does not support. Write custom PHP code to extend the engine capabilities and ensure new features interact properly with existing template tags. Test these additions in a safe staging environment before deploying them to the live production site to prevent downtime or data loss.
  • Troubleshoot runtime errors by examining logs and inspecting the engine/cache and templates directories for configuration mismatches. Resolve integration issues that arise when updating the DLE script or when third-party services fail to communicate with the CMS. Apply fixes to the core code or template logic and restart the server cache to confirm the changes take effect immediately.
  • Perform version upgrades by following official guidance to keep the database schema synchronized with the updated application files. Check compatibility between the new version and existing custom modules or templates before applying the update to the live environment. Back up the current site state and restore it if the upgrade process introduces critical bugs or breaks existing functionality.

How to hire a DataLife engine specialist on Upwork

Step 1: Post a job

Define your project needs clearly to attract qualified candidates who understand the PHP and MySQL architecture of DataLife Engine. Use the Job Post Generator powered by Umaโ„ข, Upwork's Mindful AI to draft a precise description in seconds. Describe your requirements in a few sentences, and Uma creates a structured post tailored to this role. You can write a new post, update a saved draft, or reuse an existing post to save time.

  • Specify whether you need template customization using DLE tags like {include} or backend module development for specific site features.
  • List required technical tasks such as configuring writable folder permissions, managing cache directories, or performing version upgrades.
  • Include details about your current DLE version and any specific integration challenges with third-party scripts or custom themes.

Step 2: Evaluate candidates

Review portfolios for evidence of customized DataLife Engine deployments and successful troubleshooting of template or database issues. Uma can run instant video interviews and build shortlists with side-by-side comparisons to help you identify top performers quickly.

  • Look for examples of custom mainpage.tpl edits that demonstrate clean code structure and proper use of DLE template syntax.
  • Check for experience with safe staging-to-production workflows that preserve database integrity during script updates.
  • Verify familiarity with the Administration Panel and ability to resolve common errors in engine data or cache folders.

Step 3: Interview your top choices

Discuss specific technical scenarios to gauge their depth of knowledge regarding DLE architecture and PHP compatibility. Schedule and conduct interviews within Upwork Messages, which generates an immediate transcript and summary after each session.

  • Ask how they handle CHMOD permission settings for critical directories like engine/data during initial installation.
  • Request examples of custom modules they built and how they ensured compatibility with core DLE updates.
  • Discuss their approach to debugging template tag conflicts when integrating external JavaScript or CSS assets.

Step 4: Agree on scope and begin work

Set clear milestones for deliverables such as theme implementation, plugin development, or system upgrades. Use Upwork Messages and the contract workroom for all communication and project management to keep records organized.

  • Define specific outputs like installed DLE sites with correct permissions or customized themes ready for production.
  • Rely on identity verification, payment protection, hourly tracking, and project funds for security and transparent billing.
  • Establish a testing protocol where the freelancer verifies site functionality via admin.php before marking tasks complete.

Upwork is not affiliated with and does not sponsor or endorse any of the tools or services discussed in this article. These tools and services are provided only as potential options, and each reader and company should take the time needed to adequately analyze and determine the tools or services that would best fit their specific needs and situation.

The rates and information provided in this article are based on current data and industry sources available at the time of publication. Freelance rates can vary depending on factors such as experience, location, project scope, and market conditions. Readers are encouraged to conduct their own research to confirm current rates and trends, as this information may change over time.

How much does hiring a DataLife engine specialist cost?

Hiring a DataLife engine specialist typically costs $500-$2,500 per project, depending on scope and experience. Final pricing depends on technical complexity, required integrations, template customization depth, and the freelancer's experience level.

Installation and configuration

$500-$1,000/project

Entry-level to mid-level
  • Installed DLE with correct folder permissions
  • Configured MySQL database and admin panel access
  • Tested site load and administration login

Template customization

$1,000-$2,000/project

Mid-level
  • Customized mainpage.tpl and theme assets
  • Included external scripts and styles via template tags
  • Validated responsive design across devices

Module development

$2,000-$4,000/project

Mid-level to senior-level
  • Developed custom PHP modules for new features
  • Implemented template tags for dynamic content
  • Verified module functionality in staging environment

Version upgrade

$4,000-$6,500/project

Senior-level
  • Updated script files and aligned database structure
  • Resolved conflicts in templates and custom code
  • Deployed changes to production with cache clearance

Full site migration

$6,500-$10,000/project

Expert-level
  • Migrated content and users to new DLE instance
  • Configured caching and server settings for performance
  • Documented custom changes and maintenance procedures

Frequently asked questions

Is hiring a DataLife engine specialist worth it?

For most businesses, yes: hiring a DataLife engine specialist is worthwhile. This expert handles the specific PHP and MySQL configurations required to keep your news or content site running smoothly. They resolve template conflicts and manage safe upgrades so your publication stays online without manual troubleshooting.

How do I evaluate DataLife engine specialist candidates?

Look for candidates who describe specific experiences with DLE template tags and module development. Ask them to explain how they handle CHMOD permissions for engine folders during installation or how they troubleshoot a broken admin.php access after an update.

What tasks does a DataLife engine specialist handle?

A DataLife engine specialist installs the CMS, customizes templates like mainpage.tpl, and develops plugins. They also manage database updates and fix integration errors within the engine and cache directories.

How much does it cost to hire a DataLife engine specialist?

Freelancers typically charge between $19 and $34 per hour for this specialized work. The final rate depends on the complexity of your template customization or plugin development needs.