Operational Data & Observability Engineer

Posted 4 days ago

Worldwide

Summary

We are seeking an experienced Operational Data & Observability Engineer to design, build, and improve modern observability solutions across cloud infrastructure, applications, and production environments. This role focuses on enhancing system reliability, improving incident response capabilities, and creating scalable monitoring, logging, and telemetry platforms. The ideal candidate has strong experience in SRE, DevOps, Platform Engineering, Observability Engineering, and cloud infrastructure, with hands-on expertise building monitoring solutions for large-scale distributed systems. Responsibilities - Design and implement enterprise observability strategies across infrastructure, applications, and services - Build and maintain monitoring dashboards, alerts, and Service Level Objectives (SLOs) - Develop centralized logging and log analysis solutions - Implement distributed tracing and application performance monitoring solutions - Establish performance baselines and develop anomaly detection approaches - Deploy and maintain telemetry collection systems for metrics, logs, and events - Build operational data pipelines for monitoring, analytics, and reporting - Develop APIs and integrations for operational data consumption - Ensure data quality, consistency, retention, and cost optimization for observability platforms - Troubleshoot production incidents using monitoring, logging, and tracing data - Participate in incident response and on-call rotations - Create operational documentation, runbooks, and troubleshooting guides - Partner with engineering teams to improve reliability, scalability, and operational readiness - Optimize observability infrastructure for performance, security, and resilience - Manage and enhance observability platforms such as Datadog, Grafana, Prometheus, ELK Stack, New Relic, or similar tools - Evaluate new observability technologies and recommend improvements - Automate monitoring deployments, instrumentation, and platform configurations - Perform platform upgrades, maintenance, and lifecycle management Required Skills & Experience - 3+ years of experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, Operations Engineering, or Observability Engineering - Strong experience with modern observability platforms including: Prometheus Grafana Datadog New Relic ELK / Elastic Stack Splunk CloudWatch - Strong programming or scripting skills with: Python Go Bash Similar automation languages Deep understanding of observability concepts: Metrics Logging Distributed tracing Application Performance Monitoring (APM) - Experience with cloud platforms: AWS Azure Google Cloud Platform - Hands-on experience with: Kubernetes Container orchestration Microservices architectures Infrastructure monitoring Network monitoring Database monitoring Storage performance monitoring - Experience with Infrastructure as Code tools: Terraform Ansible Similar technologies - Experience supporting production environments with incident management and root cause analysis - Ability to build and maintain operational data pipelines and telemetry systems - Understanding of security monitoring, audit logging, and compliance requirements - Familiarity with eBPF and Linux performance monitoring is a plus Preferred Qualifications - Experience designing observability solutions for large-scale distributed systems - Experience improving system reliability through automation and proactive monitoring - Experience creating custom telemetry collection and data processing solutions - Strong troubleshooting and analytical skills - Excellent communication and documentation abilities - Ability to collaborate effectively with engineering, infrastructure, and security teams What We Are Looking For We are looking for someone who is passionate about reliability, automation, and operational excellence. The ideal candidate enjoys solving complex infrastructure problems, improving visibility into systems, and helping engineering teams build more reliable and scalable platforms.

  • More than 30 hrs/week
    Hourly
  • 6+ months
    Duration
  • Expert
    Experience Level
  • $10.00

    -

    $25.00

    Hourly
  • Remote Job
  • Ongoing project
    Project Type
Skills and Expertise
Mandatory skills
Data Engineering
ETL Pipeline
Activity on this job
  • Proposals:20 to 50
  • Last viewed by client:3 days ago
  • Interviewing:
    12
  • Invites sent:
    12
  • Unanswered invites:
    8
About the client
Member since Aug 10, 2026
  • USA
    Justin6:51 PM
  • $160 total spent
    2 hires, 1 active
  • 13 hours
  • Tech & IT
    Small company (2-9 people)

Explore similar jobs on Upwork

Varicent/SPM Consultant for GuidanceHourly‐ Posted 3 weeks ago
SAP
ETL
Data Engineering Training SpecialistFixed-price‐ Posted 4 weeks ago
SAS
Data Analysis
Machine Learning
Data Modeling

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo