Observability Engineer / SRE - Datadog Experience - Full time Remote

Posted 2 weeks ago

Worldwide

Summary

We are seeking an Observability Engineer with a strong **SRE and troubleshooting mindset** to lead and evolve the organization's observability strategy across cloud and on-premises environments. This role will serve as the primary owner and subject matter expert for the organization's APM/observability platform (currently **Datadog**), building, scaling, and operating a comprehensive 24/7 monitoring solution across external websites, key applications, and critical infrastructure components such as firewalls and OpenShift. The ideal candidate is, first and foremost, a strong SRE-minded troubleshooter - someone who can dig into a production issue, reason through root cause across logs, metrics, and traces, and drive it to resolution. Platform tool depth (Datadog, Dynatrace, New Relic, AppDynamics, or similar) is important so the candidate can speak credibly to the technicals and get hands-on quickly, but it is secondary to that core troubleshooting instinct. Equally critical is the ability to communicate clearly and work directly with developers — this role is regularly **face-to-face with engineering teams**, explaining findings, negotiating instrumentation changes, and translating observability data into action. This is an individual contributor role with strong engineering and scripting expectations, **not a pure administration role**. ### Key Responsibilities - **Troubleshooting and root cause analysis:** Serve as a hands-on troubleshooter for production issues, using logs, metrics, and traces to drive incidents to resolution and identify systemic reliability risks. - **Platform ownership:** Serve as the primary owner of the organization's APM/observability platform (Datadog), architecting, building, and maintaining scalable observability solutions across cloud and on-prem environments, including Windows and Linux/Unix systems. - **Developer-facing partnership:** Work directly and regularly with development teams to translate reliability goals and incident findings into actionable monitoring strategies, dashboards, SLOs, and alerting frameworks — communicating clearly to both technical and non-technical stakeholders. - **Monitoring strategy:** Partner with application, DevOps, SRE/operations, infrastructure, and security teams to define monitoring coverage and reliability standards. - **Platform migration:** Lead and execute the migration from OpenView to Datadog while maintaining monitoring continuity and improving monitoring fidelity across migrated services and infrastructure. - **Automation and configuration as code:** Develop and automate processes using Datadog APIs, the Datadog Terraform provider, Python, PowerShell, and Bash to manage monitors, dashboards, alerts, and telemetry configuration at scale. - **Full-stack observability:** Implement and optimize APM, distributed tracing, log management, infrastructure monitoring, and Network Performance Monitoring (NPM). - **User experience monitoring:** Build and evolve RUM and Synthetic Monitoring capabilities to track end-user experience and proactively validate availability of external-facing services and critical workflows. - **SLOs and alert quality:** Define and operationalize SLOs and error budgets; reduce alert noise through correlation, enrichment, threshold tuning, and monitor dependency mapping. - **Incident management integration:** Integrate the observability platform with ServiceNow for incident/problem ticket routing and escalation; produce runbooks, post-incident reviews, and reliability dashboards. - **Standards and enablement:** Champion OpenTelemetry adoption and drive consistent logging, metrics, and tracing standards across the engineering organization. - **Application onboarding:** Onboard new applications and services into the platform and guide engineering teams on instrumentation, agent deployment, and observability best practices. - **Cost and governance:** Collaborate on platform cost optimization, data governance, and scaling strategies so the platform remains performant and cost-effective as the environment grows.

  • More than 30 hrs/week
    Hourly
  • 6+ months
    Duration
  • Intermediate
    Experience Level
  • Remote Job
  • Ongoing project
    Project Type

Contract-to-hire opportunity

This lets talent know that this job could become full time.
Learn more
Skills and Expertise
Mandatory skills
datadog
Activity on this job
  • Proposals:15 to 20
  • Last viewed by client:last week
  • Interviewing:
    11
  • Invites sent:
    10
  • Unanswered invites:
    3
About the client
Member since Aug 25, 2023
  • United States
    2:22 PM
  • $19K total spent
    34 hires, 7 active
  • 978 hours

Explore similar jobs on Upwork

Kubernetes
Red Hat
Apache and PHP-FPM Configuration ExpertHourly‐ Posted 2 months ago
Ubuntu
Apache HTTP Server

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo