Observability Engineer / SRE - Datadog Experience - Full time Hybrid NYC

Posted 3 days ago

Only freelancers located in the U.S. may apply.U.S. located freelancers only

Summary

OVERVIEW We are looking for a hands-on Observability Engineer with a strong SRE and production troubleshooting mindset to help lead and improve observability across cloud and on-premises environments. This role is not just Datadog administration. We need someone who can investigate real production issues, reason across logs, metrics, traces, infrastructure, and application behavior, and work directly with developers to turn findings into better monitoring, instrumentation, dashboards, alerts, and reliability practices. The current platform is Datadog, but experience with similar observability/APM platforms such as Dynatrace, New Relic, AppDynamics, or OpenTelemetry-based stacks is also valuable. Datadog depth is preferred, but strong troubleshooting instincts and SRE thinking are the most important requirements. WHAT YOU'LL DO Own and improve the organization’s Datadog observability/APM platform across cloud and on-prem environments Troubleshoot production issues using logs, metrics, traces, infrastructure signals, and application behavior to identify root cause and reliability risks Build and maintain monitoring coverage for external websites, key applications, critical infrastructure, firewalls, OpenShift, Windows, Linux, and Unix systems Work directly with application developers and engineering teams to define instrumentation, dashboards, SLOs, alerting, and reliability improvements Lead or support migration from OpenView to Datadog while maintaining monitoring continuity and improving signal quality Automate observability configuration using Datadog APIs, Terraform, Python, PowerShell, and Bash Implement and optimize APM, distributed tracing, log management, infrastructure monitoring, and Network Performance Monitoring Build and improve RUM and Synthetic Monitoring for external-facing services and critical user workflows Define SLOs, error budgets, alerting standards, and monitor dependency mapping to reduce noise and improve incident response Integrate observability workflows with ServiceNow for incident routing, escalation, problem management, and reporting Create runbooks, post-incident review inputs, reliability dashboards, and documentation for operational teams Promote OpenTelemetry and consistent logging, metrics, and tracing standards across engineering teams Onboard new applications and services into the observability platform Help with platform cost optimization, data governance, and scaling strategy as usage grows WHAT WE'RE LOOKING FOR Strong SRE, production support, or reliability engineering background Excellent troubleshooting and root cause analysis skills across applications, infrastructure, networks, and cloud/on-prem systems Hands-on experience with Datadog or a comparable APM/observability platform such as Dynatrace, New Relic, AppDynamics, Splunk Observability, or similar Experience with logs, metrics, traces, dashboards, alerts, SLOs, synthetic monitoring, and incident response Ability to work directly with developers and explain technical findings clearly Experience with scripting or automation using Python, PowerShell, Bash, Terraform, or APIs Familiarity with Windows and Linux/Unix environments Understanding of cloud infrastructure, OpenShift/Kubernetes, firewalls, networking, and application monitoring patterns Ability to operate as an individual contributor with strong engineering ownership NICE TO HAVE Strong Datadog experience, including APM, Logs, Infrastructure Monitoring, RUM, Synthetic Monitoring, NPM, dashboards, monitors, and API automation Experience migrating from legacy monitoring platforms such as OpenView into Datadog or another modern observability platform OpenTelemetry implementation experience ServiceNow integration experience Experience supporting regulated, enterprise, financial services, or large hybrid environments Prior work creating observability standards, onboarding playbooks, runbooks, or SLO frameworks HOW TO APPLY Describe your most relevant Datadog, APM, or observability platform experience Share an example of a difficult production issue you helped troubleshoot, including how you used logs, metrics, traces, or infrastructure signals to find root cause Tell us about your experience working directly with developers to improve instrumentation, alerting, or reliability Mention any experience with OpenTelemetry, ServiceNow, OpenShift/Kubernetes, Terraform, Python, PowerShell, or Bash Please include a Loom or similar short video where you either walk through relevant work from your profile/portfolio or briefly introduce yourself and explain why you are a strong fit for this project.

  • More than 30 hrs/week
    Hourly
  • 6+ months
    Duration
  • Intermediate
    Experience Level
  • Remote Job
  • Ongoing project
    Project Type

Contract-to-hire opportunity

This lets talent know that this job could become full time.
Learn more
Skills and Expertise
Mandatory skills
datadog
Activity on this job
  • Proposals:Less than 5
  • Last viewed by client:2 days ago
  • Interviewing:
    2
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since Aug 25, 2023
  • United States
    5:00 PM
  • $20K total spent
    35 hires, 8 active
  • 989 hours

Explore similar jobs on Upwork

Senior DevOps & SRE / Infrastructure EngineerFixed-price‐ Posted 4 weeks ago
DevOps
Microsoft Azure
Automated Monitoring
Configuration Management
Infrastructure Management
Google Cloud Platform
Amazon Web Services
Cloud Computing
Cloudflare

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo