Observability Engineer / SRE - Datadog Experience - Full time Hybrid NYC
Only freelancers located in the U.S. may apply.U.S. located freelancers only
OVERVIEW We are looking for a hands-on Observability Engineer with a strong SRE and production troubleshooting mindset to help lead and improve observability across cloud and on-premises environments. This role is not just Datadog administration. We need someone who can investigate real production issues, reason across logs, metrics, traces, infrastructure, and application behavior, and work directly with developers to turn findings into better monitoring, instrumentation, dashboards, alerts, and reliability practices. The current platform is Datadog, but experience with similar observability/APM platforms such as Dynatrace, New Relic, AppDynamics, or OpenTelemetry-based stacks is also valuable. Datadog depth is preferred, but strong troubleshooting instincts and SRE thinking are the most important requirements. WHAT YOU'LL DO Own and improve the organization’s Datadog observability/APM platform across cloud and on-prem environments Troubleshoot production issues using logs, metrics, traces, infrastructure signals, and application behavior to identify root cause and reliability risks Build and maintain monitoring coverage for external websites, key applications, critical infrastructure, firewalls, OpenShift, Windows, Linux, and Unix systems Work directly with application developers and engineering teams to define instrumentation, dashboards, SLOs, alerting, and reliability improvements Lead or support migration from OpenView to Datadog while maintaining monitoring continuity and improving signal quality Automate observability configuration using Datadog APIs, Terraform, Python, PowerShell, and Bash Implement and optimize APM, distributed tracing, log management, infrastructure monitoring, and Network Performance Monitoring Build and improve RUM and Synthetic Monitoring for external-facing services and critical user workflows Define SLOs, error budgets, alerting standards, and monitor dependency mapping to reduce noise and improve incident response Integrate observability workflows with ServiceNow for incident routing, escalation, problem management, and reporting Create runbooks, post-incident review inputs, reliability dashboards, and documentation for operational teams Promote OpenTelemetry and consistent logging, metrics, and tracing standards across engineering teams Onboard new applications and services into the observability platform Help with platform cost optimization, data governance, and scaling strategy as usage grows WHAT WE'RE LOOKING FOR Strong SRE, production support, or reliability engineering background Excellent troubleshooting and root cause analysis skills across applications, infrastructure, networks, and cloud/on-prem systems Hands-on experience with Datadog or a comparable APM/observability platform such as Dynatrace, New Relic, AppDynamics, Splunk Observability, or similar Experience with logs, metrics, traces, dashboards, alerts, SLOs, synthetic monitoring, and incident response Ability to work directly with developers and explain technical findings clearly Experience with scripting or automation using Python, PowerShell, Bash, Terraform, or APIs Familiarity with Windows and Linux/Unix environments Understanding of cloud infrastructure, OpenShift/Kubernetes, firewalls, networking, and application monitoring patterns Ability to operate as an individual contributor with strong engineering ownership NICE TO HAVE Strong Datadog experience, including APM, Logs, Infrastructure Monitoring, RUM, Synthetic Monitoring, NPM, dashboards, monitors, and API automation Experience migrating from legacy monitoring platforms such as OpenView into Datadog or another modern observability platform OpenTelemetry implementation experience ServiceNow integration experience Experience supporting regulated, enterprise, financial services, or large hybrid environments Prior work creating observability standards, onboarding playbooks, runbooks, or SLO frameworks HOW TO APPLY Describe your most relevant Datadog, APM, or observability platform experience Share an example of a difficult production issue you helped troubleshoot, including how you used logs, metrics, traces, or infrastructure signals to find root cause Tell us about your experience working directly with developers to improve instrumentation, alerting, or reliability Mention any experience with OpenTelemetry, ServiceNow, OpenShift/Kubernetes, Terraform, Python, PowerShell, or Bash Please include a Loom or similar short video where you either walk through relevant work from your profile/portfolio or briefly introduce yourself and explain why you are a strong fit for this project.
- More than 30 hrs/weekHourly
- 6+ monthsDuration
- IntermediateExperience Level
- Remote Job
- Ongoing projectProject Type
Skills and Expertise
Activity on this job
- Proposals:Less than 5
- Last viewed by client:2 days ago
- Interviewing:2
- Invites sent:0
- Unanswered invites:0
About the client
- United States5:00 PM
- $20K total spent35 hires, 8 active
- 989 hours
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by