обновлено 8 дней назад
Operational Data & Observability Engineer
145 000 - 180 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Operational Data & Observability Engineer (Observability/Cloud Infrastructure): Building monitoring, logging, tracing, and operational data platforms for reliable, scalable, and performant production environments with an accent on metrics pipelines, dashboards, alerting, SLOs, and centralized telemetry. Focus on troubleshooting production incidents, reducing alert noise and MTTR, automating platform administration, and optimizing highly available observability infrastructure.
Location: US; hybrid or remote work arrangements are available depending on business needs.
Salary: $145,000–$180,000 USD per year, plus potential bonus, equity, and/or commission programs.
Company
Nscale builds and operates platforms that require reliable, scalable, and observable production infrastructure.
What you will do
- Design and implement observability strategies across infrastructure, services, and applications.
- Build dashboards, alerts, SLOs, centralized logging pipelines, distributed tracing, performance baselines, and anomaly detection.
- Deploy and maintain metrics, logs, events, and telemetry collection systems, including operational data pipelines, APIs, and integrations.
- Troubleshoot production issues, participate in on-call rotations, and support incident response.
- Administer and automate observability platforms such as Datadog, Grafana, Prometheus, ELK Stack, and New Relic.
- Partner with DevOps, SRE, platform, and software engineering teams to improve reliability, scalability, operational readiness, and cost efficiency.
Requirements
- 3+ years of experience in DevOps, SRE, Operations Engineering, Platform Engineering, or Observability Engineering.
- Hands-on experience with monitoring platforms such as Prometheus, Grafana, Datadog, or New Relic, and centralized logging platforms such as ELK/Elastic Stack, Splunk, or CloudWatch.
- Proficiency in Python, Go, Bash, or equivalent scripting and programming languages.
- Strong understanding of metrics, logging, distributed tracing, APM, and performance monitoring across applications, infrastructure, networking, databases, and storage.
- Experience with AWS, Azure, or Google Cloud Platform and Kubernetes or other container orchestration technologies.
- Strong analytical, troubleshooting, communication, and documentation skills with a collaborative approach.
Nice to have
- Experience with microservices, incident management, root cause analysis, and post-incident reviews.
- Infrastructure as Code experience with Terraform, Ansible, or similar tools.
- Familiarity with eBPF, low-level Linux performance monitoring, custom telemetry, ETL, security monitoring, audit logging, or compliance requirements.
Culture & Benefits
- Hybrid or remote work arrangements, depending on business needs.
- Rotating on-call schedule supporting production environments.
- Occasional after-hours or incident response responsibilities for mission-critical systems.
- Medical, dental, and vision benefits, flexible paid time off, parental leave, and retirement plan participation may be available.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
13 дней назад
Site Reliability Engineer (AWS)
120 000 - 185 000$
9 дней назад
Sr. Site Reliability Engineer
160 000 - 180 000$
12 дней назад
Site Reliability Engineer - Vice President (Kubernetes)
130 000 - 160 000$
10 дней назад
Site Reliability Engineer (Kubernetes)
180 000 - 220 000$
13 дней назад
Senior Staff Site Reliability Engineer
232 338 - 290 422$
Phantom
12 дней назад
Staff Software Engineer (SRE) (Crypto)
200 000 - 250 000$