updated 22 days ago
Senior Site Reliability Engineer (AI)
141 800 - 195 000$
Match & Cover letter
Plus required for matching with this vacancy
Job description
Text:
TL;DR
Senior Site Reliability Engineer (AI): Building and operating reliable cloud-based telemetry and observability platforms with an accent on availability, latency, scalability, and resilience. Focus on designing observability systems, automating toil, driving incident response improvements, and maintaining high availability across production services.
Location: Remote within the United States
Base salary: $141,800–$195,000 USD annually, depending on geographic location, knowledge, skills, and experience.
Company
builds telemetry infrastructure and observability software that helps enterprise IT and Security teams manage and analyze real-time data for humans and AI agents.
What you will do
- Improve service delivery and reliability across the full lifecycle of cloud-based services.
- Measure and monitor production systems for availability, latency, and overall system health.
- Investigate errors and instability in production cloud services and drive operational excellence.
- Partner with product and platform teams to improve reliability, resilience, and observability.
- Identify and reduce operational toil through automation and creative engineering.
- Participate in standby, on-call, and off-hours support duties.
Requirements
- Senior-level experience designing, implementing, and operating observability systems for complex cloud platforms.
- Experience with configuration management and infrastructure as code, including Terraform or Ansible.
- Knowledge of AWS or Azure, containers, orchestration technologies, cloud security, and cloud design patterns for scale and resiliency.
- Experience with APM and observability tools such as New Relic, Splunk, CloudWatch, Prometheus, Grafana, Kibana, and Sentry.
- Experience with enterprise-scale continuous delivery, Linux systems engineering, and JavaScript, Node, or TypeScript development.
- Experience with sustainable incident response in a blameless environment and tools such as PagerDuty, FireHydrant, or Blameless.
Culture & Benefits
- Remote-first work environment with a distributed team and a high level of autonomy.
- Health, dental, vision, short-term disability, and life insurance.
- Paid holidays, paid time off, fertility treatment benefits, 401(k), and equity.
- Participation in the Corporate Bonus Program for eligible non-sales roles.
- Blameless collaboration culture focused on quality, inclusion, and continuous improvement.
Be careful: if the employer asks you to log into their system using iCloud/Google, send codes/passwords, or run code/software, don't do it - these are scammers. Always click "Report" or contact support. More in guide →
Similar vacancies
11 days ago
Senior Site Reliability Engineer (Azure)
6 days ago
Site Reliability Engineer (SRE)
9 days ago
Staff Site Reliability Engineer (GCP/Kubernetes)
CloudLinux
2 days ago
Senior Database Reliability Engineer (Dbre)
5 000 - 10 000$
6 days ago
Site Reliability Engineer II (AWS)
100 000 - 110 000$
11 days ago
Senior Site Reliability Engineer (Observability)
160 000 - 200 000$