обновлено 8 дней назад
Principal Production Engineer (AI)
164 500 - 235 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Production Engineer (AI/Cloud Infrastructure): Designing and operating highly available, scalable infrastructure across AWS, GCP, and bare-metal environments with an accent on automation, observability, and reliability engineering. Focus on building self-healing systems, reducing Mean Time to Mitigate, leading incident response, and scaling globally distributed multi-cloud services.
Location: Remote in California, USA, or hybrid in San Jose, California, with three days per week in the office
Salary: $164,500–$235,000 USD base salary per year, excluding bonus, equity, and benefits
Company
provides the Zero Trust Exchange, a cloud security platform that protects users, devices, and applications from cyberattacks and data loss.
What you will do
- Design and implement highly available, scalable infrastructure across AWS, GCP, and bare-metal environments.
- Write Python and Go code to eliminate manual toil and build automation-first, self-healing systems.
- Develop observability using Prometheus, Grafana, and OpenTelemetry; define SLIs, SLOs, and error budgets.
- Lead incident response as an Incident Commander, maintain response playbooks, and conduct post-incident analyses.
- Partner with engineering teams on operability reviews and improve the reliability and scalability of a globally distributed platform processing more than 200 billion transactions daily.
Requirements
- 10+ years of experience managing reliability, scalability, and availability for large-scale production services.
- Deep programming expertise in Python, Go, or C/C++.
- Strong knowledge of networking protocols, Linux/RHEL systems, and distributed architectures.
- Experience with high-stakes incident management and participation in a 24/7 on-call rotation.
- Experience using ITIL frameworks, incident data, problem management, and technical operability reviews.
- Foundational AI/ML knowledge and experience leveraging, securing, or positioning AI-driven solutions.
Nice to have
- Experience with AWS, Azure, GCP, and Infrastructure-as-Code tools such as Ansible, Terraform, Helm, and Temporal.
- Experience with chaos engineering and large-scale disaster recovery planning.
- Expertise in BGP, GRE, IPSec, HAProxy, DNS at scale, and OS networking internals.
Culture & Benefits
- Ownership, collaboration, trust through outcomes, and a challenge culture with ongoing feedback.
- Health plans, vacation and sick leave, parental leave, retirement options, and education reimbursement.
- In-office perks and an inclusive workplace focused on collaboration and belonging.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
13 дней назад
Site Reliability Engineer (AWS)
120 000 - 185 000$
Nscale
8 дней назад
Operational Data & Observability Engineer
145 000 - 180 000$
10 дней назад
Site Reliability Engineer (Kubernetes)
180 000 - 220 000$
12 дней назад
Site Reliability Engineer - Vice President (Kubernetes)
130 000 - 160 000$
9 дней назад
Sr. Site Reliability Engineer
160 000 - 180 000$
9 дней назад
Sr Software Development Engineer, SRE (US Federal)
163 800 - 245 800$