1 день назад
Site Reliability Engineer, AI Infrastructure (AI)
129 960 - 246 240$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer, AI Infrastructure (AI): Building and operating globally distributed, fault-tolerant infrastructure for recommendation and search systems with an accent on production ownership, observability, automation, and service reliability. Focus on designing distributed systems, leading infrastructure migrations, managing capacity and SLAs/SLOs, and resolving systemic causes of incidents in a highly secure environment.
Location: Seattle, United States; fully in-person schedule up to 5 days a week
Salary: $129,960–$246,240 annually, with potential additional bonuses, incentives, and restricted stock units.
Company
A technology joint venture focused on data privacy, cybersecurity, national security, and secure U.S. user data and applications.
What you will do
- Partner with engineering and product teams across system design, architecture reviews, deployment, operations, and continuous service refinement.
- Build tools, platforms, and automation that improve reliability, scalability, R&D efficiency, and operational workflows.
- Monitor service health, latency, and key metrics for large-scale, multi-region systems.
- Lead infrastructure migrations and architecture upgrades in a highly restricted compliance environment.
- Manage capacity, resource allocation, stability optimization, error attribution, and SLA/SLO compliance.
- Drive incident management, blameless postmortems, and systemic root-cause resolution.
Requirements
- Bachelor’s degree or equivalent practical experience in computer science, software engineering, or a related technical field.
- At least 1 year of hands-on SRE, DevOps, or systems engineering experience with large-scale, highly reliable systems.
- Strong knowledge of Linux, system performance, networking fundamentals, and Bash or shell scripting.
- Programming experience in at least one of Go, Python, C/C++, or Java.
- Experience designing, troubleshooting, and maintaining complex distributed systems in dynamic production environments.
- Familiarity with CI/CD practices and automated deployment pipelines.
Nice to have
- Experience with Kubernetes, service mesh architectures, cloud platforms, and Infrastructure-as-Code such as Terraform.
- Experience with Prometheus, Grafana, distributed tracing, or other observability tools.
- Experience building self-service tools and automation, or applying LLMs and agentic AI to operational workflows.
Culture & Benefits
- On-site collaboration in a highly secure and strictly isolated infrastructure environment.
- Medical, dental, and vision insurance from day one, plus a 401(k) plan with company match.
- Paid parental leave, disability coverage, life insurance, and wellbeing benefits.
- 10 paid holidays, 10 paid sick days, and 17 days of paid personal time, with increasing accruals by tenure.
- Commitment to inclusive recruitment and reasonable accommodations.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
19 часов назад
Site Reliability Engineer, Tech Infra - USDS (Cloud Infrastructure)
136 800 - 259 200$
9 часов назад
Site Reliability Engineer, Platform Responsibility - USDS (AI/LLM)
129 960 - 246 240$
7 часов назад
Site Reliability Engineer, Edge Services - USDS
136 800 - 359 720$
Databricks
5 дней назад
Sr Technical Solutions Engineering (AI)
130 200 - 178 950$
21 час назад
Platform Engineer (AI)
140 000 - 180 000$
3 дня назад
Sr. DevOps Engineer II (AI)
119 000 - 221 000$