2 дня назад
Staff Site Reliability Engineer (AI/ML)
112 500 - 187 500$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Site Reliability Engineer (AI/ML): Driving reliability strategy and operating high-volume, highly available cloud-native platforms on GCP and Kubernetes with an accent on observability, Linux performance, capacity planning, security, and AI/ML integration. Focus on designing production-grade AI capabilities, leading high-risk maintenance and incident response, and maintaining 99.999% reliability targets.
Location: Chicago, Illinois; hybrid work with in-person responsibilities at an assigned office at least two days per week
Salary: $112,500–$187,500 annually, with potential eligibility for an annual bonus and other incentives.
Company
provides data and technology solutions that support customers and communities through products and services built around trust and innovation.
What you will do
- Define reliability strategy and contribute to architectural and strategic decisions for major platform components.
- Research, test, implement, and continuously improve systems and engineering tooling.
- Perform capacity planning, load testing, security improvements, and high-impact platform engineering.
- Participate in on-call rotations and lead calm, blameless incident response and problem resolution.
- Plan and lead high-risk maintenance events while minimizing customer impact.
- Raise engineering standards through tooling, procedures, communication, and cross-functional collaboration.
Requirements
- 5+ years of experience in cloud architecture, site reliability engineering, platform engineering, or a related field at enterprise scale.
- Deep hands-on expertise with Google Cloud Platform and Kubernetes for high-volume, highly available workloads.
- Expertise with monitoring, observability, and alerting platforms such as Datadog, Prometheus, Grafana, and PagerDuty.
- Advanced Linux knowledge, including kernel internals, system performance tuning, hardening, and production troubleshooting.
- Hands-on experience integrating AI/ML solutions into cloud-native platforms, including LLM orchestration, vector databases, model serving, and AI observability.
- Ability to work in the hybrid Chicago office arrangement with at least two in-person days per week.
Culture & Benefits
- Medical, dental, and vision coverage with day-one eligibility, plus HSA and FSA options.
- Company-paid life and disability insurance, with optional additional coverage and legal, pet, and travel insurance.
- Adoption assistance, fertility planning coverage, caregiver support, dependent care benefits, and up to 12 weeks of paid parental leave.
- 401(k) with employer match, Employee Stock Purchase Plan, financial wellness resources, and career coaching.
- Tuition reimbursement, flexible or paid time off, up to 12 paid holidays, commuter benefits, employee discounts, and paid volunteer time.
- Wellness resources including therapy, coaching, meditation, and emotional, physical, social, and financial support.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Senior Site Reliability Engineer (Temporal)
180 000 - 200 000$
9 дней назад
Senior Manager, Site Reliability Engineering (AI Ops)
222 000 - 300 500$
Okta
6 дней назад
Staff Site Reliability Engineer (Splunk)
194 000 - 267 000$
Okta
5 дней назад
Staff Site Reliability Engineer, Federal (TS/SCI)
174 000 - 238 000$
Nscale
5 дней назад
Senior Site Reliability Engineer (AI Infrastructure Operations)
170 000 - 265 000$
3 дня назад
Staff Site Reliability Engineer (AI)
252 000 - 308 000$