5 часов назад
Staff Site Reliability Engineer (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Site Reliability Engineer (Kubernetes): Designing, deploying, and operating reliable infrastructure and applications across hybrid cloud and on-premises environments with an accent on infrastructure as code, distributed systems, and regulated operations. Focus on building immutable platforms, improving observability and incident response, and coordinating complex technical initiatives across engineering teams.
Location: Remote, but must reside near Miami, Florida or Austin, Texas. Occasional travel to data center sites may be required. General availability around U.S. time zones and participation in an on-call rotation are expected.
Company
provides high-performance computing infrastructure for AI, HPC, and digital asset mining workloads across data center campuses in North America.
What you will do
- Lead complex technical initiatives from problem definition and design through implementation, rollout, and operation.
- Design, implement, and operate reliable systems across hybrid cloud and on-premises environments.
- Build and deploy infrastructure and applications using automation and infrastructure as code.
- Implement secure, immutable infrastructure with Terraform, Kubernetes, Helm, Ansible, and related tooling.
- Improve observability, monitoring, alerting, and incident response practices.
- Coordinate work across engineering teams, influence system design, and mentor engineers.
Requirements
- Bachelor’s degree in Computer Science or a related field, 7+ years of experience, or equivalent demonstrated impact in SRE, DevOps, or Infrastructure Engineering.
- Broad experience with infrastructure, distributed systems, networking, service communication, production failure modes, scalability, and reliability.
- Experience in regulated, compliant, or change-controlled environments.
- Experience with hybrid environments, including AWS and required on-premises infrastructure.
- Strong experience with Terraform, Helm, Kustomize, Ansible, Kubernetes, virtualization, and configuration management.
- Experience with observability platforms such as Datadog and build and release systems such as GitHub Actions, Makefiles, and Python tooling.
Culture & Benefits
- Remote work with required proximity to Miami or Austin.
- Monday–Friday schedule, generally from 8:00 a.m. to 5:00 p.m.
- Participation in an on-call rotation currently scheduled for one week every five weeks.
- Work may involve professional office environments and occasional data center conditions such as noise and construction.
- Collaborative, change-controlled environment with opportunities to influence technical standards and mentor engineers.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
12 часов назад
DevOps/SRE Engineer
Okta
1 день назад
Staff Site Reliability Engineer, Federal (TS/SCI)
174 000 - 238 000$
6 дней назад
Senior Site Reliability Engineer (Cloud-Native Infrastructure)
142 800 - 178 500$
Okta
2 дня назад
Staff TDI Site Reliability Engineer (Cloud Infrastructure)
174 000 - 239 000$
3 дня назад
Platform Site Reliability Engineer (SRE)
Okta
2 дня назад
Staff TDI Site Reliability Engineer (AWS)
174 000 - 239 000$