5 дней назад
Staff Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Site Reliability Engineer (AI): Designing, deploying, and operating reliable infrastructure and applications across hybrid AWS and on-premises environments with an accent on Kubernetes, infrastructure as code, observability, and regulated operations. Focus on leading complex technical initiatives, solving distributed-systems reliability and scaling challenges, and coordinating hands-on delivery across engineering teams.
Location: Must reside near Miami, FL or Austin, TX; generally available around U.S. time zones. The role includes occasional travel to data center sites and an on-call rotation of one week every five weeks.
Company
provides high-performance compute infrastructure for AI, HPC, and digital asset mining across data center campuses in North America.
What you will do
- Lead complex technical initiatives from problem definition and system design through implementation, rollout, and operation.
- Design, implement, and operate reliable, scalable systems across hybrid cloud and on-premises environments.
- Build and deploy infrastructure and applications using automation, infrastructure as code, Kubernetes, Helm, Ansible, and related tooling.
- Improve observability, monitoring, alerting, and incident response practices.
- Coordinate execution across engineering teams, delegate effectively, and influence application architecture and technical decisions.
- Establish reliability, security, and operational best practices while mentoring engineers.
Requirements
- Bachelor's degree in Computer Science or a related field, 7+ years of experience, or equivalent demonstrated impact in SRE, DevOps, or infrastructure engineering.
- Broad experience with infrastructure, distributed systems, networking, service communication, production failure modes, scaling, and reliability.
- Experience in regulated, compliant, or change-controlled environments.
- Experience with hybrid environments, including AWS and required on-premises infrastructure.
- Strong experience with infrastructure as code, configuration management, and orchestration tools including Terraform, Helm, Kustomize, and Ansible.
- Experience with Kubernetes, virtualization, observability platforms such as Datadog, and build and release systems such as GitHub Actions, Makefiles, and Python tooling.
Culture & Benefits
- Work in an entrepreneurial, collaborative, and results-driven environment.
- Operate infrastructure supporting AI, HPC, and digital asset mining workloads.
- Work is generally scheduled Monday through Friday, 8:00 a.m. to 5:00 p.m.
- The role operates in a professional office environment and may involve data center conditions including loud noise and construction.
- Physical duties may include sitting, standing, walking, using hands, and lifting up to 25 pounds.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
1 час назад
Sr. Site Reliability Engineer (Healthcare)
125 000 - 145 000$
Deimos
7 дней назад
Senior Site Reliability Engineer
4 дня назад
Senior/Lead Site Reliability Engineer – Federal (AI)
159 000 - 230 000$
3 дня назад
Senior Site Reliability Engineer (AI)
6 дней назад
Staff Site Reliability Engineer (AWS/Kubernetes)
140 000 - 155 000CAD
2 дня назад
Senior Site Reliability Engineer (Kubernetes)
125 000 - 145 000$