5 часов назад
Site Reliability Engineer (AI/ML)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI/ML): Designing, deploying, and operating large-scale distributed systems across compute, storage, networking, and AI/ML environments with an accent on Kubernetes, service mesh, infrastructure as code, and observability. Focus on automating infrastructure workflows, troubleshooting complex performance issues, and building resilient systems for model training and data pipelines.
Location: Reston, VA, United States
Company
provides strategic database, cloud, analytics, and AI services for mid-sized and large organizations, partnering with major technology platforms including Google Cloud, AWS, Microsoft, Oracle, SAP, and Snowflake.
What you will do
- Operate and optimize Kubernetes clusters, Istio service mesh, and Linux-based systems.
- Automate infrastructure workflows using Go, Python, and Shell scripting.
- Build monitoring and observability solutions with Prometheus, Grafana, and Loki.
- Troubleshoot complex networking, storage, and system performance issues.
- Partner with AI/ML teams to prepare infrastructure for model training and data pipelines.
- Participate in on-call rotations and postmortem reviews to improve system resilience.
Requirements
- Experience with Google Cloud and Terraform.
- Strong knowledge of microservices, containers, Kubernetes, Docker, and networking.
- Hands-on experience with PKI, service mesh, and Linux systems administration.
- SRE mindset focused on automation, scalability, and reliability.
- Active TS/SCI with DoW and TS/SCI with Full Scope Polygraph security clearances are required.
- Ability to fulfill the requirements for a background check.
Nice to have
- Experience with Golang.
Culture & Benefits
- Competitive total rewards package.
- Training allowance, professional development days, training opportunities, and certification support.
- Equipment for working from home, including a laptop with a choice of operating system.
- Annual budget for personalizing the work environment.
- Annual wellness budget, paid vacation and sick days, and paid time off for volunteering.
Hiring process
- Applicants undergo background-check screening and may request accommodations during the selection process.
- AI tools may assist with application and resume review, while final decisions remain with human reviewers.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Site Reliability Engineer (AI Infrastructure)
175 000 - 265 000$
5 дней назад
Principal Site Reliability Engineer (GCP)
151 000 - 244 200$
6 дней назад
Site Reliability Engineer (Cloud Infrastructure)
122 574 - 259 200$
Okta
5 дней назад
Staff Site Reliability Engineer (Splunk)
194 000 - 267 000$
6 дней назад
Site Reliability Engineering Manager (Kubernetes)
175 000 - 238 000$
6 дней назад
Senior Site Reliability Engineer (AI)
75 000 - 85 000$