5 часов назад
Senior Site Reliability Engineer (AI Infrastructure)
215 000 - 275 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (AI Infrastructure) (Kubernetes/Go/Python): Designing and scaling control-plane and data-plane infrastructure for distributed AI workloads with an accent on Kubernetes, cloud-native systems, scheduling, and reliability. Focus on optimizing Ray cluster orchestration, integrating heterogeneous accelerators, and solving complex observability, security, and performance challenges across cloud and on-prem environments.
Location: Hybrid in San Francisco or Palo Alto, United States
Salary: $215,000–$275,000 annual base salary, plus equity and benefits
Company
develops a cloud platform based on the open-source Ray project to help developers and data scientists scale distributed machine learning applications.
What you will do
- Design, build, and scale services for orchestrating Ray clusters across cloud and on-premises environments.
- Optimize control-plane components for large-scale distributed AI and machine learning workloads.
- Build intelligent scheduling and resource management systems for heterogeneous compute clusters.
- Improve the reliability, performance, scalability, and observability of managed Ray workloads.
- Develop accelerator integrations for GPUs and TPUs, as well as container image management and dependency resolution.
- Participate in architecture discussions, code reviews, on-call support, and infrastructure troubleshooting with customer-facing teams.
Requirements
- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
- 3+ years of experience writing high-quality production code.
- Hands-on experience building and maintaining highly available, scalable, and performant distributed systems.
- Expertise with cloud-native technologies such as AWS, Azure, or GCP, and Kubernetes-based deployments.
- Strong understanding of networking, security, and authentication mechanisms in cloud environments.
- Proficiency in Go and Python, with knowledge of Linux kernel foundations, file systems, and containers.
Nice to have
- Experience with observability stacks such as Prometheus and Grafana.
- Experience contributing to open-source Ray or working with distributed systems and machine learning infrastructure.
Culture & Benefits
- Stock options and participation in 's equity program.
- Healthcare premiums covered at 95%.
- 401(k) retirement plan, wellness and education stipend, and paid parental leave.
- Fertility benefits, paid time off, commute reimbursement, and free office lunches.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 часа назад
Senior Production Engineer (AI Infrastructure)
170 000 - 205 000$
2 дня назад
Senior Software / Site Reliability Lead Engineer (AI)
142 696 - 158 303$
58 минут назад
Senior Software Engineer, Site Reliability Engineering (AWS)
153 000 - 210 000$
3 часа назад
Production Engineer (AI Infrastructure)
172 000 - 209 000$
5 часов назад
Senior Site Reliability Engineer (SRE)
170 000 - 196 000$
2 часа назад