Назад
Company hidden
5 часов назад

Senior Site Reliability Engineer (AI Infrastructure)

215 000 - 275 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (AI Infrastructure) (Kubernetes/Go/Python): Designing and scaling control-plane and data-plane infrastructure for distributed AI workloads with an accent on Kubernetes, cloud-native systems, scheduling, and reliability. Focus on optimizing Ray cluster orchestration, integrating heterogeneous accelerators, and solving complex observability, security, and performance challenges across cloud and on-prem environments.

Location: Hybrid in San Francisco or Palo Alto, United States

Salary: $215,000–$275,000 annual base salary, plus equity and benefits

Company

hirify.global develops a cloud platform based on the open-source Ray project to help developers and data scientists scale distributed machine learning applications.

What you will do

  • Design, build, and scale services for orchestrating Ray clusters across cloud and on-premises environments.
  • Optimize control-plane components for large-scale distributed AI and machine learning workloads.
  • Build intelligent scheduling and resource management systems for heterogeneous compute clusters.
  • Improve the reliability, performance, scalability, and observability of managed Ray workloads.
  • Develop accelerator integrations for GPUs and TPUs, as well as container image management and dependency resolution.
  • Participate in architecture discussions, code reviews, on-call support, and infrastructure troubleshooting with customer-facing teams.

Requirements

  • Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
  • 3+ years of experience writing high-quality production code.
  • Hands-on experience building and maintaining highly available, scalable, and performant distributed systems.
  • Expertise with cloud-native technologies such as AWS, Azure, or GCP, and Kubernetes-based deployments.
  • Strong understanding of networking, security, and authentication mechanisms in cloud environments.
  • Proficiency in Go and Python, with knowledge of Linux kernel foundations, file systems, and containers.

Nice to have

  • Experience with observability stacks such as Prometheus and Grafana.
  • Experience contributing to open-source Ray or working with distributed systems and machine learning infrastructure.

Culture & Benefits

  • Stock options and participation in hirify.global's equity program.
  • Healthcare premiums covered at 95%.
  • 401(k) retirement plan, wellness and education stipend, and paid parental leave.
  • Fertility benefits, paid time off, commute reimbursement, and free office lunches.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →