Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (AI): Building a GPU management and scheduling platform for nearly 30 models across heterogeneous hardware with an accent on utilization metrics, admission control, autoscaling, and reliable inference operations. Focus on developing Python and Go orchestration systems, automating cloud infrastructure, and operating secure, fault-tolerant production platforms for healthcare AI.
Location: Menlo Park, California, United States; on-site
Company
Hippocratic AI builds healthcare-focused, safety-oriented AI systems and production-grade language model infrastructure.
What you will do
- Design and build a GPU management and scheduling platform for approximately 30 models running on heterogeneous hardware.
- Develop metrics pipelines that interpret GPU utilization and load data to drive admission control, load shedding, and autoscaling decisions.
- Build cloud orchestration systems and operators in Python and Go for managing model fleets.
- Architect and operate scalable, fault-tolerant, secure production systems across AWS, GCP, or Azure.
- Develop Terraform-based infrastructure automation, CI/CD pipelines, monitoring, logging, and alerting.
- Partner with engineers and research scientists on complex infrastructure and operational issues, while mentoring other engineers.
Requirements
- 10+ years of professional experience across site reliability, DevOps, and software engineering.
- Computer Science degree from a top CS program.
- Strong software engineering experience building orchestration and scheduling systems with Python and/or Go.
- Experience designing metric-driven control systems such as autoscaling, load shedding, or admission control.
- Production experience with AWS, GCP, or Azure; infrastructure automation and CI/CD using Terraform, GitLab CI/CD, or similar tools.
- Strong knowledge of Docker, Kubernetes, monitoring and logging platforms, secrets management, security, problem-solving, and communication.
Nice to have
- Experience managing GPU fleets, heterogeneous accelerators, or HPC environments.
- Familiarity with ML inference serving and model deployment tools such as Triton, KServe, or Ray Serve.
- Experience with Kubernetes autoscaling internals, custom metrics, or custom controllers.
- Experience implementing HIPAA and SOC 2 compliance.
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
Culture & Benefits
- Work on healthcare AI systems designed with safety and clinical reliability as core priorities.
- Collaborate with physicians, hospital leaders, AI researchers, and engineers from leading technology and healthcare institutions.
- Join a company backed by major healthcare and AI investors.
- Equal opportunity employer committed to an inclusive workplace and accommodations during the hiring process.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →