1 день назад
DevOps – AI Factory
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
DevOps – AI Factory (AWS/GCP/Kubernetes): Designing and operating highly available distributed AI production systems with an accent on GPU inference, multicloud infrastructure, and platform reliability. Focus on building AI Factory infrastructure with Kubernetes, Ray or KubeFlow, automating delivery, and troubleshooting complex application, networking, operating system, and data-layer issues.
Location: Hybrid and flexible work environment; specific office location is not stated.
Company
is a global video technology company providing live, on-demand, and real-time video solutions for organizations.
What you will do
- Lead DevOps and platform initiatives from design through implementation and production rollout.
- Design, build, and operate highly available distributed systems across AWS, GCP, and NeoCloud providers.
- Develop and maintain the platform stack with Kubernetes, Docker, Helm, CI/CD, automation, data, caching, and messaging layers.
- Build and operate infrastructure for AI workloads, including GPU inference platforms.
- Modernize and consolidate technology stacks into a unified platform.
- Drive reliability, monitoring, observability, and complex issue resolution across application, infrastructure, networking, operating system, and data layers.
Requirements
- At least 2 years of hands-on experience in DevOps, Platform, or SRE roles operating distributed production systems at scale.
- At least 2 years of experience with cloud inference services, NeoClouds, and AI Factory systems using KubeFlow, Ray, or similar technologies.
- At least 2 years of hands-on experience with GPU inference technologies such as vLLM, SGLang, or Nvidia Triton.
- Deep expertise in multicloud AWS/GCP/NeoCloud environments and Infrastructure as Code with Terraform or similar tools.
- Strong Kubernetes and Helm expertise, including Kubernetes internals, production operations, Docker, networking, Linux, TCP/IP, DNS, and load balancing.
- Strong experience with GitHub Actions, CI/CD, automation, monitoring, observability, Prometheus, and Grafana, plus proficiency with current AI coding tools.
Nice to have
- Familiarity with AI evaluation and AI observability systems such as Opik, OpenEval, or LangSmith.
Culture & Benefits
- Hybrid and flexible work environment.
- Extended private health insurance, including mental health coverage.
- Personal and professional development programs.
- Occasional cross-company long weekends.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Senior/Staff DevOps Engineer (AI Platform Infrastructure)
180 000 - 240 000$
3 дня назад
Senior DevOps Engineer (AI)
130 000 - 196 500$
6 дней назад
Senior AI Platform Engineer (AI)
3 дня назад
DevOps Engineer (AWS)
Ramp
6 дней назад
Production Engineer (AI Infrastructure)
240 000 - 330 000$
6 дней назад