Назад
Company hidden
24 часа назад

AI Platform Support Engineer (AI)

115 000 - 140 000$
Формат работы
hybrid
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Platform Support Engineer (ML infrastructure): Supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms with an accent on distributed systems troubleshooting, customer-facing technical guidance, and platform reliability. Focus on diagnosing Kubernetes scheduling, GPU orchestration, PyTorch, networking, storage, and inference performance issues while building automation, observability, documentation, and operational improvements.

Location: Hybrid in Seattle, San Francisco, or New York, with at least 2 days per week in the office; Monday–Friday, 8:00 AM–5:00 PM PST. Occasional team and company offsites are required.

Annual base salary: $115,000–$140,000 USD, plus discretionary bonus, equity, and benefits.

Company

hirify.global builds an end-to-end platform for developing, training, and deploying AI systems, combining developer-focused software with large-scale AI compute.

What you will do

  • Partner directly with customer ML engineering teams running production training and inference workloads.
  • Diagnose distributed systems and ML infrastructure issues involving Kubernetes, GPU allocation, networking, storage, and inference serving.
  • Troubleshoot PyTorch, CUDA, NCCL, containerized workloads, and multi-node GPU systems.
  • Analyze logs, metrics, traces, and system behavior to identify root causes and performance bottlenecks.
  • Drive reliability improvements through post-incident reviews, observability, documentation, runbooks, and automation.
  • Collaborate with infrastructure, networking, and platform engineering teams to improve customer troubleshooting workflows.

Requirements

  • Strong software engineering and systems troubleshooting experience.
  • Experience with Kubernetes, containerized environments, Linux, cloud infrastructure, and distributed systems.
  • Knowledge of Linux networking, storage, process management, performance tuning, and observability tools such as Prometheus, Grafana, or OpenTelemetry.
  • Hands-on experience operating machine learning workloads and troubleshooting distributed ML systems using tools such as PyTorch, CUDA, or NCCL.
  • Experience with GPU infrastructure, orchestration, and ML infrastructure reliability, performance, or scaling issues.
  • Visa sponsorship is not available for this role.

Nice to have

  • Experience with large-scale model training, distributed inference, Ray, Kubeflow, Slurm, or similar scheduling platforms.
  • Experience with InfiniBand, RDMA, high-performance networking, bare-metal infrastructure, or ML storage systems.
  • Background in AI infrastructure, cloud, MLOps, developer tooling, or platform engineering.
  • Experience writing Python automation, tooling, or scripts.

Culture & Benefits

  • Health, dental, and vision coverage for employees and eligible dependents.
  • RSUs, U.S. 401(k) matching, unlimited PTO, company holidays, and a two-week winter break.
  • Paid parental and family leave, wellness and work-from-home stipends, and an annual learning allowance.
  • Four weeks of paid sabbatical leave after four years of service.
  • Flexible schedules, hybrid work, and complimentary meals at office hubs.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →