24 часа назад
AI Platform Support Engineer (AI)
115 000 - 140 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AI Platform Support Engineer (ML infrastructure): Supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms with an accent on distributed systems troubleshooting, customer-facing technical guidance, and platform reliability. Focus on diagnosing Kubernetes scheduling, GPU orchestration, PyTorch, networking, storage, and inference performance issues while building automation, observability, documentation, and operational improvements.
Location: Hybrid in Seattle, San Francisco, or New York, with at least 2 days per week in the office; Monday–Friday, 8:00 AM–5:00 PM PST. Occasional team and company offsites are required.
Annual base salary: $115,000–$140,000 USD, plus discretionary bonus, equity, and benefits.
Company
builds an end-to-end platform for developing, training, and deploying AI systems, combining developer-focused software with large-scale AI compute.
What you will do
- Partner directly with customer ML engineering teams running production training and inference workloads.
- Diagnose distributed systems and ML infrastructure issues involving Kubernetes, GPU allocation, networking, storage, and inference serving.
- Troubleshoot PyTorch, CUDA, NCCL, containerized workloads, and multi-node GPU systems.
- Analyze logs, metrics, traces, and system behavior to identify root causes and performance bottlenecks.
- Drive reliability improvements through post-incident reviews, observability, documentation, runbooks, and automation.
- Collaborate with infrastructure, networking, and platform engineering teams to improve customer troubleshooting workflows.
Requirements
- Strong software engineering and systems troubleshooting experience.
- Experience with Kubernetes, containerized environments, Linux, cloud infrastructure, and distributed systems.
- Knowledge of Linux networking, storage, process management, performance tuning, and observability tools such as Prometheus, Grafana, or OpenTelemetry.
- Hands-on experience operating machine learning workloads and troubleshooting distributed ML systems using tools such as PyTorch, CUDA, or NCCL.
- Experience with GPU infrastructure, orchestration, and ML infrastructure reliability, performance, or scaling issues.
- Visa sponsorship is not available for this role.
Nice to have
- Experience with large-scale model training, distributed inference, Ray, Kubeflow, Slurm, or similar scheduling platforms.
- Experience with InfiniBand, RDMA, high-performance networking, bare-metal infrastructure, or ML storage systems.
- Background in AI infrastructure, cloud, MLOps, developer tooling, or platform engineering.
- Experience writing Python automation, tooling, or scripts.
Culture & Benefits
- Health, dental, and vision coverage for employees and eligible dependents.
- RSUs, U.S. 401(k) matching, unlimited PTO, company holidays, and a two-week winter break.
- Paid parental and family leave, wellness and work-from-home stipends, and an annual learning allowance.
- Four weeks of paid sabbatical leave after four years of service.
- Flexible schedules, hybrid work, and complimentary meals at office hubs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Customer Engineer (ML/AI)
170 000 - 199 000$
Nscale
3 дня назад
Senior AI Product Engineer
180 000 - 260 000$
CoreWeave
23 часа назад
Account Solutions Architect (AI)
182 000 - 242 000$
3 дня назад
Software Engineer, ML Infrastructure Platform (AI)
160 360 - 240 540$
3 дня назад
Software Engineer (AI)
130 000 - 170 000$
3 дня назад
Software Engineer, AI Platform (AI)
145 000 - 170 000$