Назад
Company hidden
2 дня назад

AI Systems Engineer

90 000 - 100 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Systems Engineer (GPU Infrastructure): Design, build, and operate the platform layer powering large-scale AI training and inference workloads with an accent on GPU clusters, distributed training, scheduling, storage performance, and developer experience. Focus on building reliable and cost-efficient AI infrastructure, optimizing interactions between hardware, kernels, schedulers, and ML frameworks, and operating production systems at scale.

Location: 100% remote within the United States

Salary: $90,000–$100,000 annually

Company

hirify.global is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

What you will do

  • Design, build, and operate the platform layer for large-scale AI training and inference workloads.
  • Operate and optimize GPU clusters, distributed training frameworks, scheduling systems, and high-performance storage.
  • Improve reliability, infrastructure efficiency, and cost control for AI workloads.
  • Build developer experience capabilities for ML engineers and researchers.
  • Apply software engineering practices including testing, CI/CD, and code review.
  • Collaborate with ML engineers, researchers, and cross-functional stakeholders.

Requirements

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • 6+ years of experience in infrastructure, platform, or HPC engineering.
  • Hands-on experience operating GPU clusters or large-scale ML training infrastructure.
  • Strong proficiency in Python and at least one systems language such as Go or C++.
  • Deep understanding of distributed training, accelerator architectures, collective communication, Linux internals, networking, and high-performance storage.
  • Experience with Kubernetes, Slurm, Ray, or similar ML workload scheduling systems, plus a major cloud provider’s ML infrastructure offerings.

Nice to have

  • Experience operating InfiniBand or RDMA networking at scale.
  • Contributions to open-source ML infrastructure projects.
  • Familiarity with custom orchestrators or research-grade training stacks.
  • Exposure to frontier model training operations.
  • Experience with FinOps for AI workloads.

Culture & Benefits

  • Full-time direct W-2 employment.
  • Career growth opportunities within an established organization.
  • Equal employment opportunity for employees and applicants.

Hiring process

  • Submit a resume by email for consideration.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →