Назад
Company hidden
5 часов назад

Training Infrastructure Engineer (AI)

210 000 - 320 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Training Infrastructure Engineer (AI): Designing and optimizing scalable infrastructure and distributed pipelines for large-scale LLM and multimodal model training with an accent on multi-GPU systems, data storage, orchestration, and operational reliability. Focus on improving training performance and cost efficiency, troubleshooting distributed computing issues, and collaborating with AI researchers on training methodologies.

Location: Hybrid in San Mateo or New York, United States

Salary: $210K–$320K annually, plus equity

Company

hirify.global provides infrastructure for building, training, and serving specialized AI models across text, image, embedding, audio, and multimodal workloads.

What you will do

  • Design and implement scalable infrastructure for large-scale model training workloads.
  • Develop and maintain distributed training pipelines for LLMs and multimodal models.
  • Optimize training performance across GPUs, nodes, and data centers.
  • Build monitoring, logging, debugging, storage, provisioning, scaling, and orchestration systems for training operations.
  • Collaborate with AI researchers to implement and optimize training methodologies.
  • Analyze system efficiency, scalability, reliability, and cost-effectiveness while troubleshooting complex distributed performance issues.

Requirements

  • Bachelor’s degree in Computer Science, Computer Engineering, or a related field, or equivalent practical experience.
  • 3+ years of experience with distributed systems and ML infrastructure.
  • Experience with PyTorch and distributed training techniques such as data parallelism, model parallelism, and FSDP.
  • Proficiency with AWS, GCP, or Azure cloud platforms.
  • Experience with containerization and orchestration using Kubernetes and Docker.

Nice to have

  • Master’s or PhD in Computer Science or a related field.
  • Experience training large language models or multimodal AI systems.
  • Experience with ML workflow orchestration tools and ML DevOps practices.
  • Background in optimizing high-performance distributed computing systems.
  • Open-source contributions to ML infrastructure or related projects.

Culture & Benefits

  • Work on challenging AI infrastructure problems, including scalable model serving and low-latency inference.
  • Build production technology used by businesses and developers globally.
  • Operate with a high level of ownership in a fast-growing, collaborative environment.
  • Collaborate with experienced engineers and AI researchers.
  • Equity is included in the compensation package.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →