Назад
4 часа назад

ML Infrastructure Engineer (AI)

180 000 - 440 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
middle/senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

ML Infrastructure Engineer (AI/GPU Infrastructure): Designing and scaling GPU compute infrastructure and training frameworks to power recommendations on X with an accent on high-performance ML platforms and data pipelines. Focus on optimizing distributed systems, ensuring scalability of large-scale ML systems, and productionizing complex models.

Location: Palo Alto, California, United States (Onsite)

Salary: $180,000 – $440,000 USD

Company

xAI is a small, highly motivated team focused on creating AI systems that understand the universe to aid humanity in its pursuit of knowledge.

What you will do

  • Design, build, and scale GPU compute infrastructure, training frameworks, and experimentation tools.
  • Develop data pipelines and integrate large-scale data, training, and inference systems.
  • Collaborate with ML teams to productionize models and ensure seamless integration across the stack.
  • Ensure the scalability, reliability, and efficiency of large-scale machine learning systems.
  • Solve complex technical problems independently across the full stack.
  • Mentor junior engineers and contribute to the overall growth of the engineering team.

Requirements

  • Degree in computer science, machine learning, or a quantitative discipline (or equivalent experience).
  • 2+ years of industry experience with high-traffic production environments, distributed systems, or GPU infrastructure.
  • 2+ years of experience with ML platforms, training infrastructure, or collaboration with modeling engineers.
  • Strong proficiency in Python and experience with compiled languages like C++ or Rust.
  • Must be based in or able to work from Palo Alto, California.

Nice to have

  • Deep familiarity with modern ML frameworks such as JAX or PyTorch.
  • Low-level understanding of compute systems, NVIDIA drivers, CUDA toolkits, and networking.
  • Experience with Linux systems and orchestration tools.
  • Experience with job schedulers (Slurm) or configuration management (Puppet/Ansible).

Culture & Benefits

  • Flat organizational structure emphasizing initiative and hands-on contribution.
  • Competitive compensation including base salary and equity.
  • Comprehensive medical, vision, and dental coverage.
  • Access to a 401(k) retirement plan and life insurance.
  • Short and long-term disability insurance and various other perks.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →