Назад
Company hidden
1 день назад

Member of Technical Staff, Pre-Training Infrastructure (AI)

142 800 - 274 800$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff, Pre-Training Infrastructure (AI): Building and optimizing distributed training infrastructure for frontier-scale AI models across massive GPU clusters, high-throughput storage, and supercomputing systems with an accent on performance profiling, collective communication, and accelerator optimization. Focus on implementing distributed training in Python and C++, optimizing NCCL across NVLink and InfiniBand topologies, and ensuring the reliability of thousands of GPUs.

Location: Mountain View, United States. Employees within the applicable distance of a designated hirify.global office are expected to work onsite at least four days per week.

Salary: USD $142,800–$274,800 base pay per year across the U.S.; USD $188,000–$304,200 per year in the San Francisco Bay Area and New York City metropolitan area.

Company

hirify.global AI develops Copilot, Bing, Edge, generative AI research, and consumer AI products.

What you will do

  • Design, implement, test, and optimize distributed training infrastructure in Python and C++ for large-scale GPU clusters.
  • Profile, benchmark, and debug compute, memory, networking, and storage bottlenecks.
  • Optimize collective communication libraries such as NCCL for NVLink and InfiniBand topologies.
  • Collaborate with hardware teams on next-generation NVIDIA, AMD, and other accelerators.
  • Develop data-driven pretraining compute roadmaps and contribute to AI models powering Copilot and other products.
  • Drive architectural improvements and deliver infrastructure capabilities to users iteratively.

Requirements

  • Bachelor's degree in Computer Science or a related technical field and 6+ years of technical engineering experience, or equivalent experience.
  • Professional coding experience in languages including C, C++, C#, Java, JavaScript, or Python.
  • Experience with distributed computing and large-scale systems.
  • Experience profiling, benchmarking, and optimizing performance-critical systems.
  • Ability to work from the designated Mountain View office at least four days per week when within the applicable distance.

Nice to have

  • Master's degree and 8+ years of experience, or bachelor's degree and 12+ years of experience.
  • GPU programming experience with CUDA and NCCL, plus PyTorch.
  • Experience building infrastructure for large-scale machine learning or generative AI workloads.
  • Experience with InfiniBand, NVLink, storage systems, or distributed training parallelism.
  • Experience leading technical projects and supporting architectural decisions with data.

Culture & Benefits

  • Startup-like hirify.global AI organization focused on safety, responsibility, and human values.
  • Collaboration across engineering, research, hardware, and product teams.
  • Potential eligibility for benefits and additional compensation.
  • Inclusive workplace grounded in respect, integrity, accountability, and a growth mindset.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →