Назад
Company hidden
3 дня назад

Senior Machine Learning Engineer (ML Training Infrastructure)

170 000 - 240 800$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Machine Learning Engineer (ML Training Infrastructure) (Python/PyTorch): Designing and building scalable, reliable, high-performance infrastructure for distributed ML model training with an accent on observability, debuggability, and heterogeneous hardware utilization. Focus on optimizing training performance, scaling large foundational models, and integrating advanced capabilities such as FSDP and pipeline parallelism.

Location: Sunnyvale, California, United States; hybrid attendance at least 3 times per week. Travel required up to 25%; relocation benefits may be available.

Salary: $170,000–$240,800 per year, plus potential incentive compensation.

Company

hirify.global develops automotive technologies focused on safer, cleaner, and more equitable transportation.

What you will do

  • Design and develop scalable, reliable, high-performance ML training frameworks.
  • Improve system observability, debuggability, operational excellence, and user experience.
  • Analyze and optimize model-training performance across distributed workflows and heterogeneous hardware.
  • Maximize resource utilization and reduce infrastructure costs.
  • Collaborate with machine learning engineers, research scientists, and cross-functional partners to integrate platform features and technologies.

Requirements

  • Bachelor’s degree or higher in Computer Science or an equivalent field, or equivalent relevant experience.
  • 5+ years of professional software engineering experience.
  • 1+ years of specialized experience in AI/ML infrastructure and distributed training for large models.
  • Strong Python programming skills and proficiency with PyTorch, TensorFlow, or a similar framework.
  • Experience with distributed computing, GPU computing, and cloud environments such as AWS, GCP, or Azure.
  • Ability to report to the Sunnyvale location at least 3 times per week and willingness to travel as needed.

Nice to have

  • Extensive experience with PyTorch 2.x+ and distributed training frameworks.
  • Experience designing training frameworks supporting FSDP, pipeline parallelism, and other scalable approaches for foundational models.
  • Experience profiling, debugging, and optimizing training and data-loading performance.
  • Strong communication skills for consensus building, risk communication, and constructive feedback.

Culture & Benefits

  • Medical, dental, and vision benefits.
  • Health Savings Account and Flexible Spending Accounts.
  • Retirement savings plan, life insurance, and sickness and accident benefits.
  • Paid vacation and holidays, tuition assistance, and an employee assistance program.
  • GM vehicle discounts and potential performance-based incentive pay.

Hiring process

  • Applicants may complete role-related assessments and pre-employment screening where applicable.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →