Назад
4 дня назад

Staff Machine Learning Engineer (ML Efficiency)

230 000 - 322 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Machine Learning Engineer (ML Efficiency): Building infrastructure, tooling, and optimization systems for efficient large-scale ML training and inference with an accent on distributed systems, GPU utilization, performance engineering, and model serving. Focus on designing benchmarking frameworks, optimizing resource scheduling and distributed training, and improving platform scalability, reliability, and cost efficiency.

Location: Remote from the United States or Canada

Base salary: $230,000–$322,000 USD per year, plus potential equity and commission depending on the position offered.

Company

Reddit operates a large online community platform with more than 100,000 active communities and approximately 130 million daily active unique visitors.

What you will do

  • Design and build systems that improve the efficiency of machine learning training and inference workloads.
  • Develop tooling for debugging, profiling, optimization, and monitoring of model performance.
  • Improve GPU and general resource utilization through scheduling, resource management, caching, and workload optimization.
  • Build benchmarking frameworks and performance dashboards for training and serving systems.
  • Optimize distributed training infrastructure, data pipelines, and model serving architectures.
  • Lead cross-functional initiatives and define technical strategy for scalable, reliable, and cost-efficient ML platforms.

Requirements

  • BS, MS, or PhD in Computer Science or a related field.
  • 5+ years of software engineering experience.
  • Strong proficiency in Python.
  • Experience building distributed systems at scale and working with ML infrastructure, training systems, or model serving platforms.
  • Deep understanding of performance engineering, systems optimization, debugging, and profiling.
  • Proficiency in at least one systems language such as Go, C++, Rust, or Java is preferred.

Nice to have

  • Experience with large-scale recommendation, ranking, generative AI, or foundation model systems.
  • Experience with PyTorch Distributed, Ray, TensorFlow, or Spark.
  • Familiarity with GPU architectures and performance analysis tools.
  • Experience optimizing cloud infrastructure costs across large ML workloads.
  • Experience building real-time ML inference applications or internal platforms used by multiple ML teams.

Culture & Benefits

  • Flexible remote-first workforce.
  • Workspace, professional development, caregiving, family planning, and mental health support programs.
  • Private medical and dental coverage, pension or 401(k) programs with employer matching, and income replacement programs.
  • Flexible vacation, paid volunteer time off, and generous paid parental leave.
  • Interviews may be recorded, transcribed, and summarized by AI, with the option to opt out before scheduled interviews.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →