Назад
Company hidden
7 дней назад

Machine Learning Engineer, Performance Tooling (AI)

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
UK/US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Machine Learning Engineer, Performance Tooling (AI): Building self-service tools that measure, monitor, predict, and advise on AI model performance across cloud and embedded hardware with an accent on profiling, efficiency analysis, and performance prediction. Focus on modeling theoretical versus achieved performance, identifying layer- and operator-level bottlenecks, and forecasting latency, memory, utilization, and compute cost before model changes are run.

Location: London, United Kingdom or Sunnyvale, California, USA

Company

hirify.global develops AI performance tooling for model training and inference across cloud and embedded hardware.

What you will do

  • Design and build reliable, self-service performance tools for models, hardware targets, and development workflows.
  • Define measurement standards and performance methodologies used across AI development.
  • Model theoretical platform peak performance, compare it with achieved results, and identify efficiency losses at layer and operator level.
  • Predict latency, memory, utilization, and compute cost before model or recipe changes consume compute resources.
  • Own monitoring and regression alerting across model builds and training runs.
  • Collaborate with model, compiler, runtime, platform, hardware, training, and infrastructure engineers to establish performance targets and prioritize improvements.

Requirements

  • Deep hands-on performance engineering experience in complex systems, including profiling, roofline analysis, latency and throughput optimization, and root-cause analysis.
  • Experience owning a tool or service end to end, from design and delivery through adoption by other teams.
  • Strong Python skills and experience profiling and instrumenting large production codebases.
  • Hands-on experience developing deep learning models with PyTorch.
  • Ability to analyze noisy measurements and turn them into defensible conclusions.
  • Strong quantitative communication and judgment for defining measurable performance questions and prioritizing impactful work.

Culture & Benefits

  • Cross-functional collaboration across model, compiler, runtime, platform, hardware, training, and infrastructure teams.
  • Work focused on making AI training and inference faster, more efficient, and more predictable.
  • Performance decisions are supported by profiling data and quantitative analysis.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →