Назад
Company hidden
2 дня назад

Senior Machine Learning Engineer, ML Infrastructure- Online (Machine Learning)

165 600 - 273 400$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Machine Learning Engineer, ML Infrastructure- Online (Machine Learning): Building and operating large-scale online inference infrastructure for production machine learning models with an accent on low-latency serving, distributed systems, and observability. Focus on optimizing GPU and CPU utilization, enabling safe model rollouts, and designing reliable packaging, validation, monitoring, and deployment workflows.

Location: Remote, Washington, USA

Base salary: $165,600–$273,400 annually, depending on geographic zone and qualifications.

Company

hirify.global develops a leading game engine and 3D development platform used across games, XR, web, automotive, manufacturing, and healthcare.

What you will do

  • Design and operate large-scale online inference infrastructure for production machine learning models.
  • Build infrastructure for distributed training workflows using PyTorch, Ray Data, and Ray Train.
  • Integrate machine learning pipelines with workflow orchestration systems such as Flyte and Airflow.
  • Optimize inference through dynamic batching, model compilation, GPU/CPU utilization improvements, kernel fusion, request scheduling, and runtime tuning.
  • Improve observability, reliability, reproducibility, model packaging, artifact validation, compatibility testing, and deployment automation.
  • Lead architectural improvements and partner with ML engineers, platform teams, and product stakeholders on safe, scalable, and cost-efficient model iteration.

Requirements

  • Experience building and operating production-grade online ML inference systems and model-serving frameworks such as NVIDIA Triton Inference Server, TorchServe, Ray Serve, or TensorFlow Serving.
  • Strong experience with distributed systems, Kubernetes, autoscaling, service reliability, and production observability.
  • Strong Python programming skills and practical experience with production ML systems and high-scale services.
  • Experience with PyTorch, model packaging, validation, deployment workflows, safe rollouts, canary testing, A/B experimentation, and automated rollback.
  • Ability to reason about latency, throughput, reliability, scalability, and cost tradeoffs and influence architectural decisions across teams.
  • Professional English proficiency is required for frequent communication with colleagues and partners worldwide.

Culture & Benefits

  • Health, life, and disability insurance options, with eligibility varying by country and employment status.
  • Employee stock ownership and competitive retirement or pension plans.
  • Generous vacation and personal days, plus parental leave and family-care programs.
  • Mental health and wellbeing support, employee resource groups, and a global employee assistance program.
  • Training and development programs, volunteering support, and donation matching.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →