Назад
Company hidden
4 дня назад

Research Platform Engineer (ML)

200 000 - 300 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Research Platform Engineer (ML): Developing a multi-tenant research compute platform for dynamically orchestrating large-scale machine learning workloads across hybrid GPU and CPU infrastructure with an accent on HPC scheduling, Kubernetes-native engineering, and multi-cloud portability. Focus on building fault-tolerant research pipelines, optimizing high-performance data paths, and delivering observability and cost attribution for distributed workloads.

Location: New York, United States; hybrid working opportunities available

Annual base salary: $200,000–$300,000, plus eligibility for a discretionary bonus

Company

Quantitative trading firm developing high-performance electronic trading infrastructure and supporting independent trading teams with market access, data, compute, research infrastructure, risk management, compliance, and business services.

What you will do

  • Develop a multi-tenant research compute platform for large-scale machine learning workloads across GPU and CPU infrastructure.
  • Design platform abstractions, APIs, scheduler wrappers, and a durable job control plane for simulations, distributed training, and research pipelines.
  • Build multi-tenant scheduling and isolation with HPC-grade batch scheduling, topology awareness, fair-share allocation, and cloud-native flexibility.
  • Enable portable execution across on-premise datacenters and public cloud providers, and integrate workflow orchestration engines for feature, training, and backtesting pipelines.
  • Implement failure detection, retries, checkpointing, high-performance data paths, and observability for distributed workloads.
  • Provide compute utilization, queue pressure, and GPU/CPU cost attribution visibility for portfolio managers and senior management.

Requirements

  • Deep experience administering HPC job schedulers such as Slurm, including gang scheduling, fair-share priority trees, topology-aware allocation, and containerized execution.
  • Advanced understanding of Kubernetes architecture, CRDs, HPC-focused operators, admission controllers, and GPU-native schedulers.
  • Experience with distributed computing and deep learning frameworks such as Ray, PyTorch, and JAX, including multi-node scaling.
  • Hands-on experience evaluating, architecting, and operating job graph orchestration frameworks.
  • Experience designing vendor-agnostic infrastructure, cloud-bursting strategies, and compute execution across multiple datacenters and public clouds.
  • Knowledge of high-performance storage, high-speed network fabrics, cluster-wide telemetry, and cost attribution for multi-tenant environments.

Culture & Benefits

  • Collaborative, results-oriented environment with no unnecessary hierarchy or ego.
  • Generous paid time off and regional savings and financial wellness plans.
  • Hybrid working opportunities.
  • Daily breakfast, lunch, and snacks, plus in-office wellness experiences and selected wellness expense reimbursement.
  • Company-sponsored sports teams, fitness events, volunteer opportunities, charitable giving, and social events.
  • Workshops and continuous learning opportunities.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →