4 дня назад
Research Platform Engineer (ML)
200 000 - 300 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Research Platform Engineer (ML): Developing a multi-tenant research compute platform for dynamically orchestrating large-scale machine learning workloads across hybrid GPU and CPU infrastructure with an accent on HPC scheduling, Kubernetes-native engineering, and multi-cloud portability. Focus on building fault-tolerant research pipelines, optimizing high-performance data paths, and delivering observability and cost attribution for distributed workloads.
Location: New York, United States; hybrid working opportunities available
Annual base salary: $200,000–$300,000, plus eligibility for a discretionary bonus
Company
Quantitative trading firm developing high-performance electronic trading infrastructure and supporting independent trading teams with market access, data, compute, research infrastructure, risk management, compliance, and business services.
What you will do
- Develop a multi-tenant research compute platform for large-scale machine learning workloads across GPU and CPU infrastructure.
- Design platform abstractions, APIs, scheduler wrappers, and a durable job control plane for simulations, distributed training, and research pipelines.
- Build multi-tenant scheduling and isolation with HPC-grade batch scheduling, topology awareness, fair-share allocation, and cloud-native flexibility.
- Enable portable execution across on-premise datacenters and public cloud providers, and integrate workflow orchestration engines for feature, training, and backtesting pipelines.
- Implement failure detection, retries, checkpointing, high-performance data paths, and observability for distributed workloads.
- Provide compute utilization, queue pressure, and GPU/CPU cost attribution visibility for portfolio managers and senior management.
Requirements
- Deep experience administering HPC job schedulers such as Slurm, including gang scheduling, fair-share priority trees, topology-aware allocation, and containerized execution.
- Advanced understanding of Kubernetes architecture, CRDs, HPC-focused operators, admission controllers, and GPU-native schedulers.
- Experience with distributed computing and deep learning frameworks such as Ray, PyTorch, and JAX, including multi-node scaling.
- Hands-on experience evaluating, architecting, and operating job graph orchestration frameworks.
- Experience designing vendor-agnostic infrastructure, cloud-bursting strategies, and compute execution across multiple datacenters and public clouds.
- Knowledge of high-performance storage, high-speed network fabrics, cluster-wide telemetry, and cost attribution for multi-tenant environments.
Culture & Benefits
- Collaborative, results-oriented environment with no unnecessary hierarchy or ego.
- Generous paid time off and regional savings and financial wellness plans.
- Hybrid working opportunities.
- Daily breakfast, lunch, and snacks, plus in-office wellness experiences and selected wellness expense reimbursement.
- Company-sponsored sports teams, fitness events, volunteer opportunities, charitable giving, and social events.
- Workshops and continuous learning opportunities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
ML Platform Engineer (AI)
100 000 - 160 000$
4 дня назад
AI Platform Engineer, Staff
172 500 - 255 875$
5 дней назад
Senior AI/ML Platform Engineer
148 500 - 221 000$
10 дней назад
AI Platform Engineer
130 000 - 180 000$
5 дней назад
Senior ML Infrastructure & MLOps Engineer (Core Platform)
126 900 - 185 100$
Baseten
8 дней назад
Software Engineer (AI Training Infrastructure)
165 000 - 330 000$