Назад
Company hidden
обновлено 5 дней назад

Software Engineer – ML Platform

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer – ML Platform (Kubernetes/Ray): Building and scaling an ML platform for autonomous-driving training and data processing with an accent on workflow orchestration, distributed execution, and multi-tenant resource governance. Focus on optimizing training throughput, debugging complex production workloads, and developing reliable platform services across GPU, CPU, memory, storage, and networking.

Location: Austin, Texas, United States; onsite. Candidates must be authorized to work in the U.S. Remote work is not available, and relocation and visa sponsorship are not offered.

Company

hirify.global develops autonomous-driving technology and the infrastructure required to train and operate machine-learning systems at scale.

What you will do

  • Build and scale the ML compute platform on Kubernetes using Argo Workflows for training, evaluation, and data-processing orchestration.
  • Design platform capabilities including a Ray-based SDK for distributed execution and multi-tenant resource governance for GPU, CPU, memory, and I/O.
  • Optimize training throughput and platform efficiency through improved data access, caching, and reduced storage, network, and resource bottlenecks.
  • Work with ML engineers to debug complex workloads, perform root-cause analysis, and convert recurring issues into platform-level fixes.
  • Evaluate, integrate, and extend open-source tooling across the Kubernetes ecosystem.

Requirements

  • Strong proficiency in Python or Go.
  • Experience designing and building scalable, maintainable systems and services.
  • Experience operating production services end to end, including APIs, reliability practices, and observability.
  • Deep knowledge of Kubernetes scheduling, resource management, controllers, and pod lifecycle behavior.
  • Solid Linux and systems-debugging skills covering performance, networking, and storage/I/O.
  • Ability to troubleshoot production issues across logs, metrics, and traces and drive them to resolution.

Nice to have

  • C++ experience.
  • Experience with Argo Workflows, Ray, MLflow, or comparable distributed ML tooling.
  • Experience building or operating large-scale ML training systems, including GPU scheduling, distributed training, and training-data pipelines.
  • Experience optimizing resource usage and performance in distributed environments.

Culture & Benefits

  • Work directly with ML teams across the company to improve experimentation and model training at scale.
  • Build production-grade tooling for the full machine-learning model lifecycle.
  • Reasonable accommodations are available for qualified applicants and employees with disabilities.
  • Equal-opportunity employment practices apply.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →