обновлено 5 дней назад
Software Engineer – ML Platform
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Software Engineer – ML Platform (Kubernetes/Ray): Building and scaling an ML platform for autonomous-driving training and data processing with an accent on workflow orchestration, distributed execution, and multi-tenant resource governance. Focus on optimizing training throughput, debugging complex production workloads, and developing reliable platform services across GPU, CPU, memory, storage, and networking.
Location: Austin, Texas, United States; onsite. Candidates must be authorized to work in the U.S. Remote work is not available, and relocation and visa sponsorship are not offered.
Company
develops autonomous-driving technology and the infrastructure required to train and operate machine-learning systems at scale.
What you will do
- Build and scale the ML compute platform on Kubernetes using Argo Workflows for training, evaluation, and data-processing orchestration.
- Design platform capabilities including a Ray-based SDK for distributed execution and multi-tenant resource governance for GPU, CPU, memory, and I/O.
- Optimize training throughput and platform efficiency through improved data access, caching, and reduced storage, network, and resource bottlenecks.
- Work with ML engineers to debug complex workloads, perform root-cause analysis, and convert recurring issues into platform-level fixes.
- Evaluate, integrate, and extend open-source tooling across the Kubernetes ecosystem.
Requirements
- Strong proficiency in Python or Go.
- Experience designing and building scalable, maintainable systems and services.
- Experience operating production services end to end, including APIs, reliability practices, and observability.
- Deep knowledge of Kubernetes scheduling, resource management, controllers, and pod lifecycle behavior.
- Solid Linux and systems-debugging skills covering performance, networking, and storage/I/O.
- Ability to troubleshoot production issues across logs, metrics, and traces and drive them to resolution.
Nice to have
- C++ experience.
- Experience with Argo Workflows, Ray, MLflow, or comparable distributed ML tooling.
- Experience building or operating large-scale ML training systems, including GPU scheduling, distributed training, and training-data pipelines.
- Experience optimizing resource usage and performance in distributed environments.
Culture & Benefits
- Work directly with ML teams across the company to improve experimentation and model training at scale.
- Build production-grade tooling for the full machine-learning model lifecycle.
- Reasonable accommodations are available for qualified applicants and employees with disabilities.
- Equal-opportunity employment practices apply.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Platform Engineer - AI Engineering (AI)
130 000 - 190 000$
7 дней назад
Senior/Staff Software Engineer, Platform Infrastructure (AI)
180 000 - 280 000$
11 дней назад
Engineering Manager (AI)
210 000 - 300 000$
7 дней назад
Platform Engineer (AWS)
105 600 - 190 450$
7 дней назад
Senior Platform Engineer (DevX)
165 000 - 200 000$
12 дней назад
Software Engineer, Platform Operations
127 000 - 158 700$