2 дня назад
Software Engineer (SE / Sr SE), Data & ML Platform
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Software Engineer (SE / Sr SE), Data & ML Platform (Kubernetes/GPU): Operating and evolving Kubernetes infrastructure for petabyte-scale data processing, simulation, auto-labeling, scenario mining, and model training with an accent on reliability, self-service workload onboarding, and multi-tenant resource efficiency. Focus on building GitOps delivery, distributed batch and workflow platforms, GPU scheduling, resource isolation, and reusable Spark-based processing capabilities.
Location: Santa Clara, CA; hybrid workplace
Company
operates a compute platform for large-scale data processing, simulation, auto-labeling, scenario mining, and model training.
What you will do
- Operate and evolve production Kubernetes clusters, including bare-metal provisioning, highly available control planes, node lifecycle, GPU container runtime, networking, and storage.
- Build GitOps-based delivery for platform services and user applications with Argo CD, Helm, and Kustomize.
- Develop multi-tenant capabilities for scheduling, resource isolation, storage, networking, access control, secrets, and observability.
- Improve CPU/GPU utilization and cost efficiency across the compute platform.
- Build reusable distributed batch and workflow platforms for Spark processing and GPU-based replay and simulation.
- Follow Quality Management System requirements and contribute to continuous improvement.
Requirements
- BS, MS, or PhD in Computer Science or a related technical field, or equivalent practical experience.
- Hands-on experience operating production Kubernetes clusters, including node lifecycle management, upgrades, and troubleshooting.
- Experience with GitOps and infrastructure as code.
- Experience with GPU or ML workload scheduling, queueing and priorities, fractional GPU sharing, autoscaling, or multi-tenant resource management.
- Strong ownership, self-direction, curiosity, and ability to drive projects end to end.
- Level is determined by experience, technical depth, scope of ownership, and demonstrated impact.
Nice to have
- Experience with Ray or Kubeflow.
- Experience with Delta Lake or Apache Iceberg.
- Experience operating large-scale distributed data-processing and workflow systems.
- Hands-on experience with Apache Spark and working knowledge of Argo Workflows or an equivalent orchestrator.
Culture & Benefits
- Full-time employment in a hybrid workplace.
- Strong fundamentals, ownership, and the ability to learn are valued over experience with every technology in the stack.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Principal Software Engineer (Infra/Platform)
180 000 - 235 000$
2 дня назад
Senior Cloud Engineer (AWS/Kubernetes)
215 000 - 240 000$
5 часов назад
Software Engineer in Deployment (Kubernetes)
135 000 - 231 000$
2 дня назад
Foundational Software Engineer (Infrastructure)
180 000 - 300 000$
1 день назад
DevOps Engineer (Kubernetes)
135 000 - 200 000$
1 день назад
Platform Engineer (AWS)
111 662 - 145 000$