Назад
Company hidden
2 дня назад

Software Engineer (SE / Sr SE), Data & ML Platform

Формат работы
hybrid
Тип работы
fulltime
Грейд
middle/senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer (SE / Sr SE), Data & ML Platform (Kubernetes/GPU): Operating and evolving Kubernetes infrastructure for petabyte-scale data processing, simulation, auto-labeling, scenario mining, and model training with an accent on reliability, self-service workload onboarding, and multi-tenant resource efficiency. Focus on building GitOps delivery, distributed batch and workflow platforms, GPU scheduling, resource isolation, and reusable Spark-based processing capabilities.

Location: Santa Clara, CA; hybrid workplace

Company

hirify.global operates a compute platform for large-scale data processing, simulation, auto-labeling, scenario mining, and model training.

What you will do

  • Operate and evolve production Kubernetes clusters, including bare-metal provisioning, highly available control planes, node lifecycle, GPU container runtime, networking, and storage.
  • Build GitOps-based delivery for platform services and user applications with Argo CD, Helm, and Kustomize.
  • Develop multi-tenant capabilities for scheduling, resource isolation, storage, networking, access control, secrets, and observability.
  • Improve CPU/GPU utilization and cost efficiency across the compute platform.
  • Build reusable distributed batch and workflow platforms for Spark processing and GPU-based replay and simulation.
  • Follow Quality Management System requirements and contribute to continuous improvement.

Requirements

  • BS, MS, or PhD in Computer Science or a related technical field, or equivalent practical experience.
  • Hands-on experience operating production Kubernetes clusters, including node lifecycle management, upgrades, and troubleshooting.
  • Experience with GitOps and infrastructure as code.
  • Experience with GPU or ML workload scheduling, queueing and priorities, fractional GPU sharing, autoscaling, or multi-tenant resource management.
  • Strong ownership, self-direction, curiosity, and ability to drive projects end to end.
  • Level is determined by experience, technical depth, scope of ownership, and demonstrated impact.

Nice to have

  • Experience with Ray or Kubeflow.
  • Experience with Delta Lake or Apache Iceberg.
  • Experience operating large-scale distributed data-processing and workflow systems.
  • Hands-on experience with Apache Spark and working knowledge of Argo Workflows or an equivalent orchestrator.

Culture & Benefits

  • Full-time employment in a hybrid workplace.
  • Strong fundamentals, ownership, and the ability to learn are valued over experience with every technology in the stack.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →