Назад
Company hidden
13 дней назад

Member of Technical Staff (Platform)

Формат работы
remote (только Brazil)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Brazil
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff (Platform) (Kubernetes/AI): Operating model inference and training workloads across cloud and customer-hosted Kubernetes with an accent on reliability, performance, and cost efficiency. Focus on building Kubernetes controllers, optimizing GPU-backed inference and data-heavy Python pipelines, and solving deterministic scheduling, recovery, isolation, and profiling challenges.

Location: Remote from São Paulo, Brazil

Company

hirify.global builds a model platform for governed inference, training, post-training, and fine-tuning across cloud and customer-hosted Kubernetes environments.

What you will do

  • Evolve Sophos, the Kubernetes-based online and batch inference runtime.
  • Run large-scale batch inference on ephemeral jobs using multi-dimensional CPU, memory, and GPU admission control through Kueue.
  • Build and extend the Sophos controller and Kubernetes custom resources.
  • Optimize model inference engines and feature-processing pipelines using vectorized and columnar operations.
  • Own training, post-training, and fine-tuning job execution across hirify.global’s cloud and customer dataplanes.
  • Drive autoscaling, GPU serving, telemetry, performance, reliability, and cost optimization.

Requirements

  • Experience running model serving or large-scale batch compute on Kubernetes.
  • Experience building Kubernetes controllers or operators.
  • Strong Python profiling and optimization skills for data-heavy pipelines.
  • Ability to write production-quality code, participate in reviews, and operate systems in production.
  • Strong focus on compute efficiency, serving availability, inference latency, batch throughput, and job completion reliability.

Nice to have

  • Production experience with Ray, Ray Serve, or KubeRay.
  • Experience with Kueue or other batch scheduling and admission-control systems.
  • GPU serving and performance optimization experience.
  • Experience with Arrow, Parquet, Lance, or other columnar data formats.
  • Experience shipping software to customer-hosted Kubernetes, including GCP/AWS, GKE/EKS, or regulated financial-services environments.

Culture & Benefits

  • Technical ICs operate as Members of Technical Staff with ownership of systems and outcomes.
  • Work remotely from São Paulo.
  • Success is measured through 99.9% serving availability, p95/p99 online inference latency, batch throughput, cost per prediction, GPU utilization, and reliable job completion.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →