Назад
Company hidden
2 месяца назад

Member of Technical Staff, Distributed Systems (AI)

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US/Japan
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff, Distributed Systems (AI): Designing and operating GPU clusters, cloud environments, scheduling systems, and storage infrastructure for large-scale biological model training and inference with an accent on distributed systems, performance engineering, and reliability. Focus on building unified compute interfaces, optimizing thousands of accelerators, and developing fault-tolerant infrastructure for frontier biological world models.

Location: On-site in San Francisco or Tokyo

Company

hirify.global is an AI research lab developing generative genomics and biological world models to advance scientific understanding, healthcare, and biological defense.

What you will do

  • Operate and automate large GPU clusters, including provisioning, imaging, capacity planning, uptime, utilization, and cost efficiency.
  • Build a unified software interface for cluster management, model training, and inference.
  • Extend Kubernetes or Slurm for topology-aware scheduling, preemption, quotas, and fair-share multi-tenancy.
  • Profile and optimize communication, memory usage, custom kernels, compilation paths, and instrumentation across the systems stack.
  • Develop monitoring, fault tolerance, checkpointing, recovery, and incident-analysis mechanisms for long-running distributed workloads.
  • Design durable storage and artifact paths for datasets, checkpoints, logs, retention, and experiment lineage while collaborating with researchers and training scientists.

Requirements

  • Track record building distributed systems or operating large-scale GPU clusters and container orchestration systems such as Kubernetes or Slurm.
  • Proficiency in performant, maintainable software development with Python or Rust.
  • Strong systems knowledge spanning Linux, networking, and infrastructure as code.
  • Strong understanding of deep learning systems and internals, including PyTorch, Triton, CUDA, and C++.
  • Ability to debug complex distributed-training, memory, performance, and reliability issues across large codebases.
  • Excellent written and verbal communication across technical and scientific domains; U.S. work authorization is required.

Nice to have

  • Familiarity with CUDA/NCCL, distributed-training performance profiling, ML runtimes, compilers, numerics, communication libraries, and custom kernel development.
  • Experience supporting frontier or foundation model training and improving researcher productivity through infrastructure or developer tooling.
  • Open-source contributions to ML systems such as PyTorch, Torchtitan, or Megatron-LM.
  • Background in applied mathematics, systems, computational biology, or related quantitative sciences.

Culture & Benefits

  • Work on distributed training, architecture, and numerics problems for real biological applications.
  • Collaborate across AI labs, biotechs, hospital systems, and research institutes.
  • Join a culture focused on rigor, creativity, and cross-disciplinary partnership.
  • Competitive compensation, comprehensive benefits, and support for continual learning.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →