Назад
Company hidden
9 дней назад

Tech Lead Machine Learning Infrastructure Engineer - Recommendations and Search (AI)

187 040 - 438 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Tech Lead Machine Learning Infrastructure Engineer - Recommendations and Search (AI): Building and scaling distributed real-time machine learning training and inference platforms for TikTok recommendations and search with an accent on large-scale parallelism, GPU cluster efficiency, and low-latency model serving. Focus on designing 3D parallelism architectures, resilient multi-thousand-GPU infrastructure, generative recommendation systems, and secure compliant data processing.

Location: San Jose R&D; fully in-person schedule up to 5 days a week

Salary: $187,040–$438,000 annually, with potential discretionary bonuses, incentives, and restricted stock units.

Company

hirify.global operates data privacy, cybersecurity, trust and safety, and U.S. user data protection programs for TikTok applications.

What you will do

  • Drive the technical roadmap for billion-parameter distributed real-time machine learning training and inference platforms supporting recommendations and search.
  • Architect and scale multi-node distributed training systems using data, tensor, and pipeline parallelism across large GPU clusters.
  • Build low-latency, high-throughput model serving infrastructure and optimize inference pipelines for high-volume live traffic.
  • Partner with applied machine learning researchers to develop production-ready generative recommendation systems and reusable infrastructure components.
  • Design automated fault detection, asynchronous checkpointing, SLI/SLO frameworks, and resilient systems for multi-thousand-GPU environments.
  • Engineer secure multimodal data processing and storage systems with encryption, access controls, and tenant isolation.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related technical discipline.
  • 5+ years of software engineering experience with Python, C++ or Java, and large-scale distributed systems.
  • 3+ years of experience building and maintaining enterprise-scale machine learning infrastructure with GPU clusters, Kubernetes, or Slurm.
  • Deep knowledge of PyTorch or TensorFlow, GPU memory management, CUDA, and InfiniBand or RoCE networking.
  • Experience with production-grade LLM training and inference tools, profiling, and eliminating I/O, compute, or network bottlenecks.
  • Ability to work onsite in San Jose up to five days per week.

Nice to have

  • Experience with Megatron, DeepSpeed, advanced pipeline scheduling, and Model FLOPs Utilization optimization.
  • Experience scaling LLM workloads across hundreds of GPUs, including KV cache management, prefill-decoding separation, and quantization.
  • Background in regulated industries, sovereign cloud environments, or federal data security compliance.
  • Experience with CUDA, Triton, CUTLASS, TensorRT, Triton Inference Server, MoE routing, or specialized hardware optimization.
  • Strong technical leadership, mentoring, architecture documentation, and stakeholder communication skills.

Culture & Benefits

  • Medical, dental, and vision insurance from the first day of employment.
  • 401(k) savings plan with company match, life insurance, disability coverage, and wellbeing benefits.
  • Paid parental leave, 10 paid holidays, 10 paid sick days, and 17 days of paid personal time.
  • Inclusive workplace focused on creativity, curiosity, humility, resilience, and continuous innovation.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →