Назад
Company hidden
9 дней назад

Senior Machine Learning Infrastructure Engineer, Recommendations and Search (AI)

187 040 - 359 720$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Machine Learning Infrastructure Engineer, Recommendations and Search (AI): Building large-scale distributed training, inference, and model-serving platforms for TikTok recommendations and search with an accent on GPU clusters, low-latency serving, and generative recommendation systems. Focus on designing 3D parallelism, optimizing LLM workloads across large GPU clusters, and building resilient, secure infrastructure for massive live traffic.

Location: San Jose, United States; fully in-person schedule up to 5 days a week

Salary: $187,040–$359,720 annually, plus potential discretionary bonuses, incentives, and restricted stock units.

Company

hirify.global operates TikTok-related applications and focuses on protecting U.S. user data, national security, cybersecurity, trust and safety, and content moderation.

What you will do

  • Implement and own large-scale distributed real-time machine learning training and inference platforms for TikTok recommendations and search.
  • Design and scale multi-node training systems using data, tensor, and pipeline parallelism across GPU clusters.
  • Build low-latency, high-throughput model-serving infrastructure and optimize inference pipelines for massive live traffic.
  • Co-design generative recommendation systems with applied machine learning researchers and develop reusable production components.
  • Build automated fault detection and asynchronous checkpointing for hardware, data-corruption, and network failures.
  • Develop secure multimodal data systems and enforce encryption, access control, and tenant-isolation requirements.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related technical discipline.
  • 4+ years of professional software engineering experience with deep expertise in Python and C++ or Java.
  • 2+ years of experience building and maintaining enterprise-scale machine learning infrastructure.
  • Expertise in PyTorch or TensorFlow internals, GPU memory management, CUDA, and InfiniBand or RoCE networking.
  • Experience with production-grade LLM training and inference tools, profiling, and system-level troubleshooting.
  • Ability to collaborate across teams, write technical design documents, and mentor team members.

Nice to have

  • Experience with Megatron, DeepSpeed, MFU optimization, communication-computation overlap, or zero-bubble pipeline scheduling.
  • Experience scaling LLM workloads across hundreds of GPUs, including KV cache management, prefill-decoding separation, and model quantization.
  • Experience in regulated industries, sovereign cloud environments, or federal data-security compliance.
  • Experience with CUDA, Triton, Cutlass, TensorRT, Triton Inference Server, graph compilation, or low-level kernel development.
  • Interest in MLOps, mixture-of-experts routing infrastructure, and specialized hardware optimization.

Culture & Benefits

  • On-site work is emphasized to support speed, alignment, team development, and integrated execution.
  • Medical, dental, and vision insurance are available from day one.
  • Benefits include a 401(k) savings plan with company match, paid parental leave, disability coverage, life insurance, and wellbeing benefits.
  • Benefits include 10 paid holidays, 10 paid sick days, and 17 days of paid personal time, with accrual increasing by tenure.
  • Inclusive workplace with reasonable accommodations available during recruitment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →