9 дней назад
Tech Lead Machine Learning Infrastructure Engineer - Recommendations and Search (AI)
187 040 - 438 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Tech Lead Machine Learning Infrastructure Engineer - Recommendations and Search (AI): Building and scaling distributed real-time machine learning training and inference platforms for TikTok recommendations and search with an accent on large-scale parallelism, GPU cluster efficiency, and low-latency model serving. Focus on designing 3D parallelism architectures, resilient multi-thousand-GPU infrastructure, generative recommendation systems, and secure compliant data processing.
Location: San Jose R&D; fully in-person schedule up to 5 days a week
Salary: $187,040–$438,000 annually, with potential discretionary bonuses, incentives, and restricted stock units.
Company
operates data privacy, cybersecurity, trust and safety, and U.S. user data protection programs for TikTok applications.
What you will do
- Drive the technical roadmap for billion-parameter distributed real-time machine learning training and inference platforms supporting recommendations and search.
- Architect and scale multi-node distributed training systems using data, tensor, and pipeline parallelism across large GPU clusters.
- Build low-latency, high-throughput model serving infrastructure and optimize inference pipelines for high-volume live traffic.
- Partner with applied machine learning researchers to develop production-ready generative recommendation systems and reusable infrastructure components.
- Design automated fault detection, asynchronous checkpointing, SLI/SLO frameworks, and resilient systems for multi-thousand-GPU environments.
- Engineer secure multimodal data processing and storage systems with encryption, access controls, and tenant isolation.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related technical discipline.
- 5+ years of software engineering experience with Python, C++ or Java, and large-scale distributed systems.
- 3+ years of experience building and maintaining enterprise-scale machine learning infrastructure with GPU clusters, Kubernetes, or Slurm.
- Deep knowledge of PyTorch or TensorFlow, GPU memory management, CUDA, and InfiniBand or RoCE networking.
- Experience with production-grade LLM training and inference tools, profiling, and eliminating I/O, compute, or network bottlenecks.
- Ability to work onsite in San Jose up to five days per week.
Nice to have
- Experience with Megatron, DeepSpeed, advanced pipeline scheduling, and Model FLOPs Utilization optimization.
- Experience scaling LLM workloads across hundreds of GPUs, including KV cache management, prefill-decoding separation, and quantization.
- Background in regulated industries, sovereign cloud environments, or federal data security compliance.
- Experience with CUDA, Triton, CUTLASS, TensorRT, Triton Inference Server, MoE routing, or specialized hardware optimization.
- Strong technical leadership, mentoring, architecture documentation, and stakeholder communication skills.
Culture & Benefits
- Medical, dental, and vision insurance from the first day of employment.
- 401(k) savings plan with company match, life insurance, disability coverage, and wellbeing benefits.
- Paid parental leave, 10 paid holidays, 10 paid sick days, and 17 days of paid personal time.
- Inclusive workplace focused on creativity, curiosity, humility, resilience, and continuous innovation.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
ML Platform Engineer (AI)
100 000 - 160 000$
9 дней назад
Engineer II, AI/ML
88 800 - 133 200$
7 дней назад
Machine Learning Engineer, Infra, AI for Drug Discovery
147 600 - 274 000$
7 дней назад
Senior AI/ML Platform Engineer (AI)
160 000 - 271 000$
7 дней назад
Lead Data Scientist
10 дней назад
Staff Machine Learning Engineer (Fintech)
208 000 - 260 000$