Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
ML Infrastructure Engineer (AI/GPU Infrastructure): Designing and scaling GPU compute infrastructure and training frameworks to power recommendations on X with an accent on high-performance ML platforms and data pipelines. Focus on optimizing distributed systems, ensuring scalability of large-scale ML systems, and productionizing complex models.
Location: Palo Alto, California, United States (Onsite)
Salary: $180,000 – $440,000 USD
Company
xAI is a small, highly motivated team focused on creating AI systems that understand the universe to aid humanity in its pursuit of knowledge.
What you will do
- Design, build, and scale GPU compute infrastructure, training frameworks, and experimentation tools.
- Develop data pipelines and integrate large-scale data, training, and inference systems.
- Collaborate with ML teams to productionize models and ensure seamless integration across the stack.
- Ensure the scalability, reliability, and efficiency of large-scale machine learning systems.
- Solve complex technical problems independently across the full stack.
- Mentor junior engineers and contribute to the overall growth of the engineering team.
Requirements
- Degree in computer science, machine learning, or a quantitative discipline (or equivalent experience).
- 2+ years of industry experience with high-traffic production environments, distributed systems, or GPU infrastructure.
- 2+ years of experience with ML platforms, training infrastructure, or collaboration with modeling engineers.
- Strong proficiency in Python and experience with compiled languages like C++ or Rust.
- Must be based in or able to work from Palo Alto, California.
Nice to have
- Deep familiarity with modern ML frameworks such as JAX or PyTorch.
- Low-level understanding of compute systems, NVIDIA drivers, CUDA toolkits, and networking.
- Experience with Linux systems and orchestration tools.
- Experience with job schedulers (Slurm) or configuration management (Puppet/Ansible).
Culture & Benefits
- Flat organizational structure emphasizing initiative and hands-on contribution.
- Competitive compensation including base salary and equity.
- Comprehensive medical, vision, and dental coverage.
- Access to a 401(k) retirement plan and life insurance.
- Short and long-term disability insurance and various other perks.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
13 дней назад
Tech Lead Machine Learning Infrastructure Engineer - Recommendations and Search (AI)
187 040 - 438 000$
13 дней назад
ML Infrastructure Engineer (AI)
11 дней назад
Staff ML Infrastructure Engineer (Embodied AI)
171 700 - 335 300$
11 дней назад
Senior ML Infrastructure Engineer (Autonomous Driving)
128 700 - 261 300$
13 дней назад
AI/ML Engineer
107 000 - 233 000$
3 дня назад
AI Platform Engineer
130 000 - 180 000$