5 часов назад
Training Infrastructure Engineer (AI)
210 000 - 320 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Training Infrastructure Engineer (AI): Designing and optimizing scalable infrastructure and distributed pipelines for large-scale LLM and multimodal model training with an accent on multi-GPU systems, data storage, orchestration, and operational reliability. Focus on improving training performance and cost efficiency, troubleshooting distributed computing issues, and collaborating with AI researchers on training methodologies.
Location: Hybrid in San Mateo or New York, United States
Salary: $210K–$320K annually, plus equity
Company
provides infrastructure for building, training, and serving specialized AI models across text, image, embedding, audio, and multimodal workloads.
What you will do
- Design and implement scalable infrastructure for large-scale model training workloads.
- Develop and maintain distributed training pipelines for LLMs and multimodal models.
- Optimize training performance across GPUs, nodes, and data centers.
- Build monitoring, logging, debugging, storage, provisioning, scaling, and orchestration systems for training operations.
- Collaborate with AI researchers to implement and optimize training methodologies.
- Analyze system efficiency, scalability, reliability, and cost-effectiveness while troubleshooting complex distributed performance issues.
Requirements
- Bachelor’s degree in Computer Science, Computer Engineering, or a related field, or equivalent practical experience.
- 3+ years of experience with distributed systems and ML infrastructure.
- Experience with PyTorch and distributed training techniques such as data parallelism, model parallelism, and FSDP.
- Proficiency with AWS, GCP, or Azure cloud platforms.
- Experience with containerization and orchestration using Kubernetes and Docker.
Nice to have
- Master’s or PhD in Computer Science or a related field.
- Experience training large language models or multimodal AI systems.
- Experience with ML workflow orchestration tools and ML DevOps practices.
- Background in optimizing high-performance distributed computing systems.
- Open-source contributions to ML infrastructure or related projects.
Culture & Benefits
- Work on challenging AI infrastructure problems, including scalable model serving and low-latency inference.
- Build production technology used by businesses and developers globally.
- Operate with a high level of ownership in a fast-growing, collaborative environment.
- Collaborate with experienced engineers and AI researchers.
- Equity is included in the compensation package.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 часа назад
Member of Technical Staff, Training Infra (AI)
200 000 - 350 000$
9 часов назад
ML Training Infrastructure Engineer (AI)
220 000 - 320 000$
Baseten
21 час назад
Forward Deployed Engineers (AI)
200 000 - 400 000$
8 часов назад
AI Engineer
159 000 - 195 000$
5 часов назад
Technical Staff
200 000 - 350 000$
4 часа назад