Member of Technical Staff Research Infrastructure Engineer (GPU)
Мэтч & Сопровод
Покажет вашу совместимость и напишет письмо
Описание вакансии
Member of Technical Staff Research Infrastructure Engineer
Company
Black Forest Labs
Conditions
1 week ago
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will maintain and scale research infrastructure for large-scale training workloads. You will optimize application and infrastructure performance, work with research teams on cost-efficient solutions, diagnose distributed-system bottlenecks, build telemetry and monitoring, and participate in on-call incident response.
Requirements
- Experience building or operating large-scale training platforms
- Experience with large-scale GPU compute clusters
- Ability to debug performance and reliability issues across distributed fleets
- Knowledge of Kubernetes, Infrastructure as Code, AWS, and GCP
- Experience with SLURM
- Knowledge of Python, Bash, Go, NVIDIA GPU drivers and operators, OpenTelemetry, and Prometheus
Responsibilities
- Maintain and optimize research infrastructure
- Scale infrastructure while maintaining reliability and performance
- Design cost-efficient infrastructure solutions with research teams
- Identify and resolve distributed-system performance bottlenecks
- Build and evolve telemetry and monitoring systems
- Participate in on-call rotations and incident response
Benefits
- Equity
- Reasonable travel costs covered
- Relocation encouraged but not required
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Текст вакансии взят без изменений