13 дней назад
Director, Deployment Engineering — Systems Engineering (AI Infrastructure)
240 000 - 353 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Director, Deployment Engineering — Systems Engineering (AI Infrastructure): Leading the deployment, validation, and operation of core compute platforms for AI infrastructure with an accent on GPU qualification, systems reliability, and production readiness. Focus on scaling systems engineering teams, designing validation standards, automating deployment operations, and improving observability across high-availability infrastructure.
Location: AMER; travel up to 50% required for deployments, site readiness, vendor collaboration, and operational execution.
Salary: $240,000–$353,000 USD annually, plus potential bonus, equity, and/or commission.
Company
Nscale builds AI infrastructure and the compute platforms that support production AI workloads.
What you will do
- Lead and scale the systems engineering team responsible for platform systems, deployment readiness, and compute operations.
- Define multi-quarter initiatives that improve deployment velocity, platform reliability, validation quality, and operational performance.
- Establish validation standards for servers, GPUs, networking, storage, and supporting infrastructure before production acceptance.
- Lead GPU burn-in and qualification testing covering thermal, power, stress, and performance requirements.
- Partner with infrastructure, network engineering, deployment, product, and operations leaders on platform architecture and execution.
- Build scalable approaches for automation, monitoring, telemetry, incident learning, documentation, and continuous platform improvement.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical field.
- 10+ years of experience in systems engineering, compute operations, infrastructure, or engineering management.
- Experience in a large cloud provider, hyperscale data center, or similarly complex infrastructure environment.
- Strong knowledge of Linux/Unix administration, OS-level tuning, server architecture, and GPU hardware.
- Experience with virtualization, containerization, distributed systems, Kubernetes, Docker, infrastructure as code, and configuration management.
- Strong scripting, automation, data center systems design, observability, telemetry, fault tolerance, and distributed-system reliability experience.
Culture & Benefits
- Leadership role spanning systems engineering, compute operations, deployment validation, and cross-functional infrastructure execution.
- Medical, dental, and vision benefits.
- Flexible paid time off and parental leave.
- Retirement plan participation.
- Potential bonus, equity, and/or commission eligibility.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →