Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Deployment Engineering Director, Systems Engineering (AI cloud infrastructure): Leading and scaling engineering teams responsible for core platform systems for a next-generation GPU cloud with an accent on systems validation, reliability, platform architecture, and operational excellence. Focus on designing execution plans for ambiguous infrastructure challenges, aligning teams across distributed systems and developer tooling, and scaling high-performance data centre operations.
Location: UK; travel required up to 50% of the time
Company
Nscale provides high-performance GPU cloud infrastructure for AI start-ups and enterprise customers, focusing on cost-efficient AI development and sustainable data centre operations.
What you will do
- Lead and scale an engineering team responsible for core platform systems.
- Define and execute multi-quarter initiatives improving delivery velocity, system reliability, validation, and product quality.
- Partner with Product Management and engineering leaders to evolve the platform and balance architecture with near-term business needs.
- Turn ambiguous challenges involving scalability, vendor experience, and platform abstractions into clear execution plans.
- Coordinate work across internal platforms, infrastructure, and developer tooling teams.
- Raise engineering quality, operational excellence, and execution standards across the platform organization.
Requirements
- Bachelor’s degree in Computer Science or a related engineering field.
- 10+ years of Systems Engineering Management experience.
- Experience in a large cloud provider or hyperscale data centre environment.
- Experience in systems or compute operations leadership.
- Strong knowledge of Linux/Unix administration, OS-level tuning, and server/GPU hardware architecture.
- Excellent organizational, verbal, written communication, and judgment skills.
Nice to have
- Experience with virtualization, Kubernetes, Docker, and distributed systems design.
- Infrastructure-as-code and configuration management experience with Ansible, Terraform, Puppet, or Chef.
- Experience with scripting, automation, data centre design, storage, networking, or computing.
- Knowledge of system architecture, data synchronization, fault tolerance, state management, monitoring, observability, and telemetry.
- Experience with GPU burn-in and validation testing at scale, including thermal, power, and stress qualification.
Culture & Benefits
- Opportunity to shape operating standards for a next-generation AI cloud platform.
- Ownership of complex infrastructure challenges with direct business and technical impact.
- Work focused on scaling high-performance and sustainable data centre operations.
- Culture centered on innovation, ownership, accountability, openness, and transparency.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
Engineering Manager - Cloud Platform & Operations
5 дней назад
Senior Engineering Manager (AI)
126 000 - 136 000GBP
4 дня назад
Technical Lead (.NET)
70 000 - 85 000GBP
CrowdStrike
7 дней назад
Sr. Network Engineer - Data Center and Cloud (Hybrid, London)
8 дней назад
Director Embedded Systems
200 000 - 250 000CAD
9 дней назад