Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Centre of Excellence Senior Engineer (AI Infrastructure): Defining and standardising infrastructure operations across a global data centre estate with an accent on operational excellence, incident management, capacity planning, and infrastructure optimisation. Focus on leading complex infrastructure initiatives, solving systemic reliability issues, and building scalable standards, runbooks, and operating models for AI infrastructure.
Location: EMEA; London; UK. Remote-first collaboration is available with an option for office-based work; occasional travel to data centre sites across multiple countries is required.
Company
Nscale provides high-performance GPU cloud infrastructure for AI start-ups and enterprise customers.
What you will do
- Lead complex infrastructure projects, including hardware installations, upgrades, deployments, and operational improvement initiatives.
- Standardise infrastructure operations across the global data centre estate through repeatable processes, best practices, documentation, and runbooks.
- Own complex infrastructure incidents, lead major incident response, conduct root cause analysis, and implement corrective actions.
- Partner with engineering, deployment, platform, network, and operations teams on capacity planning, infrastructure optimisation, and operational readiness.
- Provide technical leadership and mentorship across multiple operational teams and data centre locations.
- Travel to data centre sites to support deployments, resolve technical issues, and improve operational execution.
Requirements
- 5+ years of experience in technical infrastructure, data centre, or operations environments.
- Extensive experience in large-scale data centre environments and solving systemic infrastructure issues.
- Experience leading technical initiatives across multiple sites or operational teams.
- Strong analytical, troubleshooting, problem-solving, communication, and stakeholder management skills.
- Advanced technical certifications relevant to infrastructure operations.
- Ability to work in fast-paced, ambiguous environments and travel to data centre sites when required.
Nice to have
- Experience with hyperscale, HPC, AI, or NVIDIA GPU infrastructure.
- Hands-on server deployment, maintenance, and troubleshooting experience.
- Experience supporting mission-critical infrastructure with a strong customer service mindset.
Culture & Benefits
- Collaborative, supportive, and innovation-focused working environment.
- Flexible workplace with autonomy over the working day and remote-first collaboration.
- Highly competitive package with reviews every 12 months.
- Dynamic progression plan focused on strategic infrastructure initiatives and operational excellence.
- Equal-opportunity environment with accommodations available for individual circumstances.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Anthropic
8 дней назад
Staff Software Engineer (AI Reliability Engineering)
325 000 - 390 000GBP
Anthropic
8 дней назад
Incident Response Manager (AI)
290 000 - 365 000$
4 дня назад
Staff Service Reliability and Operational Intelligence Engineer (AI Ops)
152 000 - 228 000$
6 дней назад
Senior Site Reliability Engineer (AI)
75 000 - 85 000$
Replit
10 дней назад
Engineering Manager (SRE)
250 000 - 325 000$
Datadog
4 дня назад
Senior Software Engineer - Incident Insights & Readiness (SRE)
192 000 - 240 000$