3 дня назад
Director, Platform Engineering (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Director, Platform Engineering (AI): Owning end-to-end reliability for the cloud, accelerated compute, DevOps, and application support platform behind enterprise AI initiatives with an accent on incident management, distributed infrastructure, and production operations. Focus on designing reliable GPU and Kubernetes environments, enforcing SLOs and safe delivery pipelines, and building observability and governance for LLM-based and agentic systems.
Location: Onsite in Kraków, Poland; ability to travel up to 20%.
Company
is a global science and technology company operating across life sciences, diagnostics, and biotechnology.
What you will do
- Own availability, performance, recovery, and end-to-end reliability for the platform supporting enterprise AI initiatives.
- Lead incident detection, triage, mitigation, post-incident reviews, SLOs, error budgets, and on-call practices.
- Design and operate GPU and accelerator fleets, cloud infrastructure, storage, networking, scheduling, and capacity management for AI workloads.
- Own application support, deployment, configuration, runtime health, CI/CD, infrastructure as code, environment management, and release engineering.
- Partner with AI teams in molecular design, autonomous labs, supply chain, and professional services to translate reliability and capacity needs into platform improvements.
- Define reliability strategies for LLM-based and agentic systems, including observability, evaluation, guardrails, governance, and safe execution patterns.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
- 10+ years of engineering experience, including substantial ownership of production reliability, SRE, or platform operations for large-scale distributed systems.
- Experience leading and scaling multidisciplinary engineering organizations across distributed time zones.
- Hands-on incident management experience, including on-call design, escalation paths, blameless post-incident reviews, and permanent remediation.
- Production experience with AWS, Azure, or GCP; Docker, Kubernetes; Terraform, Pulumi, or similar infrastructure-as-code tools; and ML/AI infrastructure at scale.
- Strong observability, automation, communication, and internal customer management skills.
Nice to have
- Production experience with LLM and agentic systems, including inference serving, evaluation, guardrails, or observability for non-deterministic workloads.
- Experience with HPC, lab automation, instrument data pipelines, or other scientific computing environments.
- Experience with SOC 2, FedRAMP, GxP, multi-cloud, or on-premise infrastructure.
Culture & Benefits
- Corporate role hosted by the Cytiva operating company in Kraków.
- Work within a culture of continuous improvement and belonging.
- Opportunity to support science and technology initiatives with real-world impact.
- Travel of up to 20% is expected.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →