8 дней назад
Principal SRE, Infrastructure & Platform
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal SRE, Infrastructure & Platform (Linux/Ansible/Kubernetes): Designing, deploying, and operating foundational infrastructure for a large-scale multi-datacenter platform spanning 30+ Points of Presence, with an accent on bare-metal provisioning, virtualization, container platforms, cloud integration, and security. Focus on automating heterogeneous on-premises and cloud environments, operating PCI-DSS-compliant infrastructure, and leading incident resolution in a 24x7 global production environment.
Location: Singapore homebase; remote-first team
Company
develops cybersecurity and application delivery solutions that help organizations create, secure, and run applications.
What you will do
- Design, deploy, and operate foundational infrastructure across 30+ Points of Presence in the Americas, EMEA, and APAC.
- Automate bare-metal provisioning, configuration management, CI/CD workflows, secrets management, and infrastructure asset management.
- Build and maintain Proxmox VE hypervisor clusters on HPE hardware, virtual machines, on-premises Kubernetes clusters, and Docker workloads.
- Engineer AWS and Azure infrastructure integrated with on-premises DNS, monitoring, identity, networking, and storage.
- Operate distributed DNS, load balancing, directory, AAA, networking, monitoring, logging, and runtime security services.
- Participate in 24x7 on-call rotation, lead incident resolution, define SLOs/SLIs, conduct blameless post-mortems, and implement high-availability improvements.
Requirements
- 7+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering in production environments.
- Strong Linux administration skills, preferably with RHEL/CentOS, including systemd, networking, storage, kernel tuning, and package management.
- Proficiency with Ansible or an equivalent configuration-management tool and experience writing CI/CD pipelines.
- Hands-on experience with a hypervisor platform and production operation of on-premises Kubernetes clusters.
- Practical AWS or Azure experience covering compute, networking, IAM, DNS, security groups, and managed services.
- Experience with networking fundamentals, secrets management, PCI-DSS infrastructure requirements, and leading production incident response.
Nice to have
- Bare-metal server fleet management with HPE iLO, IPMI, or equivalent tools.
- CMDB/IPAM, LDAP, RADIUS, and SNMP-based network monitoring experience.
- Python or Bash scripting and experience with colocation or carrier-neutral data centers.
- OSTree, DDoS mitigation, BGP route reflectors, network-function virtualization, Pulp, or open-source infrastructure tooling experience.
Culture & Benefits
- Remote-first work in a globally distributed team across multiple time zones.
- Participation in a supported 24x7 on-call rotation.
- Autonomous work with a strong focus on automation and operational excellence.
- Responsibility for documentation, runbooks, post-mortems, security hardening, and audit readiness.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →