25 дней назад
Sr. GPU Cloud South-North Network SRE Expert (EVPN-VXLAN)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Sr. GPU Cloud South-North Network SRE Expert (EVPN-VXLAN): Building and operating the IP and Ethernet infrastructure connecting AI GPU cloud sites across the US, APAC, and Iceland, with an accent on EVPN-VXLAN fabrics, DCI, BGP, WAN, and network automation. Focus on expressing route policies, ACLs, VRFs, peering, and operational changes as code for AIOps-driven monitoring and remediation.
Location: Singapore or Penang, Malaysia. The role allows flexible remote work where necessary and may involve on-site exposure to heat, noise, and vibration at datacenter or job sites.
Company
develops Bitcoin mining infrastructure and AI cloud services, operating datacenters and energy infrastructure across multiple countries.
What you will do
- Design and operate EVPN-VXLAN datacenter fabrics supporting multi-tenant isolation and workload mobility.
- Build multi-site DCI overlays and underlays connecting four US datacenters with consistent Layer 2 and Layer 3 services.
- Manage BGP, OSPF, ECMP, IP transit, peering, internet edge, and WAN infrastructure across US, APAC, and Iceland sites.
- Operate Arista, Cisco, and Palo Alto network equipment, including switches, routers, and firewalls.
- Automate network configuration and operations with Ansible, Terraform, Nautobot/NetBox, and related tooling.
- Feed route changes, network health data, BGP events, and DCI degradation into AIOps monitoring and remediation workflows.
Requirements
- 5+ years of data center network engineering experience with hands-on EVPN-VXLAN deployment.
- Strong BGP expertise, including eBGP/iBGP, route policies, communities, and traffic engineering.
- Experience designing and operating multi-site DCI with EVPN multihoming.
- Proficiency with Arista EOS and/or Cisco NX-OS in spine-leaf datacenter environments.
- Experience with Palo Alto or Fortinet firewalls, IP transit, internet peering, QoS, and network capacity planning.
- Experience with network automation tools such as Ansible, Python/Netmiko, Nautobot, or NetBox, plus an intent-driven, runbook-as-code approach.
Culture & Benefits
- Inclusive environment that values authenticity and diverse perspectives.
- Startup-style environment within a fast-growing technology company.
- Autonomy, personal accountability, learning, training, and mentoring opportunities.
- Opportunity to contribute to digital asset, AI cloud, and datacenter infrastructure projects.
- Safety and noise-reduction equipment is provided for on-site work.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior Cloud Networking SRE
2 часа назад
DevOps/SRE Engineer
NEWHR
4 дня назад
Senior PostgreSQL DBA/SRE (iPaaS)
P2P.org
4 дня назад
Site Reliability Engineer (Web3)
5 000 - 6 000$
9 дней назад
Head of Site Reliability Engineering (AI)
195 000 - 285 000$
9 дней назад