System Team Lead (Infrastructure)
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
TL;DR
System Team Lead (Infrastructure): Leading and managing the infrastructure for AI-powered restaurant operations software with an accent on distributed on-prem edge fleet management and observability. Focus on designing self-service tooling, automating operational workflows via Ansible and AWS, and maintaining high system reliability across cloud and edge.
Location: Hybrid in Taipei City, Taiwan
Company
builds AI-powered operations software for fast-food/QSR restaurant chains in the US.
What you will do
- Set the technical direction for the systems team, providing mentorship while remaining hands-on.
- Manage a distributed on-prem edge fleet of in-store servers and camera hardware via secure VPN/mesh networking.
- Develop self-service tooling and internal APIs to enable independent installation and debugging by field technicians.
- Operate a self-hosted Prometheus-based observability stack (Grafana, Mimir, Loki, Vector) across edge and cloud.
- Build event-driven auto-remediation and ticketing systems using Ansible/AWX and AWS serverless components.
- Oversee Taipei office core infrastructure, including Kubernetes, virtualization, and ISO 27001 security compliance.
Requirements
- 5+ years of experience in systems, infrastructure, or DevOps engineering with strong Ubuntu Linux administration.
- Proven track record of leading an infrastructure or SRE team while contributing technically.
- Deep understanding of TCP/IP, DNS, VLANs, VPNs, and firewalls.
- Proficiency in Infrastructure-as-Code (Terraform, Pulumi) and scripting (Bash, Python).
- Fluent in both Mandarin and English for internal coordination and communication with US stakeholders.
- Must be located in Taipei City, Taiwan to support the hybrid work model.
Nice to have
- Experience with IP cameras, ONVIF, and real-time video streaming (RTSP, H.264/H.265).
- Knowledge of large-scale storage/NAS operations (ZFS, TrueNAS, RAID/HA).
- Experience with vulnerability scanning and static analysis tools like OpenVAS or SonarQube.
- Experience with low-code workflow automation tools such as n8n.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β