Senior Staff/Principal Deployment Automation Engineer (AI)
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Location: Onsite in San Francisco, Sunnyvale, or Bellevue, United States
Salary: $250,000β$300,000 annually plus bonus; restricted stock units included.
Company
is an AI infrastructure company operating an integrated stack from energy and data centers to cloud services for large-scale AI workloads.
What you will do
- Own deployment and integration-testing automation for bare-metal, on-premise systems across the AI Cloud stack.
- Build CI/CD platforms and GitLab tooling for reliable testing, iteration, and infrastructure releases across multiple data centers.
- Design large-scale validation tests for multi-node virtualized GPU and CPU clusters, including scaling, stability, performance, and tenant-isolation testing.
- Maintain and scale bare-metal Linux configurations using Ansible, AWX, osquery, and related tools.
- Develop deployment orchestration for canary releases, Blue/Green testing, and automated rollback on production systems.
- Build Python or Go automation frameworks to provision, configure, stress-test, and observe virtualized environments.
Requirements
- 12+ years of experience and a bachelor's or master's degree in Computer Science, Electrical Engineering, or a related technical field.
- Experience building automated integration testing for AI Cloud environments, from low-level Linux systems through distributed control planes.
- Working knowledge of Kubernetes, Docker, Terraform, and Postgres.
- Advanced Python and/or Bash skills, with intimate knowledge of CI/CD pipelines and GitLab tooling.
- Experience with one or more configuration-management systems, including Ansible, Puppet, Chef, or SaltStack.
- Knowledge of Linux kernel internals, PCIe topology, VFIO, HugePages, IOMMU, distributed GPU stacks, RDMA, RoCE, and InfiniBand.
Nice to have
- Experience with MNNVL or specialized AI fabric architectures.
- Familiarity with NVIDIA Nsight, AMD Omniperf, and hardware-level debugging or performance-profiling tools.
- Knowledge of Kubernetes GPU orchestration and specialized device plugins.
Culture & Benefits
- Health, dental, and vision insurance with employer HSA contributions.
- Paid time off, holidays, parental leave, leave-of-absence programs, and volunteer time off.
- 401(k) plan with company matching up to 4% of salary.
- Professional development, tuition reimbursement, and mental health and wellness support.
- Life insurance, disability coverage, global travel insurance, commuter benefits, meals allowance, and a cell phone stipend.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β