Назад
Company hidden
2 дня назад

Senior Staff/Principal Deployment Automation Engineer (AI)

250Β 000 - 300Β 000$
Π€ΠΎΡ€ΠΌΠ°Ρ‚ Ρ€Π°Π±ΠΎΡ‚Ρ‹
onsite
Π’ΠΈΠΏ Ρ€Π°Π±ΠΎΡ‚Ρ‹
fulltime
Π“Ρ€Π΅ΠΉΠ΄
senior
Английский
b2
Π‘Ρ‚Ρ€Π°Π½Π°
US
Вакансия ΠΈΠ· списка Hirify.GlobalВакансия ΠΈΠ· Hirify Global, списка ΠΌΠ΅ΠΆΠ΄ΡƒΠ½Π°Ρ€ΠΎΠ΄Π½Ρ‹Ρ… tech-ΠΊΠΎΠΌΠΏΠ°Π½ΠΈΠΉ
Для мэтча ΠΈ ΠΎΡ‚ΠΊΠ»ΠΈΠΊΠ° Π½ΡƒΠΆΠ΅Π½ Plus

ΠœΡΡ‚Ρ‡ & Π‘ΠΎΠΏΡ€ΠΎΠ²ΠΎΠ΄

Для мэтча с этой вакансиСй Π½ΡƒΠΆΠ΅Π½ Plus

ОписаниС вакансии

ВСкст:
/
TL;DR
Senior Staff/Principal Deployment Automation Engineer (AI): Building deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters across an AI Cloud environment with an accent on bare-metal Linux, distributed systems, and infrastructure reliability. Focus on validating cluster scaling, orchestrating canary and Blue/Green deployments, automating rollback, and testing GPU performance and multi-tenant isolation.

Location: Onsite in San Francisco, Sunnyvale, or Bellevue, United States

Salary: $250,000–$300,000 annually plus bonus; restricted stock units included.

Company

hirify.global is an AI infrastructure company operating an integrated stack from energy and data centers to cloud services for large-scale AI workloads.

What you will do

  • Own deployment and integration-testing automation for bare-metal, on-premise systems across the AI Cloud stack.
  • Build CI/CD platforms and GitLab tooling for reliable testing, iteration, and infrastructure releases across multiple data centers.
  • Design large-scale validation tests for multi-node virtualized GPU and CPU clusters, including scaling, stability, performance, and tenant-isolation testing.
  • Maintain and scale bare-metal Linux configurations using Ansible, AWX, osquery, and related tools.
  • Develop deployment orchestration for canary releases, Blue/Green testing, and automated rollback on production systems.
  • Build Python or Go automation frameworks to provision, configure, stress-test, and observe virtualized environments.

Requirements

  • 12+ years of experience and a bachelor's or master's degree in Computer Science, Electrical Engineering, or a related technical field.
  • Experience building automated integration testing for AI Cloud environments, from low-level Linux systems through distributed control planes.
  • Working knowledge of Kubernetes, Docker, Terraform, and Postgres.
  • Advanced Python and/or Bash skills, with intimate knowledge of CI/CD pipelines and GitLab tooling.
  • Experience with one or more configuration-management systems, including Ansible, Puppet, Chef, or SaltStack.
  • Knowledge of Linux kernel internals, PCIe topology, VFIO, HugePages, IOMMU, distributed GPU stacks, RDMA, RoCE, and InfiniBand.

Nice to have

  • Experience with MNNVL or specialized AI fabric architectures.
  • Familiarity with NVIDIA Nsight, AMD Omniperf, and hardware-level debugging or performance-profiling tools.
  • Knowledge of Kubernetes GPU orchestration and specialized device plugins.

Culture & Benefits

  • Health, dental, and vision insurance with employer HSA contributions.
  • Paid time off, holidays, parental leave, leave-of-absence programs, and volunteer time off.
  • 401(k) plan with company matching up to 4% of salary.
  • Professional development, tuition reimbursement, and mental health and wellness support.
  • Life insurance, disability coverage, global travel insurance, commuter benefits, meals allowance, and a cell phone stipend.

Π‘ΡƒΠ΄ΡŒΡ‚Π΅ остороТны: Ссли Ρ€Π°Π±ΠΎΡ‚ΠΎΠ΄Π°Ρ‚Π΅Π»ΡŒ просит Π²ΠΎΠΉΡ‚ΠΈ Π² ΠΈΡ… систСму, ΠΈΡΠΏΠΎΠ»ΡŒΠ·ΡƒΡ iCloud/Google, ΠΏΡ€ΠΈΡΠ»Π°Ρ‚ΡŒ ΠΊΠΎΠ΄/ΠΏΠ°Ρ€ΠΎΠ»ΡŒ, Π·Π°ΠΏΡƒΡΡ‚ΠΈΡ‚ΡŒ ΠΊΠΎΠ΄/ПО, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡ‚Π΅ этого - это мошСнники. ΠžΠ±ΡΠ·Π°Ρ‚Π΅Π»ΡŒΠ½ΠΎ ΠΆΠΌΠΈΡ‚Π΅ "ΠŸΠΎΠΆΠ°Π»ΠΎΠ²Π°Ρ‚ΡŒΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡˆΠΈΡ‚Π΅ Π² ΠΏΠΎΠ΄Π΄Π΅Ρ€ΠΆΠΊΡƒ. ΠŸΠΎΠ΄Ρ€ΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β†’