Senior Infrastructure Software Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Location: Hybrid role based in one of the New York City, San Francisco, Seattle, or London hubs, with a minimum of two in-office days per week. Occasional team and company offsites are required. Visa sponsorship is not available.
Annual base salary: $180,000–$220,000 USD, plus discretionary bonus, equity, and benefits.
Company
builds an end-to-end platform for developing, training, and deploying AI systems, combining developer-focused software with large-scale GPU and bare-metal infrastructure.
What you will do
- Design, build, and operate production services, APIs, tooling, and automation for large-scale bare-metal and GPU infrastructure.
- Develop systems for server discovery, provisioning, configuration, validation, capacity deployment, and lifecycle management.
- Integrate software with hardware management and provisioning systems to reduce manual operational work.
- Build telemetry, logging, observability, and diagnostic tools for infrastructure and hardware health.
- Translate recurring operational issues and hardware failure modes into improvements in software, tooling, and automation.
- Collaborate with Networking, Infrastructure Operations, Data Center, and Platform Engineering teams on architecture and technical direction.
Requirements
- 8+ years of professional software engineering, infrastructure engineering, or related experience.
- Strong software engineering fundamentals and production backend development experience in Python or a similar object-oriented language.
- Strong experience with Linux in production environments.
- Experience building APIs, tooling, or automation for infrastructure at scale.
- Familiarity with containerization and orchestration concepts, plus HPC and bare-metal infrastructure fundamentals.
- Ability to make pragmatic architecture decisions and work with a high degree of ownership and autonomy in a fast-paced environment.
Nice to have
- Experience with PXE/iPXE, BMC, Redfish, IPMI, Dell hardware, GPU servers, or bare-metal troubleshooting and provisioning.
- Experience with network switches, routers, and firewalls, including SONiC, Palo Alto, or Juniper Networks.
- Experience with high-performance storage systems such as VAST.
- Experience supporting AI/ML or HPC infrastructure at scale.
Culture & Benefits
- Ownership-focused environment emphasizing urgency, open communication, continuous improvement, and long-term thinking.
- Medical, dental, and vision coverage for employees and eligible dependents.
- RSUs, retirement matching in the U.S. or pension contributions in the U.K., and flexible time off.
- Two-week company-wide winter break, paid parental and family leave, and a four-week paid sabbatical after four years.
- Annual learning and development allowance, wellness and work-from-home stipends, flexible schedules, and a hybrid work model.
- Complimentary meals at office hubs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →