Назад
Company hidden
2 часа назад

Software Engineer (AI Infrastructure)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer (AI Infrastructure) (Python/Kubernetes): Building and operating the platform layer behind engineering infrastructure, including CI/CD systems, deployment automation, developer environments, and observability with an accent on distributed workflows, reliability, and scalable infrastructure. Focus on designing Kubernetes platforms, debugging failures across system boundaries, and implementing durable fixes for cloud, on-premises, and specialized hardware environments.

Location: Hybrid, US and Canada offices

Company

AI hardware and software company building large-scale systems for high-speed AI training and inference.

What you will do

  • Design, build, and maintain CI/CD systems for build, testing, integration, qualification, and release workflows.
  • Build and operate Kubernetes-based platforms and services for engineering teams.
  • Develop deployment systems, internal tools, and self-service workflows that make infrastructure changes repeatable and safe.
  • Improve infrastructure reliability, capacity, performance, cost efficiency, monitoring, and operational readiness.
  • Debug failures across CI pipelines, Kubernetes, networking, storage, authentication, operating systems, and distributed applications.
  • Perform root-cause analysis and partner with software, IT, security, networking, release, and developer-productivity teams on scalable infrastructure solutions.

Requirements

  • 5+ years of professional experience in platform engineering, DevOps, infrastructure engineering, site reliability engineering, or software engineering.
  • Hands-on experience with CI/CD pipelines, automated software delivery, Kubernetes, and containerized environments.
  • Experience with a major cloud platform, preferably AWS, and programmatic infrastructure provisioning.
  • Strong Linux or Unix fundamentals and understanding of DNS, routing, load balancing, proxies, ports, TLS, and service connectivity.
  • Proficiency in Python, Shell, or another language used for infrastructure automation and operational tooling.
  • Experience with monitoring, logging, alerting, dashboards, incident investigation, and cross-system debugging.

Nice to have

  • Experience with Terraform, Kubernetes controllers or operators, custom resources, Helm, Argo CD, or similar platform technologies.
  • Experience managing artifact repositories, package registries, build caches, or software-distribution infrastructure.
  • Familiarity with build systems, dependency management, reproducible builds, and internal developer platforms.
  • Experience supporting hybrid cloud, on-premises, and specialized-hardware environments.
  • Experience with identity and access management, secrets, certificates, TLS, or mTLS; a BS/MS in Computer Science or equivalent practical experience.

Culture & Benefits

  • Work on an AI platform designed to go beyond GPU limitations.
  • Contribute to cutting-edge AI research and open-source projects.
  • Work with one of the fastest AI supercomputers in the world.
  • Combine startup vitality with job stability and a non-corporate work culture.
  • Work in an inclusive environment focused on continuous learning, growth, and support.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →