2 часа назад
Software Engineer (AI Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Software Engineer (AI Infrastructure) (Python/Kubernetes): Building and operating the platform layer behind engineering infrastructure, including CI/CD systems, deployment automation, developer environments, and observability with an accent on distributed workflows, reliability, and scalable infrastructure. Focus on designing Kubernetes platforms, debugging failures across system boundaries, and implementing durable fixes for cloud, on-premises, and specialized hardware environments.
Location: Hybrid, US and Canada offices
Company
AI hardware and software company building large-scale systems for high-speed AI training and inference.
What you will do
- Design, build, and maintain CI/CD systems for build, testing, integration, qualification, and release workflows.
- Build and operate Kubernetes-based platforms and services for engineering teams.
- Develop deployment systems, internal tools, and self-service workflows that make infrastructure changes repeatable and safe.
- Improve infrastructure reliability, capacity, performance, cost efficiency, monitoring, and operational readiness.
- Debug failures across CI pipelines, Kubernetes, networking, storage, authentication, operating systems, and distributed applications.
- Perform root-cause analysis and partner with software, IT, security, networking, release, and developer-productivity teams on scalable infrastructure solutions.
Requirements
- 5+ years of professional experience in platform engineering, DevOps, infrastructure engineering, site reliability engineering, or software engineering.
- Hands-on experience with CI/CD pipelines, automated software delivery, Kubernetes, and containerized environments.
- Experience with a major cloud platform, preferably AWS, and programmatic infrastructure provisioning.
- Strong Linux or Unix fundamentals and understanding of DNS, routing, load balancing, proxies, ports, TLS, and service connectivity.
- Proficiency in Python, Shell, or another language used for infrastructure automation and operational tooling.
- Experience with monitoring, logging, alerting, dashboards, incident investigation, and cross-system debugging.
Nice to have
- Experience with Terraform, Kubernetes controllers or operators, custom resources, Helm, Argo CD, or similar platform technologies.
- Experience managing artifact repositories, package registries, build caches, or software-distribution infrastructure.
- Familiarity with build systems, dependency management, reproducible builds, and internal developer platforms.
- Experience supporting hybrid cloud, on-premises, and specialized-hardware environments.
- Experience with identity and access management, secrets, certificates, TLS, or mTLS; a BS/MS in Computer Science or equivalent practical experience.
Culture & Benefits
- Work on an AI platform designed to go beyond GPU limitations.
- Contribute to cutting-edge AI research and open-source projects.
- Work with one of the fastest AI supercomputers in the world.
- Combine startup vitality with job stability and a non-corporate work culture.
- Work in an inclusive environment focused on continuous learning, growth, and support.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 часов назад
Senior Platform Engineer (AI)
2 дня назад
Senior AI DevOps Developer (AI)
150 000 - 206 000$
1 час назад
Data Platform Engineer, Infrastructure (Robotics)
3 часа назад
Production Engineer (AI Infrastructure)
172 000 - 209 000$
Windsurf
6 дней назад
Site Reliability Engineer (AI)
5 часов назад
Senior Site Reliability Engineer (AI Infrastructure)
215 000 - 275 000$