11 часов назад
Site Reliability Engineer (Kubernetes)
140 000 - 170 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Kubernetes): Deploying and operating Blitzy's self-hosted AI software development platform inside a secure, customer-controlled cloud environment with an accent on Kubernetes operations, observability, and infrastructure reliability. Focus on hardening restricted deployments, designing monitoring and incident response, benchmarking compute-intensive AI workloads, and partnering directly with customer infrastructure and security teams.
Location: Remote within the U.S., with occasional travel for key customer workshops. U.S. citizenship is required for customer badging.
Salary: $140,000–$170,000 base salary annually, plus bonus and equity commensurate with experience.
Company
is an AI software development platform that autonomously converts enterprise requirements into production-ready code.
What you will do
- Deploy, operate, and maintain the self-hosted platform in a secure customer-controlled cloud environment.
- Own Kubernetes deployments, releases, upgrades, capacity planning, and performance benchmarking for compute-intensive AI workloads.
- Design observability systems covering logging, metrics, tracing, alerting, and incident response within the customer security boundary.
- Partner with customer infrastructure, security, governance, and platform teams on provisioning, reviews, documentation, and escalations.
- Handle sensitive customer data according to security requirements and promote security best practices.
- Feed operational lessons into the product and infrastructure roadmap for future public sector deployments.
Requirements
- U.S. citizenship and ability to complete the customer background and badging process.
- 3+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering.
- Strong Kubernetes and container orchestration experience, including deployments in customer-controlled or restricted environments.
- Experience operating in isolated, restricted, or highly regulated environments such as defense, government, or financial services.
- Hands-on infrastructure-as-code experience with Terraform, Pulumi, or equivalent, plus at least one major cloud platform.
- Expertise in observability, incident management, on-call practices, and scripting with Python, Go, Bash, or similar tools.
Nice to have
- Experience with government-accredited or similarly certified cloud environments.
- Familiarity with sensitive data and security frameworks for regulated industries.
- Experience supporting AI/ML workloads or their infrastructure.
- Forward-deployed, residency, or embedded-engineer experience at an enterprise customer site.
- Experience in a high-growth startup environment.
Culture & Benefits
- Greenfield reliability work shaping the deployment playbook for government and defense customers.
- Direct influence on architecture and close collaboration with engineering, infrastructure, security, and customer teams.
- Fast-paced, innovation-focused environment with strong customer ownership.
- Bonus and equity are provided in addition to the base salary.
- Emphasis on sustainable performance through sleep, movement, and restorative activities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Site Reliability Engineer (Kubernetes)
130 000 - 160 000$
5 часов назад
Infrastructure Engineer (AI)
148 000 - 230 000$
8 часов назад
Senior Software Engineer (AWS/Kubernetes)
160 000 - 215 000$
13 часов назад
Principal Site Reliability Engineer (AWS/Kubernetes)
163 620 - 212 710$
5 часов назад
Senior DevOps Engineer / Site Reliability Engineer (AI)
170 000 - 220 000$
5 часов назад
Platform Engineer (AI)
160 000 - 240 000$