обновлено 8 дней назад
Site Reliability Engineer (Infrastructure Platforms)
126 400 - 314 400$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Infrastructure Platforms): Building and operating highly scalable production infrastructure for GitLab's global services with an accent on automation, infrastructure-as-code, and reliability engineering. Focus on designing robust Kubernetes-based systems, optimizing observability, and driving incident response practices to ensure platform stability at scale.
Location: Remote in Canada or the United States only
United States base salary range: $126,400–$314,400 USD annually. The range excludes bonuses, equity, and benefits.
Company
provides an intelligent orchestration platform for DevSecOps that helps organizations improve developer productivity, operational efficiency, security, compliance, and software delivery.
What you will do
- Keep user-facing services and production systems reliable, scalable, and efficient.
- Build automation, infrastructure tooling, and infrastructure-as-code workflows that reduce operational toil.
- Operate and troubleshoot Kubernetes production systems, including deployments, rollouts, and scaling.
- Ship changes safely through CI/CD and GitOps while contributing to observability through metrics, logs, alerts, and SLOs.
- Participate in on-call rotations, incident response, post-incident reviews, and continuous improvement.
- Document runbooks, architecture decisions, and repeatable operational practices.
Requirements
- Experience operating reliable production systems with strong software engineering fundamentals.
- Experience building new infrastructure tooling and automation, such as Terraform modules, Kubernetes operators or controllers, or production services.
- Experience with infrastructure as code and Kubernetes at a depth appropriate to the level being assessed.
- Hands-on experience with at least one major cloud provider: AWS or GCP.
- Ability to read, debug, and reason about code; most teams use Go, with some using Ruby.
- Comfort with on-call work, incident response, observability practices, async collaboration, and manager-of-one responsibilities.
Nice to have
- Experience using AI and automation to reduce toil and improve engineering workflows.
- Experience shaping reliability across multiple teams or setting technical direction at organizational scale.
Culture & Benefits
- Remote-first, globally distributed organization working asynchronously.
- Focus on automation, clear ownership, monitoring, metrics, and continuous reliability improvements.
- Flexible paid time off, parental leave, health and well-being benefits, and equity compensation.
- Growth and Development Fund and Team Member Resource Groups.
Hiring process
- Recruiter screen covering background, expectations, level, and team fit.
- Core technical assessment, hiring manager interview, peer technical interview, and skip-level interview.
- Final level and team placement are calibrated according to interview performance, experience scope, and current hiring needs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Baseten
8 дней назад
Site Reliability Engineer (AI)
165 000 - 330 000$
13 дней назад
Senior Site Reliability Engineer (Healthcare)
200 000 - 240 000$
Baseten
8 дней назад
Site Reliability Engineer (AI)
165 000 - 330 000$
Replit
10 дней назад
Engineering Manager (SRE)
250 000 - 325 000$
11 дней назад
Senior SRE (Kubernetes)
150 000 - 170 000$
12 дней назад
Lead Site Reliability Engineer (Cybersecurity)
145 000 - 200 000$