Lead Site Reliability Engineer (Cybersecurity)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Lead Site Reliability Engineer (Cybersecurity): Defining and scaling reliable infrastructure for a secure, mission-critical collaboration platform with an accent on Kubernetes, infrastructure-as-code, observability, and regulated cloud environments. Focus on designing high-availability systems, leading incident management, automating operations, and meeting defense and government compliance requirements.
Location: Remote within the United States; the role supports a remote-first environment and may require U.S. government security clearance and eligibility to access export-controlled information.
Salary: $145,000–$200,000 per year
Company
Secure collaborative workflow software supports defense, intelligence, security, and critical infrastructure organizations through on-premises and private-cloud deployments.
What you will do
- Define the SRE strategy, architecture, and roadmap in alignment with product and business goals.
- Lead the design and optimization of containerized workloads, infrastructure-as-code, and compliant cloud environments.
- Establish observability, monitoring, alerting, capacity planning, and performance optimization practices.
- Drive incident management, on-call operations, root cause analysis, and systemic reliability improvements.
- Partner with security and compliance teams on data sovereignty and regulatory requirements.
- Mentor SRE engineers and build a developer platform that enables secure, reliable software delivery.
Requirements
- 5+ years of experience in SRE, DevOps, or cloud infrastructure roles, with a relevant technical degree or equivalent experience.
- Expertise with container orchestration, ideally Kubernetes, and infrastructure-as-code, ideally Terraform.
- Strong experience with cloud platforms, ideally AWS, plus monitoring, alerting, and distributed-system troubleshooting.
- Proficiency in at least one scripting or programming language for automation.
- Excellent communication skills and experience influencing cross-functional teams.
- Must meet U.S. federal eligibility requirements; the role may require obtaining and maintaining a security clearance and access to export-controlled information.
Nice to have
- Experience with Grafana, Prometheus, GCP, or Azure.
- Experience designing high-availability, disaster-recovery, and scaling architectures.
- Experience in defense, finance, critical infrastructure, or other highly regulated industries.
- Knowledge of FedRAMP, DoD ATO, NIST 800-53, and cloud marketplaces.
- Open-source contributions or certifications such as CKA, CKAD, or AWS Solutions Architect.
Culture & Benefits
- Remote-first work environment.
- Open-source development model.
- Collaboration with globally distributed teams.
- Focus on learning, technical excellence, operational efficiency, and secure software delivery.
- Equal-opportunity workplace with interview accommodations available on request.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →