2 часа назад
Site Reliability Engineer (SRE) II
124 800 - 187 200$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (SRE) II (AWS/Kubernetes): Maintaining 24/7 availability and security of hybrid infrastructure spanning bare-metal RHEL servers, AWS GovCloud, and EKS with an accent on observability, incident response, and FedRAMP compliance. Focus on troubleshooting production telemetry, securing and patching systems, operating GitOps deployments, and resolving complex cloud, Kubernetes, networking, and hardware issues.
Location: San Jose; hybrid schedule with 24/7/365 day, evening, night, weekend, holiday, and on-call rotations. US citizenship is required due to FedRAMP and AWS GovCloud security requirements.
Annual base pay: $124,800–$187,200.
Company
develops and secures application delivery solutions that help organizations run applications across evolving digital environments.
What you will do
- Monitor production telemetry across bare-metal RHEL servers, AWS environments, and EKS clusters using Prometheus, Alertmanager, and Grafana.
- Triage incidents, investigate root causes through Elasticsearch and Kibana, coordinate Slack incident bridges, and document post-incident reviews.
- Administer on-premises RHEL hardware, including networking, RAID, storage, lifecycle maintenance, and vendor hardware replacements.
- Operate AWS and Kubernetes infrastructure, troubleshoot L4/L7 load balancing, manage certificates and traffic routing, and apply Terraform changes.
- Execute GitLab CI/CD, ArgoCD, and Argo Workflows deployments, including Helm and Kustomize-based releases and compliance validation.
- Maintain runbooks and automate recurring operational work with Bash and Python.
Requirements
- 3–5 years of experience in SRE, DevOps, systems administration, or hybrid infrastructure support.
- Availability for rotating 24/7 shifts, including nights, weekends, holidays, and on-call escalations.
- Hands-on administration of bare-metal RHEL servers, including IPMI, iDRAC, iLO, RAID, LVM, hardware diagnostics, and NIC bonding.
- Experience with AWS services, AWS GovCloud, Amazon EKS/Kubernetes, secure ALB/NLB and ingress troubleshooting, and Linux systems administration.
- Experience with Kibana, Elasticsearch, Prometheus, Grafana, GitLab CI/CD, ArgoCD, Argo Workflows, and Slack ChatOps.
- Knowledge of FedRAMP, NIST SP 800-53, CVE remediation, DISA STIG hardening, FIPS 140, RBAC, and SSH key management.
Nice to have
- Python or Bash automation experience.
- Terraform experience and RHCSA, RHCE, AWS SysOps Administrator, or CKA certification.
Culture & Benefits
- Hybrid work in a structured 24/7 operations environment.
- Exposure to on-premises data centers, AWS Commercial, AWS GovCloud, and Kubernetes infrastructure.
- Opportunity to work on cybersecurity, compliance, reliability, and continuous delivery operations.
- Eligible compensation may include incentive compensation, bonuses, restricted stock units, and benefits.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Site Reliability Engineer (AI Infrastructure)
175 000 - 265 000$
5 часов назад
Sr. Staff Site Reliability Engineer-Federal, Security Clearance
164 000 - 205 000$
5 дней назад
Principal Site Reliability Engineer (GCP)
151 000 - 244 200$
Okta
5 дней назад
Staff Site Reliability Engineer (Splunk)
194 000 - 267 000$
6 дней назад
Sr Software Development Engineer, SRE (US Federal)
163 800 - 245 800$
Okta
5 дней назад
Staff Site Reliability Engineer (Splunk)
194 000 - 267 000$