3 дня назад
Site Reliability Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AWS/DevOps): Building and operating reliable, secure, high-availability delivery infrastructure for AI defense platforms with an accent on observability, incident response, and deployment automation. Focus on designing monitoring and self-service tooling, scaling production systems, automating CI/CD pipelines, and resolving complex reliability issues in high-side environments.
Location: Remote from anywhere in the United States; willingness and ability to work on-site in San Diego, California is required.
Company
is a defense technology company building agentic AI systems that model adversary behavior, simulate campaigns, and support decision-making.
What you will do
- Monitor dashboards and system telemetry to identify health issues, performance degradation, and reliability risks.
- Own debugging and incident response from the initial alert through resolution, including escalation and post-incident learning.
- Build logging, monitoring, and observability tooling while improving SRE practices.
- Maintain platform health, scaling, capacity planning, and secure high-availability delivery pipelines.
- Improve deployment processes, automate build and deployment pipelines, and develop self-service engineering tools.
- Communicate system status, technical trade-offs, and incident findings with engineering teams and stakeholders.
Requirements
- 5+ years of experience in SRE, DevOps, or software engineering.
- Production monitoring, incident response, on-call, and post-mortem experience.
- Experience with PLG, Datadog, or comparable enterprise monitoring and observability tools.
- Knowledge of AWS, infrastructure as code such as Terraform or Pulumi, and scripting with Python, Bash, or similar languages.
- Ability to work independently in an agile Scrum environment and communicate clearly during live troubleshooting.
- U.S. citizenship and an active TS/SCI clearance are required; the role also requires the ability to work on-site in San Diego, California.
Nice to have
- Experience with SLOs, SLIs, error budgets, CI/CD automation, Docker, AWS GovCloud, and modern web services architectures.
- Experience with relational databases, SQL, Elasticsearch, or OpenSearch.
- Experience collaborating across engineering teams on cross-functional projects.
Culture & Benefits
- Remote-first culture with work available anywhere in the United States and WeWork access for full-time employees.
- Health, dental, and vision insurance.
- Unlimited PTO with vacation and holiday schedules.
- Monthly mental health, wellness, and fitness stipends, plus home-office and family-planning assistance.
- Paid parental leave, military reserve duty salary top-up, and child and pet care reimbursement during travel.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Site Reliability Engineer, Tech Lead (AI)
10 дней назад
Staff Site Reliability Engineer (Kubernetes)
Phantom
6 дней назад
Staff Software Engineer (SRE) (Crypto)
200 000 - 250 000$
5 дней назад
Site Reliability Engineer (Kubernetes)
180 000 - 220 000$
4 дня назад
Sr Software Development Engineer, SRE (US Federal)
163 800 - 245 800$
5 дней назад
Senior DevOps / Site Reliability Engineer (SRE) (Cybersecurity)
165 000 - 215 000$