Senior Software / Site Reliability Lead Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Senior Software / Site Reliability Lead Engineer (AI): Building and defining the SRE practice for production AI services with an accent on reliability standards, observability, SLOs, and incident response. Focus on designing production-readiness criteria, detecting AI-specific failure modes, and automating operational work across large-scale systems.
Location: 100% remote within the United States; U.S. citizenship and eligibility to obtain a Department of Defense Secret security clearance are required.
Salary: $142,696–$158,303 per year
Company
develops high-technology solutions, products, and services for defense and scientific missions.
What you will do
- Define cross-project reliability standards, SLOs, error budgets, and production-readiness criteria for AI services.
- Build and maintain observability, monitoring, alerting, logging, metrics, tracing, and dashboard systems.
- Own on-call procedures, escalation paths, incident response, post-incident reviews, and reliability improvement backlogs.
- Identify and automate repetitive operational work using scripting and infrastructure-as-code.
- Collaborate with Functional SREs and software development teams to connect reliability metrics with business outcomes and architectural decisions.
- Apply SRE principles to AI-specific risks such as model drift, token budget exhaustion, prompt injection, and upstream data-quality degradation.
Requirements
- Bachelor’s degree in Computer Science, Software Engineering, or a related STEM field with at least 8 years of relevant experience, or a master’s degree with at least 6 years.
- Production SRE or DevOps experience owning the reliability of systems used by real customers.
- Hands-on experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, ELK, or CloudWatch.
- Strong scripting and automation skills with Python, Bash, and infrastructure-as-code tools such as Terraform or CloudFormation.
- Experience with Docker, Kubernetes, container orchestration at scale, SLOs, error budgets, and production incident response.
- U.S. citizenship and the ability to obtain a Department of Defense Secret security clearance are required.
Nice to have
- Experience shipping production AI systems used by real users.
- Software engineering, API design, architecture, and design review experience.
- Defense industry experience.
Culture & Benefits
- 100% telework with a flexible work environment.
- 9/80 work schedule.
- Competitive benefits and recognition for contributions.
- Work focused on defense, scientific, and high-technology missions.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →