Назад
Company hidden
21 час назад

Senior Software / Site Reliability Lead Engineer (AI)

142 696 - 158 303$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior/lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Senior Software / Site Reliability Lead Engineer (AI): Building and defining the SRE practice for production AI services with an accent on reliability standards, observability, SLOs, and incident response. Focus on designing production-readiness criteria, detecting AI-specific failure modes, and automating operational work across large-scale systems.

Location: 100% remote within the United States; U.S. citizenship and eligibility to obtain a Department of Defense Secret security clearance are required.

Salary: $142,696–$158,303 per year

Company

hirify.global develops high-technology solutions, products, and services for defense and scientific missions.

What you will do

  • Define cross-project reliability standards, SLOs, error budgets, and production-readiness criteria for AI services.
  • Build and maintain observability, monitoring, alerting, logging, metrics, tracing, and dashboard systems.
  • Own on-call procedures, escalation paths, incident response, post-incident reviews, and reliability improvement backlogs.
  • Identify and automate repetitive operational work using scripting and infrastructure-as-code.
  • Collaborate with Functional SREs and software development teams to connect reliability metrics with business outcomes and architectural decisions.
  • Apply SRE principles to AI-specific risks such as model drift, token budget exhaustion, prompt injection, and upstream data-quality degradation.

Requirements

  • Bachelor’s degree in Computer Science, Software Engineering, or a related STEM field with at least 8 years of relevant experience, or a master’s degree with at least 6 years.
  • Production SRE or DevOps experience owning the reliability of systems used by real customers.
  • Hands-on experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, ELK, or CloudWatch.
  • Strong scripting and automation skills with Python, Bash, and infrastructure-as-code tools such as Terraform or CloudFormation.
  • Experience with Docker, Kubernetes, container orchestration at scale, SLOs, error budgets, and production incident response.
  • U.S. citizenship and the ability to obtain a Department of Defense Secret security clearance are required.

Nice to have

  • Experience shipping production AI systems used by real users.
  • Software engineering, API design, architecture, and design review experience.
  • Defense industry experience.

Culture & Benefits

  • 100% telework with a flexible work environment.
  • 9/80 work schedule.
  • Competitive benefits and recognition for contributions.
  • Work focused on defense, scientific, and high-technology missions.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →