2 дня назад
Senior AI Site Reliability Engineer (AI.x)
170 000 - 220 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior AI Site Reliability Engineer (AI.x) (GenAI/SRE): Building reliable, scalable, and secure production infrastructure for GenAI applications and AI platforms with an accent on automation, CI/CD, observability, and infrastructure as code. Focus on designing self-healing systems, minimizing MTTD and MTTR, managing SLOs and error budgets, and supporting 24/7 incident response.
Location: Austin, Texas, United States; hybrid work with regular in-person collaboration
Salary: USD $170,000–$220,000 per year, plus bonus or incentive opportunities
Company
is a financial services company whose AI Strategy & Transformation team, AI.x, develops enterprise AI platforms and next-generation generative AI solutions.
What you will do
- Lead automation-first reliability initiatives, including self-healing systems and toil reduction across AI.x platforms.
- Design CI/CD pipelines with automated testing, validation, rollback, and one-touch deployment capabilities.
- Build observability frameworks for AI services using metrics, logs, traces, intelligent alerting, and automated diagnostics.
- Provide 24/7 production support through on-call rotations, incident response, root cause analysis, and resolution.
- Define and manage SLOs, SLIs, error budgets, runbooks, monitoring, and incident management practices.
- Champion infrastructure as code, cloud-native operations, capacity planning, and collaboration with AI Engineering teams.
Requirements
- 8+ years of software engineering experience, including 4+ years as a hands-on Site Reliability Engineer.
- Bachelor’s degree in Computer Science or a related field, or equivalent experience.
- 5+ years building products from scratch, operating them in production, and ensuring reliability.
- 3+ years working with containers, public cloud, infrastructure as code, and CI/CD pipelines.
- 3+ years of experience in high-availability hybrid-cloud environments.
- Experience with observability, incident management, distributed systems, and reliability engineering principles.
Nice to have
- Experience deploying and maintaining LLM-powered applications using platforms or models such as Gemini, Claude, or OpenAI.
- Experience with Terraform and Google Cloud Platform.
- Strong communication skills and the ability to troubleshoot complex problems with incomplete or ambiguous data.
Culture & Benefits
- Hybrid work and flexibility balanced with regular in-person collaboration.
- 401(k) with company match and employee stock purchase plan.
- Health, dental, and vision insurance.
- Paid vacation, volunteering time, parental leave, and family-building benefits.
- Tuition reimbursement and a 28-day sabbatical after five years for eligible positions.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Senior DevOps Platform Engineer (AI)
140 000 - 180 000$
5 дней назад
Senior Software / Site Reliability Lead Engineer (AI)
142 696 - 158 303$
2 дня назад
Senior Site Reliability Engineer (GovCloud)
117 000 - 209 330$
5 дней назад
Senior AI DevOps Developer (AI)
150 000 - 206 000$
3 дня назад
Sr. DevOps Engineer (AI)
175 000 - 195 000$
3 дня назад
Developer Experience Engineer (AI/HPC)
150 000 - 275 000$