5 дней назад
Senior Site Reliability Engineer I (AI)
95 300 - 158 800$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer I (AI) (AWS/Terraform): Building and operating resilient, scalable infrastructure and services for mission-critical platforms with an accent on observability, automation, incident response, and AI tooling deployment. Focus on managing multi-account, multi-region AWS environments, designing reliable Terraform and CI/CD workflows, and testing service availability, recoverability, and disaster readiness.
Location: Philadelphia, PA (Market St), United States
Salary: $95,300–$158,800 annual base pay, plus eligibility for an annual incentive bonus. Geographic differentials may apply.
Company
provides scientific, technical, and medical research content, publishing, analytics, and digital platforms such as ScienceDirect, Scopus, and Mendeley.
What you will do
- Build monitoring queries and service-level baselines for critical platforms and services.
- Automate operational tasks and execute production changes to improve reliability and reduce toil.
- Support incident response, root cause analyses, post-mortems, rollback decisions, and operational runbooks.
- Participate in disaster recovery and availability, reliability, and recoverability testing.
- Deploy, monitor, and maintain services integrating AI tools and APIs.
- Support infrastructure topology design, deployment workflows, documentation, handover, and developer enablement.
Requirements
- Advanced Terraform expertise, including modules, providers, state management, lifecycle controls, drift detection, safe refactoring, remote state, locking, and cross-stack dependencies.
- Hands-on production experience with multi-account, multi-region AWS environments, including ECS, RDS, ALB, VPC, IAM, Route53, ECR, S3, Lambda, DynamoDB, SQS, Secrets Manager, KMS, and CloudWatch.
- Experience building and troubleshooting reusable GitHub Actions CI/CD workflows, OIDC authentication, approval gates, runners, Terraform deployments, application deployments, and migration pipelines.
- Knowledge of Docker, ECS Fargate, ECR, task definitions, IAM roles, health checks, autoscaling, ALB integration, and deployment rollbacks.
- Strong Linux and Git fundamentals with Bash or Python scripting for AWS CLI automation, CI/CD, and operational tooling.
- Experience operating AI services and APIs in production, including monitoring, reliability, security, and cross-layer troubleshooting.
Culture & Benefits
- Work in cross-functional Embedded Innovation Teams that turn internal AI experimentation into validated, reusable solutions.
- Collaborate directly with segment and function teams to improve internal workflows and customer outcomes.
- Country-specific benefits are available.
- The hiring process provides accommodations for candidates with disabilities or other needs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Site Reliability Engineer (AI Infrastructure)
175 000 - 265 000$
7 дней назад
Senior Site Reliability Engineer, Infrastructure (Kubernetes)
128 000 - 160 000$
11 дней назад
Site Reliability Engineering (SRE) Manager
106 000 - 130 600$
5 дней назад
Senior Site Reliability Engineer (AI)
75 000 - 85 000$
10 дней назад
Sr. Site Reliability Engineer (AI)
10 дней назад
Senior Manager - Site Reliability Engineering (SRE)
9 458 - 16 551$