5 дней назад
Senior Site Reliability Engineer, Platform Infrastructure (AWS/AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer, Platform Infrastructure (AWS/AI): Owning and scaling critical AWS infrastructure with an accent on infrastructure-as-code, observability, and AI-assisted development. Focus on designing reliable platform systems, managing incidents and SLOs, evaluating AI/ML infrastructure, and mentoring engineers across a 24/7 reliability operation.
Location: South Jordan, Utah, USA; in-office at least 4–5 days per week. Relocation assistance is available.
Company
develops smart cutting machines, design applications, and materials that help people create personalized products and home décor.
What you will do
- Own the architecture, reliability, scalability, security, and cost management of critical AWS infrastructure.
- Shape the platform infrastructure roadmap and technical strategy with software engineers and lead engineers.
- Expand infrastructure-as-code practices using tools such as Terraform and CloudFormation.
- Use AI-assisted development tools and evaluate AI/ML infrastructure, including model serving, vector databases, and LLM tooling.
- Lead production monitoring, incident response, blameless postmortems, and SLO/SLI management.
- Consult with engineering teams, mentor infrastructure-focused engineers, and participate in the on-call rotation for 24/7 reliability coverage.
Requirements
- 4+ years of experience in software engineering or site reliability engineering.
- Strong software engineering background, including backend microservices experience; .NET is a plus.
- Deep hands-on experience with AWS services such as EC2, S3, RDS, Kinesis, VPC, and IAM.
- Experience with infrastructure-as-code tools such as Terraform or CloudFormation.
- Experience with SRE principles, observability, monitoring, and logging platforms such as Datadog and OpenTelemetry.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent industry experience; experience with AI-assisted development, prompt engineering, context management, and AI/ML infrastructure.
Culture & Benefits
- Face-to-face collaboration in an in-office environment.
- Medical, dental, and vision coverage.
- 401(k) match, generous paid time off, tuition reimbursement, and an annual lifestyle stipend.
- Employee discounts and a creative, collaborative workplace.
- Employment is contingent on successfully completing a criminal background check; participates in E-Verify.
Hiring process
- Submit a resume, cover letter, portfolio, or links to relevant social profiles.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Site Reliability Engineer Technical Lead (Kubernetes)
100 000 - 150 000$
11 дней назад
Sr. Site Reliability Engineer (AI)
6 дней назад
Site Reliability Engineering Team Lead (Principal SRE, Automotive AI)
132 000 - 211 400$
10 дней назад
Staff Site Reliability Engineer (Cybersecurity)
199 750 - 270 000$
7 дней назад
Senior Site Reliability Engineer, Infrastructure (Kubernetes)
128 000 - 160 000$
6 дней назад
Senior Site Reliability Engineer (AI)
75 000 - 85 000$