3 дня назад
Site Reliability Engineer III (AI/FinOps)
159 500 - 176 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer III (AI/FinOps): Building telemetry, profiling, CI/CD guardrails, and automated rollout systems that optimize GCP and AI/SaaS costs while preserving application performance and reliability with an accent on GKE, BigQuery, OpenTelemetry, and SLO-aware FinOps. Focus on validating AI-generated optimization playbooks, designing canary analysis and rollback mechanisms, and solving latency, memory, reliability, and cloud-spending regressions.
Location: Boston, Massachusetts, United States
Annual base pay: USD 159,500–176,000
Company
is an e-commerce technology company operating a large online destination for home products.
What you will do
- Map application code paths, microservices, middleware, databases, and infrastructure to cloud costs and service-level objectives.
- Use APM, continuous profiling, distributed tracing, and OpenTelemetry to measure latency, throughput, error rates, and optimization impact.
- Validate AI-generated pull requests and optimization playbooks from Devin, Cursor, and custom agents for performance, reliability, and infrastructure safety.
- Build SLO-aware CI/CD guardrails, Terraform and Helm linting rules, canary analysis, staged rollouts, and automated rollback triggers.
- Investigate performance regressions and cloud-spending anomalies through root-cause analysis.
- Measure cost-efficiency gains, developer adoption, GCP and SaaS savings, and SLO stability.
Requirements
- 5+ years of experience in SRE, production engineering, systems or performance engineering, or software engineering in scaled production environments.
- Deep hands-on experience with Google Cloud Platform, including GKE, Kubernetes, BigQuery, and cloud resource tuning.
- Proficiency with Terraform, Helm, and automated deployment pipelines such as GitHub Actions, Buildkite, or Jenkins.
- Experience with APM, continuous profiling, distributed tracing, OpenTelemetry, Datadog, Dynatrace, or Google Cloud Operations, plus SLO and SLI definition.
- Strong Python, Go, or Bash scripting skills for operational automation, telemetry parsing, and API integrations.
- Practical experience using AI coding assistants and understanding of cloud cost drivers, resource allocation, and availability-versus-cost trade-offs.
Culture & Benefits
- Work with FinOps and Technology Cost Value Management systems across a multi-million-dollar GCP and AI/SaaS environment.
- Access a comprehensive package of medical, financial, and additional employee benefits.
- Equity, bonuses, commissions, or other compensation may be included depending on the position offered.
- supports equal opportunity and provides reasonable accommodations during the application and interview process.
Hiring process
- Live interview evaluation covers system architecture and dependency mapping, automation and coding, AI-agentic code review, and FinOps trade-off analysis.
- Evaluation includes Python or Go scripting, distributed tracing, SLO design, canary analysis, rollback strategies, and identification of unsafe AI-generated changes.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior Site Reliability Engineer (Systems Engineer III) (GCP)
140 000 - 150 000$
8 дней назад
Senior Staff Site Reliability Engineer
181 000 - 263 000$
5 дней назад
Site Reliability Engineer (AI/ML)
5 дней назад
Senior Engineer Site Reliability (AWS/Data Operations)
105 700 - 149 275$
Okta
10 дней назад
Staff Site Reliability Engineer (Splunk)
194 000 - 267 000$
3 дня назад
Senior Site Reliability Engineer (SRE)
120 000 - 175 000$