Назад
Company hidden
3 дня назад

Senior DevOps Engineer, AI Platform

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior DevOps Engineer, AI Platform (AWS/AI infrastructure): Operating multi-region AI runtime infrastructure for customer-facing accounting products with an accent on AWS Bedrock, sandboxed code execution, observability, cost attribution, and tenant isolation. Focus on designing reliable serving and delivery systems across US, EU, and AU regions, enforcing Terraform-based infrastructure, and detecting model, queue, and silent quality failures.

Location: San Jose, California; hybrid workplace

Company

hirify.global develops SaaS products for accounting workflows, including Transform, AI Matching, and AutoBuilder.

What you will do

  • Own the AWS Bedrock and Bedrock AgentCore runtime across US, EU, and AU regions, including model access, throughput, quotas, throttling, retries, and regional availability.
  • Operate sandboxed execution environments for AI-generated code with lifecycle limits, network controls, least-privilege IAM, and tenant isolation.
  • Build and review Terraform for a multi-account, multi-region AWS estate running on ECS/Fargate and Lambda.
  • Extend Grafana observability with token usage, latency, throttling, retries, tool-call failures, sandbox outcomes, generation success, and end-to-end traces.
  • Manage AI cost attribution, capacity planning, SLOs, CI/CD gates, feature-flagged releases, and rollback workflows.
  • Join the DevOps on-call rotation and create runbooks for model throttling, sandbox exhaustion, queue failures, and silent degradation.

Requirements

  • 5+ years of experience in DevOps, SRE, platform, or infrastructure engineering with production on-call ownership.
  • Deep hands-on AWS experience with ECS/Fargate, Lambda, SQS, S3, IAM, VPC, networking, and ALB/NLB.
  • Production-scale Terraform experience, including modules, state management, multi-region, and multi-account infrastructure.
  • Experience operating an LLM-backed or ML-serving workload in production, with practical knowledge of tokens, latency, throttling, and cost.
  • Experience with CI/CD, Docker, GitHub Actions, Grafana, Prometheus or OpenTelemetry, distributed tracing, and user-journey SLOs.
  • Working fluency in Python or TypeScript/Node.js and the ability to read and debug the other language.

Nice to have

  • Multi-region infrastructure under US/EU/AU data-residency constraints.
  • AI cost and performance optimization, sandboxed untrusted-code execution, FinOps, or per-tenant attribution.
  • Atmos or comparable Terraform orchestration, NX/Turborepo/Bazel, and Harness feature flags.
  • Exposure to MongoDB, PostgreSQL, Snowflake, EMR/Spark, or regulated SaaS domains such as fintech, accounting, or healthcare.
  • SOC 2 or ISO 27001 audit evidence experience.

Culture & Benefits

  • Embedded collaboration with the Transform and Close AI engineering pods.
  • Production systems are operated with an emphasis on reliability, observability, security, compliance, and cost control.
  • AI infrastructure is deployed across US, EU, and AU regions and defined entirely in Terraform.
  • Model and prompt changes are versioned, evaluated in CI, feature-flagged, and reversible.

Hiring process

  • 30-minute recruiter screen followed by approximately four hours of interviews.
  • Hiring manager discussion, technical deep dive with Terraform and service diagnosis, and AI infrastructure design exercise.
  • Team panel with engineering, security or compliance, and engineering-values discussions.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →