Назад
Company hidden
6 часов назад

Platform Support Engineer (AI)

Формат работы
hybrid
Тип работы
fulltime
Грейд
middle/senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Platform Support Engineer (AI): Supporting and improving hybrid and self-hosted AI observability deployments across AWS, Azure, and GCP with an accent on Kubernetes, Terraform, backend debugging, and customer-facing incident response. Focus on diagnosing infrastructure and performance failures, shipping fixes to backend and deployment tooling, and building diagnostics that improve reliability and customer self-service.

Location: Hybrid in San Francisco, New York City, or Seattle

Company

hirify.global provides an agent observability platform that helps teams understand agent behavior in production, identify critical patterns, and improve agents through tracing and evaluations.

What you will do

  • Own customer-facing support for hybrid and self-hosted deployments across AWS, Azure, and GCP.
  • Debug Kubernetes workloads, Terraform state, networking, VPC configuration, IAM, TLS, and cloud-provider issues.
  • Diagnose backend performance and reliability problems using logs, metrics, and traces.
  • Lead incident response for customer-impacting issues and communicate clearly through resolution.
  • Ship fixes to backend services, Terraform modules, and deployment tooling.
  • Build diagnostics, health checks, preflight validation, self-service tools, runbooks, and deployment documentation.

Requirements

  • Experience in a customer-facing technical role such as Support Engineering, SRE, DevOps, Solutions Architecture, or Infrastructure Engineering.
  • Strong Kubernetes fundamentals, including deploying, debugging, and scaling workloads.
  • Hands-on Terraform experience and depth in at least one major cloud platform; AWS is strongly preferred.
  • Backend development experience with Python, TypeScript, or Go sufficient to reproduce, trace, and fix bugs.
  • Fluency with observability tooling and clear, calm communication during customer-impacting incidents.
  • Willingness to participate in an on-call rotation for critical customer issues.

Nice to have

  • Experience supporting self-hosted or on-premises enterprise software in regulated environments.
  • Multi-cloud experience, particularly with Azure or GCP alongside AWS.
  • Experience with Postgres, ClickHouse, or similar analytical data stores.
  • Background in observability, ML infrastructure, developer platforms, or production agent development.
  • Experience building support or diagnostic tooling that reduced ticket volume.

Culture & Benefits

  • Work on complex infrastructure problems supporting AI product teams.
  • Join an early-stage team and help shape its standards, tooling, and operating practices.
  • Work closely with customers and engineering teams, with ownership to fix problems in either direction.
  • Medical, dental, and vision insurance.
  • Flexible time off, daily lunch and refreshments, Wi-Fi and cellphone stipend, salary, and equity.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →