6 часов назад
Platform Support Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Platform Support Engineer (AI): Supporting and improving hybrid and self-hosted AI observability deployments across AWS, Azure, and GCP with an accent on Kubernetes, Terraform, backend debugging, and customer-facing incident response. Focus on diagnosing infrastructure and performance failures, shipping fixes to backend and deployment tooling, and building diagnostics that improve reliability and customer self-service.
Location: Hybrid in San Francisco, New York City, or Seattle
Company
provides an agent observability platform that helps teams understand agent behavior in production, identify critical patterns, and improve agents through tracing and evaluations.
What you will do
- Own customer-facing support for hybrid and self-hosted deployments across AWS, Azure, and GCP.
- Debug Kubernetes workloads, Terraform state, networking, VPC configuration, IAM, TLS, and cloud-provider issues.
- Diagnose backend performance and reliability problems using logs, metrics, and traces.
- Lead incident response for customer-impacting issues and communicate clearly through resolution.
- Ship fixes to backend services, Terraform modules, and deployment tooling.
- Build diagnostics, health checks, preflight validation, self-service tools, runbooks, and deployment documentation.
Requirements
- Experience in a customer-facing technical role such as Support Engineering, SRE, DevOps, Solutions Architecture, or Infrastructure Engineering.
- Strong Kubernetes fundamentals, including deploying, debugging, and scaling workloads.
- Hands-on Terraform experience and depth in at least one major cloud platform; AWS is strongly preferred.
- Backend development experience with Python, TypeScript, or Go sufficient to reproduce, trace, and fix bugs.
- Fluency with observability tooling and clear, calm communication during customer-impacting incidents.
- Willingness to participate in an on-call rotation for critical customer issues.
Nice to have
- Experience supporting self-hosted or on-premises enterprise software in regulated environments.
- Multi-cloud experience, particularly with Azure or GCP alongside AWS.
- Experience with Postgres, ClickHouse, or similar analytical data stores.
- Background in observability, ML infrastructure, developer platforms, or production agent development.
- Experience building support or diagnostic tooling that reduced ticket volume.
Culture & Benefits
- Work on complex infrastructure problems supporting AI product teams.
- Join an early-stage team and help shape its standards, tooling, and operating practices.
- Work closely with customers and engineering teams, with ownership to fix problems in either direction.
- Medical, dental, and vision insurance.
- Flexible time off, daily lunch and refreshments, Wi-Fi and cellphone stipend, salary, and equity.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Member of Technical Staff, Infrastructure Engineer (AI)
175 000 - 240 000$
5 дней назад
Staff IT Engineer (AI)
200 000 - 240 000$
Writer
5 дней назад
Infrastructure Engineer (AI)
155 400 - 273 700$
6 дней назад
Principal AI Infrastructure & Cloud Engineer
5 дней назад
Software Engineer (Cloud)
175 000 - 240 000$
7 дней назад
DevOps Engineer (AI)
163 000 - 204 000$