4 дня назад
Platform Support Engineers (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Platform Support Engineers (AI): Supporting and improving hybrid and self-hosted agent observability deployments across AWS, Azure, and GCP with an accent on Kubernetes, Terraform, cloud infrastructure, and backend reliability. Focus on diagnosing complex performance and networking issues, leading customer-impacting incidents, shipping infrastructure fixes, and building self-service diagnostics.
Location: Singapore; remote
Company
provides an agent observability platform for tracing agents, running evaluations, and improving their production performance.
What you will do
- Support hybrid and self-hosted deployments across AWS, Azure, and GCP from installation through ongoing operation.
- Debug Kubernetes workloads, Terraform state, networking, VPC configuration, IAM permissions, TLS, and cloud-provider issues.
- Diagnose backend performance and reliability problems using logs, metrics, and traces.
- Lead incident response for customer-impacting infrastructure issues and participate in the on-call rotation.
- Submit fixes to backend services, Terraform modules, and deployment tooling.
- Build diagnostics, health checks, preflight validation, self-service tools, runbooks, and deployment documentation.
Requirements
- Experience in customer-facing technical support, SRE, DevOps, solutions architecture, infrastructure engineering, or backend/infrastructure engineering.
- Strong Kubernetes fundamentals, including deploying, debugging, and scaling workloads.
- Hands-on Terraform experience and depth in at least one major cloud platform, preferably AWS.
- Comfort working in Python, TypeScript, or Go backend codebases to reproduce and fix issues.
- Fluency with observability tooling and clear communication during high-pressure incidents.
- Ownership of customer problems through resolution.
Nice to have
- Experience supporting self-hosted or on-premises enterprise software in regulated environments.
- Multi-cloud experience with Azure or GCP alongside AWS.
- Experience with Postgres, ClickHouse, or similar analytical data stores.
- Background in observability, ML infrastructure, developer platforms, LLM APIs, or production agent evaluation.
- Experience building support or diagnostic tooling that reduced ticket volume.
Culture & Benefits
- Work on complex infrastructure problems supporting AI product teams.
- Join an early team and help shape operating standards, tooling, and technical quality.
- Work closely with customers and engineering teams with authority to resolve issues in either direction.
- Medical, dental, and vision insurance.
- Flexible time off, daily meals and beverages, a Wi-Fi and cellphone stipend, salary, and equity.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Infrastructure Engineer (GCP/Kubernetes)
9 дней назад
Staff IT Engineer (AI)
200 000 - 240 000$
4 дня назад
Senior Platform Engineering Manager (AI)
Writer
9 дней назад
Infrastructure Engineer (AI)
155 400 - 273 700$
10 дней назад
Senior Manager, Engineering (DevOps, Infrastructure, and Release Engineering)
9 дней назад