обновлено 15 часов назад
Senior DevOps / Platform Reliability Engineer (AI Platform)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior DevOps / Platform Reliability Engineer (AWS/AI Platform): Building and operating the CI/CD, infrastructure, observability, and security backbone for an agentic CX platform with an accent on Kubernetes, AWS reliability, and AI-native operations. Focus on designing safe auto-remediation agents, operating production data and event systems, and implementing guardrails, disaster recovery, compliance, and cost controls.
Location: East Coast, United States; remote role
Salary: Competitive compensation package; exact amount not specified
Company
develops an intelligent process automation platform that improves customer experience operations and supports enterprise self-service and agent productivity.
What you will do
- Own and evolve CI/CD pipelines for microservices and agentic workloads using GitHub Actions and OIDC-based authentication.
- Provision infrastructure with Terraform and CloudFormation, and operate production EKS and Argo CD environments.
- Manage AWS networking, Cloudflare, CloudFront, API Gateway, load balancers, DNS, security controls, and multi-tenant platform isolation.
- Operate Aurora MySQL, Redis, S3, Kafka/MSK, Lambda, and event-driven workloads, including backups, point-in-time recovery, and multi-AZ disaster recovery.
- Build observability with Prometheus, Grafana, and OpenTelemetry, including telemetry for LLM and agentic systems.
- Design AI-native DevOps capabilities such as auto-remediation agents, AI-assisted incident analysis, MCP integrations, anomaly detection, and operational guardrails.
Requirements
- 5+ years of experience in DevOps, SRE, or Platform Engineering operating production systems on AWS.
- Strong experience with CI/CD, GitHub Actions, Terraform, production EKS, Kubernetes operations, and Argo CD.
- Hands-on AWS networking experience covering VPCs, subnets, routing, security groups, NACLs, Route 53, ACM, and load balancers.
- Experience with Aurora or RDS MySQL, Redis, S3, Kafka/MSK, Lambda, backups, migrations, lifecycle management, and disaster recovery.
- Strong observability and security experience with Prometheus, Grafana, OpenTelemetry, IAM, KMS, secrets management, vulnerability scanning, and software supply chain security.
- Comfortable working with Python, Bash, and Linux systems.
Nice to have
- Experience operating LLM or ML workloads in production, including LiteLLM, Bedrock, pgvector, prompt caching, or evaluation systems.
- Experience building MCP servers or deploying LangGraph or CrewAI agent frameworks in production.
Culture & Benefits
- Automation is preferred over repetitive operational toil.
- Small team with high ownership and responsibility for defining platform standards.
- Human review is maintained for risky AI-assisted actions, with blameless incident reviews and documented decisions.
- Flexible remote work from anywhere, with a co-working reimbursement of up to $200 per month.
- Health, dental, and vision benefits with 100% of employee premiums and 75%–80% of most dependent premiums covered.
- 401(k), paid parental leave, unlimited PTO, and home office support of up to $500 plus $100 per month for internet, phone, and related expenses.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Staff Site Reliability Engineer (Linux/Network Troubleshooting/Scripting) (AI)
3 дня назад
DevOps Engineer (AI)
15 часов назад
Tech Lead, DevOps Engineer (AWS)
160 000 - 180 000$
2 дня назад
Staff Platform Infrastructure Engineer (AI)
180 000 - 220 000$
1 день назад
Senior Site Reliability Engineer (Cloud-Native Infrastructure)
142 800 - 178 500$
3 дня назад