Назад
Company hidden
обновлено 15 часов назад

Senior DevOps / Platform Reliability Engineer (AI Platform)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior DevOps / Platform Reliability Engineer (AWS/AI Platform): Building and operating the CI/CD, infrastructure, observability, and security backbone for an agentic CX platform with an accent on Kubernetes, AWS reliability, and AI-native operations. Focus on designing safe auto-remediation agents, operating production data and event systems, and implementing guardrails, disaster recovery, compliance, and cost controls.

Location: East Coast, United States; remote role

Salary: Competitive compensation package; exact amount not specified

Company

hirify.global develops an intelligent process automation platform that improves customer experience operations and supports enterprise self-service and agent productivity.

What you will do

  • Own and evolve CI/CD pipelines for microservices and agentic workloads using GitHub Actions and OIDC-based authentication.
  • Provision infrastructure with Terraform and CloudFormation, and operate production EKS and Argo CD environments.
  • Manage AWS networking, Cloudflare, CloudFront, API Gateway, load balancers, DNS, security controls, and multi-tenant platform isolation.
  • Operate Aurora MySQL, Redis, S3, Kafka/MSK, Lambda, and event-driven workloads, including backups, point-in-time recovery, and multi-AZ disaster recovery.
  • Build observability with Prometheus, Grafana, and OpenTelemetry, including telemetry for LLM and agentic systems.
  • Design AI-native DevOps capabilities such as auto-remediation agents, AI-assisted incident analysis, MCP integrations, anomaly detection, and operational guardrails.

Requirements

  • 5+ years of experience in DevOps, SRE, or Platform Engineering operating production systems on AWS.
  • Strong experience with CI/CD, GitHub Actions, Terraform, production EKS, Kubernetes operations, and Argo CD.
  • Hands-on AWS networking experience covering VPCs, subnets, routing, security groups, NACLs, Route 53, ACM, and load balancers.
  • Experience with Aurora or RDS MySQL, Redis, S3, Kafka/MSK, Lambda, backups, migrations, lifecycle management, and disaster recovery.
  • Strong observability and security experience with Prometheus, Grafana, OpenTelemetry, IAM, KMS, secrets management, vulnerability scanning, and software supply chain security.
  • Comfortable working with Python, Bash, and Linux systems.

Nice to have

  • Experience operating LLM or ML workloads in production, including LiteLLM, Bedrock, pgvector, prompt caching, or evaluation systems.
  • Experience building MCP servers or deploying LangGraph or CrewAI agent frameworks in production.

Culture & Benefits

  • Automation is preferred over repetitive operational toil.
  • Small team with high ownership and responsibility for defining platform standards.
  • Human review is maintained for risky AI-assisted actions, with blameless incident reviews and documented decisions.
  • Flexible remote work from anywhere, with a co-working reimbursement of up to $200 per month.
  • Health, dental, and vision benefits with 100% of employee premiums and 75%–80% of most dependent premiums covered.
  • 401(k), paid parental leave, unlimited PTO, and home office support of up to $500 plus $100 per month for internet, phone, and related expenses.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →