Назад
Company hidden
4 дня назад

AI Infrastructure Engineer

180 000 - 400 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Infrastructure Engineer (AI/Cloud Infrastructure): Building and operating the multi-cloud infrastructure, deployment systems, and observability stack that power autonomous AI agents with an accent on Terraform, Kubernetes, AWS, Azure, GCP, and regulated production environments. Focus on defining SRE practices for non-deterministic agentic systems, automating customer environments, maintaining cloud parity, and making agent behavior operationally observable.

Location: New York City, United States; on-site

Salary: $180,000–$400,000 base annually, plus equity

Company

hirify.global applies AI to transform critical institutions across industries such as healthcare, manufacturing, and energy, combining embedded engineering, product, and research expertise with an in-house toolkit for agentic workflows.

What you will do

  • Own infrastructure, deployment, operational reliability, SLOs, alert routing, on-call operations, incident response, and postmortems for autonomous AI systems.
  • Design and evolve Terraform infrastructure across AWS, Azure, and GCP, including tested modules, policy-as-code, drift management, and environment blueprints.
  • Automate customer-environment provisioning, including networking, DNS, identities, tagging policies, certificates, and firewall exceptions.
  • Maintain cloud parity across managed Kubernetes environments and support BYOC, single-tenant, on-premises, and air-gapped deployments.
  • Own the observability stack, including Grafana, Loki, Mimir, Tempo, Alloy, Langfuse, and LiteLLM, with a focus on agent decisions, data usage, and cost.
  • Build paved deployment paths for engineers and embed HIPAA and SOC 2 controls directly into infrastructure and delivery pipelines.

Requirements

  • 5+ years operating production infrastructure in SRE, platform engineering, or production engineering.
  • Production experience with AWS, Azure, and GCP, including networking, IAM, and EKS, AKS, or GKE.
  • Deep Terraform expertise covering module design, state, testing, drift, and blast-radius management.
  • Experience operating production Kubernetes across multiple cloud providers and working with environments outside direct organizational control.
  • Experience with regulated infrastructure and controls such as HIPAA, SOC 2, PCI, or FedRAMP.
  • Proficiency in Python, Go, or Bash, along with customer-facing communication skills for architecture reviews and infrastructure coordination.

Nice to have

  • Experience with vendor control planes such as Ryvn, Nuon, or Replicated; GitOps and progressive delivery across multi-tenant fleets.
  • Multi-region infrastructure experience and expertise with the Grafana stack at multi-tenant scale.
  • GPU and inference operations using Ray, SkyPilot, vLLM, SGLang, LLM gateways, or cost attribution systems.
  • MLOps or research-to-production handoff experience.
  • Experience designing observability and blast-radius controls for non-deterministic autonomous systems.

Culture & Benefits

  • Office-first environment focused on in-person collaboration and outcomes.
  • Autonomy, trust, and ownership in a small, rapidly growing team.
  • Equity, 401(k) matching, comprehensive health and family benefits, and daily lunch.
  • Equal opportunity workplace with reasonable accommodations for qualified applicants and employees.
  • Values include ambitious problem-solving, customer outcomes, decisive communication, execution intensity, and kindness.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →