Назад
Company hidden
обновлено 13 дней назад

Infrastructure Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US/Georgia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Infrastructure Engineer (AI): Maintaining and improving Lineate and customer production infrastructure across cloud and on-prem environments with an accent on reliability, automation, observability, and incident response. Focus on operating Kubernetes and Databricks workloads, coordinating root-cause analysis, and building AI-agent automation for operational support tasks.

Location: Georgian office; in-person role

Company

hirify.global is an international software development company building custom AI solutions, cloud and data infrastructure, and augmented engineering teams for FinTech, HealthTech, AdTech, and other clients.

What you will do

  • Maintain the reliability, availability, performance, security, and cost-efficiency of hirify.global and customer production systems.
  • Monitor systems, respond to alerts and support requests, diagnose issues, and operate Kubernetes workloads and Databricks resources using established playbooks.
  • Administer Linux and Windows servers and support production infrastructure across cloud and on-prem environments.
  • Coordinate incident triage, mitigation, root-cause analysis, post-incident reviews, and customer communication.
  • Improve observability through metrics, logs, traces, dashboards, alerting, and SLO tracking.
  • Build automation for ticket triage, runbook execution, log analysis, knowledge-base maintenance, and other routine operational tasks using AI agents.

Requirements

  • At least 3 years of hands-on Linux system administration experience and at least 2 years of production AWS experience.
  • Working familiarity with Kubernetes and AWS Databricks, including workload operation, job execution, cluster management, log reading, and troubleshooting with playbooks.
  • Practical experience using AI agents or LLM-based tools for operational, scripting, or support automation.
  • At least 1 year of Windows system administration experience and 1+ year of practical scripting with Python, Bash, or Go.
  • Experience with CI/CD pipelines, Git workflows, Docker, and production IT infrastructure.
  • Participation in a 24/7 shift schedule and professional customer support through Slack, email, or Zendesk.

Nice to have

  • Infrastructure as Code with Terraform, Ansible, or CloudFormation.
  • Observability tools such as Prometheus, Grafana, ELK/Loki, Datadog, or CloudWatch.
  • Networking, cloud security, database operations, incident management, or data-platform experience.
  • MCP servers, prompt-based automations, AI-driven runbooks, or additional GCP or Azure experience.

Culture & Benefits

  • Professional development opportunities and a clear career path.
  • Social benefits package, company equipment, gym membership, and English lessons.
  • Flexible vacation time and inclusive in-person and digital events.
  • International IT environment with a structured performance management system.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →