Назад
Company hidden
21 час назад

L2 Support Engineer (AI)

100 000 - 140 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
L2 Support Engineer (AI): Supporting installation, upgrades, and day-to-day operation of an agentic AI software development platform with an accent on Kubernetes, cloud infrastructure, distributed-systems debugging, and observability. Focus on diagnosing cross-service failures, reproducing issues safely in multi-tenant environments, building runbooks and alerts, and producing evidence-backed escalations during US-aligned hours and rotating on-call coverage.

Location: Cambridge, Massachusetts, United States (on-site). Customer coverage is aligned with US business hours, roughly ET–PT, with rotating on-call coverage and occasional evening, early-morning, or weekend escalations.

Salary: $100,000–$140,000 per year plus equity.

Company

hirify.global is an AI software development platform that autonomously converts enterprise requirements into production-ready software.

What you will do

  • Deploy and install the platform in customer environments and troubleshoot installation issues.
  • Support upgrades and daily operations while keeping customer environments stable.
  • Partner with L1 support to triage, diagnose, resolve, or escalate customer-reported issues.
  • Investigate failures across compute, networking, storage, and application services using logs, pod state, queue depth, traces, and database records.
  • Build dashboards, monitors, alerts, runbooks, and evidence-backed incident escalations.
  • Communicate incident status and resolutions clearly to customers.

Requirements

  • Experience debugging distributed systems across services, queues, and network hops using evidence-based investigation.
  • Hands-on Kubernetes and Docker experience.
  • Experience with at least one of GCP, AWS, or Azure, including managed Kubernetes, cloud logging, IAM, and storage behavior.
  • Strong monitoring and observability skills, including log queries, request or trace ID correlation, traces, dashboards, and alerts.
  • Knowledge of Python, Redis or message queues, WebSockets, networking, SQL/PostgreSQL, source-control platforms, CI/CD, Helm, Linux, and Windows.
  • Ability to handle secrets, credentials, certificates, and multi-tenant customer environments safely, defaulting to read-only diagnostics.

Nice to have

  • ArgoCD and Vault experience.
  • Jira or similar incident-management and ticketing experience.
  • Prior customer-facing support or SRE/on-call experience.

Culture & Benefits

  • In-person collaboration with a high-performance, customer-focused engineering culture.
  • Rotating on-call responsibilities are shared fairly.
  • On-call is compensated or followed by time off in lieu according to policy.
  • Recovery time is protected after heavy incidents.
  • Focus on sustainable performance, sleep, movement, and restorative activities.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →