Назад
Company hidden
6 дней назад

GenAI Platform Engineering Lead (AI)

116 400 - 194 000$
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
GenAI Platform Engineering Lead (AI/Azure): Leading the reliability, observability, and operational readiness of a production AI platform with an accent on Azure infrastructure, incident response, infrastructure automation, and service assurance. Focus on building observability pipelines, automating Terraform-based deployments, monitoring model latency and token costs, and solving complex production reliability and security challenges.

Location: Buffalo, New York, United States of America

Salary: $116,400–$194,000 annual USD

Company

M&T Bank is a banking organization operating under financial, risk, security, audit, and regulatory controls.

What you will do

  • Lead the reliability, observability, service assurance, and operational readiness of the Bank’s AI platform.
  • Direct incident response, troubleshooting, escalation, root-cause analysis, corrective actions, and production support.
  • Build and maintain observability pipelines with OpenTelemetry, Prometheus, Grafana, Azure Monitor, and Log Analytics.
  • Oversee Azure infrastructure, Azure API Management, Microsoft Entra ID, Terraform, and controlled CI/CD deployment pipelines.
  • Drive automation with Python, Bash, PowerShell, GitHub Actions, GitLab CI, and infrastructure-as-code practices.
  • Manage engineering teams, client relationships, staffing, project priorities, budgets, vendors, and cross-functional work with cybersecurity, risk, architecture, and application teams.

Requirements

  • At least 9 years of combined higher education and/or work experience, including 4 years in engineering or architecture and 3 years in leadership.
  • At least 5 years in site reliability engineering, platform operations, infrastructure engineering, or a related discipline, including production on-call support.
  • Hands-on Azure experience with Azure API Management, Azure Monitor, Log Analytics, Microsoft Entra ID, and Terraform.
  • Experience with metrics, traces, logs, observability pipelines, infrastructure CI/CD, scripting, automation, and production incident response.
  • Strong analytical, problem-solving, project management, presentation, written communication, and decision-making skills.
  • Experience leading and developing engineers, coordinating complex projects, and operating within risk, audit, security, change-management, and regulatory requirements.

Nice to have

  • Bachelor’s degree and 10 or more years of technology management, SRE, platform operations, or large program leadership experience.
  • Experience operating AI or large language model platforms in production.
  • Knowledge of token cost tracking, model latency monitoring, provider failover, caller-level rate limiting, AI gateway patterns, prompt logging, PII interception, and content filtering.
  • Experience integrating application infrastructure with SIEM platforms and working in regulated environments.

Culture & Benefits

  • Work supports collaboration across technology, business, risk, cybersecurity, architecture, and vendor partners.
  • Operational practices emphasize reliability, auditability, security, cost awareness, documentation, and continuous improvement.
  • The role includes leadership development, staffing oversight, performance management, and budget responsibility.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →