6 дней назад
GenAI Platform Engineering Lead (AI)
116 400 - 194 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
GenAI Platform Engineering Lead (AI/Azure): Leading the reliability, observability, and operational readiness of a production AI platform with an accent on Azure infrastructure, incident response, infrastructure automation, and service assurance. Focus on building observability pipelines, automating Terraform-based deployments, monitoring model latency and token costs, and solving complex production reliability and security challenges.
Location: Buffalo, New York, United States of America
Salary: $116,400–$194,000 annual USD
Company
M&T Bank is a banking organization operating under financial, risk, security, audit, and regulatory controls.
What you will do
- Lead the reliability, observability, service assurance, and operational readiness of the Bank’s AI platform.
- Direct incident response, troubleshooting, escalation, root-cause analysis, corrective actions, and production support.
- Build and maintain observability pipelines with OpenTelemetry, Prometheus, Grafana, Azure Monitor, and Log Analytics.
- Oversee Azure infrastructure, Azure API Management, Microsoft Entra ID, Terraform, and controlled CI/CD deployment pipelines.
- Drive automation with Python, Bash, PowerShell, GitHub Actions, GitLab CI, and infrastructure-as-code practices.
- Manage engineering teams, client relationships, staffing, project priorities, budgets, vendors, and cross-functional work with cybersecurity, risk, architecture, and application teams.
Requirements
- At least 9 years of combined higher education and/or work experience, including 4 years in engineering or architecture and 3 years in leadership.
- At least 5 years in site reliability engineering, platform operations, infrastructure engineering, or a related discipline, including production on-call support.
- Hands-on Azure experience with Azure API Management, Azure Monitor, Log Analytics, Microsoft Entra ID, and Terraform.
- Experience with metrics, traces, logs, observability pipelines, infrastructure CI/CD, scripting, automation, and production incident response.
- Strong analytical, problem-solving, project management, presentation, written communication, and decision-making skills.
- Experience leading and developing engineers, coordinating complex projects, and operating within risk, audit, security, change-management, and regulatory requirements.
Nice to have
- Bachelor’s degree and 10 or more years of technology management, SRE, platform operations, or large program leadership experience.
- Experience operating AI or large language model platforms in production.
- Knowledge of token cost tracking, model latency monitoring, provider failover, caller-level rate limiting, AI gateway patterns, prompt logging, PII interception, and content filtering.
- Experience integrating application infrastructure with SIEM platforms and working in regulated environments.
Culture & Benefits
- Work supports collaboration across technology, business, risk, cybersecurity, architecture, and vendor partners.
- Operational practices emphasize reliability, auditability, security, cost awareness, documentation, and continuous improvement.
- The role includes leadership development, staffing oversight, performance management, and budget responsibility.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Director, AI Engineering & Automation
134 000 - 248 000$
7 дней назад
Engineering Manager (AI)
144 000 - 181 000CAD
6 дней назад
Forward Deployed Team Lead (AI)
130 000 - 194 000$
6 дней назад
Forward Deployed Team Lead, Federal (AI)
130 000 - 194 000$
7 дней назад
Engineering Manager (AI)
128 000 - 183 000$
7 дней назад
Senior Manager, AI Engineering (Agent OS Platform)
239 300 - 358 900$