Назад
Company hidden
10 часов назад

Senior Principal Engineer (AI Infrastructure)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Principal Engineer (AI Infrastructure): Define and deliver platform-scale orchestration architecture for distributed AI infrastructure across cloud services and on-premise edge appliances with an accent on control planes, observability, fleet management, and autonomous infrastructure agents. Focus on building reliable multi-tenant systems, validating architecture through POCs, and solving complex problems across telemetry, data storage, device lifecycle management, and intent-driven change execution.

Location: United States; on-site at headquarters

Company

hirify.global builds high-performance infrastructure for demanding artificial intelligence workloads across silicon, systems, and networking.

What you will do

  • Define cross-cutting architecture for platform-scale orchestration initiatives from vision through delivery.
  • Build distributed control planes connecting cloud services with on-premise edge appliances through mTLS gRPC streams.
  • Design observability and telemetry platforms using OpenTelemetry, Kafka, ClickHouse, VictoriaMetrics, real-time state views, and fleet-wide health aggregation.
  • Build autonomous infrastructure agents for onboarding, upgrades, drift remediation, failure recovery, and multi-step workflow orchestration.
  • Develop fleet orchestration, device lifecycle management, zero-touch provisioning, multi-tenant SaaS, RBAC, and enterprise identity integrations.
  • Own data architecture and intent compilation workflows across relational, graph, time-series, key-value, streaming, and Git-backed storage systems.

Requirements

  • 15–18 years of experience building and operating distributed systems for enterprise customers across cloud and on-premise environments.
  • Deep proficiency in Go or a comparable systems language, with strong distributed systems fundamentals.
  • Production experience with microservices, gRPC, Protocol Buffers, OpenAPI, Kubernetes, Helm, and hybrid cloud/on-premise deployments.
  • Experience designing observability platforms, stream-processing systems, time-series databases, alerting pipelines, and real-time dashboards at scale.
  • Experience with multi-tenant platforms, tenant isolation, RBAC, identity federation, API versioning, backward compatibility, and contract-first development.
  • Ability to lead cross-team technical initiatives, communicate architecture clearly, build POCs, write code, and own delivery outcomes.

Nice to have

  • AI/ML agent architectures for infrastructure operations, including LangGraph, AutoGen, MCP, human-in-the-loop gating, and autonomous workflow orchestration.
  • Network automation, infrastructure management platforms, datacenter networking, OpenConfig, gNMI, or fabric management at scale.
  • Hub-and-spoke, edge computing, or control-plane/data-plane separation architectures.
  • OpenTelemetry contributor experience, device enrollment systems, zero-touch provisioning, open-source contributions, or published distributed-systems work.

Culture & Benefits

  • Work in a talent-dense, high-performing environment focused on ownership, technical rigor, and speed.
  • Partner cross-functionally with Product, QA, and Customer Engineering to shape roadmaps and quality gates.
  • Work on foundational AI infrastructure with immediate real-world impact.
  • Accessibility accommodations are available throughout the hiring process.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →