Назад
Company hidden
1 месяц назад

Technical Support Principal Engineer (AI Network)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Technical Support Principal Engineer (AI Network) (SONiC/AI Networking): Building and validating reliable networking platforms for GPU AI clusters with an accent on switching ASIC pipeline tuning, SONiC NOS development, and Ethernet/RoCEv2 or InfiniBand fabric performance. Focus on debugging distributed infrastructure, developing observability and automated testing tools, and solving complex deployment and customer proof-of-concept issues at scale.

Location: US - Headquarters; on-site

Company

hirify.global builds high-performance infrastructure for demanding artificial intelligence workloads, spanning silicon, systems, and networking.

What you will do

  • Own end-to-end network bring-up and performance for GPU AI clusters during customer proof-of-concepts.
  • Design, develop, and maintain SONiC NOS features and enhancements.
  • Debug software, Linux systems, networking stacks, ASICs, switching platforms, and distributed infrastructure.
  • Build tools for traffic monitoring, performance measurement, deployment efficiency, observability, telemetry, and automated testing.
  • Collaborate with hardware, software, infrastructure, QA, test, and product teams to identify root causes and deliver robust solutions.
  • Support customer deployments, integration testing, proof-of-concepts, field issue resolution, technical documentation, and engineering communication.

Requirements

  • 12+ years of professional experience and at least 3 years of hands-on SONiC or equivalent network operating system development experience preferred.
  • Bachelor’s or master’s degree in computer science, electrical engineering, or a related field.
  • Strong programming skills in Python, Go, or a similar language.
  • Strong knowledge of Linux, TCP/IP networking, routing, switching, VLANs, network troubleshooting, and Linux internals.
  • Experience with PTF and SPyTest for network validation, plus familiarity with Docker containers.
  • Mandatory knowledge of network ASICs and switch hardware architecture, with strong written and verbal communication skills.

Nice to have

  • Data-center networking, switch ASIC tuning, and platform bring-up experience.
  • RoCEv2 deployments, including ECN/PFC design, DCQCN tuning, DSCP/PCP mapping, and queue shaping.
  • SONiC QoS and buffer configuration, Enterprise-OS QoS, and deep buffer and queueing knowledge.
  • Optics and PHY expertise, including PAM4, FEC, DOM/RS-FEC counters, autonegotiation, link training, and DAC/AOC selection.
  • Experience with network telemetry, benchmarking, EVPN/VXLAN, BlueField DPUs, GPUDirect RDMA, and AI storage fabrics.

Culture & Benefits

  • High-performance startup environment focused on technical rigor, ownership, speed, and first-principles engineering.
  • Opportunity to work on infrastructure powering next-generation GPU clusters and AI data centers.
  • Collaborative work across engineering and customer-facing teams.
  • Equal opportunity employer committed to accessible hiring and accommodations.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →