Назад
Company hidden
3 часа назад

Technical Support Senior Staff Engineer (AI Networking)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Technical Support Senior Staff Engineer (AI Networking): Building and validating high-performance networking platforms for GPU AI clusters, with an accent on SONiC NOS development, switching ASIC tuning, and Ethernet/RoCEv2 or InfiniBand fabric performance. Focus on debugging distributed infrastructure, developing observability and automation tools, and solving complex platform, networking, and customer deployment issues at scale.

Location: US - Headquarters; on-site

Company

hirify.global builds high-performance infrastructure for demanding artificial intelligence workloads across silicon, systems, and networking.

What you will do

  • Own end-to-end network bring-up and performance for GPU AI clusters during customer proof-of-concepts.
  • Design, develop, and maintain features and enhancements for the SONiC network operating system.
  • Debug software, Linux systems, networking stacks, distributed infrastructure, and complex system issues.
  • Develop tools for traffic monitoring, performance measurement, deployment efficiency, observability, telemetry, and automated testing.
  • Support customer deployments, integration testing, proof-of-concepts, and field issue resolution.
  • Collaborate with hardware, software, infrastructure, QA, test, product, and customer teams on root-cause analysis and robust solutions.

Requirements

  • Bachelor’s or master’s degree in computer science, electrical engineering, or a related field.
  • 10+ years of work experience, including hands-on SONiC or equivalent network operating system development experience.
  • Strong programming skills in Python, Go, or a similar language.
  • Solid knowledge of Linux, TCP/IP networking, routing, switching, VLANs, network troubleshooting, and Docker containers.
  • Experience with PTF and SPyTest for network validation; knowledge of network ASICs and switch hardware architecture is mandatory.
  • Strong problem-solving, written communication, verbal communication, and cross-functional collaboration skills.

Nice to have

  • Data-center networking with switch ASIC tuning and platform bring-up.
  • Deep knowledge of buffers, queueing, QoS, ECN, PFC, VOQ, and shared pools.
  • Experience with optics and PHY technologies, EVPN/VXLAN, BlueField DPU offloads, GPUDirect RDMA, storage fabrics, and AI networking benchmarks.
  • Experience with tools such as ethtool, devlink, mlnx_qos, perfquery, Prometheus, gNMI, REST, perftest, nccl-tests, and iPerf3.
  • Python or Ansible automation and Git-based configuration workflows.

Culture & Benefits

  • Work in a fast-paced startup environment focused on foundational AI infrastructure.
  • Operate with strong ownership, technical rigor, and speed.
  • Collaborate with experienced engineers across hardware, software, infrastructure, QA, test, and product teams.
  • Work on technologies powering next-generation GPU clusters and AI data centers.
  • Accessibility accommodations are available throughout the hiring process.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →