10 часов назад
Forward-Deployed Engineer, AI Fabric
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Forward-Deployed Engineer, AI Fabric (AI Networking): Troubleshooting and resolving complex AI fabric and data-center networking issues in production with an accent on Ethernet forwarding, BGP, QoS, hardware, optical links, and Linux systems. Focus on reproducing L2/L3 failures in the lab, improving AI-driven diagnostics, validating classification and mitigation, and driving customer issues through root-cause analysis to verified fixes.
Location: US - Headquarters; on-site. Customer travel and on-call work are part of the role.
Company
builds high-performance infrastructure and networking systems for demanding artificial intelligence workloads.
What you will do
- Own complex AI fabric and system issues end-to-end, from intake and triage through root-cause analysis, workaround delivery, and verified fixes.
- Reproduce production failures in the lab, including packet loss, latency, retransmits, ECMP behavior, and PFC/ECN interactions.
- Serve as an escalation point for engineering and provide clear, reproducible problem statements when development support is required.
- Improve AI agents by contributing detection signatures, classification logic, resolution recommendations, labeled data, and feedback from field cases.
- Work directly with customers to understand fabric topologies, workloads, and operational constraints, communicating findings to technical and non-technical stakeholders.
- Identify product supportability gaps, improve diagnostics, maintain detection capabilities, and work with source code to understand behavior and validate fixes.
Requirements
- 5+ years of hands-on experience with high-speed data-center switching platforms at scale.
- Deep Ethernet troubleshooting experience covering L2/L3 forwarding, ECMP, packet-level analysis, BGP, route reflectors, and fabric-scale deployments.
- Experience troubleshooting QoS, including PFC, ECN, DSCP, and queue management.
- Background with ASICs, FPGAs, chassis platforms, pluggable optics, fiber links, DOM telemetry, NOS, firmware, and drivers.
- Proficiency with Linux, system administration, log analysis, streaming telemetry, counters, event-driven diagnostics, and complex lab test scenarios.
- Experience contributing to or training ML/AI systems, plus clear written, verbal, and customer-facing communication skills.
Nice to have
- Experience with SONiC or other open network operating systems.
- Familiarity with AI data-center networking, GPU clusters, NCCL, RDMA/RoCEv2, NVIDIA Spectrum, BlueField, or ConnectX.
- Understanding of Ultra Ethernet Consortium specifications, PCIe architecture, Ixia/Keysight equipment, HFT telemetry, or WJH event data.
Culture & Benefits
- Work on foundational infrastructure for high-performance AI workloads.
- Operate in a small, high-performing team that values ownership, technical rigor, and speed.
- Collaborate globally while working directly with customers and spending time on site when needed.
- Participate in an on-call rotation as the global team expands to reduce out-of-hours coverage.
Hiring process
- Accessibility accommodations are available throughout the hiring process upon request.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 часов назад
Edge AI Systems Engineer
8 часов назад
Principal Forward Deployable Autonomous Systems Engineer (Autonomous Systems)
200 000 - 240 000$
10 часов назад
Embedded Firmware Engineer (AI)
17 часов назад
Performance Library Engineer (AI/ML)
130 000 - 230 000$
9 часов назад
Staff Embedded Engineer (Renewable Energy)
180 000 - 210 000$
13 часов назад