5 дней назад
Senior Software Development Engineer in Test (AI Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Software Development Engineer in Test (AI Infrastructure) (Distributed AI Infrastructure): Designing and automating tests for large-scale AI clusters with an accent on reliability, failure scenarios, performance, security, and observability. Focus on validating distributed ML training and inference systems, debugging hardware and software across thousands of nodes, and achieving highly reliable cluster deployments.
Location: Hybrid, Toronto office, Canada
Company
builds large-scale AI infrastructure and wafer-scale accelerator systems for high-speed model training and inference.
What you will do
- Design and execute optimized test strategies and methodologies for cutting-edge AI infrastructure.
- Break down large distributed ML training and inference systems into components suitable for unit testing.
- Automate tests for cluster features covering high availability, failure scenarios, performance, stress, and security.
- Validate cluster software, including Kubernetes, Prometheus, and Grafana, as well as hardware components and high-speed interconnects.
- Debug and improve large-scale systems to support cluster reliability targets of up to 99.9999% uptime and strong observability.
Requirements
- Bachelor’s or master’s degree in computer science, electrical engineering, AI, data science, or a related field.
- 5+ years of experience testing enterprise software, distributed systems, datacenter hardware, or datacenter software.
- Strong coding skills in Python, Golang, or C/C++.
- Strong debugging skills across distributed systems, hardware, and software, with experience using tools such as pdb, gdb, strace, and network monitors.
- Strong knowledge of operating-system internals, datacenter layouts, server and device performance, memory, BIOS, PCIe, networking, and storage.
- Experience with AWS, Kubernetes, and Docker.
Nice to have
- Experience with Grafana and Prometheus monitoring.
- Understanding of ML model training and inference.
- Experience with ML hardware accelerators, including GPUs or custom accelerator ASICs.
Culture & Benefits
- Build an AI platform designed beyond the constraints of GPUs.
- Work with a high-performance AI supercomputer and contribute to cutting-edge AI research.
- Combine startup vitality with job stability.
- Work in a simple, non-corporate environment that respects individual beliefs.
- Access continuous learning, growth, and team support in an inclusive workplace.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →