4 дня назад
Member of Technical Staff (AI)
200 000 - 230 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Member of Technical Staff (AI): Designing and building schedulers, distributed storage systems, caching layers, and datacenter networks that support large-scale AI training and inference with an accent on heterogeneous accelerator fleets, low-latency model serving, and system performance measurement. Focus on turning systems research into production infrastructure, eliminating bottlenecks across the stack, and maintaining reliability under real workloads and failures.
Location: San Mateo or New York, United States
Salary: $200,000–$230,000 per year
Company
provides a platform for building, training, and serving specialized AI models across text, image, embedding, audio, and multimodal workloads.
What you will do
- Design, build, and operate core infrastructure for large-scale AI training and inference.
- Develop benchmarks, traces, and simulators to evaluate scheduling, caching, and network performance.
- Identify and eliminate bottlenecks across kernels, drivers, scheduler policies, storage, and networking.
- Turn research ideas into production systems that perform reliably under real workloads and failures.
- Collaborate with research and inference teams to align infrastructure and model design.
- Track emerging hardware, interconnects, and systems research to guide technical direction.
Requirements
- PhD completed within the last six months or expected by December 2026 in computer science, computer engineering, electrical engineering, or a related field.
- Research experience in distributed systems, operating systems, scheduling, storage, computer networks, computer architecture, or high-performance computing.
- Deep expertise in at least one area: scheduling and resource management, distributed storage and caching, or datacenter networking.
- Strong systems programming skills in C/C++, Rust, Go, or a similar language, plus working proficiency in Python.
- Experience building and evaluating real systems with rigorous performance measurement.
- Ability to communicate system designs and results clearly to diverse audiences.
Nice to have
- First-authored publications at leading systems or networking venues.
- Infrastructure, cloud, or HPC internships, or open-source contributions to systems projects.
- Experience with GPU clusters, CUDA, ROCm, Triton, Kubernetes, cloud infrastructure, or multi-tenant production clusters.
- Experience with distributed performance profiling and debugging.
- Research or engineering achievements demonstrated through grants, fellowships, patents, or systems competitions.
Culture & Benefits
- Work on AI infrastructure problems involving low-latency inference and scalable model serving.
- Collaborate with experienced engineers and AI researchers.
- Receive mentorship from a senior engineer and take ownership of a real systems problem from the start.
- Start dates can be coordinated around thesis defense timelines.
- Operate in a fast-growing environment focused on technical ownership and direct impact.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 дней назад
Senior Backend Engineer (Distributed Systems)
215 000 - 265 000$
10 дней назад
Software Engineer (Rust/C++)
Canva
7 дней назад
Staff Software Engineer (Video Performance)
262 000 - 329 000$
10 дней назад
Forward Deployed Engineer (AI)
136 000 - 216 000$
Anthropic
4 дня назад
Staff+ Software Engineer, ML Sampling Path (AI)
320 000 - 485 000$
11 дней назад
Software Developer IV
114 447 - 199 036$