3 дня назад
Architect/Staff Systems Software Engineer (AI)
337 000GBP
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Architect/Staff Systems Software Engineer (AI): Building the runtime and distributed serving stack that connects PyTorch and JAX to the DX-1 decode accelerator for rack-scale inference with an accent on distributed execution, memory-aware scheduling, and hardware/software co-design. Focus on scaling inference across disaggregated accelerator topologies, debugging pre- and post-silicon behaviour, and defining reliability, observability, and tooling standards across runtime, network, and accelerator layers.
Location: London, UK; on-site
Starting from £337K annually; additional compensation includes equity and an annual Living-Local Bonus for residences within 20 minutes of the office.
Company
is developing the DX-1, a dataflow accelerator architected specifically for large-scale AI model decoding and disaggregated inference.
What you will do
- Design, build, and extend the distributed inference and serving stack connecting PyTorch and JAX to the DX-1 accelerator.
- Define tensor, pipeline, and data parallelism, collective communication, KV-cache management and offload, and memory-aware scheduling across disaggregated accelerator topologies.
- Improve reliability across distributed failure domains through fault handling, graceful degradation, load balancing, recovery, observability, tracing, and diagnostic tooling.
- Drive pre-silicon and bring-up work using simulation, emulation, FPGA prototyping, and analytical modelling.
- Set systems, tooling, observability, and review standards across hardware, compiler, modelling, and software teams.
- Shape platform direction and solve ambiguous cross-team systems problems while developing senior technical talent.
Requirements
- Deep systems software experience with hands-on C/C++ and strong fundamentals across runtime, networking, and accelerator layers.
- Ownership of complex end-to-end systems problems, ideally extending distributed inference or serving stacks such as vLLM, SGLang, NVIDIA Dynamo, or TensorRT-LLM in production.
- Experience with distributed inference at cluster scale, including parallelism strategies, collective communication, KV-cache and memory management, and reliability across failure domains.
- Fluency at the framework boundary, connecting accelerators with PyTorch, JAX, and serving systems, plus whole-stack debugging using tracing, workload replay, and architectural analysis.
- Strong judgment on speed, cost, and quality trade-offs, excellent communication, and the ability to influence cross-functional teams without formal authority.
- Bachelor’s degree or higher in computer science, electrical engineering, mathematics, or a related field.
Nice to have
- Experience with dataflow or non-GPU accelerator architectures and pre- or post-silicon bring-up on custom ASIC or FPGA hardware.
- Production observability at scale, including hardware counters, Prometheus/Grafana-style export, and device and cluster views.
- Depth in HPC cluster design, high-speed networking, distributed systems, or heterogeneous compute platforms.
Culture & Benefits
- Meaningful stock options and ownership in the company.
- Employer-contributed retirement plans.
- Annual Living-Local Bonus for living within 20 minutes of the office.
- Competitive compensation based on experience, skills, and location.
- Eligibility is subject to U.S. export control regulations and recent citizenship or permanent residency status.
- Applicants with most recent citizenship or permanent residency in Iran, North Korea, Syria, Cuba, Russia, Belarus, China, Hong Kong, Macau, or Venezuela are generally not eligible.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →