2 дня назад
Infrastructure Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Infrastructure Engineer (AI/HPC): Building and maintaining the compute, storage, and networking foundations for on-premise silicon development workloads with an accent on Linux administration, cluster management, observability, and performance tuning. Focus on troubleshooting high-performance computing infrastructure, automating routine operations, and resolving networking, storage, and resource-utilisation bottlenecks.
Location: London or Bristol, United Kingdom
Company
is an AI hardware startup developing chips and systems designed to accelerate inference for frontier AI models.
What you will do
- Maintain and troubleshoot on-premise Rocky Linux and RHEL-family server infrastructure.
- Deploy and maintain compute infrastructure using infrastructure-as-code tools such as Ansible.
- Set up and maintain monitoring and observability for resource utilisation, service health, and machine failures.
- Diagnose networking and storage issues involving DNS, VLANs, routing, network filesystems, capacity, and performance.
- Support cluster compute and job scheduling operations, including Slurm, as well as user and identity management through FreeIPA or LDAP.
- Document infrastructure changes, maintain runbooks, and work with engineering teams to resolve infrastructure bottlenecks.
Requirements
- 2–3+ years of hands-on Linux systems administration experience.
- Working knowledge of monitoring solutions such as Prometheus, Grafana, Zabbix, or similar tools.
- Solid TCP/IP, DNS, VLAN, routing, and networking troubleshooting fundamentals.
- Understanding of network, parallel, and distributed filesystems, disk and volume management, and storage performance troubleshooting.
- Command-line proficiency and scripting skills in Bash, Python, or similar languages.
- Methodical troubleshooting skills, curiosity, independence, and ownership of problems through resolution.
Nice to have
- Exposure to HPC environments or cluster compute concepts.
- Familiarity with Ansible, Terraform, or similar infrastructure-as-code tools.
- Linux certifications such as RHCSA or LFCS.
- Data centre experience and knowledge of server installation, diagnostics, and component replacement.
Culture & Benefits
- High ownership and autonomy in driving work forward.
- Rapid collaboration with leadership and hardware, software, silicon, and modelling teams.
- Inclusive and diverse office environment focused on pragmatic execution and technical curiosity.
- Competitive salary and meaningful equity participation.
- Private medical, dental and vision coverage, contributory pension, 25 days of holiday plus bank holidays, and life and critical illness insurance.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →