Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer in Network Infrastructure (SRE/Networking): Building and operating Nebius network infrastructure with an accent on reliability targets (SLIs/SLOs, error budgets), incident response, and observability. Focus on designing safer change workflows and automating operational practices to keep network services stable while scaling quickly.
Location: Remote - Europe (Amsterdam, Netherlands)
Company
Nebius builds a cloud infrastructure platform for the global AI economy, spanning compute, storage, networking, and applied AI.
What you will do
- Define and own reliability goals for network services and critical paths using SLIs/SLOs, availability targets, and error budgets.
- Drive reliability improvements across the network, including site readiness, inter-site connectivity (DCI), and operational standards.
- Own incident response for your areas: lead investigations/postmortems and implement durable fixes.
- Build and evolve observability with actionable metrics/logs/traces, alerting, and faster debug loops.
- Design safer network change workflows with automation, CI/CD, test/staging, canarying, rollbacks, and auditability.
- Partner with network engineers and platform teams to embed operability into designs.
Requirements
- Strong production Linux fundamentals and a structured approach to debugging complex systems.
- Solid networking fundamentals and understanding of real network failures (control plane vs data plane, latency/loss, failure domains).
- Hands-on experience operating high-availability systems and improving them over time.
- Ability to write and maintain software/automation (Go is common; Python is welcome).
- Experience with modern infrastructure tooling and automating operational workflows (e.g., IaC, CI/CD, container platforms).
- Work location: must be based in Europe.
Nice to have
- Experience with high-throughput traffic processing (load balancers, tunneling/decap, NAT64, datapath-heavy systems).
- Low-level networking performance/debug background (eBPF/XDP, DPDK, perf/ftrace, kernel networking internals).
- Experience building network-safe delivery pipelines (testing labs, staged rollouts, automated verification, drift detection).
- Background with large-scale network observability/telemetry (routing/flow telemetry, regression detection at scale).
Culture & Benefits
- Competitive compensation and opportunities for career growth and learning.
- Flexibility and ownership in an engineering-first environment.
- Collaborative, innovative culture with an international team.
- Opportunity to work on impactful AI projects.
Hiring process
- Interviews focused on SRE/network reliability experience and practical problem-solving.
- Discussion of reliability practices, incident ownership, and automation/observability approach.
- Final evaluation of fit for engineering-first collaboration and ownership.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Infrastructure Engineer
7 дней назад
Site Reliability Engineer (AWS)
5 дней назад
Senior Software Engineer, Site Reliability Engineering (AWS)
153 000 - 210 000$
5 дней назад
Software Reliability Engineer (Temporary FTE)
109 250 - 163 370$
7 дней назад
Sr. Site Reliability Engineer
125 000 - 145 000$
6 дней назад