3 дня назад
B2B Systems Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
B2B Systems Site Reliability Engineer (AI): Improving the reliability of production services on AWS through service-level objectives, observability, troubleshooting, automation, and AI-assisted tooling with an accent on cross-stack diagnosis and safe agentic development. Focus on eliminating toil, building guardrails for AI-authored changes, scaling SRE practices, and coordinating reliability initiatives across engineering and business teams.
Location: Remote in Poland. Applications are accepted only from candidates based in Poland who have sponsorship to live and work in Poland. Periodic attendance at a office or collaborative work location may be required for specific events.
Company
develops Apple device management and security software for organizations, schools, and healthcare providers worldwide.
What you will do
- Partner with engineering teams to define service-level objectives, error budgets, and reliability indicators.
- Investigate complex production issues across application, data, infrastructure, and network layers using logs, metrics, traces, profilers, and AI-assisted analysis.
- Create technical documentation, runbooks, architecture notes, postmortems, and proofs of concept for technical and non-technical audiences.
- Eliminate systemic toil through automation, AI agents, infrastructure tooling, and process improvements.
- Build the context, integrations, tests, and guardrails required for safe and reliable AI-authored changes.
- Lead cross-team reliability initiatives, mentor engineers, support critical customer escalations, and improve the SRE practice.
Requirements
- At least 5 years of experience in software engineering, SRE, or production operations.
- Strong production troubleshooting skills across the stack, including profilers, heap and thread dumps, query plans, traces, logs, and metrics.
- Hands-on experience operating production services on AWS, including services such as EC2, S3, EKS, RDS/Aurora, and CloudFront.
- Experience with observability tools such as Grafana, Prometheus, or LogicMonitor; infrastructure as code; and production-grade automation in a general-purpose language.
- Experience with Agile development processes and clear technical documentation for technical and non-technical audiences.
- Strong judgement in applying AI and hands-on experience with agentic development tools such as Claude Code, Cursor, or Copilot, including scoping, verification, and safe production use.
Nice to have
- SQL query optimization and database engine tuning.
- CI/CD tooling such as GitHub Actions or Jenkins.
- Experience with chaos engineering, fault injection, disaster recovery, or FinOps practices.
- Bachelor's degree or equivalent combination of relevant experience and education.
Culture & Benefits
- Open and flexible culture based on trust, respect, ownership, and continuous improvement.
- Remote work with periodic in-person collaboration for selected events and important moments.
- Opportunity to work with a small, empowered engineering team and contribute to software used by more than 75,000 global customers.
- Supportive leadership and a clear career path for professional growth.
- Work focused on Apple platform innovation, reliability, and secure device management.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Senior Site Reliability Engineer (Kubernetes, B2B)
10 дней назад
Global IT Site Reliability Engineer Senior Manager
7 дней назад
Senior Site Reliability Engineer (SRE) – Infrastructure & Systems (B2B)
10 дней назад
Staff Site Reliability Engineer (Cybersecurity)
199 750 - 270 000$
11 дней назад
Site Reliability Engineer (AI)
5 дней назад
Senior Site Reliability Engineer - Dublin (AI)
85 000 - 115 000€