2 часа назад
SRE Engineer (AI, Platform Reliability)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
SRE Engineer (AI, Platform Reliability): Improve the reliability of ’s platform and Tier 0 critical infrastructure with an accent on distributed platform reliability, incident command and stabilization, and AI-driven observability and incident analysis. Focus on building safe deployment patterns, reducing operational toil, and driving root-cause fixes across complex systems.
Location: Remote
Company
builds technology to increase access to the global economy.
What you will do
- Build and extend reliability platforms across multiple systems and organizations.
- Standardize reliability tooling and deploy-safety patterns (progressive delivery, automated rollback, guardrails).
- Lead incident command and coordinate mitigation for sev 0–1 incidents, including structured escalation.
- Use AI-driven tooling to improve signal detection, reduce alert noise, and accelerate root cause analysis.
- Drive platform-wide reliability improvements, shared operational tooling, and evidence-based maturity assessments.
- Participate in primary platform oncall (12 hours/day, one week every few weeks depending on team size).
Requirements
- 5+ years of software development experience.
- Experience running production oncall for high-availability systems and strong incident management skills (structured triage, mitigation under pressure, blameless postmortems).
- Fluency with CI/CD pipelines, progressive rollout strategies, and rollback automation.
- Monitoring & observability expertise, including building/tuning alerts for uptime, error rates, latency regression, and resource exhaustion.
- Comfort with vendor/dependency management, including maintaining validated escalation contacts reachable within ≤ 5 minutes.
- Familiarity with AI-driven tooling for observability, incident analysis, or automation.
Nice to have
- Demonstrated technical initiative and leadership on previous backend/platform-focused projects.
- Experience with evidence-based maturity assessments using trailing 90-day data windows.
Culture & Benefits
- Remote work with medical insurance, flexible time off, and retirement savings plans.
- Work across multiple time zones; may require work outside normal business hours.
- Equal opportunity employer with an inclusive interview experience and reasonable accommodations.
- Modern family planning support.
Hiring process
- Applications may be evaluated using automated AI tools for efficiency and consistency.
- Recruiter contact available for accommodations and hiring-practice/data-usage questions.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior Site Reliability Engineer
125 000 - 145 000$
DualEntry
6 дней назад
Senior Site Reliability Engineer (SRE) (Fintech)
45 000 - 80 000$
4 дня назад
Site Reliability Engineer (Monetization)
Together AI
3 дня назад
Lead Site Reliability Engineering Manager (AI)
2 дня назад
Senior SRE Engineer (AWS/Kubernetes)
6 дней назад
Sr. DevOps Engineer (Azure)
176 362 - 293 937$