10 дней назад
Site Reliability Engineer, AML Platform - USDS (AI/ML)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer, AML Platform - USDS (AI/ML): Building and operating highly available, scalable, and fault-tolerant distributed AI/ML systems with an accent on performance analysis, automation, monitoring, and incident response. Focus on designing reliability pipelines, optimizing system capacity, resolving production incidents, and implementing preventative measures through root cause analysis.
Location: Sydney, Australia; fully in-person up to 5 days a week
Company
operates data privacy, cybersecurity, and trust and safety programs to protect U.S. user data, applications, algorithms, and the content ecosystem.
What you will do
- Design, build, and maintain highly available, scalable, and fault-tolerant systems for distributed AI/ML workloads.
- Monitor and analyze system performance and resolve issues before they affect users.
- Develop automated monitoring, alerting, and incident response systems and pipelines.
- Collaborate with software engineering teams to improve application reliability, scalability, and performance.
- Implement security best practices and maintain compliance with regulatory requirements.
- Participate in on-call rotations, conduct root cause analysis, lead post-mortems, and implement preventative measures.
Requirements
- Expertise in analyzing and troubleshooting Linux-based distributed systems.
- Bachelor's or Master's degree in Computer Science, Computer Engineering, or equivalent SRE or software engineering experience.
- Programming experience in at least one of C, C++, Python, or Go.
- Strong understanding of data structures and algorithms.
- Competent knowledge of relational database systems.
- Availability to respond to incidents during and outside normal business hours.
Nice to have
- Experience designing and maintaining large-scale systems.
- Strong understanding of code optimization and routine task automation.
- Proficiency with TensorFlow, PyTorch, MXNet, or PaddlePaddle.
Culture & Benefits
- Fully in-person work supports fast alignment, real-time decision-making, and integrated execution.
- Work takes place within a diverse and inclusive environment.
- The role provides exposure to coding, performance analysis, large-scale system operations, and hardware and capacity decisions.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
Senior Observability & Telemetry Engineer (GPU/AI Infrastructure)
Canva
3 дня назад
Principal Production Engineer
Airwallex
4 дня назад
Staff Site Reliability Engineer (Fintech)
6 дней назад
Senior Production Engineer (SRE)
5 дней назад
Senior Cloud Operations Engineer, Infrastructure
102 400 - 153 200$
Canva
3 дня назад