Назад
Company hidden
10 дней назад

Site Reliability Engineer, AML Platform - USDS (AI/ML)

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
Australia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer, AML Platform - USDS (AI/ML): Building and operating highly available, scalable, and fault-tolerant distributed AI/ML systems with an accent on performance analysis, automation, monitoring, and incident response. Focus on designing reliability pipelines, optimizing system capacity, resolving production incidents, and implementing preventative measures through root cause analysis.

Location: Sydney, Australia; fully in-person up to 5 days a week

Company

hirify.global operates data privacy, cybersecurity, and trust and safety programs to protect U.S. user data, applications, algorithms, and the content ecosystem.

What you will do

  • Design, build, and maintain highly available, scalable, and fault-tolerant systems for distributed AI/ML workloads.
  • Monitor and analyze system performance and resolve issues before they affect users.
  • Develop automated monitoring, alerting, and incident response systems and pipelines.
  • Collaborate with software engineering teams to improve application reliability, scalability, and performance.
  • Implement security best practices and maintain compliance with regulatory requirements.
  • Participate in on-call rotations, conduct root cause analysis, lead post-mortems, and implement preventative measures.

Requirements

  • Expertise in analyzing and troubleshooting Linux-based distributed systems.
  • Bachelor's or Master's degree in Computer Science, Computer Engineering, or equivalent SRE or software engineering experience.
  • Programming experience in at least one of C, C++, Python, or Go.
  • Strong understanding of data structures and algorithms.
  • Competent knowledge of relational database systems.
  • Availability to respond to incidents during and outside normal business hours.

Nice to have

  • Experience designing and maintaining large-scale systems.
  • Strong understanding of code optimization and routine task automation.
  • Proficiency with TensorFlow, PyTorch, MXNet, or PaddlePaddle.

Culture & Benefits

  • Fully in-person work supports fast alignment, real-time decision-making, and integrated execution.
  • Work takes place within a diverse and inclusive environment.
  • The role provides exposure to coding, performance analysis, large-scale system operations, and hardware and capacity decisions.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →