Назад
Company hidden
5 дней назад

Site Reliability Engineer (Bilingual SRE)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Japan
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (Bilingual SRE) (Kubernetes/AWS): Improving the reliability, performance, and scalability of high-availability microservice systems with an accent on observability, incident management, and resilience. Focus on analyzing bottlenecks, developing monitoring and alerting tools, reducing MTTR, and implementing reliable deployment and configuration practices.

Location: Hybrid workstyle with office attendance required according to organizational guidelines and team objectives; Japan-based work environment

Company

hirify.global is a fintech company operating a cashless payment and financial lifestyle platform with more than 75 million users.

What you will do

  • Analyze existing technologies and develop monitoring and notification tools to improve observability and system visibility.
  • Verify failure scenarios in advance, troubleshoot production incidents, conduct root cause analysis, and reduce MTTR.
  • Develop solutions that improve high availability, scalability, resilience, and system performance.
  • Integrate telemetry and alerting platforms to track and improve reliability.
  • Apply best practices for system development, configuration management, and deployment.
  • Document technical knowledge, share incident learnings, and collaborate across teams on reliability improvements.

Requirements

  • At least 5 years of software development experience with Python, Java, Go, or similar languages, plus strong foundations in data structures, algorithms, problem solving, and complexity analysis.
  • Experience troubleshooting and tuning high-performance microservice architectures running on Kubernetes and AWS in highly available production environments.
  • Experience with observability tools and data collection, as well as database technologies such as RDS, NoSQL, or distributed TiDB.
  • Proactive ability to identify and resolve performance bottlenecks, scalability issues, and resilience problems.
  • Ability to communicate verbally in both English and Japanese is required.
  • A coding challenge is included in the selection process.

Nice to have

  • Container image management, optimization, large-scale distributed systems, and capacity planning experience.
  • Understanding of infrastructure as code and automation tools such as Terraform and CloudFormation.
  • SRE or DevOps implementation experience.
  • Experience with CloudWatch, VictoriaMetrics, Prometheus, Snowflake, Sigma, CloudFront, or Nginx.
  • Experience designing or operating disaster recovery strategies and multi-region architectures.

Culture & Benefits

  • Hybrid workstyle with super flex time and no core hours; standard hours are 9:00–17:45 with a one-hour break.
  • Full-time employment with social insurance, employee pension, employment insurance, and compensation insurance.
  • Annual leave of up to 14 days in the first year and 5 days of personal leave annually, allocated proportionally where applicable.
  • Special paid leave can be used for illness, injury, hospital visits, and care needs involving employees, family members, or pets.
  • Visa sponsorship and relocation support are available.
  • 401K and translation, interpretation, and language-learning support are provided.

Hiring process

  • The selection process includes a coding challenge.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →