Назад
Company hidden
2 часа назад

Site Reliability Engineer (Cloud/Azure/SRE)

Формат работы
remote/onsite
Тип работы
fulltime
Грейд
middle/senior
Английский
b2
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (Cloud/Azure/SRE): Architecting, building, and evolving a core monitoring and observability platform with an accent on SRE principles, cloud infrastructure automation, and data quality monitoring. Focus on designing intelligent systems for proactive issue resolution, integrating cloud technologies, and embedding FinOps practices to optimize performance and cost efficiency.

What you will do

  • Architect, build, and evolve the core monitoring and observability platform and services using modern technologies and best practices.
  • Maintain and enhance stability, availability, and performance of critical applications by applying Site Reliability Engineering principles.
  • Develop intelligent systems for proactive identification and automated resolution of technical issues and anomalies.
  • Build capabilities to measure and report on data quality dimensions across supply chain pipelines and critical data sources.
  • Develop unified dashboards (Grafana, Power BI, Databricks) for KPIs, system health, operational metrics, and financial insights.
  • Drive automation through Infrastructure as Code, CI/CD pipelines, and scripting for provisioning and deployment workflows.
  • Integrate and optimize technologies like Fivetran, AKS, Kafka, Azure SQL Server, Databricks, and Flink within Azure cloud.
  • Embed FinOps practices to monitor cloud resource usage and spending for cost optimization.
  • Analyze operational incidents and data quality issues to drive continuous platform improvements.

Requirements

  • At least 4 years of experience in a similar role with fundamental understanding of SRE methodologies.
  • Experience with Cloud & Kubernetes (MS Azure, AKS, Kafka, HVR, Databricks).
  • Proficiency in coding and scripting (Python, Java, Shell Scripting) and strong automation skills.
  • Expertise in monitoring, alerting, and logging systems (Grafana, Datadog, Dynatrace, Prometheus).
  • Experience with CI/CD tools (Terraform, GitHub Actions) and infrastructure as code.
  • Experience building enhanced data resiliency by defining and tracking key Data Quality metrics.
  • English proficiency at least B2.

Culture & Benefits

  • Stable employment with a company established since 2008 and 1800+ employees across 7 global sites.
  • Flexible "office as an option" model allowing remote or office work.
  • Workation opportunities aligned with company policy.
  • Great Place to Work® certified employer.
  • Flexible working hours and contract forms.
  • Comprehensive onboarding with a buddy system.
  • Access to Udemy learning platform and certificate training programs.
  • Upskilling support with 110+ training opportunities yearly.
  • Internal promotion culture with 76% of managers promoted internally.
  • Diverse, inclusive, and values-driven community.
  • Autonomy in work approach and referral bonuses.
  • Well-being and health support activities.
  • Opportunities to donate to charities and support the environment.
  • Modern office equipment provided or available to borrow depending on location.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →