Назад
Company hidden
14 часов назад

Site Reliability Engineer (Aerospace)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (Aerospace): Building a production-grade observability platform for satellite, ground-station, and deep-space communication networks with an accent on metrics, logging, distributed tracing, and service reliability. Focus on designing SLOs, SLIs, and error budgets, automating Kubernetes and cloud infrastructure with IaC and GitOps, and leading monitoring and incident response.

Location: Remote (United States). Access to export-controlled information requires U.S. person status, eligibility to access the information without export authorization, or eligibility and reasonable likelihood of obtaining the necessary export authorization.

Company

hirify.global develops laser communications technology and temporospatial software-defined networking platforms for aerospace, satellite, airborne, cislunar, and deep-space communications.

What you will do

  • Design and build a centralized observability platform for metrics, logging, and distributed tracing using tools such as Prometheus, Loki, Tempo, and OpenTelemetry.
  • Define and manage SLOs, SLIs, and error budgets for core products.
  • Partner with software engineers to establish observability standards, templates, documentation, and instrumentation.
  • Automate observability deployment and operations with Terraform and GitOps tools such as ArgoCD.
  • Provide visibility into Kubernetes clusters and GCP and AWS environments.
  • Lead monitoring, alerting, incident response, and blameless post-mortems; participate in on-call rotations.

Requirements

  • 4+ years of experience in SRE or platform engineering, focused on observability for large-scale distributed compute or network systems.
  • Hands-on experience with observability platforms including Prometheus, Grafana, Loki or ELK, OpenTelemetry, Tempo or Jaeger, and similar tools.
  • Production experience with Google Cloud Platform and Kubernetes.
  • Experience with Infrastructure as Code and GitOps principles.
  • Proficiency in a systems programming language, preferably Go or Python.
  • Experience defining and managing SLOs, SLIs, and error budgets for highly available production services.

Nice to have

  • Multi-cloud experience with GCP and AWS, GitLab CI, service mesh technologies, or JVM observability.
  • Experience instrumenting Go and C++ applications.
  • Active Secret clearance or higher.

Culture & Benefits

  • Flexible working arrangements, including hybrid remote and in-office schedules.
  • Professional development and advancement opportunities.
  • Collaborative, supportive, and inclusive work environment.
  • Competitive compensation with equity options.
  • Health, dental, vision, and life insurance, 401(k), and paid time off.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →