14 часов назад
Site Reliability Engineer (Aerospace)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Aerospace): Building a production-grade observability platform for satellite, ground-station, and deep-space communication networks with an accent on metrics, logging, distributed tracing, and service reliability. Focus on designing SLOs, SLIs, and error budgets, automating Kubernetes and cloud infrastructure with IaC and GitOps, and leading monitoring and incident response.
Location: Remote (United States). Access to export-controlled information requires U.S. person status, eligibility to access the information without export authorization, or eligibility and reasonable likelihood of obtaining the necessary export authorization.
Company
develops laser communications technology and temporospatial software-defined networking platforms for aerospace, satellite, airborne, cislunar, and deep-space communications.
What you will do
- Design and build a centralized observability platform for metrics, logging, and distributed tracing using tools such as Prometheus, Loki, Tempo, and OpenTelemetry.
- Define and manage SLOs, SLIs, and error budgets for core products.
- Partner with software engineers to establish observability standards, templates, documentation, and instrumentation.
- Automate observability deployment and operations with Terraform and GitOps tools such as ArgoCD.
- Provide visibility into Kubernetes clusters and GCP and AWS environments.
- Lead monitoring, alerting, incident response, and blameless post-mortems; participate in on-call rotations.
Requirements
- 4+ years of experience in SRE or platform engineering, focused on observability for large-scale distributed compute or network systems.
- Hands-on experience with observability platforms including Prometheus, Grafana, Loki or ELK, OpenTelemetry, Tempo or Jaeger, and similar tools.
- Production experience with Google Cloud Platform and Kubernetes.
- Experience with Infrastructure as Code and GitOps principles.
- Proficiency in a systems programming language, preferably Go or Python.
- Experience defining and managing SLOs, SLIs, and error budgets for highly available production services.
Nice to have
- Multi-cloud experience with GCP and AWS, GitLab CI, service mesh technologies, or JVM observability.
- Experience instrumenting Go and C++ applications.
- Active Secret clearance or higher.
Culture & Benefits
- Flexible working arrangements, including hybrid remote and in-office schedules.
- Professional development and advancement opportunities.
- Collaborative, supportive, and inclusive work environment.
- Competitive compensation with equity options.
- Health, dental, vision, and life insurance, 401(k), and paid time off.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Deimos
2 дня назад
Senior Site Reliability Engineer
10 часов назад
Site Reliability Engineer (TypeScript)
180 000 - 220 000$
15 часов назад
Site Reliability Engineer (Azure)
7 дней назад
Site Reliability Engineer - DevSecOps Engineer (Cloud)
11 часов назад
Infrastructure and Reliability Engineer (Kubernetes/Terraform)
196 000 - 235 000$
18 часов назад