5 дней назад
SRE Monitoring Platform Software Engineer (Early Career / Temporary)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
SRE Monitoring Platform Software Engineer (Early Career / Temporary) (Observability and GPU Cloud Infrastructure): Building monitoring, automation, alerting, storage, and remediation components for a multi-region GPU rental fleet with an accent on telemetry ingestion, distributed systems, Kubernetes, and SLOs. Focus on writing production code and tests, instrumenting services with OpenTelemetry, building dashboards and runbooks, and operating components through GitOps and CI/CD under senior-engineer guidance.
Location: Remote within San Jose, California or Austin, Texas
Company
provides Bitcoin mining infrastructure, AI computational infrastructure, data center operations, and cloud capabilities for AI workloads.
What you will do
- Build collection, ingestion, querying, storage, enrichment, and monitoring components for metrics, logs, traces, and profiles.
- Develop alerting, correlation, SLO, topology, and cluster-health services for a multi-region GPU cloud platform.
- Contribute to Kubernetes, Slurm, Ray, Volcano, Kueue, and KubeRay collection plugins and platform workflows.
- Develop remediation, orchestration, inspection-probe, and job-scheduling components.
- Instrument services with metrics, logs, and traces using OpenTelemetry; build dashboards and operational runbooks.
- Write unit, integration, and contract tests and participate in chaos, soak, and on-call activities under senior-engineer guidance.
Requirements
- 0–2 years of software engineering experience; strong graduate projects or internships are acceptable.
- Proficiency in one programming language: Go preferred, or Python, Java, or Rust.
- Knowledge of data structures, algorithms, concurrency, TCP/HTTP networking, operating systems, and distributed-systems fundamentals.
- Hands-on exposure to monitoring or observability tools such as Prometheus, Grafana, or Loki; ability to write basic PromQL and instrument a service.
- Familiarity with Linux, shell tools, Kubernetes basics, Git, CI pipelines, and unit and integration testing.
- Clear written and verbal English is required.
Nice to have
- Internship or project experience with monitoring, observability, telemetry pipelines, platform tooling, or SRE tooling.
- Exposure to GPU and AI infrastructure such as DCGM, InfiniBand/RoCE, Kubernetes GPU Operator, Slurm, or Ray.
- Exposure to AIOps or ML-adjacent tools, including anomaly detection and alert correlation.
- Contributions to open-source observability or cloud-native projects.
Culture & Benefits
- Greenfield platform development within an established Plugin Framework, GitOps pipeline, and SLO framework.
- Mentorship from senior and principal engineers responsible for the platform architecture.
- Progressive on-call participation, beginning with shadowing before taking primary responsibility.
- Opportunity to develop observability and distributed-systems skills at production scale across GPU infrastructure.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Site Observability Engineer
100 000 - 150 000$
3 дня назад
Reliability Monitoring Engineer (Observability)
100 000 - 150 000$
6 дней назад
Senior Site Reliability Engineer (Observability)
CrowdStrike
2 дня назад
Sr. Platform Engineer - Kubernetes (Remote)
140 000 - 215 000$
5 дней назад
Staff Software Engineer, Compute Platform (Kubernetes)
Reddit
5 дней назад
Staff Software Engineer (Observability)
217 000 - 303 900$