3 дня назад
SRE Monitoring Platform Software Engineer (Entry Level)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
SRE Monitoring Platform Software Engineer (Entry Level) (Observability/AI Infrastructure): Building monitoring, automation, and observability components for a multi-region GPU cloud with an accent on metrics, logs, traces, alerting, Kubernetes integrations, and SLOs. Focus on writing production code and tests, instrumenting services with OpenTelemetry, building dashboards and runbooks, and progressively operating components through on-call.
Location: Remote within San Jose, CA or Austin, TX
Company
builds Bitcoin mining infrastructure, AI computational infrastructure, cloud capabilities, and GPU data center platforms across multiple countries.
What you will do
- Build collection, ingestion, query, storage, enrichment, and monitoring components for metrics, logs, traces, and profiles.
- Develop alerting, correlation, SLO, topology, cluster-health, remediation, workflow, and job-scheduling components.
- Integrate observability and platform plugins for Kubernetes, Slurm, Ray, Volcano, Kueue, and KubeRay.
- Instrument services with OpenTelemetry, build dashboards, and write operational runbooks.
- Write unit, integration, and contract tests and participate in chaos and soak testing.
- Operate shipped components with senior-engineer guidance, including shadow on-call participation before taking primary responsibility.
Requirements
- 0–2 years of software engineering experience; strong projects or internships are accepted for new graduates.
- Programming experience in Go, Python, Java, or Rust, with the ability to write clean, readable, tested code.
- Knowledge of data structures, algorithms, concurrency, TCP/HTTP, operating systems, and distributed-systems fundamentals.
- Hands-on exposure to monitoring or observability tools such as Prometheus, Grafana, or Loki, including basic PromQL and service instrumentation.
- Familiarity with Linux, shell tools, Kubernetes fundamentals, Git, CI pipelines, and unit and integration testing.
- Clear written and verbal English, strong communication skills, curiosity, and willingness to learn GPU/AI infrastructure, AIOps, distributed systems, and production observability.
Nice to have
- Internship or project experience in monitoring, observability, telemetry pipelines, platform engineering, or SRE tooling.
- Exposure to GPU and AI infrastructure, including DCGM, InfiniBand/RoCE, Kubernetes GPU Operator, Slurm, or Ray.
- Exposure to AIOps or ML-adjacent tooling such as anomaly detection and alert correlation.
- Contributions to open-source observability or cloud-native projects.
Culture & Benefits
- Greenfield platform development with an established Plugin Framework, GitOps pipeline, and SLO framework.
- Mentorship from senior and principal engineers who own the architecture.
- Opportunity to learn the full observability stack at production scale.
- Full-time employment with remote work limited to the stated locations.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
DevOps & SRE Engineer (Kubernetes)
100 000 - 150 000$
5 дней назад
Senior Site Reliability Engineer (Linux Systems & Application Observability)
180 000 - 200 000$
NDA
2 дня назад
SRE Engineer
CrowdStrike
5 дней назад
Engineer II, Site Reliability (Cybersecurity)
3 дня назад
Staff Site Reliability Engineer (AI)
252 000 - 308 000$
5 дней назад