8 часов назад
SRE Platform Software Engineer (AI Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
SRE Platform Software Engineer (AI Infrastructure): Building and operating a multi-region SRE platform for a global GPU rental fleet with an accent on microservices, GitOps automation, observability, and infrastructure reliability. Focus on developing production-ready platform services, maintaining strict SLOs, writing resilience tests, and participating in a mentored on-call rotation.
Location: Remote within San Jose, California or Austin, Texas
Company
builds Bitcoin mining solutions and AI computational infrastructure, including data centers and cloud capabilities for high-demand artificial intelligence workloads.
What you will do
- Build, test, and deploy SRE microservices such as collection agents, telemetry pipelines, alert engines, and cluster health services.
- Deliver infrastructure and application updates through GitOps, declarative configuration, and automated CI/CD pipelines.
- Monitor and optimize metrics, logs, and traces to support high availability across GPU infrastructure.
- Participate in a mentored on-call rotation, write runbooks, and prepare incident post-mortems.
- Develop unit, integration, and end-to-end tests to validate platform resilience.
Requirements
- Bachelor’s degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience and internships.
- 0–2 years of hands-on software development experience.
- Proficiency in Go, Java, or Rust, plus scripting with Python or Bash.
- Knowledge of data structures, algorithms, object-oriented design, APIs, concurrency, networking, and distributed systems.
- Hands-on exposure to Docker, Kubernetes, and Linux through coursework, projects, open-source work, or internships.
- Experience writing unit and integration tests, with strong technical writing and collaboration skills.
Nice to have
- Experience with Kubernetes Operators, Helm, ArgoCD, or Flux.
- Exposure to Prometheus, OpenTelemetry, Grafana, Loki, or time-series databases.
- Familiarity with NVIDIA DCGM, CUDA, GPU or AI infrastructure, and high-performance computing.
- Knowledge of Terraform or Ansible.
Culture & Benefits
- Full-time role with a build-and-run operating model.
- Collaboration with senior engineers and participation in a mentored on-call rotation.
- Work on infrastructure supporting multiple regions, data centers, cloud teams, and tenants.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
CrowdStrike
5 дней назад
Sr Engineer, SRE TechOps CICD (Remote)
140 000 - 215 000$
8 часов назад
Senior SRE Engineer (AI)
17 часов назад
Senior Platform Engineer (AI)
SandboxAQ
6 дней назад
Staff Platform Engineer (AI)
121 600 - 228 000$
3 дня назад
Sr. SRE (AWS/Kubernetes)
10 часов назад