Назад
Company hidden
8 часов назад

SRE Platform Software Engineer (AI Infrastructure)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
junior
Английский
b2
Страна
Singapore/US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
SRE Platform Software Engineer (AI Infrastructure): Building and operating a multi-region SRE platform for a global GPU rental fleet with an accent on microservices, GitOps automation, observability, and infrastructure reliability. Focus on developing production-ready platform services, maintaining strict SLOs, writing resilience tests, and participating in a mentored on-call rotation.

Location: Remote within San Jose, California or Austin, Texas

Company

hirify.global builds Bitcoin mining solutions and AI computational infrastructure, including data centers and cloud capabilities for high-demand artificial intelligence workloads.

What you will do

  • Build, test, and deploy SRE microservices such as collection agents, telemetry pipelines, alert engines, and cluster health services.
  • Deliver infrastructure and application updates through GitOps, declarative configuration, and automated CI/CD pipelines.
  • Monitor and optimize metrics, logs, and traces to support high availability across GPU infrastructure.
  • Participate in a mentored on-call rotation, write runbooks, and prepare incident post-mortems.
  • Develop unit, integration, and end-to-end tests to validate platform resilience.

Requirements

  • Bachelor’s degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience and internships.
  • 0–2 years of hands-on software development experience.
  • Proficiency in Go, Java, or Rust, plus scripting with Python or Bash.
  • Knowledge of data structures, algorithms, object-oriented design, APIs, concurrency, networking, and distributed systems.
  • Hands-on exposure to Docker, Kubernetes, and Linux through coursework, projects, open-source work, or internships.
  • Experience writing unit and integration tests, with strong technical writing and collaboration skills.

Nice to have

  • Experience with Kubernetes Operators, Helm, ArgoCD, or Flux.
  • Exposure to Prometheus, OpenTelemetry, Grafana, Loki, or time-series databases.
  • Familiarity with NVIDIA DCGM, CUDA, GPU or AI infrastructure, and high-performance computing.
  • Knowledge of Terraform or Ansible.

Culture & Benefits

  • Full-time role with a build-and-run operating model.
  • Collaboration with senior engineers and participation in a mentored on-call rotation.
  • Work on infrastructure supporting multiple regions, data centers, cloud teams, and tenants.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →