Назад
Company hidden
3 дня назад

SRE Monitoring Platform Software Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
junior
Английский
b2
Страна
Singapore/US/Norway +4 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
SRE Monitoring Platform Software Engineer (AI): Building monitoring, automation, and observability components for a multi-region GPU rental platform with an accent on metrics, logs, traces, alerting, SLOs, and Kubernetes infrastructure. Focus on writing production code and tests, instrumenting services with OpenTelemetry, building dashboards and runbooks, and operating components through GitOps and CI/CD.

Location: Penang, Malaysia or Singapore

Company

hirify.global is a technology company focused on Bitcoin mining, AI cloud, ASIC hardware, and large-scale data center operations.

What you will do

  • Build collection agents, metrics, logs, traces, and profiles storage services, enrichment services, and collection monitoring.
  • Develop alerting, correlation, and SLO framework components and tune default alert rules.
  • Contribute to topology and cluster-health services and collection plugins for Kubernetes, Slurm, Ray, Volcano, Kueue, and KubeRay.
  • Build remediation actuators, orchestration and workflow components, inspection probes, and job schedulers.
  • Instrument services with OpenTelemetry, create dashboards and runbooks, and participate in on-call operations.
  • Write unit, integration, and contract tests and participate in chaos and soak testing under senior-engineer guidance.

Requirements

  • 0–2 years of software engineering experience; strong graduate projects or internships are accepted.
  • Proficiency in at least one of Go, Python, Java, or Rust, with the ability to write clean, readable, tested code.
  • Knowledge of data structures, algorithms, concurrency, networking, operating systems, and distributed-systems fundamentals.
  • Exposure to Prometheus, Grafana, Loki, or similar observability tools, plus basic PromQL and service instrumentation skills.
  • Familiarity with Linux, shell tools, Kubernetes fundamentals, Git, CI pipelines, and unit and integration testing.
  • Clear written and verbal English is required.

Nice to have

  • Hands-on experience with GPU or AI infrastructure, AIOps, observability platforms, or production-scale distributed systems.
  • Experience with Kubernetes, Slurm, Ray, Volcano, Kueue, or KubeRay projects.

Culture & Benefits

  • Inclusive environment that values authenticity and diverse perspectives.
  • Startup-style environment with opportunities to work on new projects and systems.
  • Direct contribution to digital-asset and AI-cloud infrastructure.
  • Autonomy, personal accountability, rapid growth, training, mentoring, and other welfare benefits.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →