Назад
Company hidden
6 дней назад

Staff MaaS Backend Engineer (AI)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff MaaS Backend Engineer (AI): Re-architecting a Model-as-a-Service platform into a globally distributed, multi-tenant token service with an accent on Go services, high-throughput inference APIs, reliability, and invoice-grade metering. Focus on building active-active regional infrastructure, optimizing KV caching and TTFT, enforcing tenant isolation, and delivering exactly-once billing under partial failure conditions.

Location: Remote within San Jose, CA or Austin, TX

Company

hirify.global is a technology company building Bitcoin mining infrastructure and AI computational infrastructure, including data centers and cloud capabilities for high-demand AI workloads.

What you will do

  • Co-own the MaaS system architecture with the Principal Architect and deliver incremental, reversible platform improvements.
  • Own Go services and inference APIs supporting OpenAI and Anthropic compatibility, streaming, tool calling, structured output, versioning, routing, circuit breaking, and fallbacks.
  • Optimize token throughput, KV and prefix caching, time-to-first-token latency, model serving, LoRA multiplexing, and cold-start-aware autoscaling across Kubernetes GPU fleets.
  • Design globally distributed regional inference pools, active-active control planes, capacity-aware failover, load shedding, SLOs, on-call runbooks, and peak-concurrency testing.
  • Build secure multi-tenant authorization, identity, API key and OAuth credential lifecycle management, distributed quotas, rate limits, abuse controls, and zero-retention data paths.
  • Deliver exactly-once token metering, usage ledgers, billing reconciliation, safe stateful migrations, shadow traffic, dual writes, tracing, and cost telemetry.

Requirements

  • 8+ years of backend engineering experience, including 3+ years owning a high-traffic, multi-tenant API platform for paying customers.
  • Deep hands-on experience with Go-based services, production Kubernetes, Envoy, GPU-aware scheduling, PostgreSQL, Redis, and Kafka.
  • Experience scaling distributed systems through multi-region active-active deployments, caching, backpressure, performance engineering, and zero-downtime brownfield migrations.
  • Strong understanding of LLM serving, server-sent-event streaming, KV and prefix caching, TTFT, throughput, and model-serving trade-offs.
  • Experience building reliable metering and billing systems that reconcile high-volume events under partial failure conditions.
  • Ability to own observability, SLOs, on-call responsibilities, incident reviews, fail-closed authorization, and strict cross-tenant isolation.

Culture & Benefits

  • Full-time employment on a revenue-bearing AI infrastructure platform.
  • Direct collaboration with the Principal Architect and ownership of backend engineering standards.
  • Responsibility for production reliability, measurable service objectives, and operational improvements.
  • Work on infrastructure supporting AI computation and Bitcoin mining operations across multiple countries.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →