7 дней назад
Software Engineer, Platform (AI Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Software Engineer, Platform (AI Infrastructure): Building low-latency, high-reliability APIs, AI gateways, automated arena runtimes, and serving layers for large-scale real-world model evaluation with an accent on streaming reliability, provider integration, and enterprise-grade infrastructure. Focus on designing distributed systems, recovering from partial failures during model streams, instrumenting deep observability, and scaling zero-to-one AI platform products.
Location: San Francisco, United States; hybrid with a minimum of 3 days per week onsite. Fully remote candidates will only be considered with a very strong endorsement.
Company
develops infrastructure for evaluating how AI models perform in real-world workflows, including online arenas, leaderboards, and model-serving systems.
What you will do
- Build low-latency, high-reliability APIs for leaderboards, models, and arenas.
- Design streaming systems across heterogeneous LLM providers, including partial-failure recovery, mid-stream fallback, and response normalization.
- Develop enterprise infrastructure for rate limiting, authentication, usage metering, cost attribution, audit logging, and SOC 2 compliance.
- Build distributed tracing, latency analysis, token-level usage tracking, and real-time observability dashboards.
- Integrate backend services with the evaluation platform, Arena data, customer benchmarks, Leaderboards, and Evals platforms.
- Work closely with researchers, engineers, and product leadership on zero-to-one and scaling initiatives.
Requirements
- Approximately 5 or more years of backend engineering experience, including substantial work with distributed systems, infrastructure, or developer-facing platforms.
- Strong proficiency in Go; Go is the primary backend language and a must-have requirement.
- Experience with LLM provider APIs such as OpenAI, Anthropic, or Google, including streaming, token management, rate limits, and provider-specific behavior.
- Product-oriented approach to API and developer experience design, with the ability to work effectively in an ambiguous startup environment.
- Availability to work onsite in San Francisco at least 3 days per week; fully remote consideration requires a very strong endorsement.
Nice to have
- Cloud infrastructure experience with AWS, GCP, or Azure, plus Kubernetes, Terraform, Postgres, or Redis.
- Experience building API gateways, proxies, or developer tools such as Bifrost, Kong, Envoy, or Tyk.
- Background in AI/ML infrastructure, model serving, inference, or evaluation frameworks.
- Experience with SSO, RBAC, audit logs, multi-tenancy, or billing systems such as Stripe, Metronome, or Orb.
- Familiarity with vLLM, LiteLLM, LangChain, or other modern AI infrastructure tools.
Culture & Benefits
- Competitive compensation and equity aligned with the market where the employee is permanently based.
- Comprehensive medical, dental, vision, and wellness benefits.
- Opportunity to work on AI infrastructure with a small, mission-driven team.
- Culture focused on transparency, trust, craftsmanship, curiosity, and community impact.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →