Назад
Company hidden
10 часов назад

Principal Member of Technical Staff (AI Infrastructure)

200 000 - 350 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Member of Technical Staff (AI Infrastructure): Building and operating Kubernetes infrastructure for AI scientist agents and thousands of persistent scientific workloads with an accent on cluster orchestration, custom operators, scheduling, and resource efficiency. Focus on scaling stateful workloads, designing resilient storage and networking, and ensuring observability and reliability across distributed systems.

Location: On-site in San Francisco, Dogpatch

Salary: $200,000–$350,000 annually plus equity

Company

hirify.global builds and deploys AI scientist agents to accelerate scientific research and the development of new medicines.

What you will do

  • Architect, implement, and operate highly available Kubernetes clusters supporting thousands of concurrent, persistent agents, jobs, and services.
  • Develop CRDs and Kubernetes operators for AI agent lifecycles, research pipelines, and long-running compute tasks.
  • Define cluster scaling, node pool management, autoscaling, scheduling, placement, affinity, and resource quota strategies.
  • Build infrastructure-as-code and GitOps workflows for reproducible environment management.
  • Own Kubernetes storage, networking, service mesh, ingress, security, observability, monitoring, alerting, and incident response.
  • Collaborate with backend, ML, and research teams to translate workload requirements into reliable infrastructure patterns.

Requirements

  • 10+ years of professional infrastructure or platform engineering experience with deep production Kubernetes expertise.
  • Experience designing CRDs and Kubernetes operators using Kubebuilder, Operator SDK, or controller-runtime.
  • Experience operating and scaling Kubernetes clusters with thousands of persistent or long-lived resources.
  • Strong understanding of Kubernetes internals, cloud infrastructure, networking, storage, IAM, RBAC, and secrets management.
  • Proficiency in at least one systems or backend language and hands-on experience with Terraform, Pulumi, Crossplane, or similar tools.
  • Ability to work autonomously, make sound technical judgments, and drive projects from concept through production.

Nice to have

  • Experience with data-intensive platforms, scientific computing, or ML/AI infrastructure.
  • Startup or small-team experience with significant architectural ownership and ambiguity.
  • Experience scaling systems, teams, or platforms during rapid growth.

Culture & Benefits

  • Mission-driven, fast-moving environment focused on accelerating science and medicine.
  • Full healthcare coverage with premiums paid for employees and dependents.
  • Paid parental leave, fertility coverage, new parent support, and mental health support.
  • 401(k) matching, commuter benefits, and quarterly health and wellness support.
  • Daily office lunch, late-work dinners, team off-sites, and company events.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →