11 часов назад
Member of Technical Staff, Infrastructure Engineer (Kubernetes)
175 000 - 240 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Member of Technical Staff, Infrastructure Engineer (Kubernetes): Designing, scaling, and operating the Kubernetes infrastructure powering autonomous AI scientist agents with an accent on resilient scheduling, resource orchestration, and production-grade platform operations. Focus on managing heterogeneous CPU and GPU workloads, implementing observability and fault-tolerant storage and networking, and solving complex distributed infrastructure issues.
Location: On-site at the San Francisco office in the Dogpatch neighborhood
Salary: $175,000–$240,000 per year plus equity
Company
Scientific builds and deploys AI scientist agents that accelerate scientific research and the development of new medicines. The engineering team combines expertise across biology, physics, chemistry, and AI.
What you will do
- Architect, implement, and operate highly available Kubernetes clusters supporting thousands of concurrent agents, jobs, and services.
- Define cluster scaling, node pool management, autoscaling, and resource quota strategies for rapidly growing workloads.
- Design scheduling, placement, and affinity strategies for heterogeneous CPU-, GPU-, and memory-intensive workloads.
- Establish observability, monitoring, alerting, and incident response practices using Prometheus, Grafana, Datadog, or similar tools.
- Own Kubernetes storage and networking, including persistent volumes, CSI drivers, service mesh, network policies, and ingress architecture.
- Partner with backend, ML, and research teams to translate workload requirements into reliable infrastructure patterns and guide complex troubleshooting.
Requirements
- 5+ years of professional infrastructure or platform engineering experience.
- Hands-on production experience with Kubernetes and cloud platforms such as AWS EKS, GCP GKE, or Azure AKS.
- Proficiency in at least one systems or backend programming language for operator development and infrastructure tooling.
- Experience with Terraform, Pulumi, or Crossplane and GitOps workflows.
- Strong knowledge of container networking, storage, security, RBAC, Pod Security Standards, and secrets management.
- Ability to work autonomously, make sound technical judgments, and drive projects from concept through production.
Nice to have
- Experience with gVisor or Kata Containers and container snapshots.
Culture & Benefits
- Full healthcare coverage with premiums paid for employees and dependents.
- Support for growing families, including a new parent stipend, fertility coverage, and 12 weeks of paid parental leave.
- Mental health support, pet care support, commuter benefits, 401(k) matching, and a quarterly health and wellness benefit.
- Daily lunch in the office and dinner when working late.
- Regular team off-sites and company events.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
11 часов назад
Infrastructure Engineer (AI)
200 000 - 400 000$
1 час назад
Senior Cloud Engineer (AWS/Kubernetes)
215 000 - 240 000$
3 часа назад
Senior Infrastructure Engineer (AI)
168 000 - 213 000$
9 часов назад
Site Reliability Engineer (TypeScript)
180 000 - 220 000$
LangChain
10 часов назад
Database Infra Engineer (AI)
180 000 - 230 000$
12 часов назад
Infrastructure Engineer (AI)
160 000 - 245 000$