2 дня назад
Systems Engineer (AI Infrastructure)
140 000 - 225 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Systems Engineer (AI Infrastructure): Building and operating AI inference infrastructure, Kubernetes workloads, and hardware-isolated agent sandboxes with an accent on Rust systems programming, GPU orchestration, model serving, and secure execution. Focus on designing inference control planes, managing distributed stateful data systems, enforcing isolation boundaries, and improving reliability, scalability, and metering accuracy.
Location: Remote within the United States; applicants must be currently authorized to work in the United States on a full-time basis. Travel to company summits twice per year is required.
Compensation: $140K–$225K annually, plus equity and bonus eligibility.
Company
is a public benefit corporation building physics-informed machine learning and enterprise AI solutions for climate, sustainability, energy, real estate, logistics, and utilities.
What you will do
- Manage open-weight model serving, including deployments, configuration, cold-start strategies, service-level objectives, upgrades, and canary releases.
- Administer Kubernetes infrastructure for inference and sandbox workloads, including operators, CRDs, autoscaling, GPU scheduling, and node lifecycle management.
- Operate vector and graph data stores with tested backup, restore, scaling, and failover procedures.
- Own the secure sandbox runtime and host-side control plane for lifecycle management, execution, snapshots, forking, teardown, metering, and threat modeling.
- Build and direct FastAPI control-plane services, infrastructure-as-code, and observability dashboards while debugging across backend services.
- Maintain open-source engineering standards through reviewed pull requests, documentation, and reproducible builds.
Requirements
- At least five years of shipping production systems in a systems language; Rust is preferred, with deep Go, C/C++, or Zig experience also considered.
- Production experience operating Kubernetes workloads, including controllers or operators, scheduling, autoscaling, and node lifecycle management.
- Strong threat-modeling skills for isolation boundaries, including namespaces, cgroups, seccomp, hypervisors, least privilege, credential protection, audit trails, and human approval for system-of-record writes.
- Hands-on experience deploying or operating open-weight LLM serving infrastructure such as vLLM or SGLang in reliable, metered production endpoints.
- Practical experience with Rust and Tokio, Python and FastAPI, Kubernetes operators, KEDA, Karpenter, GPU device plugins or DRA, distributed-systems performance, and observability.
- Current United States work authorization is required; visa sponsorship is not available. A bachelor's degree is required.
Nice to have
- Experience with Firecracker, Kata, gVisor, or comparable isolation technologies.
- Familiarity with Vault or KMS-class secrets management and egress control.
- Experience operating pgvector, Qdrant, Neo4j, or comparable stateful vector and graph systems.
- Master's degree.
Culture & Benefits
- Fast-growing, profitable, mission-driven startup focused on AI transformation in critical industries.
- Fully remote culture with a cluster of teammates in Seattle.
- Health insurance with dependent coverage.
- Flexible paid time off, equity, and bonus eligibility.
- Opportunities to build and release infrastructure according to open-source standards.
Hiring process
- Applicants should apply to no more than two roles within a six-month period.
- Candidates who do not meet every qualification are still encouraged to apply.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →