Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Database Infra Engineer (AI): Operating and scaling a production database for AI observability and evaluation across cloud environments with an accent on Kubernetes, infrastructure as code, and distributed stateful workloads. Focus on building safe deployment pipelines, automating failover and upgrades, and ensuring reliability, disaster recovery, and cost-efficient capacity at massive scale.
Location: San Francisco, CA, United States; on-site
Salary: $180,000–$230,000 USD per year
Company
LangChain builds open-source frameworks and a platform for building, evaluating, deploying, and operating AI agents at scale, including LangSmith and SmithDB.
What you will do
- Own SmithDB deployment and operations across cloud environments, including cluster lifecycle management, upgrades, and automated failover.
- Build infrastructure tooling with Terraform, Kubernetes, Helm, or equivalent technologies to provision, configure, and scale database nodes.
- Operate Kubernetes infrastructure for multi-tenant, high-throughput, low-latency distributed database services.
- Develop deployment pipelines, rollout strategies, infrastructure as code, and CI/CD promotion from development through production.
- Drive reliability engineering, including incident response, postmortems, SLOs, disaster recovery, capacity planning, and cost efficiency.
- Collaborate with SmithDB engineers to translate database engine features into production-ready infrastructure changes and safe rollouts.
Requirements
- 5+ years of experience in infrastructure, platform engineering, or SRE.
- Hands-on experience with Kubernetes and AWS, GCP, or Azure cloud infrastructure.
- Strong scripting or systems programming ability with Go, Python, or a similar language.
- Experience with infrastructure as code and CI/CD tooling such as Terraform, Pulumi, CDK, Helm, or ArgoCD.
- Experience running stateful workloads in production, including persistent volumes, managed node groups, and cloud storage.
- Experience with on-call operations for high-traffic data systems, incident triage, runbooks, and automation.
Nice to have
- Production database ownership experience with Postgres, ClickHouse, Redis, or similar systems.
- Ability to read and reason about Rust.
- Knowledge of replication, backups, point-in-time recovery, connection pooling, and graceful degradation under load.
Culture & Benefits
- Work in a small, fast-moving, high-autonomy systems engineering team.
- Build a greenfield database system handling production AI observability data at massive scale.
- Medical, dental, and vision coverage.
- Flexible vacation, a 401(k) plan, and meals on in-office days in the US.
- Compensation includes base salary, variable compensation where relevant, equity, benefits, and perks.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 часов назад
Infrastructure and Reliability Engineer (Kubernetes/Terraform)
196 000 - 235 000$
11 часов назад
DevOps Engineer (AI)
85 000 - 180 000$
8 часов назад
Infrastructure Engineer (AI)
150 000 - 300 000CAD
7 часов назад
Infrastructure Engineer (AI)
160 000 - 245 000$
4 часа назад
Senior DevOps Engineer / Site Reliability Engineer (AI)
170 000 - 220 000$
5 часов назад
Infrastructure Engineer (AI)
158 000 - 235 000$