Назад
2 дня назад

Software Engineer — AI Infrastructure (AI)

180 000 - 240 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer — AI Infrastructure (AI): Building a shared agentic platform, LLM cost-efficiency systems, data access layers, and event-driven infrastructure with an accent on orchestration, governance, observability, and reliability. Focus on designing MCP gateways, agent memory and evaluation layers, token optimization, multi-source data access, and production SLOs across AWS and Kubernetes.

Location: Hybrid in Redwood City, California, or San Francisco, California

Salary: $180,000–$240,000 USD per year

Company

Snorkel AI helps enterprises transform expert knowledge and data into specialized, production-ready AI systems.

What you will do

  • Design and build a shared agentic platform with orchestration, memory, context, MCP gateway, evaluation, and observability layers.
  • Build LLM cost and efficiency infrastructure for queuing, throttling, batching, caching, model routing, and token attribution.
  • Develop shared Python access libraries and SDKs for Snowflake, S3, and RDS with authentication, RBAC, pagination, and query governance.
  • Implement event-driven data flows with CDC connectors, schema registries, event routing, and dead letter queues.
  • Build governance, lineage, PII handling, retention, and audit infrastructure for data, agents, and tool calls.
  • Own reliability and cost through OpenTelemetry instrumentation, SLOs, alerting, infrastructure optimization, and on-call support.

Requirements

  • 2+ years of experience building and operating platform infrastructure, data infrastructure, or backend systems with significant data components.
  • Strong Python proficiency and hands-on production experience with LLM APIs, tokens, context windows, rate limits, batching, and caching.
  • Daily use of AI-assisted development tools such as Claude Code or Cursor is required.
  • Strong SQL skills and experience with at least two of Snowflake, Redshift, and Postgres.
  • Experience with AWS services including S3, RDS, EKS, EventBridge, and IAM, plus Terraform-managed environments and Kubernetes.
  • Familiarity with Prefect, Airflow, or Dagster; dbt; and data governance concepts including RBAC, PII handling, audit logging, and lineage.

Nice to have

  • Experience with agentic systems, MCP gateways, agent memory, knowledge layers, or evaluation harnesses.
  • Experience with LLM serving and gateway infrastructure, event-driven architectures, shared SDKs, or distributed compute with Ray.
  • Experience with OpenTelemetry, ClickHouse, LLM trace observability, or regulated environments such as SOC 2, FedRAMP, or HIPAA.

Culture & Benefits

  • Work in a rapidly scaling company with market-proven solutions and robust funding.
  • Opportunity to shape technical priorities, strategic decisions, and company-wide developer velocity and product quality.
  • Support for technical growth, leadership development, and cross-functional learning.
  • Equal employment opportunity and reasonable accommodation for candidates with disabilities.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →