Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior / Staff AI Engineer (AI): Building infrastructure for large-scale agentic workloads, synthetic data generation, evaluation pipelines, simulation environments, and LLM systems with an accent on distributed execution, observability, reproducibility, and production reliability. Focus on scaling non-deterministic experiments, evaluating long-running agent trajectories, safely isolating tool execution, and turning research workflows into reusable platform capabilities.
Location: Hybrid in New York City, NY, or San Francisco, CA
Company
Snorkel AI develops an AI platform that helps enterprises transform expert knowledge and data into specialized, production-ready AI systems.
What you will do
- Design and build infrastructure for large-scale agentic workloads involving tools, external services, sandboxes, and simulated environments.
- Develop synthetic data generation, automated labeling, and evaluation systems for training and testing AI models and agents.
- Build orchestration and distributed compute systems for thousands to millions of experiments and simulations across heterogeneous environments.
- Operate LLM infrastructure for routing, retries, caching, provider failover, rate limiting, cost attribution, and efficient execution.
- Instrument workloads with traces, model interactions, tool calls, environment state, evaluation results, latency, reliability, and cost data.
- Collaborate with research, product, and engineering teams to turn experimental workflows into reliable APIs, SDKs, and reusable platform capabilities.
Requirements
- 5+ years of experience building production software systems in AI/ML infrastructure, distributed systems, data platforms, ML platforms, or backend infrastructure.
- Experience operating non-deterministic AI or ML workloads in production or at significant scale.
- Strong proficiency in Python and experience building production-quality APIs, services, and developer tooling.
- Strong background in distributed systems and cloud platforms, including compute orchestration, storage, networking, isolation, and failure handling; AWS experience is preferred.
- Experience with Prefect, Airflow, Dagster, Ray, Kubernetes, or similar workflow and distributed execution systems.
- Track record of leading complex engineering initiatives, influencing stakeholders, and delivering measurable impact.
Nice to have
- Experience with LLM or agent infrastructure, model gateways, agent runtimes, tool execution, tracing, or multi-agent systems.
- Experience building evaluation platforms, synthetic data systems, reinforcement learning environments, simulations, or benchmarks.
- Experience scaling distributed AI workloads across containers, Kubernetes, serverless compute, sandboxes, or heterogeneous environments.
- Experience with safe isolation and sandboxing for model-generated code or tool calls.
- Prior experience as a Tech Lead, Team Lead, or hands-on Engineering Manager, or in a hyper-growth startup environment.
Culture & Benefits
- Meaningful ownership over the infrastructure used to experiment with, evaluate, and productionize AI systems.
- Work on emerging challenges involving non-deterministic systems, agent behavior, long-running trajectories, and large-scale simulations.
- Environment supporting technical growth, leadership development, cross-functional learning, and shared strategic impact.
- Rapidly scaling company with market-proven solutions and robust funding.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →