Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Production Safety (AI): Building and operating safety-critical production systems for text, multimodal, and agentic AI products with an accent on distributed services, model guardrails, Kubernetes, and observability. Focus on designing reliable inference infrastructure, managing staged rollouts and failover, and detecting failures and regressions in large-scale training and evaluation pipelines.
Location: London, United Kingdom; Mountain View and New York City, United States; Zurich, Switzerland
Base salary: USD $119,800–$234,700 per year in the U.S.; USD $160,200–$261,000 per year in the San Francisco Bay Area and New York City metropolitan area.
Company
Microsoft AI develops AI systems and products, including frontier models, product engineering, and responsible AI technologies.
What you will do
- Design, build, deploy, and operate safety-critical services in production inference paths for text, multimodal, and agentic AI systems.
- Integrate model-based classifiers, policy engines, and guardrails with model APIs and serving platforms.
- Build containerized Kubernetes services with CI/CD, staged regional rollouts, rollback, and failover mechanisms.
- Define service-level objectives for availability, latency, throughput, correctness, and cost, including capacity planning and autoscaling.
- Develop monitoring and debugging capabilities for service health, safety decisions, training pipelines, and evaluation runs.
- Validate launches and contribute to on-call support, incident response, root-cause analysis, and reusable operational standards.
Requirements
- Bachelor’s degree in Computer Science, Engineering, a related technical field, or equivalent practical experience.
- Strong software engineering skills in one or more production languages such as C++, C#, Java, Go, Rust, or Python.
- Experience designing, deploying, and operating distributed services or large-scale production systems.
- Experience with containerized workloads, Kubernetes, and automated CI/CD pipelines.
- Knowledge of observability, capacity planning, failure isolation, graceful degradation, incident response, security, privacy, and system failure modes.
- Ability to collaborate across engineering, machine learning, product, and policy teams and communicate system tradeoffs clearly.
Nice to have
- Experience with online inference platforms, model serving, API gateways, policy enforcement systems, or latency-sensitive infrastructure.
- Experience operating machine learning models, training pipelines, or evaluation systems in production.
- Experience with Azure technologies, observability tooling, or capacity planning.
- Familiarity with AI safety, trust and safety, abuse prevention, content moderation, security engineering, privacy, or responsible AI.
Culture & Benefits
- Work with interdisciplinary teams across frontier model development, product engineering, responsible AI, security, and production inference.
- Collaborate with Microsoft product organizations, external partners, and customers.
- Some roles may be eligible for benefits and additional compensation.
- Applications are accepted on an ongoing basis until the position is filled, with the posting open for at least five days.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Staff Software Engineer (AI Platform)
189 300 - 333 400$
Scale AI
4 дня назад
Staff Software Engineer, Platform (AI)
4 дня назад
Lead Application Developer (AI)
130 000 - 190 000$
Baseten
57 минут назад
Distributed Systems Engineer (AI Inference)
180 000 - 360 000$
Scale AI
3 дня назад
Senior Software Engineer (Generative AI)
252 000 - 315 000$
3 дня назад
Advanced Software Engineer - AI Platform
103 000 - 155 000$