2 дня назад
Member of Technical Staff (MLOps)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Member of Technical Staff (MLOps) (Platform Infrastructure): Building internal infrastructure, developer tools, and MLOps systems that support research experiments and product deployments with an accent on cloud platforms, CI/CD, observability, and reproducible workflows. Focus on designing scalable platform architecture, automating data and model operations, and improving system reliability and developer velocity.
Location: Palo Alto, United States; on-site
Company
is an early-stage organization developing research and product environments supported by AI infrastructure and internal platforms.
What you will do
- Design, build, and maintain core infrastructure for research and product environments, including cloud compute, storage, CI/CD, observability, and security.
- Develop internal developer tools, automation systems, shared libraries, and platform services that improve engineering productivity.
- Build reliable systems for data management, model experimentation, evaluation, and deployment across research and production.
- Define and improve build, test, and release workflows from prototype through production.
- Monitor observability and performance metrics, diagnose bottlenecks, and optimize engineering workflows.
- Contribute to infrastructure strategy and architecture as the organization scales.
Requirements
- Strong software engineering background with experience in infrastructure, platform, or DevOps engineering.
- Experience with cloud environments such as AWS or GCP, Docker, Kubernetes, and CI/CD systems.
- Experience building developer tools, automation frameworks, or internal platforms.
- Familiarity with data pipelines, job scheduling, and ML experimentation workflows.
- Strong problem-solving and communication skills, with the ability to collaborate with research scientists and product engineers.
- Ability to work on-site in Palo Alto.
Nice to have
- Experience with machine learning infrastructure, training pipelines, or model evaluation tooling.
- Background in monitoring and observability tools such as Prometheus, Grafana, or Datadog.
- Knowledge of infrastructure as code and configuration management best practices.
- Interest in reproducible, safe, and efficient AI development.
- Experience at an early-stage startup or small research organization.
Culture & Benefits
- Close collaboration with research scientists and software engineers.
- Work focused on improving reproducibility, reliability, and efficiency across research and product workflows.
- Opportunity to shape infrastructure strategy and architecture as the organization grows.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
DevOps Engineer (Defense Technology)
125 000 - 160 000$
2 дня назад
Senior Staff DevOps Engineer – Orchestration (AI)
2 дня назад
Platform Engineer (AI)
1 день назад
DevOps Engineer (Robotics)
2 дня назад
Foundational Software Engineer (Infrastructure)
180 000 - 300 000$
2 дня назад