7 дней назад
Senior MLOps Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior MLOps Engineer (AI) (GCP/Vertex AI): Building and operating production infrastructure, deployment pipelines, and monitoring systems for complex AI models with an accent on scalable GCP services, low-latency model serving, and ML observability. Focus on converting research prototypes into secure auto-scaling microservices, automating training and deployment, and detecting drift and data-quality issues in production.
Location: Remote Poland
Company
develops cybersecurity solutions that help customers monitor, manage, and protect risks associated with digital identities and personal information.
What you will do
- Architect and manage scalable ML infrastructure on GCP using Vertex AI, GKE, Cloud Storage, Cloud Run, and GPU/TPU compute.
- Own the end-to-end deployment and serving lifecycle for machine learning models using containers, Triton Inference Server, vLLM, and MLflow.
- Build automated, reproducible pipelines for model training, testing, evaluation, and deployment with Airflow, Vertex AI Pipelines, GitHub Actions, and ArgoCD.
- Implement production monitoring for system health and ML metrics, including latency, throughput, prediction accuracy, feature drift, and data distribution shifts.
- Provide scalable training environments and standardized deployment templates for AI and research engineers.
- Transform AI prototypes and notebooks into secure, resilient, auto-scaling microservices while supporting feature stores, dataset versioning, and stream or batch processing.
Requirements
- At least 5 years of hands-on experience designing, deploying, and maintaining production ML workloads in cloud environments.
- Deep experience with GCP, including Vertex AI, Cloud Storage, GKE, Cloud Run, IAM, and VPC configurations.
- Expertise in Docker, Kubernetes/GKE, Triton Inference Server, vLLM, and MLflow.
- Experience with Airflow, Vertex AI Pipelines, GitHub Actions, ArgoCD, and Terraform.
- Proficiency in Python and SQL for scripting, automation, API development, and data manipulation.
- Hands-on experience with logging, telemetry, and ML observability using Grafana, Prometheus, GCP Cloud Monitoring, or comparable tools.
Nice to have
- Experience with large-scale LLM or deep learning inference and training workloads.
- GCP Professional Machine Learning Engineer or Professional Cloud Architect certification.
- Familiarity with feature stores such as Feast or Vertex AI Feature Store.
Culture & Benefits
- Work on cybersecurity products addressing real customer protection needs.
- Contribute in a nimble organization where individual impact is visible and valued.
- Access opportunities to learn new technologies, products, and markets in a growth-oriented environment.
- Join an inclusive workplace committed to preventing discrimination and harassment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →