3 дня назад
Senior MLOps Engineer (GCP)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior MLOps Engineer (GCP/AI): Architecting and operating production infrastructure, pipelines, and tooling for deploying, scaling, and monitoring AI models on Google Cloud Platform with an accent on Vertex AI, Kubernetes, model serving, and ML observability. Focus on building low-latency inference services, automating training and deployment workflows, and converting AI prototypes into secure, resilient, auto-scaling microservices.
Location: Remote Lithuania
Company
develops cybersecurity solutions that help customers monitor, manage, and protect risks associated with digital identities and personal information.
What you will do
- Architect and manage scalable machine learning infrastructure on Google Cloud using Vertex AI, GKE, GCS, Cloud Run, and GPU/TPU compute.
- Own the full model deployment lifecycle and build high-throughput, low-latency inference services with Docker and serving tools such as Triton Inference Server, vLLM, and MLflow.
- Build reproducible pipelines for model training, testing, evaluation, and deployment using Airflow, Vertex AI Pipelines, GitHub Actions, and ArgoCD.
- Implement system and machine learning observability, including latency, throughput, uptime, feature drift, prediction accuracy, and data distribution monitoring.
- Provide scalable training environments and standardized deployment templates for AI and research engineers.
- Turn AI prototypes and notebooks into secure, resilient, auto-scaling microservices while supporting feature stores, dataset versioning, and stream or batch processing.
Requirements
- At least 5 years of hands-on experience designing, deploying, and maintaining production machine learning workloads in cloud environments.
- Deep practical experience with GCP, including Vertex AI, Cloud Storage, GKE, Cloud Run, IAM, and VPC configurations.
- Expertise with containerization, Kubernetes/GKE, and specialized model serving tools such as Triton, vLLM, and MLflow.
- Experience with Airflow, Vertex AI Pipelines, modern CI/CD, and Terraform for infrastructure as code.
- Proficiency in Python and SQL for scripting, automation, API development, and data manipulation.
- Hands-on experience with logging, telemetry, and drift detection using Grafana, Prometheus, GCP Cloud Monitoring, or comparable ML observability frameworks.
Nice to have
- Experience running large-scale LLM or deep learning inference and training workloads.
- GCP Professional Machine Learning Engineer or Professional Cloud Architect certification.
- Familiarity with feature stores such as Feast or Vertex AI Feature Store.
Culture & Benefits
- Work on cybersecurity products addressing real customer protection needs.
- Contribute directly to a nimble organization where individual impact is visible.
- Work with cross-functional AI research, data engineering, backend, and software engineering teams.
- Develop new technologies, products, and markets in a fast-paced, growth-oriented environment.
- Inclusive workplace committed to preventing discrimination and harassment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →