15 дней назад
Senior MLOps Engineer (GCP)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior MLOps Engineer (GCP/AI): Architecting and operating GCP infrastructure, deployment pipelines, and serving systems for production machine learning models with an accent on scalability, low-latency inference, and observability. Focus on building automated training and deployment workflows, monitoring model and system performance, and converting AI prototypes into secure, auto-scaling microservices.
Location: Remote Latvia
Company
develops cybersecurity solutions that help customers monitor, manage, and protect risks associated with digital identities and personal information.
What you will do
- Architect and manage scalable machine learning infrastructure on Google Cloud Platform using Vertex AI, GKE, Cloud Storage, Cloud Run, and GPU/TPU instances.
- Own the end-to-end deployment and serving lifecycle for machine learning models, including high-throughput and low-latency inference services.
- Build reproducible CI/CD/CT pipelines for model training, testing, evaluation, and deployment with Airflow, Vertex AI Pipelines, and GitHub Actions.
- Implement production observability for system health and ML metrics, including latency, throughput, prediction accuracy, feature drift, and data distribution shifts.
- Provide standardized training environments and deployment templates for AI researchers and engineers.
- Turn AI prototypes and notebooks into secure, resilient, auto-scaling microservices while integrating feature stores, dataset versioning, and stream or batch processing.
Requirements
- At least 5 years of hands-on experience designing, deploying, and maintaining production machine learning workloads in cloud environments.
- Deep experience with GCP, including Vertex AI, Cloud Storage, GKE, Cloud Run, IAM, and VPC configurations.
- Expertise with Docker, Kubernetes/GKE, Triton Inference Server, vLLM, and MLflow.
- Experience with Airflow, Vertex AI Pipelines, GitHub Actions, ArgoCD, and Terraform.
- Proficiency in Python and SQL for scripting, automation, API development, and data manipulation.
- Hands-on experience with logging, telemetry, and ML observability tools such as Grafana, Prometheus, and GCP Cloud Monitoring.
Nice to have
- Experience running large-scale LLM or deep learning inference and training workloads.
- GCP Professional Machine Learning Engineer or Professional Cloud Architect certification.
- Familiarity with feature stores such as Feast or Vertex AI Feature Store.
Culture & Benefits
- Work on cybersecurity products addressing customers’ evolving digital security needs.
- Operate in a nimble, growth-oriented environment where individual contributions have visible impact.
- Collaborate with AI researchers, data engineers, backend teams, and other technical specialists.
- Access opportunities to learn new technologies, products, and markets as the organization expands.
- Inclusive workplace committed to equal opportunity and free from discrimination and harassment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →