2 Π΄Π½Ρ Π½Π°Π·Π°Π΄
ML Platform Engineer (AI)
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Π’Π΅ΠΊΡΡ:
TL;DR
ML Platform Engineer (AI) (Python/PyTorch/JAX): Building and operating the infrastructure behind AI products, from model training and evaluation to deployment, inference, and observability, with an accent on reliability, scalability, latency, and cost efficiency. Focus on designing distributed ML platforms, optimizing high-throughput model serving, and creating reproducible pipelines and evaluation systems for rapidly evolving AI workloads.
Company
is building A1, a proactive AI assistant for everyday conversations, errands, organization, and workflows.
What you will do
- Build and operate ML infrastructure and platforms powering AI products.
- Design systems for model training, evaluation, deployment, inference, and experimentation.
- Optimize model-serving infrastructure for high throughput, low latency, reliability, and cost efficiency.
- Develop reproducible pipelines for data preparation, training, evaluation, model release, and continuous improvement.
- Build evaluation, benchmarking, observability, monitoring, tracing, and alerting infrastructure for AI/ML workloads.
- Collaborate with AI engineers, researchers, and product engineers to turn model requirements into production-ready systems.
Requirements
- Strong software engineering fundamentals and experience building production systems.
- Experience with ML infrastructure, platforms, or production machine learning systems.
- Experience with model deployment, inference, evaluation, or data pipelines.
- Strong understanding of distributed systems and system reliability.
- Production-quality Python development and the ability to work effectively in ambiguous, fast-moving environments.
- Ownership, experimentation, and continuous improvement mindset.
Culture & Benefits
- Work in a fast-moving environment focused on experimentation and continuous improvement.
- Build reusable platform primitives instead of duplicating infrastructure for each AI product.
- Help evolve the AI stack as new models, architectures, and inference techniques emerge.
- Improve the reliability, scalability, observability, and maintainability of production AI systems.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β