6 ΠΌΠ΅ΡΡΡΠ΅Π² Π½Π°Π·Π°Π΄
Engineer, Supercomputing & Distributed Systems (AI)
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Π’Π΅ΠΊΡΡ:
TL;DR
Engineer, Supercomputing & Distributed Systems (AI): Building and operating the infrastructure for Krea's research and inference, including distributed training, Kubernetes GPU clusters, and petabyte-scale data pipelines with an accent on custom distributed datastores and job orchestration systems. Focus on scaling workloads and research between clusters in multiple datacenters and building fault tolerance systems for large-scale pretraining.
Location: On-site in San Francisco
Company
Krea AI is building next-generation AI creative tools, dedicated to making AI intuitive and controllable for creatives.
What you will do
- Design multi-stage pipelines that turn petabytes of raw data into clean, annotated datasets.
- Manage distributed training and inference on 1000+ GPU Kubernetes clusters.
- Profile and optimize dataloaders streaming thousands of images per second.
- Customize and train models to filter billions of images.
- Build fault tolerance systems for large-scale pretraining.
Requirements
- Experience with Python, PyArrow, DuckDB, SQL, PyTorch, Pandas, NumPy.
- Experience with Kubernetes.
- Fundamental knowledge of containerization, operating systems, file-systems, and networking.
- Intuition for distributed systems and a great mental model of how systems interact and function under different conditions.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β