Senior AI Platform & Agentic Infrastructure Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Senior AI Platform & Agentic Infrastructure Engineer (AI): Inherit, operate, and progressively migrate an internal multi-agent platform into a regulator-grade production standard with an accent on resilient cloud infrastructure, agent runtime + evaluation harnesses, and governed data/Responsible-AI controls. Focus on building end-to-end systems that remain trustworthy under regulated audit constraints, including HA/DR, observability, provenance/lineage, and threat modeling for agent access to untrusted content.
Company
OKX builds crypto exchange and wallet products, including an AI-native internal audit multi-agent platform.
What you will do
- Operate and migrate a Google Workspace–native prototype estate (Apps Script/Drive/Docs-Sheets-Slides APIs and Claude Code tooling) to the target platform without interrupting daily and board-cycle workflows.
- Re-architect Hive Mind into a resilient AWS or GCP production platform with HA, DR, SLOs, and full observability.
- Build the agentic runtime and evaluation harness: orchestration, multi-model routing, tool/MCP/plugin integration, and red-team + regression gates before and after production.
- Implement the Responsible-AI and model-governance layer: hallucination/bias/drift controls, output validation, guardrails, and complete logging.
- Engineer governed data infrastructure with provenance and lineage, plus data protection (encryption, secrets, least-privilege, retention, and residency).
- Stand up CAAT and continuous-monitoring data foundations and own the cloud foundation (IaC, CI/CD, identity/networking, observability, and cost controls) so the platform is examinable.
Requirements
- 7+ years building and operating resilient backend or platform systems in production, including on-call ownership over time.
- Proven brownfield migrations: taking a founder-built/prototype system to production grade while it stays in daily use.
- Strong engineering fundamentals: fluent Python, strong SQL and data modeling, plus one additional systems language (TypeScript/Node or Go) and ability to connect infrastructure end to end.
- Proven agentic runtime and harness engineering: built model-agnostic agent orchestration, model routing and evaluation across providers, agent SDKs (Claude Agent SDK or equivalent), MCP servers, skill/hook-based tooling, and evaluation + red-team harnesses with Responsible-AI controls.
- Data engineering with provenance/lineage and security/data-protection engineering by default (encryption, IAM, secrets, retention) with systems designed to withstand external audit.
- Cloud resilience engineering on AWS or GCP: Terraform, CI/CD, HA/DR, SLOs, observability, automated testing, and documentation; ability to ship inside locked-down enterprise security environments.
Culture & Benefits
- Work on an internal AI-native audit capability with a small two-person engineering team responsible for deployment, debugging, and hotfixes.
- Production-grade reliability expectations: recovery measured in hours and board-cycle windows treated as critical.
- Competitive total compensation, including performance bonus and long-term incentives.
- L&D programs and education subsidy for growth.
- Comprehensive healthcare coverage for employees and dependants; wellness and meal allowances.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →