Назад
14 часов назад

Manager, ML Solutions Architecture (Token Factory)

228 000 - 285 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US/Netherlands
Вакансия из списка Hirify.GlobalВакансия из Hirify RU Global, списка компаний с восточно-европейскими корнями
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Manager, ML Solutions Architecture (Token Factory): Leading US Solutions Architecture teams delivering and supporting production deployments of open-source LLMs through serverless inference and fine-tuning with an accent on benchmarking, serving-stack optimization, and customer delivery outcomes. Focus on managing technical teams, reviewing inference methodologies, improving delivery operations, and solving complex production feasibility and escalation challenges.

Location: Remote - United States

Salary: $228,000–$285,000 USD per year

Company

Builds a full-stack AI cloud platform for data and model training, inference, and production deployment, including the Token Factory serverless platform for open-source LLM inference and fine-tuning.

What you will do

  • Manage and grow a regional team of Solutions Architects, including goal setting, performance reviews, promotion cases, onboarding, and individual development plans.
  • Own PoC delivery and post-sales technical support outcomes, including optimization timelines, success-criteria attainment, and production handoffs.
  • Review benchmarking methods, serving configurations, results, and closure documentation before customer delivery.
  • Coordinate staffing, escalations, account coverage, and delivery across teams and time zones.
  • Maintain documentation, runbooks, onboarding materials, definitions of done, closure templates, and ticket-tracking standards.
  • Translate recurring customer issues into platform priorities and provide leadership with clear assessments of feasibility, account health, and capacity.

Requirements

  • 3+ years of experience managing technical teams, including performance management and difficult conversations.
  • Experience managing a customer-facing team working on customer timelines and escalations.
  • Strong knowledge of LLM architectures, fine-tuning, evaluation design, inference internals, quantization, KV-cache management, batching, routing, and speculative decoding.
  • Experience with or working knowledge of Python, vLLM, SGLang, TensorRT-LLM, and the ability to review benchmark methodology and code.
  • Excellent communication skills for technical and executive audiences, including enterprise customers.
  • Comfort with documentation, process design, reporting, ambiguity, and distributed teams.

Nice to have

  • Experience scaling a team from 5 to 15+ while maintaining delivery quality.
  • Customer-facing technical experience at a cloud, inference, or AI infrastructure provider.
  • Hands-on experience running LLMs in production and debugging inference workloads.
  • Experience with multimodal AI models, Docker, Kubernetes, and infrastructure-as-code.

Culture & Benefits

  • Competitive compensation and benefits.
  • Career growth and learning opportunities.
  • Flexibility, ownership, and an international working environment.
  • Opportunity to work on impactful AI projects with distributed teams.
  • Applicants must be authorized to work in the country in which they apply and provide proof of employment eligibility.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →