Назад
Company hidden
1 месяц назад

Senior AI Engineer (LLM)

200 000 - 300 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK/Singapore/US +1 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior AI Engineer (LLM): Building DV-specific model capabilities with an accent on fine-tuning and distilling open-weight models, on-prem inference infrastructure, and intelligent routing across open and closed providers. Focus on model evaluation, GPU workload management, and balancing cost, quality, latency, and data sensitivity in production.

Location: Chicago, Hong Kong, London, New York, or Singapore

Base salary: $200,000–$300,000 USD per year, plus discretionary bonus eligibility.

Company

hirify.global is a proprietary trading firm within the DV Group, providing liquidity and hedging opportunities across worldwide financial markets.

What you will do

  • Build and operate a model gateway routing inference across open and closed models while tracking cost, latency, and quality.
  • Design and run distillation pipelines that generate training data from frontier model outputs for task-specific open models.
  • Fine-tune and evaluate open-weight models such as Llama, Qwen, and Mistral for firm-specific tasks.
  • Deploy and maintain on-premise LLM inference infrastructure using vLLM, TGI, or equivalent tools on Kubernetes.
  • Build production evaluation and regression frameworks covering model quality, cost, and latency.
  • Partner with agent engineering to align the model layer with agent workloads and define criteria for selecting open or closed models.

Requirements

  • 5+ years of software engineering experience and strong Python skills.
  • Production experience fine-tuning or distilling open-weight models, beyond building inference API wrappers.
  • Experience serving LLMs on-premise with vLLM, TGI, Triton, or equivalent technology.
  • Production experience managing GPU infrastructure, including provisioning, scheduling, and utilization monitoring.
  • Experience with model evaluation and regression testing in production.
  • Experience with Kubernetes, GPU workload management, and tradeoffs between open and closed models across cost, quality, latency, and data sensitivity.

Nice to have

  • Experience with quantization, PEFT/LoRA, or other efficient training techniques.
  • Experience designing model gateways or inference proxies with routing, fallback, and rate limiting.
  • Experience in financial services or other regulated and sensitive-data environments.
  • Familiarity with the open model ecosystem, including Hugging Face, model cards, and licensing.

Culture & Benefits

  • Discretionary bonus eligibility.
  • Medical, dental, and vision insurance.
  • HSA, FSA, and dependent care options.
  • Employer-paid group term life and AD&D insurance, with voluntary LTD and additional life and AD&D options.
  • Flexible vacation policy.
  • Retirement plan with employer match.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →