Назад
Company hidden
обновлено 2 дня назад

AI Researcher (Efficient AI)

84 - 91$
Формат работы
hybrid
Тип работы
project
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Researcher (Efficient AI) (LLMs, Model Compression, On-Device AI): Developing efficient methods and prototypes that make LLMs, VLMs, multimodal models, and AI agents faster, smaller, and deployable on constrained devices with an accent on compression, quantization, inference optimization, and emerging architectures. Focus on solving long-context and KV-cache challenges, implementing low-latency inference and kernel optimizations, and evaluating methods across language, vision, reasoning, and agentic benchmarks.

Location: Hybrid in Santa Clara, California, United States

Salary: $84.13–$91.34 USD per hour

Company

hirify.global is a global technology corporation developing products and platforms across home appliances, media, vehicles, and eco solutions.

What you will do

  • Research, prototype, and implement methods that improve AI model efficiency, inference performance, and deployment on constrained devices.
  • Optimize LLMs, SLMs, VLMs, multimodal models, and agentic workloads across post-training, inference, and deployment workflows.
  • Develop and evaluate compression techniques including PTQ, QAT, pruning, and low-rank approximation.
  • Address long-context inference and KV-cache compression challenges for reasoning and agentic applications.
  • Implement efficient architectures and inference optimizations such as MoE, SSMs, speculative decoding, constrained decoding, and kernel-level optimization.
  • Build experimental pipelines, run standardized benchmarks, and contribute to publications, technical reports, open-source releases, and IP submissions.

Requirements

  • M.S. or Ph.D. in Computer Science, Computer Engineering, Machine Learning, Mathematics, or a related technical field.
  • Research or engineering experience in machine learning, efficient AI, model optimization, or AI systems.
  • Strong Python programming skills and experience with PyTorch or a comparable deep learning framework.
  • Hands-on experience with LLMs, SLMs, VLMs, multimodal models, or generative AI systems.
  • Ability to read research papers, implement technical methods, run experiments, and communicate results clearly.
  • Strong written and verbal communication skills for reports, presentations, demonstrations, and technical documentation.

Nice to have

  • Publications in reputable machine learning or systems venues.
  • Experience with llama.cpp, GGUF, vLLM, SGLang, TensorRT-LLM, or related inference and deployment systems.
  • Experience with PTQ, QAT, LoRA, distillation, instruction tuning, DPO, OPD, RLVR, or reasoning-oriented adaptation.
  • Experience with low-level kernel implementations, on-device acceleration, multi-agent systems, or hardware/software co-design.
  • Familiarity with MoE, SSMs, hybrid attention, or Looped Transformers.

Culture & Benefits

  • One-year contract with potential extension based on business needs and performance.
  • Collaboration with experienced researchers in LG's Emerging Technology Lab.
  • Opportunity to contribute to research publications, open-source projects, and intellectual property.
  • Contractors are eligible for relevant benefit programs offered through partner agencies.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →