Назад
Company hidden
4 дня назад

Edge AI/Model Optimization Engineer

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Edge AI/Model Optimization Engineer (LLMs/Edge Computing): Deploying, tuning, benchmarking, and sustaining AI models and inference services on constrained GPU-enabled edge platforms such as the X9 Spider Mission Computer with an accent on quantization, runtime configuration, and operational reliability. Focus on building stress-testing frameworks, optimizing latency and resource utilization in disconnected and low-bandwidth environments, and validating model-serving deployments for tactical missions.

Location: Aberdeen, Maryland, United States; Workplace: Hybrid

Company

hirify.global provides advanced software solutions and professional services for mission and business support, including the ReadiChat agentic AI platform.

What you will do

  • Evaluate, tune, benchmark, and operationalize LLMs, embedding models, and inference services on GPU-enabled edge platforms, including the X9 Spider Mission Computer.
  • Optimize quantization, batching, context windows, caching, inference scheduling, and GPU memory utilization for constrained hardware.
  • Benchmark agentic AI workflows, inference pipelines, and model-serving architectures against latency, throughput, reliability, and resource constraints.
  • Package, deploy, validate, troubleshoot, and sustain local model-serving components in tactical, airborne, edge, and disconnected environments.
  • Build repeatable performance and stress-testing frameworks covering failover, degraded resources, tool-call overhead, and disconnected operations.
  • Collaborate with AI engineers, systems integrators, mission stakeholders, and operational users; train technical personnel and maintain deployment and optimization documentation.

Requirements

  • Bachelor’s degree in Computer Science, Electrical Engineering, Computer Engineering, Data Science, Artificial Intelligence, or a related technical discipline.
  • 5+ years of experience in AI/ML deployment, model optimization, edge computing, GPU acceleration, or AI inference operations.
  • Experience deploying and optimizing LLMs, embedding models, or inference pipelines in resource-constrained environments.
  • Experience with GPU-enabled systems and inference technologies such as CUDA, TensorRT, ONNX Runtime, vLLM, Ollama, or equivalent platforms.
  • Experience with Linux, Docker, Kubernetes, Python, AI/ML deployment frameworks, benchmarking, and runtime performance optimization.
  • Active security clearance is required.

Nice to have

  • Experience with tactical, airborne, mission-command, or DDIL environments.
  • Familiarity with X9 Spider or similar embedded GPU-enabled mission systems and operational AI ecosystems.
  • Experience with INT8, FP16, GGUF, GPTQ, AWQ, or similar model quantization approaches.
  • Experience conducting hardware evaluations and performance trade studies for edge compute systems.

Culture & Benefits

  • Full-time employment in a hybrid work environment.
  • Work focused on operational AI capabilities and mission-supporting technologies.
  • Culture emphasizing fairness, respect, accountability, adaptability, and employee contributions.
  • Environment that encourages participation at all levels, innovation, and responsible risk-taking.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →