4 дня назад
Edge AI/Model Optimization Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Edge AI/Model Optimization Engineer (LLMs/Edge Computing): Deploying, tuning, benchmarking, and sustaining AI models and inference services on constrained GPU-enabled edge platforms such as the X9 Spider Mission Computer with an accent on quantization, runtime configuration, and operational reliability. Focus on building stress-testing frameworks, optimizing latency and resource utilization in disconnected and low-bandwidth environments, and validating model-serving deployments for tactical missions.
Location: Aberdeen, Maryland, United States; Workplace: Hybrid
Company
provides advanced software solutions and professional services for mission and business support, including the ReadiChat agentic AI platform.
What you will do
- Evaluate, tune, benchmark, and operationalize LLMs, embedding models, and inference services on GPU-enabled edge platforms, including the X9 Spider Mission Computer.
- Optimize quantization, batching, context windows, caching, inference scheduling, and GPU memory utilization for constrained hardware.
- Benchmark agentic AI workflows, inference pipelines, and model-serving architectures against latency, throughput, reliability, and resource constraints.
- Package, deploy, validate, troubleshoot, and sustain local model-serving components in tactical, airborne, edge, and disconnected environments.
- Build repeatable performance and stress-testing frameworks covering failover, degraded resources, tool-call overhead, and disconnected operations.
- Collaborate with AI engineers, systems integrators, mission stakeholders, and operational users; train technical personnel and maintain deployment and optimization documentation.
Requirements
- Bachelor’s degree in Computer Science, Electrical Engineering, Computer Engineering, Data Science, Artificial Intelligence, or a related technical discipline.
- 5+ years of experience in AI/ML deployment, model optimization, edge computing, GPU acceleration, or AI inference operations.
- Experience deploying and optimizing LLMs, embedding models, or inference pipelines in resource-constrained environments.
- Experience with GPU-enabled systems and inference technologies such as CUDA, TensorRT, ONNX Runtime, vLLM, Ollama, or equivalent platforms.
- Experience with Linux, Docker, Kubernetes, Python, AI/ML deployment frameworks, benchmarking, and runtime performance optimization.
- Active security clearance is required.
Nice to have
- Experience with tactical, airborne, mission-command, or DDIL environments.
- Familiarity with X9 Spider or similar embedded GPU-enabled mission systems and operational AI ecosystems.
- Experience with INT8, FP16, GGUF, GPTQ, AWQ, or similar model quantization approaches.
- Experience conducting hardware evaluations and performance trade studies for edge compute systems.
Culture & Benefits
- Full-time employment in a hybrid work environment.
- Work focused on operational AI capabilities and mission-supporting technologies.
- Culture emphasizing fairness, respect, accountability, adaptability, and employee contributions.
- Environment that encourages participation at all levels, innovation, and responsible risk-taking.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →