pythonfastapidockertransformersraghugging facepytorchnlpkazakh languagellm fine tuning
Вакансия из Telegram канала - Название доступно после авторизации
Пожаловаться
Зарплата и рынок
По рынку СНГ данных пока нет
На международном рынке: $119к/год ($82к - $163к)
85
Хорошая вакансия
развернуть
Роль четко определена с ясным акцентом на NLP и LLM, а зарплата достойная для региона, что делает это хорошей возможностью для специалистов в области ИИ.
Кликните для подробной информации
Четкая рольДостойная зарплатаВлиятельная работа
Оценка от Hirify AI
Мэтч & Сопровод
Покажет вашу совместимость и напишет письмо
Создать профиль и узнать мэтч
Описание вакансии
TL;DR
AI Engineer (NLP): Building Kazakh-language NLP tools, text corpora, data pipelines, and fine-tuned open-source LLMs with an accent on corpus engineering, model adaptation, and containerized deployment. Focus on fine-tuning Llama, Mistral, Qwen, and Gemma with LoRA/QLoRA, developing morphological analyzers and tokenizers, and optimizing inference services.
#вакансия #астана #ai #ml #nlp #python AI Engineer
Til-Qazyna National Scientific and Practical Center
Astana | Office: | Mon–Fri, 08:00–17:30
700,000 KZT gross
About Us & The Team
Til-Qazyna is a national scientific and practical center dedicated to building the core digital and AI infrastructure for the Kazakh language. We develop foundational language resources, linguistic tools, and open-source models, including custom text corpora, morphological analyzers, RAG pipelines, and fine-tuned Large Language Models.
We are looking for an AI Engineer to focus exclusively on natural language processing, corpus engineering, and LLM fine-tuning.
Key Responsibilities
• LLM Adaptation & Fine-Tuning: Fine-tune, evaluate, and benchmark open-source LLMs (Llama, Mistral, Qwen, Gemma) for Kazakh language understanding and generation using parameter-efficient methods (LoRA, QLoRA, PEFT).
• NLP & Linguistic Tooling: Develop, maintain, and optimize rule-based and neural NLP components, including morphological analyzers, lemmatizers, tokenizers, and terminology parsers.
• Corpus & Data Pipelines: Design automated pipelines for large-scale text crawling, data cleaning, deduplication, synthetic data generation, and dataset curation.
• Inference & Deployment: Package NLP models and LLMs into containerized microservices (Docker, FastAPI) and optimize inference using modern serving frameworks (vLLM, Ollama, TensorRT-LLM, ONNX).
Requirements
• 1+ years of hands-on experience as NLP / AI / Machine Learning engineer.
• Solid understanding of the Transformer architecture, attention mechanisms, embeddings, and tokenization.
• Experience with core ML/NLP libraries: PyTorch, Hugging Face (transformers, datasets, accelerate, peft), Scikit-learn.
• Experience building data preprocessing pipelines and working with structured/unstructured text data.
• Practical knowledge of Docker, Linux, and Git.
• Degree in Computer Science, Data Science, Applied Mathematics.
Nice to Have
• Hands-on experience with modern LLM serving and optimization engines (vLLM, TGI, llama.cpp, quantization techniques like AWQ/GPTQ).
• Experience with RAG stacks (LangChain, LlamaIndex, ChromaDB, Qdrant, FAISS).
• Understanding of the morphological, agglutinative, or grammatical characteristics of the Kazakh language.
• Familiarity with alignment techniques (DPO, RLHF) or synthetic data generation.
What We Offer
• Direct impact on sovereign AI development and language technology used at scale.
• Official employment in full compliance with the Labor Code of the Republic of Kazakhstan.
• Access to high-performance GPU compute clusters for model training and experimentation.
To Apply
Send your CV to Показать контакты on Telegram.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Текст вакансии взят без изменений
Источник - Telegram канал. Название доступно после авторизации