1 день назад
Junior Data Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Junior Data Engineer (AI): Building and maintaining cloud-based data pipelines and AI-ready datasets for analytics, machine learning, and GenAI workloads with an accent on data warehousing, orchestration, data quality, and feature-oriented modeling. Focus on designing scalable batch and near-real-time pipelines, preparing training and inference data, and enabling search and retrieval with embeddings and metadata.
Location: San Antonio, TX, United States; full-time onsite role
Company
is expanding its Data team to build data solutions that support analytics, machine learning, and AI workloads.
What you will do
- Design, develop, and maintain scalable, secure, and cost-efficient pipelines for structured and unstructured data.
- Build and manage cloud-native data warehouse, lakehouse, and streaming architectures.
- Implement ingestion and ELT pipelines using Openflow, Snowpipe-style services, third-party frameworks, and Lambda-based pipelines.
- Collaborate with data scientists, ML engineers, and business stakeholders on feature, training, and inference data requirements.
- Develop AI-ready datasets, including feature tables, historical snapshots, time-aware datasets, text datasets, embeddings, and metadata.
- Improve data quality, observability, governance, security, performance, reliability, scalability, and cost efficiency.
Requirements
- Bachelor’s degree or higher in Computer Science, Engineering, Data Science, or a related field, with one to three years of experience.
- At least one year of experience with modern data warehousing best practices and advanced SQL.
- Experience with PostgreSQL, MySQL, Snowflake, Redshift, or similar database platforms.
- Hands-on experience building cloud-based data pipelines and using orchestration tools such as Airflow or AWS Step Functions.
- Knowledge of data modeling, ETL/ELT patterns, data quality frameworks, machine learning data pipelines, feature engineering, and training data preparation.
- Proficiency in Python, Java, or Scala, plus strong problem-solving, communication, collaboration, and Agile/Scrum skills.
Nice to have
- Familiarity with feature stores, vector databases, MLOps, data versioning, data lineage, and reproducible pipelines.
- Experience with DBT, AWS Glue, SSIS, Fivetran, or similar ingestion and transformation tools.
- Cloud certification in AWS, Azure, or GCP.
Culture & Benefits
- Mentorship and a supportive, collaborative environment focused on continuous learning.
- Career growth opportunities, a Leadership Academy, and a Mentor Program.
- Continuing education and career certification opportunities.
- Healthcare coverage options, traditional and Roth 401(k) plans, and a wellness program.
- Work/life balance, employee engagement activities, recognition awards, and years-of-service awards.
- Pre-employment drug testing is required; tobacco users are not hired where permitted by law.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →