Back
Company hidden
updated 25 days ago

Big Data Engineer - Intern (AWS/Spark)

Work format
onsite
Work type
fulltime
Grade
trainee
English
b2
Country
China
This vacancy is from Hirify.Global listVacancy from Hirify Global, list of international tech companies
Plus is required to make matches and apply

Match & Cover letter

Plus required for matching with this vacancy

Job description

Text:
/
TL;DR
Big Data Engineer - Intern (AWS/Spark): Developing and maintaining big data platforms, data lakes, data warehouses, and end-to-end data pipelines with an accent on Spark, AWS, Airflow, and machine learning platforms. Focus on designing scalable ETL and machine learning pipelines, administering high-scale data infrastructure, and supporting analytics for high-tech manufacturing.

Location: Onsite in Wuxi, China. Travel: None.

Company

hirify.global develops storage solutions and data technologies for managing large-scale data growth, serving customers and businesses globally.

What you will do

  • Help develop and maintain big data platforms, including data lakes, data warehouses, and data integration systems.
  • Apply big data architecture and administration expertise across AWS EMR, Hadoop, AWS S3, Databricks, and related technologies.
  • Develop and manage Spark ETL frameworks, orchestrate data workflows with Airflow, and support Presto/Trino query development.
  • Design, scale, and deploy machine learning pipelines using platforms such as Spark ML, H2O, and KNIME.
  • Collaborate with application architects and business subject-matter experts on end-to-end data pipelines and supporting infrastructure.
  • Build productive relationships with peer organizations, partners, and software vendors.

Requirements

  • Excellent coding skills in one or more programming languages and willingness to learn new technologies.
  • Experience or strong skills in large-scale data engineering, cloud technologies, and the Hadoop ecosystem.
  • Knowledge of Spark, Hadoop, Hive, Kafka, and EMR.
  • Experience with cloud-based big data solutions, data warehouse appliances, data lakes, and machine learning or data science platforms.
  • Proficiency in Python, Java, or Scala, along with familiarity with DevOps, continuous delivery, and Agile development.
  • Strong communication, collaboration, problem-solving, and learning skills.

Nice to have

  • Understanding of microservices and container-based development with Docker and Kubernetes.
  • Experience in a software product development environment.

Culture & Benefits

  • Collaborative work with business groups and peer engineering organizations.
  • Onsite canteen, grab-and-go market, and coffee shop.
  • Basketball, badminton, yoga, and group exercise activities.
  • Music, dance, photography, literature, and Toastmasters clubs.
  • Onsite festivals, celebrations, and community volunteering opportunities.

Be careful: if the employer asks you to log into their system using iCloud/Google, send codes/passwords, or run code/software, don't do it - these are scammers. Always click "Report" or contact support. More in guide →