Tech Stack
Job Description, Responsibilities & Requirements
About the Position
Data Engineer with Bigdata, Spark, Hive and Airflow | Basic knowledge working with Kubernetes IRC274090
We are seeking a highly skilled Data Engineer with expertise in Big Data technologies such as Spark, Hive, and Airflow to join our team in Noida, India. This role involves working on batch and streaming data generated from Roku TVs and other streaming devices.
Responsibilities
- Design and optimize Big Data pipelines and data sets ranging from data ingestion to processing to data visualization.
- Write and optimize Spark Jobs, Spark SQL, and Hive and SQL queries to process large data.
- Experience in using Kafka or other message brokers.
- Configure, monitor, and schedule jobs using Oozie and/or Airflow.
- Process streaming data directly from Kafka using Spark jobs.
- Handle different file formats (ORC, AVRO, and Parquet) and unstructured data.
- Experience with No SQL databases like Amazon S3 and data warehouse tools like AWS Redshift, Snowflake, or BigQuery.
- Work experience on cloud platforms such as AWS, GCP, or Azure.
Requirements
- 8+ years of hands-on experience working on Big Data Platforms.
- Good experience in any one programming language - Scala/Python, Python preferred.
- Experience in writing and optimizing complex Hive and SQL queries.
- Experience in using Kafka or any other message brokers.
- Configuring, monitoring, and scheduling of jobs using Oozie and/or Airflow.
- Processing streaming data directly from Kafka using Spark jobs.
- Should be able to handle different file formats (ORC, AVRO, and Parquet) and unstructured data.
- Should have experience with any one No SQL databases like Amazon S3.
- Should have worked on any of the Data warehouse tools like AWS Redshift, Snowflake, or BigQuery.
- Work experience on any one cloud AWS, GCP, or Azure.
Nice to Have
- Experience in AWS cloud services like EMR, S3, Redshift, EKS/ECS, etc.
- Experience in GCP cloud services like Dataproc, Google storage, etc.
- Experience in working with huge Big data clusters with millions of records.
- Experience in working with ELK stack, especially Elasticsearch.
- Experience in Hadoop MapReduce, Apache Flink, Kubernetes, etc.
We Offer
- Exciting Projects: Focus on industries like High-Tech, communication, media, healthcare, retail, and telecom.
- Collaborative Environment: Collaborate with a diverse team of highly talented people in an open, laidback environment.
- Work-Life Balance: Flexible work schedules, opportunities to work from home, and paid time off and holidays.
- Professional Development: Regular training in communication skills, stress management, professional certifications, and technical and soft skill training.
- Excellent Benefits: Competitive salaries, family medical insurance, Group Term Life Insurance, Group Personal Accident Insurance, NPS (National Pension Scheme), periodic health awareness programs, extended maternity leave, annual performance bonuses, and referral bonuses.
- Fun Perks: Sports events, cultural activities, food subsidies, corporate parties, and discounts for popular stores and restaurants.
About the Company
GlobalLogic, a Hitachi Group Company, is a trusted digital engineering partner to the world’s largest and most forward-thinking companies. Since 2000, we’ve been at the forefront of the digital revolution – helping create some of the most innovative and widely used digital products and experiences. Today we continue to collaborate with clients in transforming businesses and redefining industries through intelligent products, platforms, and services.