Live opening · Posted 14 days ago

Azure Data Engineer

IT OPENDOORS LLC · Hyderabad, Telangana, India (Hybrid)
Linkedin No
You are 14 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 14 days ago
CompanyIT OPENDOORS LLC
LocationHyderabad, Telangana, India (Hybrid)
Work modeNo
SourceLinkedin
Listed14 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
5 min from Linkedin publishing this role to us finding it
2 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
71,224 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Job Description :
We are currently looking for a AZURE Data Engineer for the client PepsiCo. This is a full-time and Hybrid position .
Required Technical Skill Set and Tools:
Programming:
1. Programming Languages: Python (OR) Scala (OR) Java
2. Database Languages: SQL (OR) HQL (OR) PL/SQL (OR) Any SQL like (TSQL , Teradata SQL so on)
3. Scripting Languages: Shell Scripting (OR) UNIX Shell Scripting (OR) Power Shell Scripting
4. ETL data pipeline building experience
ETL:
5. ETL Tools: Informatica (OR) Infoworks (OR) DataStage (OR) Abinitio
6. ETL Data Pipelines: Build data pipelines Using PySpark (Spark or Scala or Java) for data processing, Cleansing, transformation and loading
Tools:
7. Spark (OR) PYSpark (OR) ScalaSpark (OR) Databricks Spark
8. Hive (OR) Any SQL related databases
9. Sqoop
10. Kafka (OR) Ni-Fi (Streaming Technologies)
11. No-SQL Databases like Mongo DB, Cosmo DB so on
12. Informatica,
13. Infoworks
14. Oozie
15. SPARK on GCP is called DATAPROC
Cloud Technologies:
AZURE Databricks/Azure HDI (OR) Google Cloud Platform (GCP) (OR) AWS Databricks/AWS EMR
Cloud SQL databases: Snowflake Cloud (OR) Teradata Cloud
Machine Learning (ML)/ Gen AI Job Duties:
1. Performing in development language environments- e.g. Python, Java, Scala, R, SQL, etc. and applying analytical methods to large and complex datasets leveraging one of those languages
2. Experience in machine learning, natural language processing and deep learning.
3. Candidate should have a working exposure to Generative AI based projects that includes designing and implementing solutions based on Langchain framework and designing efficient prompt for LLM’s.
4. Candidate should have good experience in pre-training and fine-tuning Large Language Models (LLM’s) on HuggingFace models & other Large Language Models.
5. Proven ability with NLP and text-based extraction techniques.
6. Familiarity with deep learning architectures used for text analysis, computer vision and signal processing.
7. Understanding of not only how to develop data science analytic models but how to operationalize these models so they can run in an automated context
8. Utilizing and applying knowledge commonly used data science packages including Spark, Pandas, SciPy, and NumPy.
9. Experience manipulating and analyzing complex, high-volume, high-dimensionality data from varying sources
10. Applying techniques such as multivariate regressions, Bayesian probabilities, clustering algorithms, machine learning, dynamic programming, stochastic-processes, queuing theory, algorithmic
knowledge to efficiently research and solve complex development problems and application of engineering methods to define, predict and evaluate the results obtained.
11. Developing and deploying A.I. solutions as part of a larger automation pipeline
12. Demonstrates extensive abilities and/or a proven record of success in the application of statistical modelling, algorithms, data mining and machine learning algorithms problem solving
Cloud Data Engineer Job Duties:
1. Analysis of Business requirement to understand what users need, clarify with users on open questions and resolutions.
2. Design and develop spark Dataframes/RDD, real time streaming application using python and Scala programming languages.
3. Design and develop Hadoop Map Reduce, real time streaming application using Pytho/Scala/Java programing language, Kafka and Flume.
4. Design and Develop Apache Hive tables, hive external tables, partitioning, bucketing, performance tuning.
5. Analyze data volume and design / develop various file formats like parquet, Avro etc.
6. Develop and implemented Databricks, Hadoop and Big data applications and processes utilizing tools and technologies such as Databricks(Azure/AWS) Streaming, Spark, Scala, Python, HDFS, Hive, Sqoop, Oozie, Kafka, Flume, Zookeeper, etc.
7. Schedule Spark and hive jobs in scheduler using Tidal, Control-M and Autosys.
8. Write SparkSQL queries on source data to analyses data in Google Big Query, Apache Hive and Google Big Table databases.
9. Write Spark Dataframe scripts as required, Issue analysis and providing resolutions by implementing technical solution.
10. Maintain in-depth knowledge of IT industry best practices, technologies, architectures and emerging technologies.
11. Design of Conceptual, Logical and Physical Datamodels for OLTP Databases and OLAP Data Warehouses across various data domains like lifesciences, Retail, Insurance so on.
12. Proficient in Data Modeling techniques using Star Schema, Snowflake Schema, Fact and Dimension tables.
13. Expert in Business Modeling Techniques using Process Flow Modeling and Data Flow Modeling.
14. Extensive experience of using data modelling tools like ER Studio Data Architect and ERwin in Relational and Dimensional Data modeling for creating Logical and Physical Design of Database and ER Diagrams using ER Studio data modeling tool, this includes creating the base tables, associative tables, and views, and choosing and creating all indexes.
13. Work and deliver forward and reverse engineering processes. Create DDL scripts for implementing Data modeling changes. Created ER Studio reports in HTML, RTF format depending upon the requirement, published Data model in model mart, created naming convention files and replicated these data model changes in the database.
14. Work on writing functional specifications, translating business requirements to technical specifications, created/maintained/modified data base design document with detailed description of logical entities and physical tables.

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App