Data Engineering with Hadoop

7h 3m
English
Paid

Big Data is not just a buzzword, but a real phenomenon. Every day, companies around the world collect and process vast amounts of data at high speeds. This data is often unstructured and inconsistent, making it nearly impossible to process using traditional methods.

One of the platforms that has proven itself for working with big data is Apache Hadoop. This is an open-source framework in Java that allows processing and storing large volumes of data in clusters using simple programming models. Hadoop is a flexible, fast, and affordable architecture capable of detecting and handling failures at the application level.

Read more about the course

What You Will Learn

In this course led by Suyog Nagaokar, you will gain a comprehensive understanding of the Hadoop architecture and its components:

  • HDFS
  • YARN
  • MapReduce
  • Hive
  • Sqoop

The course includes theoretical foundations and practical lab exercises. You will learn to:

  • Understand the concept of the Hadoop ecosystem
  • Use basic Hadoop commands
  • Implement solutions based on each Hadoop component to solve real business problems

You will install and configure a full Hadoop environment using Cloudera Quickstart VM right on your computer. In practice, you will learn to:

  • Store and query data using Sqoop, Hive, and MySQL
  • Write Hive queries to analyze data on Hadoop
  • Work with data clusters using HDFS, MapReduce, and YARN
  • Manage clusters using Hue

Requirements

  • A PC with a 64-bit version of Windows or Linux and internet access
  • At least 8 GB of free (not total) RAM to complete practical tasks (if less, you can follow along with the training but without practice)
  • Basic programming skills, preferably in Python
  • Familiarity with the Linux command line will be a big plus

The course is suitable for both beginners and those who want to deepen their knowledge in Big Data and learn to work with one of the most popular frameworks in the industry.

Watch Online Data Engineering with Hadoop

Join premium to watch
Go to premium
# Title Duration
1 What can you expect from this course? 02:10
2 Introduction to Big Data 14:50
3 What is Hadoop? Why Hadoop? 05:38
4 Hadoop Architecture – Overview 02:39
5 Hadoop Architecture – Key services 07:13
6 Storage/Processing characteristics 07:51
7 Store and process data in HDFS 03:56
8 Handling failures - Part 1 05:10
9 Handling failures - Part 2 07:33
10 Rack Awareness 05:59
11 Hadoop 1 v/s Hadoop 2 12:51
12 Hadoop Ecosystem 03:36
13 Vanilla/HDP/CDH/Cloud distributions 10:12
14 Install Cloudera Quickstart Docker 07:19
15 Hands-on with Linux and Hadoop commands 05:49
16 Hive Overview 04:54
17 How Hive works 05:57
18 Hive query execution flow 04:59
19 Creating a Data Warehouse & Loading data 05:10
20 Creating a Hive Table 21:19
21 Load data from local & HDFS 17:19
22 Internal tables vs External tables 17:20
23 Partitioning & Bucketing. (Cardinality concept) 16:24
24 Static Partitioning - Lab 14:58
25 Dynamic Partitioning - Lab 13:55
26 Bucketting - Lab 22:32
27 Storing Hive query output 11:34
28 Hive SerDe 14:26
29 ORC File Format 14:10
30 Sqoop overview 03:52
31 Sqoop list-databases and list-tables 06:31
32 Scoop Eval? 03:59
33 Import RDBMS table with Sqoop 11:40
34 Handling parallelism in Sqoop 09:02
35 Import table without primary key 11:01
36 Custom Query for Sqoop Import 08:48
37 Incremental Sqoop Import - Append 09:52
38 Incremental Sqoop Import - Last Modified 13:55
39 Scoop Job 08:01
40 Sqoop Import to a Hive table 10:59
41 Sqoop Import all tables - Part 1 06:20
42 Sqoop Import all tables - Part 2 14:03
43 Sqoop Export 06:14
44 Export Hive table 04:36
45 Export with Staging table 06:24

Similar courses to Data Engineering with Hadoop

Machine Learning with Python : COMPLETE COURSE FOR BEGINNERS

Machine Learning with Python : COMPLETE COURSE FOR BEGINNERSudemy

Category: Python, Data processing and analysis
Duration 13 hours 12 minutes 31 seconds
Mathematical Foundations of Machine Learning

Mathematical Foundations of Machine Learningudemy

Category: Python, Data processing and analysis
Duration 16 hours 25 minutes 26 seconds
Introduction to Data Engineering 2025

Introduction to Data Engineering 2025Andreas Kretz

Category: Data processing and analysis
Duration 44 minutes 26 seconds
Data Engineering on GCP

Data Engineering on GCPAndreas Kretz

Category: Data processing and analysis
Duration 1 hour 17 minutes 33 seconds
Deep Learning A-Z™: Hands-On Artificial Neural Networks

Deep Learning A-Z™: Hands-On Artificial Neural Networksudemy

Category: Python, Data processing and analysis
Duration 22 hours 36 minutes 30 seconds
Apache Spark Certification Training

Apache Spark Certification TrainingFlorian Roscheck

Category: Python, Data processing and analysis
Duration 15 hours 13 minutes 1 second
Getting Started with Embedded AI | Edge AI

Getting Started with Embedded AI | Edge AIudemy

Category: Data processing and analysis
Duration 3 hours 33 minutes 42 seconds
MongoDB Fundamentals

MongoDB FundamentalsAndreas Kretz

Category: MongoDB, Data processing and analysis
Duration 1 hour 23 minutes 19 seconds
Case Study in Product Data Science

Case Study in Product Data ScienceLunarTech

Category: Data processing and analysis
Duration 1 hour 4 minutes 47 seconds
Deep Learning: Advanced Computer Vision

Deep Learning: Advanced Computer Visionudemy

Category: Data processing and analysis
Duration 15 hours 10 minutes 54 seconds