Skip to main content
CF

Apache Spark Certification Training

15h 13m 1s
English
Free

Apache Spark Certification Training is a 99-lesson 15 hours 13 minutes self-paced course by Florian Roscheck. Master Apache Spark and showcase your skills with the Databricks Associate Developer for Apache Spark certification .

Course facts

Lessons
99
Duration
15 hours 13 minutes
Level
All levels
Language
English
Updated
Instructor
Florian Roscheck
Price
Free

Master Apache Spark and showcase your skills with the Databricks Associate Developer for Apache Spark certification. This course is designed to transform you into a PySpark professional and prepare you to ace the Databricks Spark certification exam.

Join us for an engaging and easy-to-understand journey into Apache Spark and elevate your big data career to new heights!

What Will You Learn?

The aim of this course is to teach you fundamental PySpark skills and equip you to achieve certification as a Databricks Certified Associate Developer for Apache Spark. The course is comprised of 18 comprehensive modules that will guide you through Apache Spark's internal workings and practical usage.

Course Highlights:

  • Develop expertise in coding with Spark DataFrames
  • Gain confidence with the Databricks certification exam content
  • Understand Spark's distributed and fault-tolerant data processing
  • Master the use of Spark in Databricks
  • Learn about the Spark cluster architecture
  • Discover when and how Spark evaluates code
  • Explore Spark's efficient memory management mechanisms
  • Resolve common Spark issues like out-of-memory errors
  • Understand how Spark handles complex operations such as joins
  • Become proficient in navigating the Spark UI
  • ...and much more – check out the full list below!

Who Is This Course For?

This course is designed for individuals with basic Python skills eager to advance their big data processing abilities through PySpark. It also targets those aiming to pass the Databricks Certified Associate Developer for Apache Spark certification.

Ideal Participants:

  • Those interested in using Apache Spark with Python and PySpark, rather than Scala
  • Data analysts and developers seeking to enhance their portfolio with verified big data skills and Databricks experience
  • Data engineers desiring certification to verify their Apache Spark skills and advance their careers
  • Data scientists aiming to work efficiently with large data sets in Apache Spark
  • Organizations seeking to empower their data professionals with effective Apache Spark skills
  • Anyone looking to strengthen their understanding of Apache Spark's inner workings

Who teaches Apache Spark Certification Training? Florian Roscheck

Florian Roscheck thumbnail

Florian Roscheck is a data engineer and educator focused on the Apache Spark ecosystem and the certification preparation around it.

His CourseFlix listing carries Apache Spark Certification Training — a structured preparation course for the Databricks-administered Spark certification, covering Spark SQL, DataFrames, partitioning, and the performance-tuning material the cert exam tests on.

Material is paid and aimed at data engineers preparing for the Spark certification or otherwise picking up Spark for production data pipelines. For broader data-engineering content, see CourseFlix's Data processing and analysis category page.

What lessons are included in Apache Spark Certification Training?

  • Space or K: play or pause
  • J: rewind 10 seconds
  • L: forward 10 seconds
  • Left Arrow: rewind 5 seconds
  • Right Arrow: forward 5 seconds
  • Up Arrow: volume up
  • Down Arrow: volume down
  • M: mute or unmute
  • F: toggle fullscreen
  • T: toggle theater mode
  • I: toggle mini player
  • 0 to 9: seek to 0 to 90 percent of the video
  • Shift plus N: next video
  • Shift plus P: previous video
0:00 0:00
#Lesson TitleDuration
101. Introduction 09:48
202. Certification Exam Overview 05:03
303. Signing up for Databricks Community Edition 01:44
404. Loading Data Into Databricks 02:43
505. Overview of the Spark Cluster Architecture and its Components 08:24
606. Getting to Know the Spark Driver 11:55
707. Getting to Know Executors 07:37
808. Discovering Execution Modes 17:33
909. Overview 05:38
1010. Internal Types, DataFrames, Datasets, RDDs, and the Spark SQL API 19:10
1111. Hands-on Session_ Exploring Data APIs on Databricks Community Edition 08:26
1212. Intro to Labs 01:12
1313. Intro & Creating DataFrames 06:58
1414. Exercise_ Creating a DataFrame 01:07
1515. Exercise_ Creating a DataFrame - Solution 01:59
1616. Working with Schemas 26:10
1717. Exercise_ Building a Simple Schema 01:46
1818. Exercise_ Building a Simple Schema - Solution 05:13
1919. Exercise_ Building a Complex Schema 02:28
2020. Exercise_ Building a Complex Schema - Solution 05:53
2121. Type Conversion of DataFrame Columns 07:20
2222. Exercise_ Changing the Type of a Column 01:50
2323. Exercise_ Changing the Type of a Column - Solution 04:20
2424. Overview 09:18
2525. Shuffles 07:52
2626. Data Skew 13:15
2727. Spark Configurations for Partitions 03:47
2828. Hands-on Session_ The Power of Partitions 30:18
2929. Storage Layout 17:39
3030. Caching and Storage Levels 10:29
3131. Memory in Action 30:59
3232. Hands-on Session_ Executor Memory Management - Part 1 10:27
3333. Hands-on Session_ Executor Memory Management - Part 2 13:08
3434. Intro & How to Get Help in PySpark 04:01
3535. Partitioning Recap 09:45
3636. Exercise_ Repartitioning 01:31
3737. Exercise_ Repartitioning - Solution 06:08
3838. Caching Recap 03:27
3939. Exercise_ Caching 01:13
4040. Exercise_ Caching - Solution 03:20
4141. Overview 07:40
4242. Hands-On Session_ Actions vs. Transformations 06:47
4343. Intro & Reading Data 18:36
4444. Exercise_ Reading Parquet Files 02:20
4545. Exercise_ Reading Parquet Files - Solution 03:44
4646. Reading from CSV Files 17:18
4747. Exercise_ Reading CSV Files 02:29
4848. Exercise_ Reading CSV Files - Solution 03:55
4949. Reading from JSON Files 05:16
5050. Writing Data 10:57
5151. Exercise_ Writing to Parquet Files 02:08
5252. Exercise_ Writing to Parquet Files - Solution 04:27
5353. Writing to CSV Files 02:53
5454. Exercise_ Writing to CSV Files 02:16
5555. Exercise_ Writing to CSV Files - Solution 03:12
5656. Writing to JSON Files 01:58
5757. Using PySpark with SQL 05:01
5858. Exercise_ SQL in PySpark 00:46
5959. Exercise_ SQL in PySpark - Solution 02:16
6060. Overview 16:33
6161. Hands-on Session_ Discovering the Spark UI 12:27
6262. Intro & Removing Data 16:58
6363. Exercise_ Removing Data 00:59
6464. Exercise_ Removing Data - Solution 03:16
6565. Modifying Data 30:49
6666. Exercise_ Modifying Data 02:08
6767. Exercise_ Modifying Data - Solution 07:22
6868. Analyzing Data 18:14
6969. Exercise_ Analyzing Data 01:39
7070. Exercise_ Analyzing Data - Solution 06:30
7171. The Catalyst Optimizer 18:32
7272. Adaptive Query Execution 15:32
7373. Dynamic Partition Pruning 10:08
7474. The DAG_ Achieving Fault Tolerance 12:25
7575. Intro & Working With Dates and Times 33:30
7676. Exercise_ Working With Dates and Times 02:10
7777. Exercise_ Working With Dates and Times - Solution 08:00
7878. Working With Strings 15:30
7979. Exercise_ Working With Strings 03:20
8080. Exercise_ Working With Strings - Solution 07:47
8181. Working with Arrays 14:38
8282. Exercise_ Working With Arrays 05:17
8383. Exercise_ Working With Arrays - Solution 13:19
8484. Accumulator and Broadcast Variables 11:14
8585. Joins 34:02
8686. Hands-on Session_ Cross-Cluster Communication 42:39
8787. Intro & Grouping and Aggregating 19:16
8888. Exercise_ Grouping and Aggregating 01:43
8989. Exercise_ Grouping and Aggregating - Solution 07:19
9090. Joining 15:06
9191. Exercise_ Joining 03:58
9292. Exercise_ Joining - Solution 03:58
9393. User-Defined Functions (UDFs) 20:29
9494. Exercise_ UDFs 04:06
9595. Exercise_ UDFs - Solution 17:51
9696. Signing up for the Exam 02:24
9797. Last Minute Preparations 01:34
9898. Introduction 04:36
9999. Congratulations! 00:50

Books

Read Book Apache Spark Certification Training

#TitleTypeOpen
11-Proposed Timeline PDF
2Apache Spark Certification Exam Guide PDF
3Mastery Map 1 - Cluster Components PDF
4Mastery Map 2 - Spark Execution Modes PDF
5Mastery Map 3 - Spark Data APIs PDF
6Mastery Map 4 - Executor Memory Layout PDF
7Mastery Map 5 - PySpark Storage Levels PDF
8Mastery Map 6 - Executor Out-of-Memory Errors PDF
9Mastery Map 7 - Actions Vs. Transformations PDF
10Mastery Map 8 - Execution Hierarchy PDF
11Mastery Map 9 - A Query, From Plan to Execution PDF
12Mastery Map 10 - Adaptive Query Execution Strategies PDF
13Mastery Map 11 - Dynamic Partition Pruning PDF
14Mastery Map 12 - Joins PDF

What courses are similar to Apache Spark Certification Training?

Frequently asked questions

What prerequisites should I have before taking this course?
Before enrolling, you should have basic Python skills, as the course builds on this foundation to teach PySpark. Familiarity with data processing concepts would be beneficial, but the course starts with introductory modules such as 'Introduction' and 'Certification Exam Overview' to help you get up to speed.
What projects or hands-on exercises will I complete during the course?
The course includes various hands-on sessions and exercises to reinforce learning. You will work on creating and manipulating DataFrames, building simple and complex schemas, changing column types, managing memory, and working with different data formats like Parquet and CSV. These practical exercises ensure you gain real-world experience with PySpark.
Who is the target audience for this course?
This course is intended for individuals with a desire to enhance their big data processing skills through PySpark. It's well-suited for those aiming to achieve the Databricks Certified Associate Developer for Apache Spark certification, especially if they have a foundational knowledge of Python.
How does this course compare to other big data processing courses?
The course is specifically focused on preparing students for the Databricks Certified Associate Developer for Apache Spark certification. Unlike general big data courses, this program offers targeted training on Spark's internal workings, practical usage, and Databricks integration, making it ideal for those pursuing certification as a Spark developer.
What specific tools or platforms will I learn to use in the course?
The course emphasizes the use of Databricks and Apache Spark. You'll learn how to sign up for the Databricks Community Edition and perform various tasks like loading data, creating DataFrames, and navigating the Spark UI. These tools are integral to learning PySpark and preparing for the certification exam.
What topics are not covered in this course?
The course does not cover topics outside the scope of PySpark and the Databricks platform. It focuses on Spark's architecture, functions, and practical data processing techniques. Topics like machine learning with Spark MLlib or integration with other big data tools are not included.
What is the expected time commitment to complete the course?
The course consists of 99 lessons, including hands-on sessions and exercises, with a focus on practical application and exam preparation. Although the total runtime is listed as 00:00:00, students should anticipate dedicating several weeks to complete the modules comprehensively, especially if they are preparing for the certification exam.