Banner

Spark

Apache Spark is the next generation successor to MapReduce. Spark is a powerful, open-source processing engine for data in the Hadoop cluster, optimized for speed, ease of use, and sophisticated analytics. The Spark framework supports streaming data processing and complex, iterative algorithms, enabling applications to run up to 100x faster than traditional Hadoop MapReduce programs.

The 2 day Spark course is aimed at developers who are encountering Spark for the first time and want to understand how to build Big Data Products with Spark. The course would enable participants to build complete, unified Big Data applications combining batch, streaming, and interactive analytics on all their data.

Developers would be able to write sophisticated parallel applications to execute faster decisions, better decisions, and real-time actions, applied to a wide variety of use cases, architectures, and industries.

The course has a practical focus, mixing presentation with in-depth hands-on labs and exercises.

Prerequisites

To benefit from this course you should have programming experience with Scala or with Python. The language of instruction is Scala. Basic Linux knowledge is expected.

Day 1

First Brush

Big Data Why and What?

Introduction to Spark.

Spark shell.

Programming with Spark.

Running Spark

Jobs, Stages, Task

Web UI

Stand-alone cluster

Building and running.

RDDs

Resilient Distributed Datasets.

RDD Operations.

Map Reduce.

Key value pair.

Parallel Programming

Partitions and Data Locality.

Executing parallel operations.

Caching and Persistence

Caching Overview.

Distributed Persistence.

Day 2

Spark Streaming

Overview

Streaming operations.

Sliding window operations.

Streaming Applications.

Stateful Transformations.

Good to Know

Spark Context.

Spark Properties.

Logging.

Iterative Algorithms.

Graph Analysis.

Machine Learning.

Spark SQL.

Tunning Spark Application

Spark and Hadoop

HDFS

Using HDFS with Spark.

Spark and MapReduce.

Academy

Spark

2 Days

2 Days

Instructor-Led Course

Instructor-Led Course

Beginner

Beginner

This Spark includes:

Maximum Class Size of 15

Access to Course Materials

Certificate of Completion

Access to a Private Channel with Trainers in the Academy Slack

A Q&A session one week post-course

A pre-and-post meeting with our trainers

Let's have a conversation

Schedule a meeting