Deep Dive Into Spark

Apache Spark is the next-generation successor to MapReduce. Spark is a powerful, open-source processing engine for data in the Hadoop cluster, optimized for speed, ease of use, and sophisticated analytics. The Spark framework supports streaming data processing and complex, iterative algorithms, enabling applications to run up to 100x faster than traditional Hadoop MapReduce programs.

The 5 day Spark course is aimed at developers who are encountering Spark for the first time and want to understand how to build Big Data Products with Spark. The course would enable participants to build complete, unified Big Data applications combining batch, streaming, and interactive analytics on all their data.

Developers would be able to write sophisticated parallel applications to execute faster decisions, better decisions, and real-time actions, applied to a wide variety of use cases, architectures, and industries.

The course has a practical focus, mixing presentation with in-depth hands-on labs and exercises.

Proposed Structure

Day 1

Big Data Why and What?
Introduction to Spark
Spark Installation and Modes of Operation
Spark shell
Spark Fundamentals
Role of Spark Context
MapReduce in Spark
Transformations in RDD
Actions in RDD

Day 2

RDD API In Detail.
Types of RDD (Pair RDD, Numeric RDD, JDBC RDD, Key-Value etc).
Creating RDD From Different File Formats (Parquet, Avro, JSON, JDBC).
Partitions and Data Locality.
Executing parallel operations
Caching Overview
Distributed Persistence
Accumulators and Broadcast Variables
RDD Internals
RDD Lineage

Day 3

Role of SQLContext
Running Spark SQL in Spark shell
Introduction to Data Frames
Creating Data Frames
Transformations and Operations on Data Frames
Interoperating with RDDs
Creating Datasets
Difference between Data Frames and Data Sets.
Conversion from Data Frame to Dataset and vice versa.
Scheduling Across Applications
Scheduling Within Application

Day 4

Role of StreamingContext
Streaming Applications
Operations in DStreams
Sliding Window Operations
Performance Tuning of DStreams
Stateful and Stateless Transformations in DStreams.
Data Types
Basic Statistics

Day 5

Configuration of SQLContext
Web UI
Data Serialization
Memory Management
Broadcasting Large Variables
Event Logging
SSL Configuration
Standalone mode
Submitting Applications
Spark Standalone
Amazon EC2

Course Prerequisites

To benefit from this course you should have programming experience with Scala or with Python. The language of instruction is Scala. Basic Linux knowledge is expected.

For more information on the course or a discussion on your custom need, send a mail to


Thank you for showing your interest. Someone from Knoldus Inc. will get in touch with you soon!

Please fill up the form, to begin the training of leading technologies:

Thank you for your interest. Someone from Knoldus Inc. will get in touch with you soon!
Trusted by innovative organizations, big and small.
Awards and Recognitions
Upcoming Webinar:

Leveraging Lagom to Build Microservices Register Now!