Knoldus Inc

Advanced Spark (Scala/Java)

Apache Spark is the next generation successor to MapReduce. Spark is a powerful, open-source processing engine for data in the Hadoop cluster, optimized for speed, ease of use, and sophisticated analytics. The Spark framework supports streaming data processing and complex, iterative algorithms, enabling applications to run up to 100x faster than traditional Hadoop MapReduce programs.

The 2 day Spark course is aimed at developers who are encountering Spark for the first time and want to understand how to build Big Data Products with Spark. The course would enable participants to build complete, unified Big Data applications combining batch, streaming, and interactive analytics on all their data.

Developers would be able to write sophisticated parallel applications to execute faster decisions, better decisions, and real-time actions, applied to a wide variety of use cases, architectures, and industries.

The course has a practical focus, mixing presentation with in-depth hands-on labs and exercises.

Prerequisities

Day 1

First Brush

RDD Fundamentals

Programming with Spark

Day 2

RDDs

Coaching and Persistence

Parallel Programming

Advanced Concepts of RDD

Day 3

Spark SQL

Datasets

Data Frames

Spark Schedulers

Day 4

Spark Streaming

Spark MLLib

DStreams

Day 5

Clustering

Monitoring

Tunning and Debugging

Security

Deployment

Advanced Spark (Scala/Java)

This Advanced Spark includes:

STAY UPDATED ON UPCOMING EVENTS

STAY UPDATED ON UPCOMING EVENTS