Conceptualizing the Processing Model for Apache Spark Structured Streaming

Add to wishlistAdded to wishlistRemoved from wishlist 0

Add to compare+

Duration	2h 57m
level	Intermediate
Course Creator	Janani Ravi
Last Updated	18-Sep-20

Pluralsight

Category: Data Engineering

Much real-world data is available in streams; from self-driving car sensors to weather monitors. Apache Spark 2 is a strong analytics engine with first-class support for streaming operations using micro-batch and continuous processing.

Add your review

Description
Reviews (0)

Structured Streaming in Spark 2 is a unified model that treats batch as a prefix of stream. This allows Spark to perform the same operations on streaming data as on batch data, and Spark takes care of the details involved in incrementalizing the batch operation to work on streams. In this course, Conceptualizing the Processing Model for Apache Spark Structured Streaming, you will use the DataFrame API as well as Spark SQL to run queries on streaming sources and write results out to data sinks. First, you will be introduced to streaming DataFrames in Spark 2 and understand how structured streaming in Spark 2 is different from Spark Streaming available in earlier versions of Spark. You will also get a high level understanding of how Spark’s architecture works, and the role of drivers, workers, executors, and tasks. Next, you will execute queries on streaming data from a socket source as well as a file system source. You will perform basic operations on streaming data using Data frames and register your data as a temporary view to run SQL queries on input streams. You will explore the append, complete, and update modes to write data out to sinks. You will then understand how scheduling and checkpointing works in Spark and explore the differences between the micro-batch mode of execution and the new experimental continuous processing mode that Spark offers. Finally, you will discuss the Tungsten engine optimizations which make Spark 2 so much faster than Spark 1, and discuss the stages of optimization in the Catalyst optimizer which works with SQL queries. At the end of this course, you will be able to build and execute streaming queries on input data, write these out to reliable storage using different output modes, and checkpoint your streaming applications for fault tolerance and recovery.
Author Name: Janani Ravi
Author Description:
Janani has a Masters degree from Stanford and worked for 7+ years at Google. She was one of the original engineers on Google Docs and holds 4 patents for its real-time collaborative editing framework. After spending years working in tech in the Bay Area, New York, and Singapore at companies such as Microsoft, Google, and Flipkart, Janani finally decided to combine her love for technology with her passion for teaching. She is now the co-founder of Loonycorn, a content studio focused on providing … more

Course Overview
2mins
Getting Started with Structured Streaming
55mins
Executing Streaming Queries
44mins
Understanding Scheduling and Checkpointing
26mins
Configuring Processing Models
20mins
Understanding Query Planning
28mins

User Reviews

0.0 out of 5

★★★★★

Write a review

There are no reviews yet.

Be the first to review “Conceptualizing the Processing Model for Apache Spark Structured Streaming” Cancel reply

Conceptualizing the Processing Model for Apache Spark Structured Streaming

Description
Reviews (0)

Start Course

All Categories

Conceptualizing the Processing Model for Apache Spark Structured Streaming

Table of Contents

User Reviews

Be the first to review “Conceptualizing the Processing Model for Apache Spark Structured Streaming” Cancel reply

COURSE PROVIDERS

CATEGORIES

Quick Links

Contact Us

Compare items

All Categories

Conceptualizing the Processing Model for Apache Spark Structured Streaming

Table of Contents

User Reviews

Be the first to review “Conceptualizing the Processing Model for Apache Spark Structured Streaming” Cancel reply

Related Products

Transfer and transform data with Azure Synapse Analytics pipelines

Use Apache Spark in Azure Databricks

Data Streaming

Explore fundamentals of large-scale analytics

Introduction to Azure Data Lake Storage Gen2

Organizational Privacy Engineering

COURSE PROVIDERS

CATEGORIES

Quick Links

Contact Us

Compare items