Apache Spark for Data Engineering

Let us dive into learning what is apache spark and why has it become really important in the field of data engineering?
What is Apache Spark?
Basically, Apache Spark is an Open Source Analytical Process Engine developed for large scale distributed data processing and machine learning applications.
Spark was developed at the University of California at Berkley, and later donated it to Apache Software Foundation.
Implementations of Spark:
Spark can be implemented in various interfaces:
| Spark | Interface for Scala and Java |
| PySpark | Python Interface for Spark |
| SparklyR | Spark interface for R language |
Features of Apache Spark:
Immutable in nature
Distributed processing
Lazy Evaluation
Cache
Inbuilt optimization for DataFrames
In memory computing
Architecture:

img ref: https://spark.apache.org/
Apache Spark Driver: It is the central program that is responsible for the execution of an application across a cluster. Main responsibility includes creating a Spark Context, which immune manages communication between the Cluster Manager and Executors.
Apache Spark Context: It is nothing but a representation of a connection to a Spark Cluster and can be used to created RDDs.
Cluster Manager: It is a platform/service that is used to provide some resources to Spark worker nodes which allows Spark to run in a distributed computing environment. It has resources like Memory, CPU, Network etc. Responsible for load balancing, memory utilization, Task Scheduling etc. (Standalone, Hadoop YARN, Apache Mesos, Kubernetes)
Worker Node: Which are also known as Compute Node are responsible for executing data processing tasks assigned by the driver. They are mainly responsible for performing computations on the data stored in the memory and communicate with the driver to receive commands and send desired results.
Executor: In Spark, an executor is a JVM process that runs on the worker nodes inside a cluster and is responsible for executing tasks assigned by the spark driver. Responsible for handling computations and data processing. (Parallel execution, Resource Allocation etc.)
Role of Apache Spark in Data Engineering:

img ref: Google
These are the main components of Apache Spark.
Spark Core: It is nothing but a base engine for large scale parallel and distributed data processing. Responsibilities are memory management, scheduling fault recovery etc.
Spark SQL: Spark component that is used to querying data either by using SQL or HIVEQuery Language. It provides support for various data sources and makes it possible to query SQL queries with code transformations leading it to be a very powerful tool.
Spark Streaming: Supports a real time processing of streaming data, such a production website server, log files, social media sites like X etc. It receives input data stream and divides the data into batches which then gets processed by Spark Engine and it results in generations a dream of batches.

MLib: It is a machine learning library that provides various algorithms designed to scale out on cluster for classification, regression, clustering, filtering to name a few. Some algorithms also work on streaming data like linear regression k-means clustering etc.
GraphX: Is a library that is used to manipulate and transform graphs and perform graph parallel operation. it has tool for ETL, exploratory analysis, graph computations etc. It also has libraries of common graph algorithms like PageRank etc.
Conclusion:
To summarize, Spark helps in neutralizing the intensive task of processing high volumes of real time data which is in both structured and unstructured form and integrates machine learning and graph algorithms seamlessly. Its complex capabilities is what makes it a hero in the field of data engineering and machine learning.
References: Google, Medium and the right owners