Data platformProcessingApache SparkSpark Application and Cluster Architecture

Spark Application and Cluster Architecture

Sau khi viết Spark program bằng API, chuyện gì xảy ra khi submit/run application lên cluster?

From User Entry Point to Spark Application

Spark Programming Model, chúng ta đã đứng nhìn từ phase đầu tiên: hiểu Spark đưa ra abstraction gì cho user thông qua Spark APIs. User chọn một entry point như SparkSession, nhìn data qua RDD/DataFrame/Dataset/SQL, rồi mô tả computation bằng transformations và actions.

Nhưng viết code xong chưa có nghĩa là computation đã chạy trên cluster. Để một Spark program thật sự trở thành distributed execution, nó phải được run/submit thành một Spark application. Đây là phase thứ hai: một Spark program được đưa vào runtime bằng đường nào, và cuối cùng execution được tạo ra ở đâu.

Cách classic và dễ tìm thấy nhất từ official documentations là spark-submit . User đưa JAR/Python file, dependencies, config và arguments cho Spark; spark-submit launch application lên cluster thông qua một interface chung cho các cluster managers.

Nhưng Spark không chỉ có mỗi batch-submit path. Với pyspark, spark-shell hoặc sparkR, user chạy Spark theo kiểu interactive shell; phía sau các shell này vẫn dựa vào cơ chế submit/runtime của Spark, nhưng trải nghiệm của user là một REPL/session đang sống để thử từng command. Với PySpark cài bằng pip, một file Python cũng có thể chạy trực tiếp bằng python app.py và tạo SparkSession bằng SparkSession.builder.getOrCreate(). Trong unit test, code có thể tạo SparkContext với master = "local" ngay trong process test rồi stop khi xong.

Ở Kubernetes, Spark Operator lại đưa thêm một entry path khác: user khai báo một SparkApplication CRD, còn operator chuyển khai báo đó thành driver/executor pods. Với Spark Connect, client còn tách khỏi Spark driver/server side: user tạo session bằng .remote("sc://..."), gửi logical plan qua gRPC, còn server side giữ phần Spark runtime/execution.

Nhìn chung, dù entry path khác nhau, Spark vẫn cần biến phần logic user viết thành một đơn vị runtime có thể chạy được. Nó được đặt vào một boundary mà Spark có thể quản lý, cấp resources, điều phối driver/executors, theo dõi jobs/tasks, và kết thúc khi runtime không còn cần nữa. Boundary đó chính là Spark application.

Spark's Distributed Execution

Khi chạy trên cluster, một Spark application không phải là một process đơn lẻ. Official Spark docs mô tả Spark applications như những sets of processes độc lập trên cluster, được điều phối bởi SparkContext trong main program — tức driver program.

Nói theo flow runtime: driver tạo hoặc giữ SparkContext/SparkSession, rồi dùng context đó để kết nối tới cluster manager. Cluster manager cấp resources cho application. Khi resources đã có, Spark acquire executors trên các worker nodes; executors là những processes chạy computations và giữ data/cache/shuffle data cho application. Sau đó driver gửi application code và tasks xuống executors để chạy.

Apache Spark components and architecture

Hình này là mental map cho phase này: một Spark application có driver program và SparkSession/SparkContext ở phía điều phối; cluster manager nằm ở boundary cấp resources; executors chạy trên worker nodes và dùng các cores để thực thi tasks. Vì vậy, khi nói “Spark application”, đừng chỉ nghĩ tới file code hay command submit. Nên nghĩ tới cả runtime boundary được dựng lên để chạy logic đó.

Có vài điểm architecture cần giữ ngay từ đầu. Mỗi application có executor processes riêng và các executors này thường sống trong suốt vòng đời application, chạy nhiều tasks bằng nhiều threads. Cách này giúp cô lập applications với nhau: mỗi driver schedule tasks của application riêng, và tasks từ các applications khác nhau chạy trong các JVM/processes khác nhau. Đổi lại, data không tự share giữa hai Spark applications; muốn share thì phải ghi ra external storage.

Spark cũng không gắn cứng với một cluster manager cụ thể. Miễn là Spark acquire được executor processes và các processes đó giao tiếp được với driver, Spark có thể chạy trên Standalone, YARN, Kubernetes hoặc một platform bọc phía trên.

Điểm cuối rất quan trọng: driver phải network-addressable từ worker nodes trong suốt vòng đời application, vì executors cần kết nối ngược lại driver để nhận tasks/report status. Vì driver schedule tasks cho cluster, nó cũng nên chạy gần worker nodes, ideally cùng local network. Nếu user ở xa cluster, cách tốt hơn thường là gửi request tới một driver/server gần cluster, thay vì để driver chạy trên laptop xa rồi điều phối executors qua network chậm/khó ổn định.

Driver Program

Driver program là process chạy main() của application và tạo SparkContext/SparkSession. Đây là component điều phối chính của Spark application.

Driver có nhiều vai trò cùng lúc. Nó giữ chương trình chính của user, giao tiếp với cluster manager để xin resources, biến Spark operations thành DAG computations, schedule work, rồi phân phối tasks xuống executors. Khi resources đã được cấp, driver giao tiếp trực tiếp với executors để gửi tasks và nhận status/result.

Nói đơn giản: driver không phải nơi ôm toàn bộ data lớn để xử lý. Driver là nơi giữ control plane của application — nơi biết application đang muốn làm gì, cần chia work ra sao, và executors nào đang tham gia chạy work đó.

SparkSession and SparkContext

Ở phase viết code, user thường chạm vào SparkSession trước. spark.read, spark.sql(...), createDataFrame(...) hay nhiều thao tác DataFrame/Dataset đều đi qua SparkSession.

Ở phase runtime, SparkSession/SparkContext là entry point nối user program với Spark application. Official cluster overview nói Spark applications là các sets of processes độc lập trên cluster, được coordinated bởi SparkContext object trong main program, tức driver program.

Có thể giữ mental model tạm thời như sau:

SparkSession
  -> user-facing entry point cho SQL/DataFrame/Dataset

SparkContext
  -> lower-level runtime/context object phối hợp với cluster/executors

Chi tiết SparkSession khác SparkContext thế nào, SharedState/SessionState nằm ở đâu, và tại sao Databricks notebook/runtime lại làm câu chuyện này thú vị hơn — để sang page riêng SparkContext, SparkSession & State.

Cluster Manager and Worker Nodes

Cluster manager là external service chịu trách nhiệm acquire/allocate resources trên cluster. Nó không chạy trực tiếp từng transformation của user. Nó cấp resources để application có thể launch executors.

Spark có thể chạy với nhiều cluster managers: standalone cluster manager của Spark, YARN, Kubernetes, và trong lịch sử có Mesos. Managed platforms như Databricks lại bọc thêm một layer vận hành phía trên, nhưng mental model ban đầu vẫn là: driver cần resources, cluster manager cấp resources, executors được launch để chạy work.

Worker node là node có thể chạy application code trong cluster. Executor là process chạy trên worker node, không phải bản thân node.

cluster manager
  -> allocates resources
  -> launches executors on worker nodes

worker node
  -> machine/pod/node where executor process can run

Executors

Executor là process được launch cho một application trên worker node. Executor chạy tasks và giữ data trong memory/disk storage xuyên suốt các tasks của application đó.

Executors giao tiếp với driver program. Driver gửi tasks xuống executors; executors chạy computation trên partitions, giữ cache hoặc shuffle data trung gian nếu cần, rồi report status/result về driver.

Một điểm rất quan trọng: mỗi Spark application có executors riêng. Điều này giúp cô lập applications với nhau ở cả scheduling side và executor side, nhưng cũng có nghĩa data không tự động được share giữa hai Spark applications khác nhau nếu không ghi ra external storage.

application A
  -> its own executors

application B
  -> its own executors

share data between applications
  -> write/read external storage

Deploy Modes

Deploy mode trả lời một câu hỏi đơn giản: driver process chạy ở đâu?

Trong client mode, driver chạy ở phía submitter/client. Cluster vẫn cấp executors, nhưng driver nằm ngoài cluster workers theo nghĩa nó chạy ở process/machine submit application.

Trong cluster mode, cluster manager launch driver bên trong cluster. Submitter gửi application lên, còn driver sống trong cluster runtime.

client mode
  submitter machine/process runs driver
  cluster runs executors

cluster mode
  cluster manager launches driver inside cluster
  cluster runs executors

Ở giai đoạn này chỉ cần biết deploy mode ảnh hưởng vị trí driver. Các khác biệt cụ thể về networking, logs, dependency distribution, failure behavior để học sau khi đã chắc driver/executor lifecycle.

Spark Connect as a Variant

Một biến thể hiện đại của phase run/submit là Spark Connect. Trong classic Spark, user code thường chạy gần driver hơn: application code tạo SparkSession/SparkContext, driver giữ application control và trực tiếp coordinate execution.

Spark Connect tách client khỏi driver/server side. Client không cần chạy trong cùng process với Spark driver; nó gửi unresolved logical plan qua gRPC tới Spark Connect server, còn server side giữ session/execution với Spark cluster.

Không cần đi sâu Spark Connect ở page này. Chỉ cần ghi nhớ nó như một biến thể của boundary:

classic Spark
  user application / driver tightly coupled

Spark Connect
  client sends plan remotely
  server side owns Spark execution/session boundary

Chi tiết này để sang Spark Client/Server and Connect, nhưng nhắc ở đây giúp tránh hiểu nhầm rằng mọi Spark program đều có cùng client-driver layout.

Where Execution Units Fit

Sau khi application đã chạy, Spark mới bắt đầu sinh ra các execution units như job, stage và task khi có action cần kết quả.

Execution units

Job là parallel computation được spawn để phục vụ một action như save hoặc collect.

Stage là phần nhỏ hơn của job; mỗi job được chia thành các stages phụ thuộc nhau, gần giống ý tưởng map/reduce stages trong MapReduce.

Task là unit of work được gửi tới một executor.

Page này chỉ đặt job/stage/task vào đúng chỗ trong architecture. Chúng không phải process, cũng không phải cluster resources. Chúng là cách Spark chia work bên trong một running application. Câu hỏi chi tiết — action tạo job như thế nào, shuffle cắt stage ra sao, task gắn với partition thế nào — sẽ được giải thích trong Spark Execution Model.

My Summary

Từ góc nhìn cá nhân, page này là phase chuyển từ “mình viết Spark code” sang “Spark code đó chạy trong architecture nào”. Đây là lý do nó nên đứng giữa Programming Model và Execution Model.

Spark program sau khi submit sẽ trở thành một Spark application. Application có driver program làm nơi điều phối, SparkSession/SparkContext làm entry point/context, cluster manager cấp resources, và executors chạy tasks trên worker nodes. Deploy mode chủ yếu quyết định driver nằm ở phía client hay trong cluster.

Khi đã có vocabulary này, page Execution Model mới dễ đọc hơn. Lúc đó job/stage/task không còn là một đống term rơi từ trên trời xuống; mình biết chúng là execution units sinh ra bên trong một running application có driver và executors phối hợp với nhau.

On this page