Loading the catalog…
Loading the catalog…
Delta Lake Theory - 🌊 Delta Lake Beginner's Guide - Theory Delta Lake Practical Guide Part 1 - 🌊 Delta Lake Beginner's Guide - Practical Edition (Part 1. Local Environment) Delta Lake in Practice Part 2 - 🌊 Delta Lake Beginner's Guide - Practical Edition (Part 2. Utilizing the delta-spark library) 0. INTRO In the previous post , 🌊 Delta Lake Beginner's Guide - Theory, we covered the theoretical aspects of Delta Lake in detail. In this post, we will practice the features of Delta Lake in a Pyspark Docker Container environment. In this practice session, we use PySpark as the tool for handling data, and files of type delta are stored and managed in a local directory. hyunsoolee0506/pyspark-cloud:3.5.1 For Docker containers , I recommend using an image I have created separately, considering that you will be integrating with a cloud environment for future practice . However, for this local environment practice, it is fine to proceed using Google Colab. For the practice session, delta_data.zip we used data from two CSV files located in the file below. 👉 delta_data.zip 1️⃣ Setting up the Practice Environment ▪ 1) Create a Docker Container /workspace/spark Configure the directory to be used as a volume on the user's computer to be mapped to the directory inside the container . docker run -d \ --name pyspark \ -p 8888 :8888 \ -p 4040 :4040 \ -v [ 사용자 디렉토리 ] :/workspace/spark \ hyunsoolee0506/pyspark-cloud:3.5.1 After executing the command above, 8888 you can access the juypter lab development environment by connecting to the port. ▪ 2) Install Library Install the libraries required for the practice. hyunsoolee0506/pyspark-cloud:3.5.1 Although they are already installed in the image, for Colab, you must install them by running the code below. pip install pyspark == 3.5 .1 delta-spark == 3.2 .0 pyarrow findspark ▪ 3) Pyspark Delta Lake Environment Setup https://docs.delta.io/latest/quick-start.html#set-up-apache-spark-with-delta-lake To use Delta Lake in PySpark, configure the relevant extensions when creating a SparkSession. from delta import * from pyspark . sql import SparkSession import pyspark . sql . functions as F builder = SparkSession . builder . appName ( "DeltaLakeLocal" ) . enableHiveSupport ( ) . config ( "spark.sql.extensions" , "io.delta.sql.DeltaSparkSessionExtension" ) . config ( "spark.sql.catalog.spark_catalog" , "org.apache.spark.sql.delta.catalog.DeltaCatalog" ) spark = configure_spark_with_delta_pip ( builder ) . getOrCreate ( ) 2️⃣ Create Database and Table ▪ 1) Create Database deltalake_db Creates a new database named. spark . sql ( "CREATE DATABASE IF NOT EXISTS deltalake_db" ) spark . sql ( "SHOW DATABASES" ) . show ( ) - - - + - - - - - - - - - - - - + | namespace | + - - - - - - - - - - - - + | default | | deltalake_db | + - - - - - - - - - - - - + ▪ 2) Create CSV type table trainer_data.csv trainer Creates a table to store the file's data . - 테이블 이름 : trainer - 스키마 : - id → INT - name → STRING - age → INT - hometown → STRING - prefer_type → STRING - badge_count → INT - level → STRING query = f""" CREATE TABLE IF NOT EXISTS deltalake_db.trainer ( id INT, name STRING, age INT, hometown STRING, prefer_type STRING, badge_count INT, level STRING ) USING csv OPTIONS ( path '[trainer_data.csv 파일 경로]', header 'true', inferSchema 'true', delimiter ',' ) """ spark . sql ( query ) # 테이블 생성 확인 spark . sql ( "SHOW TABLES FROM deltalake_db" ) . show ( ) - - - + - - - - - - - - - - - - + - - - - - - - - - + - - - - - - - - - - - + | namespace | tableName | isTemporary | + - - - - - - - - - - - - + - - - - - - - - - + - - - - - - - - - - - + | deltalake_db | trainer | false | + - - - - - - - - - - - - + - - - - - - - - - + - - - - - - - - - - - + 3️⃣ Create delta type table csv Since data from an existing file delta cannot be inserted directly into a table of the same type simultaneously with its creation, delta you must create the table through two steps as follows. delta Create empty table of type delta csv Insert table data into the table ▪ 1) Create table trainer_delta delta Creates a table named . /workspace/spark/deltalake/delta_local/trainer_delta/ delta We have configured the settings so that table-related data is stored under the corresponding path . query = f""" CREATE TABLE IF NOT EXISTS deltalake_db.trainer_delta ( id INT, name STRING, age INT, hometown STRING, prefer_type STRING, badge_count INT, level STRING ) USING delta LOCATION '/workspace/spark/deltalake/delta_local/trainer_delta/' """ spark . sql ( query ) ▪ 2) Insert data Insert trainer the data from the table created above into the table. trainer_delta query = """ INSERT INTO deltalake_db.trainer_delta SELECT * FROM deltalake_db.trainer; """ spark . sql ( query ) Once data insertion is complete, _delta_log/ you can verify that a new folder and parquet file have been created in the delta table directory. 4️⃣ Reading delta type tables In PySpark, you can read tables stored in slightly different ways. ▪ 1) Basic reading in Spark LOCAL_DELTA_PATH = '/workspace/spark/deltalake/delta_local/trainer_delta' df = spark . read . format ( "delta" ) . load ( LOCAL_DELTA_PATH ) df . show ( 5 ) - - - + - - - + - - - - - - + - - - + - - - - - - - - + - - - - - - - - - - - + - - - - - - - - - - - + - - - - - - - - - - - - + | id | name | age | hometown | prefer_type | badge_count | level | + - - - + - - - - - - + - - - + - - - - - - - - + - - - - - - - - - - - + - - - - - - - - - - - + - - - - - - - - - - - - + | 1 | Brian | 28 | Seoul | Electric | 8 | Master | | 3 | Susan | 18 | Gwangju | Rock | 7 | Expert | | 6 | Vicki | 17 | Daejeon | Ice | 4 | Intermediate | | 9 | Olivia | 45 | Incheon | Psychic | 3 | Intermediate | | 10 | Mark | 16 | Gangwon | Fire | 4 | Intermediate | + - - - + - - - - - - + - - - + - - - - - - - - + - - - - - - - - - - - + - - - - - - - - - - - + - - - - - - - - - - - - + only showing top 5 rows delta. Read as ▪ 2) query = f"SELECT * FROM delta.` { LOCAL_DELTA_PATH } `" spark . sql ( query ) ▪ 3) Read from Hive Catalog spark . table ( 'deltalake_db.trainer_delta' ) 5️⃣ Save after editing the table This time, we will modify the table contents and overwrite the existing directory. This step is intended to prepare for future table change history inquiries and to practice Time Travel queries, a core feature of Delta Lake. ▪ 1) Save after excluding 'Beginner' # Beginner 제외한 dataframe 생성 df_1 = df . filter ( F . col ( 'level' ) != 'Beginner' ) # 기존 경로에 덮어쓰기 df_1 . write . format ( 'delta' ) . mode ( 'overwrite' ) . save ( LOCAL_DELTA_PATH ) # 데이터 확인 df = spark . read . format ( "delta" ) . load ( LOCAL_DELTA_PATH ) df . select ( 'level' ) . distinct ( ) . show ( ) - - - + - - - - - - - - - - - - + | level | + - - - - - - - - - - - - + | Expert | | Advanced | | Master | | Intermediate | + - - - - - - - - - - - - + ▪ 2) Save after excluding 'Advanced' # Advanced 제외한 dataframe 생성 df_2 = df_1 . filter ( F . col ( 'level' ) != 'Advanced' ) # 기존 경로에 덮어쓰기 df_2 . write . format ( 'delta' ) . mode ( 'overwrite' ) . save ( LOCAL_DELTA_PATH ) # 데이터 확인 df = spark . read . format ( "delta" ) . load ( LOCAL_DELTA_PATH ) df . select ( 'level' ) . distinct ( ) . show ( ) - - - + - - - - - - - - - - - - + | level | + - - - - - - - - - - - - + | Expert | | Master | | Intermediate | + - - - - - - - - - - - - + You can see that as the data is overwritten, a parquet file is added, and _delta_log/ metadata (files) are also added within the folder . .json 6️⃣ Change History Lookup and Time Travel Query ▪ 1) View History delta View the change history for the table. So far, the table has a structure where a total of three write operations have occurred since the initial creation (CREATE). Therefore, the versions also exist as 0, 1, 2, and 3. VERSION 0 → trainer_delta Table creation status, No data VERSION 1 → trainer Initial state after data is inserted into the table VERSION 2 → Saved state after excluding 'Beginner' rows VERSION 3 → Saved state after excluding 'Advanced' row query = "DESCRIBE HISTORY deltalake_db.trainer_delta" spark . sql ( query ) . show ( vertical = True , truncate = False ) - - - - RECORD 0 - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - version | 3 timestamp | 2025 - 03 - 21 02 : 09 : 39.085 userId | NULL userName | NULL operation | WRITE operationParameters | { mode - > Overwrite , partitionBy - > [ ] } job | NULL notebook | NULL clusterId | NULL readVersion | 2 isolationLevel | Serializable isBlindAppend | false operationMetrics | { numFiles - > 1 , numOutputRows - > 42 , numOutputBytes - > 3125 } userMetadata | NULL engineInfo | Apache - Spark / 3.5 .1 Delta - Lake / 3.2 .0 - RECORD 1 - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - version | 2 timestamp | 2025 - 03 - 21 02 : 08 : 32.646 userId | NULL userName | NULL operation | WRITE operationParameters | { mode - > Overwrite , partitionBy - > [ ] } job | NULL notebook | NULL clusterId | NULL readVersion | 1 isolationLevel | Serializable isBlindAppend | false operationMetrics | { numFiles - > 1 , numOutputRows - > 85 , numOutputBytes - > 3868 } userMetadata | NULL engineInfo | Apache - Spark / 3.5 .1 Delta - Lake / 3.2 .0 - RECORD 2 - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - version | 1 timestamp | 2025 - 03 - 21 01 : 46 : 21.446 userId | NULL userName | NULL operation | WRITE operationParameters | { mode - > Append , partitionBy - > [ ] } job | NULL notebook | NULL clusterId | NULL readVersion | 0 isolationLevel |
What RADAR observed and classified to build this opportunity. It is what the source published, not a verification that the offer is still active.
🌊 Delta Lake Beginner's Guide - Practical Edition (Part 1. Local Environment). Delta Lake Theory - 🌊 Delta Lake Beginner's Guide - Theory Delta Lake Practical Guide Part 1 - 🌊 Delta Lake Beginner's Guide - Practical Edition (Part 1. Local Environment) Delta Lake in Practice Part 2 - 🌊 Delta Lake Beginner's Guide - Practical Edition (Part 2. Utilizing the delta-spark library) 0. INTRO In the…
Open source