Loading the catalog…
Loading the catalog…
🌊 Delta Lake Beginner's Guide - Practical Edition (Part 1. Local Environment) post-thumbnail Delta Lake Theory - 🌊 Delta Lake Beginner's Guide - Theory Delta Lake Practical Guide Part 1 - 🌊 Delta Lake Beginner's Guide - Practical Edition (Part 1. Local Environment) Delta Lake in Practice Part 2 - 🌊 Delta Lake Beginner's Guide - Practical Edition (Part 2. Utilizing the delta-spark library) INTRO In the previous post , 🌊 Delta Lake Beginner's Guide - Theory, we covered the theoretical aspects of Delta Lake in detail. In this post, we will practice the features of Delta Lake in a Pyspark Docker Container environment. In this practice session, we use PySpark as the tool for handling data, and files of type delta are stored and managed in a local directory. hyunsoolee0506/pyspark-cloud:3.5.1For Docker containers , I recommend using an image I have created separately, considering that you will be integrating with a cloud environment for future practice . However, for this local environment practice, it is fine to proceed using Google Colab. For the practice session, delta_data.zipwe used data from two CSV files located in the file below. 👉 delta_data.zip 1️⃣ Setting up the Practice Environment ▪ 1) Create a Docker Container /workspace/sparkConfigure the directory to be used as a volume on the user's computer to be mapped to the directory inside the container . docker run -d --name pyspark -p 8888:8888 -p 4040:4040 -v [사용자 디렉토리]:/workspace/spark hyunsoolee0506/pyspark-cloud:3.5.1 After executing the command above, 8888you can access the juypter lab development environment by connecting to the port. ▪ 2) Install Library Install the libraries required for the practice. hyunsoolee0506/pyspark-cloud:3.5.1Although they are already installed in the image, for Colab, you must install them by running the code below. pip install pyspark==3.5.1 delta-spark==3.2.0 pyarrow findspark ▪ 3) Pyspark Delta Lake Environment Setup https://docs.delta.io/latest/quick-start.html#set-up-apache-spark-with-delta-lake To use Delta Lake in PySpark, configure the relevant extensions when creating a SparkSession. from delta import * from pyspark.sql import SparkSession import pyspark.sql.functions as F builder = SparkSession.builder.appName("DeltaLakeLocal") .enableHiveSupport() .config("spark.sql.extensions", "io.delta.sql.DeltaSparkSessionExtension") .config("spark.sql.catalog.spark_catalog", "org.apache.spark.sql.delta.catalog.DeltaCatalog") spark = configure_spark_with_delta_pip(builder).getOrCreate() 2️⃣ Create Database and Table ▪ 1) Create Database deltalake_dbCreates a new database named. spark.sql("CREATE DATABASE IF NOT EXISTS deltalake_db") spark.sql("SHOW DATABASES").show() +------------+ | namespace| +------------+ | default| |deltalake_db| +------------+ ▪ 2) Create CSV type table trainer_data.csvtrainerCreates a table to store the file's data . 테이블 이름 : trainer 스키마 : id → INT name → STRING age → INT hometown → STRING prefer_type → STRING badge_count → INT level → STRING query = f""" CREATE TABLE IF NOT EXISTS deltalake_db.trainer ( id INT, name STRING, age INT, hometown STRING, prefer_type STRING, badge_count INT, level STRING ) USING csv OPTIONS ( path '[trainer_data.csv 파일 경로]', header 'true', inferSchema 'true', delimiter ',' ) """ spark.sql(query) 테이블 생성 확인 spark.sql("SHOW TABLES FROM deltalake_db").show() +------------+---------+-----------+ | namespace|tableName|isTemporary| +------------+---------+-----------+ |deltalake_db| trainer| false| +------------+---------+-----------+ 3️⃣ Create delta type table csvSince data from an existing file deltacannot be inserted directly into a table of the same type simultaneously with its creation, deltayou must create the table through two steps as follows. deltaCreate empty table of type deltacsvInsert table data into the table ▪ 1) Create table trainer_deltadeltaCreates a table named . /workspace/spark/deltalake/delta_local/trainer_delta/deltaWe have configured the settings so that table-related data is stored under the corresponding path . query = f""" CREATE TABLE IF NOT EXISTS deltalake_db.trainer_delta ( id INT, name STRING, age INT, hometown STRING, prefer_type STRING, badge_count INT, level STRING ) USING delta LOCATION '/workspace/spark/deltalake/delta_local/trainer_delta/' """ spark.sql(query) ▪ 2) Insert data Insert trainerthe data from the table created above into the table.trainer_delta query = """ INSERT INTO deltalake_db.trainer_delta SELECT * FROM deltalake_db.trainer; """ spark.sql(query) Once data insertion is complete, _delta_log/you can verify that a new folder and parquet file have been created in the delta table directory. 4️⃣ Reading delta type tables In PySpark, you can read tables stored in slightly different ways. ▪ 1) Basic reading in Spark LOCAL_DELTA_PATH = '/workspace/spark/deltalake/delta_local/trainer_delta' df = spark.read.format("delta").load(LOCAL_DELTA_PATH) df.show(5) +---+------+---+--------+-----------+-----------+------------+ | id| name|age|hometown|prefer_type|badge_count| level| +---+------+---+--------+-----------+-----------+------------+ | 1| Brian| 28| Seoul| Electric| 8| Master| | 3| Susan| 18| Gwangju| Rock| 7| Expert| | 6| Vicki| 17| Daejeon| Ice| 4|Intermediate| | 9|Olivia| 45| Incheon| Psychic| 3|Intermediate| | 10| Mark| 16| Gangwon| Fire| 4|Intermediate| +---+------+---+--------+-----------+-----------+------------+ only showing top 5 rows delta.Read as ▪ 2) query = f"SELECT * FROM delta. {LOCAL_DELTA_PATH} " spark.sql(query) ▪ 3) Read from Hive Catalog spark.table('deltalake_db.trainer_delta') 5️⃣ Save after editing the table This time, we will modify the table contents and overwrite the existing directory. This step is intended to prepare for future table change history inquiries and to practice Time Travel queries, a core feature of Delta Lake. ▪ 1) Save after excluding 'Beginner' Beginner 제외한 dataframe 생성 df_1 = df.filter(F.col('level') != 'Beginner') 기존 경로에 덮어쓰기 df_1.write .format('delta') .mode('overwrite') .save(LOCAL_DELTA_PATH) 데이터 확인 df = spark.read.format("delta").load(LOCAL_DELTA_PATH) df.select('level').distinct().show() +------------+ | level| +------------+ | Expert| | Advanced| | Master| |Intermediate| +------------+ ▪ 2) Save after excluding 'Advanced' Advanced 제외한 dataframe 생성 df_2 = df_1.filter(F.col('level') != 'Advanced') 기존 경로에 덮어쓰기 df_2.write .format('delta') .mode('overwrite') .save(LOCAL_DELTA_PATH) 데이터 확인 df = spark.read.format("delta").load(LOCAL_DELTA_PATH) df.select('level').distinct().show() +------------+ | level| +------------+ | Expert| | Master| |Intermediate| +------------+ You can see that as the data is overwritten, a parquet file is added, and _delta_log/metadata (files) are also added within the folder ..json 6️⃣ Change History Lookup and Time Travel Query ▪ 1) View History deltaView the change history for the table. So far, the table has a structure where a total of three write operations have occurred since the initial creation (CREATE). Therefore, the versions also exist as 0, 1, 2, and 3. VERSION 0 → trainer_deltaTable creation status, No data VERSION 1 → trainerInitial state after data is inserted into the table VERSION 2 → Saved state after excluding 'Beginner' rows VERSION 3 → Saved state after excluding 'Advanced' row query = "DESCRIBE HISTORY deltalake_db.trainer_delta" spark.sql(query).show(vertical=True, truncate=False) -RECORD 0-------------------------------------------------------------------------------------------------------------- version | 3 timestamp | 2025-03-21 02:09:39.085 userId | NULL userName | NULL operation | WRITE operationParameters | {mode -> Overwrite, partitionBy -> []} job | NULL notebook | NULL clusterId | NULL readVersion | 2 isolationLevel | Serializable isBlindAppend | false operationMetrics | {numFiles -> 1, numOutputRows -> 42, numOutputBytes -> 3125} userMetadata | NULL engineInfo | Apache-Spark/3.5.1 Delta-Lake/3.2.0 -RECORD 1-------------------------------------------------------------------------------------------------------------- version | 2 timestamp | 2025-03-21 02:08:32.646 userId | NULL userName | NULL operation | WRITE operationParameters | {mode -> Overwrite, partitionBy -> []} job | NULL notebook | NULL clusterId | NULL readVersion | 1 isolationLevel | Serializable isBlindAppend | false operationMetrics | {numFiles -> 1, numOutputRows -> 85, numOutputBytes -> 3868} userMetadata | NULL engineInfo | Apache-Spark/3.5.1 Delta-Lake/3.2.0 -RECORD 2------------------------------------------------------------------------------------------------------------- version | 1 timestamp | 2025-03-21 01:46:21.446 userId | NULL userName | NULL operation | WRITE operationParameters | {mode -> Append, partitionBy -> []} job | NULL notebook | NULL clusterId | NULL readVersion | 0 isolationLevel | Serializable isBlindAppend | true operationMetrics | {numFiles -> 1, numOutputRows -> 90, numOutputBytes -> 3980} userMetadata | NULL engineInfo | Apache-Spark/3.5.1 Delta-Lake/3.2.0 -RECORD 3-------------------------------------------------------------------------------------------------------------- version | 0 timestamp | 2025-03-21 01:43:25.471 userId | NULL userName | NULL operation | CREATE TABLE operationParameters | {partitionBy -> [], clusterBy -> [], description -> NULL, isManaged -> false, properties -> {}} job | NULL notebook | NULL clusterId | NULL readVersion | NULL isolationLevel | Serializable isBlindAppend | true operationMetrics | {} userMetadata | NULL engineInfo | Apache-Spark/3.5.1 Delta-Lake/3.2.0 ▪ 2) Time Travel - Version Whenever the table is modified, the contents of the table corresponding to a specific version are retrieved based on the assigned version number. 👉 Load initial version (version 0) table df_pre = spark.read .format("delta") .option("versionAsof", 0) .load(LOCAL_DELTA_PATH) df_pre.select('level').distinct().show() +-----+ |level| +-----+ +-----+ 👉 Load version 2 table df_pre = spark.read .format("delta") .option("versionA
What RADAR observed and classified to build this opportunity. It is what the source published, not a verification that the offer is still active.
sadasdasdasd. 🌊 Delta Lake Beginner's Guide - Practical Edition (Part 1. Local Environment) post-thumbnail Delta Lake Theory - 🌊 Delta Lake Beginner's Guide - Theory Delta Lake Practical Guide Part 1 - 🌊 Delta Lake Beginner's Guide - Practical Edition (Part 1. Local Environment) Delta Lake in Practice Part 2 - 🌊 Delta Lake Beginner's Guide - Practical Edition (Part 2. Utilizing the…
Open source