{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":31254,"databundleVersionId":3103714,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 《实验报告4：PySpark大数据预处理》\n---\n### 实验内容：\n1. 使用Spark MLlib进行大数据预处理。\n2. 复习DataFrame的基础操作。\n3. 练习特征工程。\n\n### 实验目标：\n1. 掌握Spark MLlib大数据预处理。\n2. 提高DataFrame的基础操作的熟练度。\n3. 掌握特征工程的基本工具与技术。\n\n### 实验提示:\n\n\n### 实验报告作业提交要求:\n**本次实验为个人作业，计时练习，请在下课前提交.word答题纸和.ipynb文件**","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"\n\nH&M集团拥有53个在线市场和4850家商店。H&M在线商店为购物者提供了广泛的产品选择以供浏览。但是，由于选择太多，顾客可能无法迅速找到他们感兴趣的产品或他们正在寻找的产品，最终，他们可能不会进行购买。为了增强购物体验，产品推荐至关重要。更重要的是，帮助顾客做出正确的选择对可持续性也有积极的影响，因为它减少了退货，从而最小化了运输排放。\n\nH&M集团邀请你基于以往的交易数据以及客户和产品元数据来开发产品推荐系统。可用的元数据包括简单的数据，如服装类型和顾客年龄，产品描述的文本数据，以及服装图像的图像数据。\n\n数据集中哪些信息可能有用需要你来发现和决定。如果你想研究分类数据类型的算法，或者深入研究自然语言处理和图像处理的深度学习，这都取决于你。\n\n**在实验报告四中，我们使用Kaggle笔记本对H&M集团的销售信息进行探索性分析和简单的预处理。所需数据在笔记本Input栏中的。请在《实验报告纸》 中回答以下问题。每个问题1分。**","metadata":{}},{"cell_type":"code","source":"# Install pyspark\n!pip install pyspark\n# Create a spark session\nfrom pyspark.sql import SparkSession\nspark = SparkSession.builder.master('local[*]').appName('HW4').getOrCreate()\n# Import spark sql functions\nfrom pyspark.sql.functions import *","metadata":{"execution":{"iopub.status.busy":"2024-11-07T02:04:54.3914Z","iopub.execute_input":"2024-11-07T02:04:54.39184Z","iopub.status.idle":"2024-11-07T02:04:58.791559Z","shell.execute_reply.started":"2024-11-07T02:04:54.391798Z","shell.execute_reply":"2024-11-07T02:04:58.789043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## A 交易数据\n","metadata":{}},{"cell_type":"markdown","source":"**<span> Transactions data description: </span>**\n\n> `t_dat` **<span style=\"color:#023e8a;\">: A unique identifier of every customer</span>**  \n> `customer_id` **<span style=\"color:#023e8a;\">: A unique identifier of every customer </span>**  **<span style=\"color:#FF0000;\">(in </span>** `customers` **<span style=\"color:#FF0000;\"> table)</span>**  \n> `article_id` **<span style=\"color:#023e8a;\">: A unique identifier of every article</span>**  **<span style=\"color:#FF0000;\">(in </span>** `articles` **<span style=\"color:#FF0000;\"> table)</span>**  \n> `price` **<span style=\"color:#023e8a;\">: Price of purchase</span>**  \n> `sales_channel_id` **<span style=\"color:#023e8a;\">: 1 or 2</span>**  ","metadata":{"execution":{"iopub.status.busy":"2024-11-01T13:34:55.038397Z","iopub.execute_input":"2024-11-01T13:34:55.039911Z","iopub.status.idle":"2024-11-01T13:34:55.052933Z","shell.execute_reply.started":"2024-11-01T13:34:55.039858Z","shell.execute_reply":"2024-11-01T13:34:55.051232Z"}}},{"cell_type":"markdown","source":"读取transactions_train.csv数据，生成DataFrame transaction，回答：\n* 1.transaction有多少列？(      )\n* 2.transaction有多少行？(      )\n* 3.transaction第0列的名称和类型是什么？(      )","metadata":{}},{"cell_type":"code","source":"# 读取数据\ntransaction = spark.read.format('csv').option(\"header\",True) \\\n              .load(\"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-11-07T02:05:15.411387Z","iopub.execute_input":"2024-11-07T02:05:15.411846Z","iopub.status.idle":"2024-11-07T02:05:21.687096Z","shell.execute_reply.started":"2024-11-07T02:05:15.411807Z","shell.execute_reply":"2024-11-07T02:05:21.685803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 查看列数和行数\nnum_columns = len(transaction.columns)\nnum_rows = transaction.count()\n\n# 第0列的名称和数据类型\nfirst_column_name = transaction.columns[0]\nfirst_column_type = transaction.schema[0].dataType\nprint(f\"列数: {num_columns}, 行数: {num_rows}, 第0列名称: {first_column_name}, 类型: {first_column_type}\")\n\n","metadata":{"execution":{"iopub.status.busy":"2024-11-07T02:05:52.371704Z","iopub.execute_input":"2024-11-07T02:05:52.372145Z","iopub.status.idle":"2024-11-07T02:06:02.884496Z","shell.execute_reply.started":"2024-11-07T02:05:52.372104Z","shell.execute_reply":"2024-11-07T02:06:02.882073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"列数: 5, 行数: 31788324, 第0列名称: t_dat, 类型: StringType()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"从transaction中按照0.5%的比例，随机抽取一个样本，将随机种子设置为123456，命名为df，回答：\n* 4.df第一行的article_id是多少？(         )","metadata":{}},{"cell_type":"code","source":"#随机抽样\ndf = transaction.sample(0.005,123456) #用于随机抽样的函数","metadata":{"execution":{"iopub.status.busy":"2024-11-07T02:06:42.396422Z","iopub.execute_input":"2024-11-07T02:06:42.396854Z","iopub.status.idle":"2024-11-07T02:06:42.407216Z","shell.execute_reply.started":"2024-11-07T02:06:42.396814Z","shell.execute_reply":"2024-11-07T02:06:42.40551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# 获取df第一行的article_id\nfirst_article_id = df.first()[\"article_id\"]\nprint(f\"df第一行的article_id: {first_article_id}\")\n\n","metadata":{"execution":{"iopub.status.busy":"2024-11-07T02:07:05.532057Z","iopub.execute_input":"2024-11-07T02:07:05.532515Z","iopub.status.idle":"2024-11-07T02:07:05.81117Z","shell.execute_reply.started":"2024-11-07T02:07:05.532476Z","shell.execute_reply":"2024-11-07T02:07:05.809871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df第一行的article_id: 0448515001","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"使用df，计算并回答：\n* 5.数据中包含多少个不同的消费者？（）\n* 6.购买商品总数最多的第五名消费者的购买商品总数是什么？（） \n* 7.购买总额最大的第五名消费者的购买总额是多少？（）","metadata":{}},{"cell_type":"code","source":"# 不同消费者数量\nunique_customers = df.select(\"customer_id\").distinct().count()\n\n# 第五名购买商品最多的消费者及其购买总数\ntop_customers_by_items = df.groupBy(\"customer_id\").count() \\\n    .orderBy(desc(\"count\")).take(5)\nfifth_top_item_count = top_customers_by_items[4][\"count\"]\n\n# 第五名购买总额最高的消费者及其购买总额\ntop_customers_by_amount = df.groupBy(\"customer_id\").agg(sum(\"price\").alias(\"total_spent\")) \\\n    .orderBy(desc(\"total_spent\")).take(5)\nfifth_top_total_spent = top_customers_by_amount[4][\"total_spent\"]\n\nprint(f\"不同消费者数量: {unique_customers}\")\nprint(f\"第五名购买商品最多的消费者商品总数: {fifth_top_item_count}\")\nprint(f\"第五名购买总额最高的消费者总额: {fifth_top_total_spent}\")","metadata":{"execution":{"iopub.status.busy":"2024-11-07T02:07:50.102378Z","iopub.execute_input":"2024-11-07T02:07:50.102907Z","iopub.status.idle":"2024-11-07T02:09:13.005437Z","shell.execute_reply.started":"2024-11-07T02:07:50.10286Z","shell.execute_reply":"2024-11-07T02:09:13.004268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"不同消费者数量: 131832\n第五名购买商品最多的消费者商品总数: 8\n第五名购买总额最高的消费者总额: 0.433864406779661","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"将交易时间列从string类型转换为date类型，生成dfDate，回答：\n* 8.交易最多的年份是哪一年？（）\n* 9.交易最多的月份是哪个月？（）\n* 10.周几的交易最多？（）","metadata":{}},{"cell_type":"code","source":"# 将交易时间列从string转换为date类型\ndfDate = df.withColumn(\"t_dat\", to_date(\"t_dat\", \"yyyy-MM-dd\"))\n\n# 交易最多的年份\nmost_common_year = dfDate.groupBy(year(\"t_dat\").alias(\"year\")).count() \\\n    .orderBy(desc(\"count\")).first()[\"year\"]\n\n# 交易最多的月份\nmost_common_month = dfDate.groupBy(month(\"t_dat\").alias(\"month\")).count() \\\n    .orderBy(desc(\"count\")).first()[\"month\"]\n\n# 交易最多的星期几\nmost_common_day = dfDate.groupBy(date_format(\"t_dat\", \"E\").alias(\"day_of_week\")).count() \\\n    .orderBy(desc(\"count\")).first()[\"day_of_week\"]\n\nprint(f\"交易最多的年份: {most_common_year}\")\nprint(f\"交易最多的月份: {most_common_month}\")\nprint(f\"交易最多的星期几: {most_common_day}\")","metadata":{"execution":{"iopub.status.busy":"2024-11-07T02:09:39.36285Z","iopub.execute_input":"2024-11-07T02:09:39.363345Z","iopub.status.idle":"2024-11-07T02:10:58.068153Z","shell.execute_reply.started":"2024-11-07T02:09:39.363301Z","shell.execute_reply":"2024-11-07T02:10:58.066906Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"交易最多的年份: 2019\n交易最多的月份: 6\n交易最多的星期几: Sat","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## B 商品数据\n读取articles.csv，将随机种子设置为**你的学号**，从中抽取50%的样本，命名为df2回答：\n* 11.df2第一行记录的articles_id是什么？（）\n\n使用inner join，将df2连接到dfDate上，生成df3，回答：\n* 12. df3有多少行？()\n* 13. 交易数量最多的商品大类（department_name）是什么？（）\n* 14. 交易数量最多的商品类别（product_type_name）是什么？（）\n* 15. 交易数量最多的商品颜色（colour_group_name）是什么？（）\n","metadata":{}},{"cell_type":"code","source":"# 读取articles.csv并抽样\narticles = spark.read.format('csv').option(\"header\", True) \\\n    .load(\"../input/h-and-m-personalized-fashion-recommendations/articles.csv\")\ndf2 = articles.sample(0.5, seed=2211020111)  # 设置学号为种子\n\n# 获取 df2 第一行的 article_id\nfirst_article_id_df2 = df2.first()[\"article_id\"]\nprint(f\"df2 第一行的 article_id: {first_article_id_df2}\")\n\n\n# 将 df2 与 dfDate 进行 inner join 生成 df3\ndf3 = dfDate.join(df2, \"article_id\", \"inner\")\ndf3_count = df3.count()\n\n# 交易数量最多的商品大类\nmost_common_department = df3.groupBy(\"department_name\").count() \\\n    .orderBy(desc(\"count\")).first()[\"department_name\"]\n\n# 交易数量最多的商品类别\nmost_common_product_type = df3.groupBy(\"product_type_name\").count() \\\n    .orderBy(desc(\"count\")).first()[\"product_type_name\"]\n\n# 交易数量最多的商品颜色\nmost_common_colour = df3.groupBy(\"colour_group_name\").count() \\\n    .orderBy(desc(\"count\")).first()[\"colour_group_name\"]\n\nprint(f\"df3行数: {df3_count}\")\nprint(f\"交易数量最多的商品大类: {most_common_department}\")\nprint(f\"交易数量最多的商品类别: {most_common_product_type}\")\nprint(f\"交易数量最多的商品颜色: {most_common_colour}\")","metadata":{"execution":{"iopub.status.busy":"2024-11-07T02:19:39.191818Z","iopub.execute_input":"2024-11-07T02:19:39.192885Z","iopub.status.idle":"2024-11-07T02:21:40.485675Z","shell.execute_reply.started":"2024-11-07T02:19:39.192816Z","shell.execute_reply":"2024-11-07T02:21:40.484096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":" article_id: 0108775015\ndf3行数: 80308\n交易数量最多的商品大类: Swimwear\n交易数量最多的商品类别: Trousers\n交易数量最多的商品颜色: Black","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## C 用户数据\n","metadata":{}},{"cell_type":"markdown","source":"读取customers.csv，命名为df4. 使用inner join，将df4连接到df3上，生成df5，回答：\n* 16. df5有多少行？()\n* 17. 将用户年龄age按百分数分为5组，生成df6, df6中交易数量最多的用户年龄组是什么？（）\n* 18. df6中交易数量最多的用户时尚信息关注程度fashion_news_frequency是什么？（）\n\n","metadata":{}},{"cell_type":"code","source":"# 读取 customers.csv 数据\ncustomers = spark.read.format('csv').option(\"header\", True) \\\n    .load(\"../input/h-and-m-personalized-fashion-recommendations/customers.csv\")\ndf4 = customers\n\n# 将 df4 与 df3 进行 inner join 生成 df5\ndf5 = df3.join(df4, \"customer_id\", \"inner\")\ndf5_count = df5.count()\nprint(f\"df5行数: {df5_count}\")\n\n# 将 age 列转换为整数类型\ndf5 = df5.withColumn(\"age\", col(\"age\").cast(\"int\"))\n\n# 计算年龄的百分位数\nquantiles = df5.approxQuantile(\"age\", [0.2, 0.4, 0.6, 0.8, 1.0], 0)\nage_bins = [float('-inf')] + quantiles  # 创建年龄的分组边界\n\n# 按年龄分组，并找到交易最多的年龄组和时尚关注程度\ndf6 = df5.withColumn(\"age_group\", when(col(\"age\") <= age_bins[1], \"0-20%\") \\\n                                   .when(col(\"age\") <= age_bins[2], \"21-40%\") \\\n                                   .when(col(\"age\") <= age_bins[3], \"41-60%\") \\\n                                   .when(col(\"age\") <= age_bins[4], \"61-80%\") \\\n                                   .otherwise(\"81-100%\"))\n\n# 年龄组交易数量最多的分组\nmost_common_age_group = df6.groupBy(\"age_group\").count() \\\n    .orderBy(desc(\"count\")).first()[\"age_group\"]\n\n# 时尚信息关注程度最高的类别\nmost_common_fashion_news = df6.groupBy(\"fashion_news_frequency\").count() \\\n    .orderBy(desc(\"count\")).first()[\"fashion_news_frequency\"]\n\nprint(f\"交易数量最多的年龄组: {most_common_age_group}\")\nprint(f\"交易数量最多的时尚信息关注程度: {most_common_fashion_news}\")","metadata":{"execution":{"iopub.status.busy":"2024-11-07T02:31:11.052878Z","iopub.execute_input":"2024-11-07T02:31:11.053971Z","iopub.status.idle":"2024-11-07T02:33:40.89731Z","shell.execute_reply.started":"2024-11-07T02:31:11.053919Z","shell.execute_reply":"2024-11-07T02:33:40.895805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"交易数量最多的年龄组: 21-40%\n交易数量最多的时尚信息关注程度: NONE","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## D 特征工程\n在df6的基础上，将以下列组合成一个数值向量特征，组合之前需要对相应的列进行处理，转换成数值列。\n\n* day-of-week\n* price\n* colour_group_name\n* age\n* fashion_news_frequency\n\n回答\n* 19.最后生成的特征向量有多少个元素？（）\n* 20.解释第一行的特征向量的含义。（）\n\n\n","metadata":{}},{"cell_type":"code","source":"from pyspark.ml.feature import VectorAssembler\n\n# 将 price 列转换为浮点数\ndf6 = df6.withColumn(\"price\", col(\"price\").cast(\"float\"))\n\n# 组合特征向量\nassembler = VectorAssembler(inputCols=[\"day_index\", \"price\", \"colour_index\", \"age\", \"fashion_index\"], outputCol=\"features\")\ndf6 = assembler.transform(df6)\n\n# 获取第一行特征向量和特征向量的元素数量\nfirst_feature_vector = df6.select(\"features\").first()[\"features\"]\nprint(f\"第一行特征向量: {first_feature_vector}\")\nprint(f\"特征向量的元素数量: {len(first_feature_vector)}\")","metadata":{"execution":{"iopub.status.busy":"2024-11-07T02:37:04.40709Z","iopub.execute_input":"2024-11-07T02:37:04.40844Z","iopub.status.idle":"2024-11-07T02:37:49.112935Z","shell.execute_reply.started":"2024-11-07T02:37:04.408367Z","shell.execute_reply":"2024-11-07T02:37:49.111696Z"},"trusted":true},"execution_count":null,"outputs":[]}]}