{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <p style=\"background-color:#f7e9ec;font-family:newtimeroman;color:#d90b1c;font-size:150%;text-align:center;border-radius:10px 10px;border-style:solid;border-color:#d90b1c;\">Recommendation system for H and M Fashion</p>","metadata":{}},{"cell_type":"markdown","source":"**For H and M Fashion EDA please check out my notebook** https://www.kaggle.com/nadianizam/h-m-fashion-eda","metadata":{}},{"cell_type":"markdown","source":"<h1 style=\"background-color:#f7e9ec;font-family:newtimeroman;color:#d90b1c;\">Terminologies</h1>\n\nThere are certain terminologies which needs to be understood before moving forward.\n\n**Apache Spark:** Apache Spark is an open-source distributed general-purpose cluster-computing framework.It can be used with Hadoop too.\n\n**Collaborative filtering:** Collaborative filtering is a method of making automatic predictions (filtering) about the interests of a user by collecting preferences or taste information from many users. Consider example if a person A likes item 1, 2, 3 and B like 2,3,4 then they have similar interests and A should like item 4 and B should like item 1.\n\n**Alternating least square(ALS) matrix factorization:** The idea is basically to take a large (or potentially huge) matrix and factor it into some smaller representation of the original matrix through alternating least squares. We end up with two or more lower dimensional matrices whose product equals the original one.ALS comes inbuilt in Apache Spark.\n\n**PySpark:** PySpark is the collaboration of Apache Spark and Python. PySpark is the Python API for Spark.","metadata":{}},{"cell_type":"markdown","source":"<h1 style=\"background-color:#f7e9ec;font-family:newtimeroman;color:#d90b1c;\">1.Initialize spark session</h1>","metadata":{}},{"cell_type":"code","source":"!pip install pyspark","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-04-29T17:38:33.156127Z","iopub.execute_input":"2022-04-29T17:38:33.156512Z","iopub.status.idle":"2022-04-29T17:39:18.219109Z","shell.execute_reply.started":"2022-04-29T17:38:33.156402Z","shell.execute_reply":"2022-04-29T17:39:18.218335Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<h1 style=\"background-color:#f7e9ec;font-family:newtimeroman;color:#d90b1c;\">2-Load libraries</h1>","metadata":{}},{"cell_type":"code","source":"import pyspark\nfrom pyspark.sql import SparkSession\nfrom pyspark.sql.types import StructType,StructField, StringType, IntegerType \nfrom pyspark.sql.types import ArrayType, DoubleType, BooleanType\nfrom pyspark.sql.functions import col,array_contains\nfrom pyspark.sql import SQLContext \nfrom pyspark.ml.recommendation import ALS\nfrom pyspark.sql.functions import udf,col,when\nfrom pyspark.sql.functions import to_timestamp,date_format\nimport numpy as np\nimport pandas as pd\nfrom pyspark.sql.types import *\nfrom pyspark.sql.functions import *\nfrom pyspark.sql.window import *\n\nsc = SparkSession.builder.appName(\"Recommendations\").config(\"spark.sql.files.maxPartitionBytes\", 5000000).getOrCreate()\nspark = SparkSession(sc)\n\n","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-04-29T17:39:20.823610Z","iopub.execute_input":"2022-04-29T17:39:20.823866Z","iopub.status.idle":"2022-04-29T17:39:27.861202Z","shell.execute_reply.started":"2022-04-29T17:39:20.823835Z","shell.execute_reply":"2022-04-29T17:39:27.860332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<h1 style=\"background-color:#f7e9ec;font-family:newtimeroman;color:#d90b1c;\">3-Load Dataset in Apache Spark</h1>","metadata":{}},{"cell_type":"code","source":"transaction = spark.read.option(\"header\",True) \\\n              .csv(\"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\")\ntransaction.printSchema()","metadata":{"execution":{"iopub.status.busy":"2022-04-29T17:40:11.829570Z","iopub.execute_input":"2022-04-29T17:40:11.829902Z","iopub.status.idle":"2022-04-29T17:40:16.900232Z","shell.execute_reply.started":"2022-04-29T17:40:11.829862Z","shell.execute_reply":"2022-04-29T17:40:16.898219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from pyspark.sql.functions import min, max\nfrom pyspark.sql.functions import unix_timestamp, lit\nmin_date, max_date = transaction.select(min(\"t_dat\"), max(\"t_dat\")).first()\nmin_date, max_date","metadata":{"execution":{"iopub.status.busy":"2022-04-29T17:40:28.951516Z","iopub.execute_input":"2022-04-29T17:40:28.951829Z","iopub.status.idle":"2022-04-29T17:41:27.402825Z","shell.execute_reply.started":"2022-04-29T17:40:28.951793Z","shell.execute_reply":"2022-04-29T17:41:27.402039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"background-color:#f7e9ec;font-family:newtimeroman;color:#d90b1c;\">5-Select data for recommendation</h1>","metadata":{}},{"cell_type":"markdown","source":"In this transaction dataset we have 31,788,324 rows and 5 columns.Let's capture first what are the most recently bought articles.For recommendation I am selecting only date 2020-09-22 which is the last transaction date.</h1>","metadata":{}},{"cell_type":"code","source":"\nhm =  transaction.withColumn('t_dat', transaction['t_dat'].cast('string'))\nhm = hm.withColumn('date', from_unixtime(unix_timestamp('t_dat', 'yyyy-MM-dd')))\nhm = hm.withColumn('year', year(col('date')))\nhm = hm.withColumn('month', month(col('date')))\nhm = hm.withColumn('day', date_format(col('date'), \"d\"))\n\nhm = hm[hm['year'] == 2020]\nhm = hm[hm['month'] == 9]\nhm = hm[hm['day'] == 22]\ntransaction.unpersist()\n\n# Prepare the dataset\nhm = hm.groupby('customer_id', 'article_id').count()\nhm.show(5)","metadata":{"execution":{"iopub.status.busy":"2022-04-29T17:41:36.012972Z","iopub.execute_input":"2022-04-29T17:41:36.013692Z","iopub.status.idle":"2022-04-29T17:43:36.287394Z","shell.execute_reply.started":"2022-04-29T17:41:36.013653Z","shell.execute_reply":"2022-04-29T17:43:36.286720Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print((hm.count(), len(hm.columns)))","metadata":{"execution":{"iopub.status.busy":"2022-04-29T17:43:44.981888Z","iopub.execute_input":"2022-04-29T17:43:44.982162Z","iopub.status.idle":"2022-04-29T17:45:41.609544Z","shell.execute_reply.started":"2022-04-29T17:43:44.982134Z","shell.execute_reply":"2022-04-29T17:45:41.608908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Count the total number of article in the dataset\nnumerator = hm.select(\"count\").count()\n\n# Count the number of distinct customerid and distinct articleid\nnum_users = hm.select(\"customer_id\").distinct().count()\nnum_articles = hm.select(\"article_id\").distinct().count()\n\n# Set the denominator equal to the number of customer multiplied by the number of articles\ndenominator = num_users * num_articles\n\n# Divide the numerator by the denominator\nsparsity = (1.0 - (numerator *1.0)/denominator)*100\nprint(\"Sparsity: \", \"%.2f\" % sparsity + \"%.\")","metadata":{"execution":{"iopub.status.busy":"2022-04-29T17:45:58.675526Z","iopub.execute_input":"2022-04-29T17:45:58.676419Z","iopub.status.idle":"2022-04-29T17:51:42.397715Z","shell.execute_reply.started":"2022-04-29T17:45:58.676366Z","shell.execute_reply":"2022-04-29T17:51:42.397047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"userId_count = hm.groupBy(\"customer_id\").count().orderBy('count', ascending=False)\nuserId_count.show()","metadata":{"execution":{"iopub.status.busy":"2022-04-29T17:52:01.215170Z","iopub.execute_input":"2022-04-29T17:52:01.215430Z","iopub.status.idle":"2022-04-29T17:53:56.456318Z","shell.execute_reply.started":"2022-04-29T17:52:01.215399Z","shell.execute_reply":"2022-04-29T17:53:56.455225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articleId_count = hm.groupBy(\"article_id\").count().orderBy('count', ascending=False)\narticleId_count.show()","metadata":{"execution":{"iopub.status.busy":"2022-04-29T17:54:05.593015Z","iopub.execute_input":"2022-04-29T17:54:05.593265Z","iopub.status.idle":"2022-04-29T17:56:03.745179Z","shell.execute_reply.started":"2022-04-29T17:54:05.593236Z","shell.execute_reply":"2022-04-29T17:56:03.744490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"background-color:#f7e9ec;font-family:newtimeroman;color:#d90b1c;\">5-Importing important modules</h1>","metadata":{}},{"cell_type":"code","source":"from pyspark.ml.evaluation import RegressionEvaluator\nfrom pyspark.ml.recommendation import ALS","metadata":{"execution":{"iopub.status.busy":"2022-04-29T17:56:12.984187Z","iopub.execute_input":"2022-04-29T17:56:12.984820Z","iopub.status.idle":"2022-04-29T17:56:12.989312Z","shell.execute_reply.started":"2022-04-29T17:56:12.984779Z","shell.execute_reply":"2022-04-29T17:56:12.988317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<h1 style=\"background-color:#f7e9ec;font-family:newtimeroman;color:#d90b1c;\">6-Converting String to index</h1>\n\nBefore making an ALS model it needs to be clear that ALS only accepts integer value as parameters. Hence we need to convert customer_id and article_id column in index form.","metadata":{}},{"cell_type":"code","source":"from pyspark.ml.feature import StringIndexer\nfrom pyspark.ml import Pipeline\nfrom pyspark.sql.functions import col\nindexer = [StringIndexer(inputCol=column, outputCol=column+\"_index\") for column in list(set(hm.columns)-set(['count'])) ]\npipeline = Pipeline(stages=indexer)\ntransformed = pipeline.fit(hm).transform(hm)\ntransformed.show()","metadata":{"execution":{"iopub.status.busy":"2022-04-29T17:56:22.006835Z","iopub.execute_input":"2022-04-29T17:56:22.007280Z","iopub.status.idle":"2022-04-29T18:03:06.478840Z","shell.execute_reply.started":"2022-04-29T17:56:22.007241Z","shell.execute_reply":"2022-04-29T18:03:06.478177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<h1 style=\"background-color:#f7e9ec;font-family:newtimeroman;color:#d90b1c;\">7-Creating training and test data</h1>","metadata":{}},{"cell_type":"code","source":"(training,test)=transformed.randomSplit([0.8, 0.2])","metadata":{"execution":{"iopub.status.busy":"2022-04-29T18:03:22.489687Z","iopub.execute_input":"2022-04-29T18:03:22.490053Z","iopub.status.idle":"2022-04-29T18:03:22.515930Z","shell.execute_reply.started":"2022-04-29T18:03:22.490017Z","shell.execute_reply":"2022-04-29T18:03:22.515163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<h1 style=\"background-color:#f7e9ec;font-family:newtimeroman;color:#d90b1c;\">8-Creating ALS model and fitting data</h1>\n\nTo build the model explicitly specify the columns. Set nonnegative as ‘True’, since we are looking count greater than 0. The model also gives an option to select implicit ratings. Since we are working with explicit, set it to ‘False’ or by default it takes explicit.\n\nWhen using simple random splits as in Spark’s CrossValidator or TrainValidationSplit, it is actually very common to encounter users and/or items in the evaluation set that are not in the training set. By default, Spark assigns NaN predictions during ALSModel.transform when a user and/or item factor is not present in the model.We set cold start strategy to ‘drop’ to ensure we don’t get NaN evaluation metrics.","metadata":{}},{"cell_type":"code","source":"from pyspark.ml.evaluation import RegressionEvaluator\nfrom pyspark.ml.recommendation import ALS\nfrom pyspark.ml.tuning import CrossValidator, ParamGridBuilder\n\n\n#create ALS model\nals=ALS(userCol=\"customer_id_index\",itemCol=\"article_id_index\",ratingCol=\"count\",coldStartStrategy=\"drop\",nonnegative=True)\n\n#tune model using ParamGridBuilder\nparam_grid = ParamGridBuilder()\\\n            .addGrid(als.rank, [15,20,25])\\\n            .addGrid(als.maxIter,[5,10,15])\\\n            .addGrid(als.regParam,[0.09,0.14,0.19])\\\n            .build()\n#define evaluator as RMSE\nevaluator = RegressionEvaluator(metricName = \"rmse\",labelCol = 'count', predictionCol = 'prediction')\n\n#Build cross validation using CrossValidator\ncv = CrossValidator(estimator=als,estimatorParamMaps=param_grid, evaluator=evaluator,numFolds=3)\n\n\n#Fit ALS model to training data\nmodel = cv.fit(training)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-04-29T18:03:29.613902Z","iopub.execute_input":"2022-04-29T18:03:29.614160Z","iopub.status.idle":"2022-04-29T19:05:05.712355Z","shell.execute_reply.started":"2022-04-29T18:03:29.614132Z","shell.execute_reply":"2022-04-29T19:05:05.711632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\"\"\"als=ALS(maxIter=5,regParam=0.09,rank=25,userCol=\"customer_id_index\",itemCol=\"article_id_index\",ratingCol=\"count\",coldStartStrategy=\"drop\",nonnegative=True)\nmodel=als.fit(training)\"\"\"\"\"\"","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-04-25T10:33:54.896318Z","iopub.execute_input":"2022-04-25T10:33:54.896652Z","iopub.status.idle":"2022-04-25T10:37:58.043826Z","shell.execute_reply.started":"2022-04-25T10:33:54.896611Z","shell.execute_reply":"2022-04-25T10:37:58.039774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<h1 style=\"background-color:#f7e9ec;font-family:newtimeroman;color:#d90b1c;\">9-Evaluate rmse</h1>","metadata":{}},{"cell_type":"code","source":"#Extract best model from the tuning exercise using ParamGridBuilder\nbest_model = model.bestModel\n\n#Generate predictions and evaluate using RMSE\npredictions = best_model.transform(test)\nrmse = evaluator.evaluate(predictions)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-04-29T19:06:21.317333Z","iopub.execute_input":"2022-04-29T19:06:21.317622Z","iopub.status.idle":"2022-04-29T19:08:18.947865Z","shell.execute_reply.started":"2022-04-29T19:06:21.317592Z","shell.execute_reply":"2022-04-29T19:08:18.947078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#print evaluation metrics and model parameters\nprint(\"RMSE =\" + str(rmse))\nprint(\"**Best Model**\")\nprint(\"Rank : {}\".format(best_model.rank))\nprint(\"MaxIter: {}\".format(best_model._java_obj.parent().getMaxIter()))\nprint(\"RegParam: {}\".format(best_model._java_obj.parent().getRegParam()))","metadata":{"execution":{"iopub.status.busy":"2022-04-29T19:08:23.274279Z","iopub.execute_input":"2022-04-29T19:08:23.274555Z","iopub.status.idle":"2022-04-29T19:08:23.283855Z","shell.execute_reply.started":"2022-04-29T19:08:23.274523Z","shell.execute_reply":"2022-04-29T19:08:23.282990Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"background-color:#f7e9ec;font-family:newtimeroman;color:#d90b1c;\">10-Providing Recommendations by Article id</h1>","metadata":{}},{"cell_type":"code","source":"user_recs=best_model.recommendForAllItems(10).show(10)","metadata":{"execution":{"iopub.status.busy":"2022-04-29T19:08:42.261496Z","iopub.execute_input":"2022-04-29T19:08:42.261749Z","iopub.status.idle":"2022-04-29T19:08:54.227188Z","shell.execute_reply.started":"2022-04-29T19:08:42.261721Z","shell.execute_reply":"2022-04-29T19:08:54.226528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<h1 style=\"background-color:#f7e9ec;font-family:newtimeroman;color:#d90b1c;\">11-Providing Recommendations by Customer id</h1>","metadata":{}},{"cell_type":"code","source":"df_recom = best_model.recommendForAllUsers(10)\ndf_recom.show(10)","metadata":{"execution":{"iopub.status.busy":"2022-04-29T19:16:38.068545Z","iopub.execute_input":"2022-04-29T19:16:38.069090Z","iopub.status.idle":"2022-04-29T19:16:50.265156Z","shell.execute_reply.started":"2022-04-29T19:16:38.069048Z","shell.execute_reply":"2022-04-29T19:16:50.263935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_recom = df_recom.select(\"customer_id_index\",\"recommendations.article_id_index\")\ndf_recom.show(10)\ndf_recom = df_recom.toPandas()","metadata":{"execution":{"iopub.status.busy":"2022-04-29T19:17:03.941087Z","iopub.execute_input":"2022-04-29T19:17:03.941344Z","iopub.status.idle":"2022-04-29T19:17:26.320994Z","shell.execute_reply.started":"2022-04-29T19:17:03.941312Z","shell.execute_reply":"2022-04-29T19:17:26.320330Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_recom.sort_values('customer_id_index')","metadata":{"execution":{"iopub.status.busy":"2022-04-29T19:17:38.278234Z","iopub.execute_input":"2022-04-29T19:17:38.278501Z","iopub.status.idle":"2022-04-29T19:17:38.298580Z","shell.execute_reply.started":"2022-04-29T19:17:38.278454Z","shell.execute_reply":"2022-04-29T19:17:38.297941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"background-color:#f7e9ec;font-family:newtimeroman;color:#d90b1c;\">12-Converting back to string form</h1>\n\nAs seen in above image the results are in integer form we need to convert it back to its original name.The code is little bit longer given so many conversions.","metadata":{}},{"cell_type":"code","source":"from pyspark.ml.evaluation import RegressionEvaluator\nfrom pyspark.ml.recommendation import ALS\nfrom pyspark.sql import Row\nimport pandas as pd\nmd=transformed.select(transformed['article_id'],transformed['article_id_index'],transformed['customer_id'],transformed['customer_id_index'])\nmd=md.toPandas()\nmd","metadata":{"execution":{"iopub.status.busy":"2022-04-29T19:17:48.043068Z","iopub.execute_input":"2022-04-29T19:17:48.043736Z","iopub.status.idle":"2022-04-29T19:19:52.557095Z","shell.execute_reply.started":"2022-04-29T19:17:48.043695Z","shell.execute_reply":"2022-04-29T19:19:52.556348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dict1 =dict(zip(md['article_id_index'],md['article_id']))\ndict2=dict(zip(md['customer_id_index'],md['customer_id']))\ndf_recom['article_id'] = df_recom['article_id_index'].map(lambda x: [dict1[y] for y in x if y in dict1])\ndf_recom['customer_id']=df_recom['customer_id_index'].map(dict2)\ndf_recom","metadata":{"execution":{"iopub.status.busy":"2022-04-29T19:20:04.380198Z","iopub.execute_input":"2022-04-29T19:20:04.380447Z","iopub.status.idle":"2022-04-29T19:20:04.462673Z","shell.execute_reply.started":"2022-04-29T19:20:04.380419Z","shell.execute_reply":"2022-04-29T19:20:04.461867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"recom_final = df_recom.drop(['customer_id_index','article_id_index'], axis = 1)\nfinalpre=recom_final[['customer_id','article_id']]\nfinalpre","metadata":{"execution":{"iopub.status.busy":"2022-04-29T19:27:03.683254Z","iopub.execute_input":"2022-04-29T19:27:03.683519Z","iopub.status.idle":"2022-04-29T19:27:03.702640Z","shell.execute_reply.started":"2022-04-29T19:27:03.683470Z","shell.execute_reply":"2022-04-29T19:27:03.701930Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"background-color:#f7e9ec;font-family:newtimeroman;color:#d90b1c;\">13-Export the prediction</h1>","metadata":{}},{"cell_type":"code","source":"my_pred = finalpre.toPandas()\nmy_pred.to_csv('my_pred.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2022-04-29T19:27:16.893422Z","iopub.execute_input":"2022-04-29T19:27:16.893949Z","iopub.status.idle":"2022-04-29T19:27:16.911154Z","shell.execute_reply.started":"2022-04-29T19:27:16.893912Z","shell.execute_reply":"2022-04-29T19:27:16.910344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\"background-color:#f7e9ec;font-family:newtimeroman;color:#d90b1c;font-size:150%;text-align:center;border-radius:10px 10px;border-style:solid;border-color:#d90b1c;\">Please do leave your comments /suggestions and if you like this kernel greatly appreciate to UPVOTE</p>","metadata":{}}]}