{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# [H&M Fashion] GPU-accelerated RecSys Dataset Profiler by NVIDIA Merlin\n\n### Context\n\nThis notebook is based on a [*RecSys Dataset Profiler* template](https://github.com/NVIDIA-Merlin/competitions/blob/main/SIGIR_eCommerce_Challenge_2021/task1_session_based_rec/0-eda/coveo_retail_recsys_dataset_profiler.ipynb) created by [NVIDIA Merlin](https://developer.nvidia.com/nvidia-merlin) team. \nIt is useful for a first assessment of datasets for recommender systems, generating some useful statistics and plots from user interactions datasets, that are useful to decide important aspects of recommender systems like:\n\n- May Collaborative Filtering algorithms (e.g., Matrix Factorization (MF), Neural Collaborative Filtering (NCF)) be suitable for this dataset (e.g. based on the user-item matrix sparsity and on items and users long-tail distribution)?\n- Should the recommender system care about recommending or not items that users have already interacted/purchased in the past?\n- How many new items became available daily and how many new users are observed every day? Based on those answers, you can understand the level of User and Item cold-start problem in the dataset\n- How fast do the items loose relevance for users? Should we care about recommending fresher items rather than the old ones?\n\nFeel free to use this notebook for other RecSys datasets :)\n\n### How to use this notebook for a RecSys dataset?\nYou just need to set some variables from the **Config** section and run the full notebook :)\n\n### Requirements\nThis notebook uses NVIDIA RAPIDS for GPU-accelerated data analysis, which is pre-installed in Kaggle Kernels (remember to enable the GPU Accelerator).   \n\nThe `cudf` is a GPU-accelerated dataframe library equivalent to `pandas`. The `dask_cudf` allow for distributing the data frame operations across multiple GPUs, for lightining fast processing for large datasets. In Kaggle Kernels we only have access to a single GPU, but we keep `dask_cudf` dependency to enable running such EDA in a distributed fashion to take advantage of multiple GPU if available.\n\nIf you are running this notebook in another environment, you can install RAPIDS on your own environment using ```conda``` following the examples [here](https://rapids.ai/).\n\n### NVIDIA Merlin\nThis notebook was created by the NVIDIA Merlin team. \nNVIDIA Merlin is an open-source framework for building large-scale deep learning recommender system. You can find more resources about Merlin here:\n- https://developer.nvidia.com/nvidia-merlin\n- https://medium.com/nvidia-merlin\n- https://github.com/NVIDIA-Merlin","metadata":{}},{"cell_type":"markdown","source":"## Imports","metadata":{}},{"cell_type":"code","source":"import os\nimport shutil\nimport pandas as pd\nimport numpy as np\nfrom collections import OrderedDict\nfrom IPython.display import display\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:04.353709Z","iopub.execute_input":"2022-02-10T16:28:04.354076Z","iopub.status.idle":"2022-02-10T16:28:04.358829Z","shell.execute_reply.started":"2022-02-10T16:28:04.354037Z","shell.execute_reply":"2022-02-10T16:28:04.357885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# These dependencies require a GPU\nimport cupy as cp\nimport cudf\nimport dask as dask, dask_cudf","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:04.360342Z","iopub.execute_input":"2022-02-10T16:28:04.360726Z","iopub.status.idle":"2022-02-10T16:28:08.063008Z","shell.execute_reply.started":"2022-02-10T16:28:04.360691Z","shell.execute_reply":"2022-02-10T16:28:08.062279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Config","metadata":{}},{"cell_type":"markdown","source":"## Dataset metadata","metadata":{}},{"cell_type":"code","source":"#If you change the order of the creation of new keys in the \"dataset_info\" OrderedDict, \n#the order of the final CSV columns with the dataset profile will change accordingly\ndataset_info = OrderedDict()\ndataset_info['name'] = 'H&B_fashion'\ndataset_info['domain'] = 'ecommerce'\ndataset_info['description'] = ''\ndataset_info['source'] = 'https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations'\ndataset_info['event_types'] = 'purchases'","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:08.064272Z","iopub.execute_input":"2022-02-10T16:28:08.064533Z","iopub.status.idle":"2022-02-10T16:28:08.069788Z","shell.execute_reply.started":"2022-02-10T16:28:08.064498Z","shell.execute_reply":"2022-02-10T16:28:08.069133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data format and path","metadata":{}},{"cell_type":"code","source":"#Accepted data formats are: csv | tsv | parquet\nDATA_FORMAT = 'csv' \n#List of columns names to be used for CSV / TSV files without the header line\nHEADLESS_CSV_COLUMN_NAMES = None #Example: ['col1','col2']","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:08.071884Z","iopub.execute_input":"2022-02-10T16:28:08.07238Z","iopub.status.idle":"2022-02-10T16:28:08.07833Z","shell.execute_reply.started":"2022-02-10T16:28:08.072342Z","shell.execute_reply":"2022-02-10T16:28:08.077621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DATA_FOLDER = \"../input/h-and-m-personalized-fashion-recommendations/\"\nFILENAME_PATTERN = 'transactions_train.csv'\nDATA_PATH = os.path.join(DATA_FOLDER, FILENAME_PATTERN)\n\n!ls $DATA_PATH","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:08.079962Z","iopub.execute_input":"2022-02-10T16:28:08.080288Z","iopub.status.idle":"2022-02-10T16:28:08.782163Z","shell.execute_reply.started":"2022-02-10T16:28:08.080254Z","shell.execute_reply":"2022-02-10T16:28:08.78136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Dataset columns","metadata":{}},{"cell_type":"code","source":"HAS_TIMESTAMP = True\nHAS_USER = True\n\n# Set these column names from your input dataset\nCOL_ITEM_ID = 'article_id'\nCOL_USER_ID = 'customer_id'\nCOL_DATETIME = 't_dat' #Can be None if HAS_TIMESTAMP=False","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:08.783878Z","iopub.execute_input":"2022-02-10T16:28:08.78418Z","iopub.status.idle":"2022-02-10T16:28:08.789512Z","shell.execute_reply.started":"2022-02-10T16:28:08.784134Z","shell.execute_reply":"2022-02-10T16:28:08.788571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Do not change\ncols_origin = [COL_ITEM_ID]\nif HAS_USER:\n    cols_origin.append(COL_USER_ID)\nif HAS_TIMESTAMP:\n    cols_origin.append(COL_DATETIME)\n    \ndataset_info['has_timestamp'] = HAS_TIMESTAMP","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:08.791127Z","iopub.execute_input":"2022-02-10T16:28:08.791469Z","iopub.status.idle":"2022-02-10T16:28:08.798796Z","shell.execute_reply.started":"2022-02-10T16:28:08.791433Z","shell.execute_reply":"2022-02-10T16:28:08.798047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Dates config","metadata":{}},{"cell_type":"markdown","source":"### Datetime / timestamp conversion","metadata":{}},{"cell_type":"code","source":"#The rest of the notebook expects the Time column to be in 'datetime64' dtype\n#Possible values:\n# - None - For a datetime column in a parquet file, keep datetimes as they are (no conversion required)\n# - 's' - For timestamp in seconds or general date represented as string like: '2016-04-09' or '2019-10-01 02:15:47 UTC'\n# - 'ms' - For timestamp in miliseconds\nDATETIME_CONVERTION = 's'","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:08.800229Z","iopub.execute_input":"2022-02-10T16:28:08.800708Z","iopub.status.idle":"2022-02-10T16:28:08.806165Z","shell.execute_reply.started":"2022-02-10T16:28:08.800671Z","shell.execute_reply":"2022-02-10T16:28:08.805381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Dates filtering","metadata":{}},{"cell_type":"code","source":"#Whether to filter the dataset by date times (inclusive)\nMIN_DATETIME = None   #pd.Timestamp(2017, 10, 01)\nMAX_DATETIME = None #Including hour","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:08.807822Z","iopub.execute_input":"2022-02-10T16:28:08.808086Z","iopub.status.idle":"2022-02-10T16:28:08.813998Z","shell.execute_reply.started":"2022-02-10T16:28:08.808055Z","shell.execute_reply":"2022-02-10T16:28:08.813232Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Outputs","metadata":{}},{"cell_type":"code","source":"OUTPUT_PATH = '/kaggle/working'\n#Wheather to save the dataset info profile to CSV\nSAVE_DATASET_INFO_CSV = True\n#Whether to save the item and user frequency cumulative distributions (e.g. for later plotting of multiple datasets distributions in the same chart)\nSAVE_USERS_ITEMS_CUM_FREQ_DISTR = True","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:08.815358Z","iopub.execute_input":"2022-02-10T16:28:08.816757Z","iopub.status.idle":"2022-02-10T16:28:08.823404Z","shell.execute_reply.started":"2022-02-10T16:28:08.81672Z","shell.execute_reply":"2022-02-10T16:28:08.822538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## GPU config","metadata":{}},{"cell_type":"code","source":"#The list of GPU devices to be used by Dask-cuDF. Should be comma-separated, e.g., \"0,1,2,3\"\nCUDA_VISIBLE_DEVICES = \"0\"\n# Caches the dataset into GPU memory. Disable this if you have a dataset larger than GPU memory and you start getting CUDA OOM erros\nCACHE_DATASET_ON_GPU = True","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:08.824406Z","iopub.execute_input":"2022-02-10T16:28:08.824979Z","iopub.status.idle":"2022-02-10T16:28:08.832031Z","shell.execute_reply.started":"2022-02-10T16:28:08.824943Z","shell.execute_reply":"2022-02-10T16:28:08.831293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Listing the available GPUs","metadata":{}},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:08.833394Z","iopub.execute_input":"2022-02-10T16:28:08.834175Z","iopub.status.idle":"2022-02-10T16:28:09.526427Z","shell.execute_reply.started":"2022-02-10T16:28:08.834133Z","shell.execute_reply":"2022-02-10T16:28:09.525601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"os.environ[\"CUDA_VISIBLE_DEVICES\"]=CUDA_VISIBLE_DEVICES","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:09.532113Z","iopub.execute_input":"2022-02-10T16:28:09.533073Z","iopub.status.idle":"2022-02-10T16:28:09.538009Z","shell.execute_reply.started":"2022-02-10T16:28:09.533022Z","shell.execute_reply":"2022-02-10T16:28:09.537295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Uses a RAID folder if it is available (DGX), if not it uses the /tmp folder for DASK workspace\ntmp_directory = '/tmp'    \ndask_workdir = os.path.join(tmp_directory, 'dask-workdir')    \nprint('Dask dir:', dask_workdir)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:09.539103Z","iopub.execute_input":"2022-02-10T16:28:09.539441Z","iopub.status.idle":"2022-02-10T16:28:09.548649Z","shell.execute_reply.started":"2022-02-10T16:28:09.539404Z","shell.execute_reply":"2022-02-10T16:28:09.547823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Make sure we have a clean worker space for Dask\nif os.path.isdir(dask_workdir):\n    shutil.rmtree(dask_workdir)\nos.mkdir(dask_workdir)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:09.549994Z","iopub.execute_input":"2022-02-10T16:28:09.550365Z","iopub.status.idle":"2022-02-10T16:28:09.557682Z","shell.execute_reply.started":"2022-02-10T16:28:09.550301Z","shell.execute_reply":"2022-02-10T16:28:09.557018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data loading & preproc","metadata":{}},{"cell_type":"code","source":"#Do not change the dest column names\ncols_dest = ['ItemId']\nif HAS_USER:\n    cols_dest.append('UserId')\nif HAS_TIMESTAMP:\n    cols_dest.append('Time')","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:09.558973Z","iopub.execute_input":"2022-02-10T16:28:09.559291Z","iopub.status.idle":"2022-02-10T16:28:09.568765Z","shell.execute_reply.started":"2022-02-10T16:28:09.559254Z","shell.execute_reply":"2022-02-10T16:28:09.567828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if DATA_FORMAT == 'parquet':\n    ddf = dask_cudf.read_parquet(DATA_PATH)\n    \nelif DATA_FORMAT in ['csv', 'tsv']:\n    ddf = dask_cudf.read_csv(DATA_PATH,                              \n                             sep='\\t' if DATA_FORMAT == 'tsv' else ',',\n                             names=HEADLESS_CSV_COLUMN_NAMES\n                            )\nelse:\n    ValueError('Acceptable data formats are: parquet | csv | tsv')","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:09.571675Z","iopub.execute_input":"2022-02-10T16:28:09.571972Z","iopub.status.idle":"2022-02-10T16:28:11.814103Z","shell.execute_reply.started":"2022-02-10T16:28:09.571931Z","shell.execute_reply":"2022-02-10T16:28:11.81325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(ddf.dtypes)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:11.818328Z","iopub.execute_input":"2022-02-10T16:28:11.818546Z","iopub.status.idle":"2022-02-10T16:28:11.827941Z","shell.execute_reply.started":"2022-02-10T16:28:11.818519Z","shell.execute_reply":"2022-02-10T16:28:11.826735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset_info['columns'] = ','.join(list(ddf.columns))\ndataset_info['columns']","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:11.829482Z","iopub.execute_input":"2022-02-10T16:28:11.830275Z","iopub.status.idle":"2022-02-10T16:28:11.841893Z","shell.execute_reply.started":"2022-02-10T16:28:11.830233Z","shell.execute_reply":"2022-02-10T16:28:11.841141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Keep only the required columns and renaming to fixed names\nddf = ddf[cols_origin]\nddf.columns = cols_dest","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:11.843599Z","iopub.execute_input":"2022-02-10T16:28:11.844433Z","iopub.status.idle":"2022-02-10T16:28:11.855322Z","shell.execute_reply.started":"2022-02-10T16:28:11.844402Z","shell.execute_reply":"2022-02-10T16:28:11.854057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Datetime processing","metadata":{}},{"cell_type":"code","source":"if HAS_TIMESTAMP:\n    display(ddf['Time'].head())","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:11.857104Z","iopub.execute_input":"2022-02-10T16:28:11.857539Z","iopub.status.idle":"2022-02-10T16:28:15.459348Z","shell.execute_reply.started":"2022-02-10T16:28:11.857503Z","shell.execute_reply":"2022-02-10T16:28:15.458667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if HAS_TIMESTAMP:\n    #Converts date time if configured to do so\n    if DATETIME_CONVERTION is not None:\n        ddf['Time'] = ddf['Time'].astype(f'datetime64[{DATETIME_CONVERTION}]')","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:15.460457Z","iopub.execute_input":"2022-02-10T16:28:15.460855Z","iopub.status.idle":"2022-02-10T16:28:16.038546Z","shell.execute_reply.started":"2022-02-10T16:28:15.460817Z","shell.execute_reply":"2022-02-10T16:28:16.037772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if HAS_TIMESTAMP:\n    display(ddf['Time'].head())","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:16.041706Z","iopub.execute_input":"2022-02-10T16:28:16.041911Z","iopub.status.idle":"2022-02-10T16:28:16.311255Z","shell.execute_reply.started":"2022-02-10T16:28:16.041878Z","shell.execute_reply":"2022-02-10T16:28:16.310581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if HAS_TIMESTAMP:\n    #Creates a string representation of the dates\n    ddf['DateStr'] = ddf['Time'].dt.strftime(\"%Y-%m-%d\")","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:16.312372Z","iopub.execute_input":"2022-02-10T16:28:16.312951Z","iopub.status.idle":"2022-02-10T16:28:16.328858Z","shell.execute_reply.started":"2022-02-10T16:28:16.312913Z","shell.execute_reply":"2022-02-10T16:28:16.328104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if HAS_TIMESTAMP:\n    #Filtering the dataset based on minimum and maximum datetimes (inclusive)\n    if MIN_DATETIME is not None:\n        ddf = ddf[ddf['Time'] >= MIN_DATETIME]\n\n    if MAX_DATETIME is not None:\n        ddf = ddf[ddf['Time'] <= MAX_DATETIME]    ","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:16.330122Z","iopub.execute_input":"2022-02-10T16:28:16.330355Z","iopub.status.idle":"2022-02-10T16:28:16.336389Z","shell.execute_reply.started":"2022-02-10T16:28:16.330323Z","shell.execute_reply":"2022-02-10T16:28:16.335715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Dataset caching","metadata":{}},{"cell_type":"code","source":"#Caches the dataset into GPU memory (if it fits). This is a lazy op, so caching will happen in the next compute() op\nif CACHE_DATASET_ON_GPU:\n    ddf, = dask.persist(ddf)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:16.337786Z","iopub.execute_input":"2022-02-10T16:28:16.338096Z","iopub.status.idle":"2022-02-10T16:28:59.977475Z","shell.execute_reply.started":"2022-02-10T16:28:16.338059Z","shell.execute_reply":"2022-02-10T16:28:59.976746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Extracting columns from date","metadata":{}},{"cell_type":"code","source":"%%time\nif HAS_TIMESTAMP:\n    min_date = ddf['Time'].min().compute().date()\n    dataset_info['first_date'] = min_date.strftime('%Y-%m-%d')\n    print(min_date)\nelse:\n    dataset_info['first_date'] = None","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:28:59.978638Z","iopub.execute_input":"2022-02-10T16:28:59.979258Z","iopub.status.idle":"2022-02-10T16:29:00.00332Z","shell.execute_reply.started":"2022-02-10T16:28:59.979216Z","shell.execute_reply":"2022-02-10T16:29:00.002589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_TIMESTAMP:\n    max_date = ddf['Time'].max().compute().date()\n    dataset_info['last_date'] = max_date.strftime('%Y-%m-%d')\n    print(max_date)\nelse:\n    dataset_info['last_date'] = None","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:00.004582Z","iopub.execute_input":"2022-02-10T16:29:00.004856Z","iopub.status.idle":"2022-02-10T16:29:00.025206Z","shell.execute_reply.started":"2022-02-10T16:29:00.004827Z","shell.execute_reply":"2022-02-10T16:29:00.02447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if HAS_TIMESTAMP:\n    dataset_info['num_days'] = (max_date - min_date).days + 1\n    print(dataset_info['num_days'])\nelse:\n    dataset_info['num_days'] = None","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:00.026345Z","iopub.execute_input":"2022-02-10T16:29:00.027001Z","iopub.status.idle":"2022-02-10T16:29:00.032746Z","shell.execute_reply.started":"2022-02-10T16:29:00.026962Z","shell.execute_reply":"2022-02-10T16:29:00.031985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_TIMESTAMP:\n    #Creating an auxiliary table with Pandas to extract weekofyear, because it is not available yet on cudf (at 0.17 version)\n    dates_df = pd.DataFrame(pd.date_range(start=min_date, end=max_date), columns=['date'])\n    dates_df['weekofyear'] = dates_df['date'].dt.weekofyear\n    dates_df['month'] = dates_df['date'].dt.month\n    dates_df['year'] = dates_df['date'].dt.year\n    dates_df['year-week'] = dates_df['year'].astype('str') + \"-w\" + \\\n                            dates_df['weekofyear'].apply(lambda x: \"{:02d}\".format(x))\n    dates_df['year-month'] = dates_df['year'].astype('str') +\"-\" +  \\\n                             dates_df['month'].apply(lambda x: \"{:02d}\".format(x))\n    dates_df['DateStr'] = dates_df['date'].dt.strftime(\"%Y-%m-%d\")\n    del dates_df['weekofyear'], dates_df['month'], dates_df['year'], dates_df['date']\n    dates_df.set_index('DateStr', inplace=True)\n    dates_df = cudf.from_pandas(dates_df)\n    display(dates_df.head())","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:00.034036Z","iopub.execute_input":"2022-02-10T16:29:00.034814Z","iopub.status.idle":"2022-02-10T16:29:00.086632Z","shell.execute_reply.started":"2022-02-10T16:29:00.034775Z","shell.execute_reply":"2022-02-10T16:29:00.08585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Includes year-week and year-month into the main dataframe\nif HAS_TIMESTAMP:\n    ddf = ddf.merge(dates_df,  left_on='DateStr', right_index=True)\n    display(ddf.head(10))","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:00.087943Z","iopub.execute_input":"2022-02-10T16:29:00.088388Z","iopub.status.idle":"2022-02-10T16:29:00.178133Z","shell.execute_reply.started":"2022-02-10T16:29:00.088346Z","shell.execute_reply":"2022-02-10T16:29:00.177313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Basic stats","metadata":{}},{"cell_type":"markdown","source":"## Counts","metadata":{}},{"cell_type":"markdown","source":"#### Number of interactions","metadata":{}},{"cell_type":"code","source":"%%time\nnrows = len(ddf)\ndataset_info['num_interactions'] = nrows\nprint(nrows)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:00.179532Z","iopub.execute_input":"2022-02-10T16:29:00.179992Z","iopub.status.idle":"2022-02-10T16:29:00.632852Z","shell.execute_reply.started":"2022-02-10T16:29:00.179952Z","shell.execute_reply":"2022-02-10T16:29:00.632071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Number of items","metadata":{}},{"cell_type":"code","source":"%%time\nn_items = ddf['ItemId'].nunique().compute()\ndataset_info['num_items'] = n_items\nn_items","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:00.63446Z","iopub.execute_input":"2022-02-10T16:29:00.634733Z","iopub.status.idle":"2022-02-10T16:29:01.154868Z","shell.execute_reply.started":"2022-02-10T16:29:00.634697Z","shell.execute_reply":"2022-02-10T16:29:01.154179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Number of users","metadata":{}},{"cell_type":"code","source":"%%time\nif HAS_USER:\n    n_users = ddf['UserId'].nunique().compute()\nelse:\n    n_users = 0\ndataset_info['num_users'] = n_users\nn_users","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:01.156183Z","iopub.execute_input":"2022-02-10T16:29:01.156634Z","iopub.status.idle":"2022-02-10T16:29:06.143952Z","shell.execute_reply.started":"2022-02-10T16:29:01.156584Z","shell.execute_reply":"2022-02-10T16:29:06.143227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### User-Item Matrix Sparcity","metadata":{}},{"cell_type":"code","source":"%%time\nif HAS_USER:\n    n_unique_user_item_pairs = len(ddf[['UserId', 'ItemId']].drop_duplicates())\nelse:\n    n_unique_user_item_pairs = 0\nn_unique_user_item_pairs","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:06.148051Z","iopub.execute_input":"2022-02-10T16:29:06.151073Z","iopub.status.idle":"2022-02-10T16:29:25.865917Z","shell.execute_reply.started":"2022-02-10T16:29:06.151027Z","shell.execute_reply":"2022-02-10T16:29:25.86517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if HAS_USER:\n    sparsity = 1.0 - (n_unique_user_item_pairs / (n_items * n_users))\nelse:\n    sparsity = None\ndataset_info['sparsity_user_item_matrix'] = sparsity\nsparsity","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:25.8674Z","iopub.execute_input":"2022-02-10T16:29:25.867895Z","iopub.status.idle":"2022-02-10T16:29:25.874865Z","shell.execute_reply.started":"2022-02-10T16:29:25.867858Z","shell.execute_reply":"2022-02-10T16:29:25.874212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Distributions","metadata":{}},{"cell_type":"code","source":"def gini_index_cupy(array):\n    \"\"\"Calculate the Gini coefficient of a numpy array.\"\"\"\n    # based on bottom eq:\n    # http://www.statsdirect.com/help/generatedimages/equations/equation154.svg\n    # from:\n    # http://www.statsdirect.com/help/default.htm#nonparametric_methods/gini.htm\n    # All values are treated equally, arrays must be 1d:\n    array = array.flatten()\n    if cp.amin(array) < 0:\n        # Values cannot be negative:\n        array -= cp.amin(array)\n    # Values cannot be 0:\n    array += 0.0000001\n    # Values must be sorted:\n    array = cp.sort(array)\n    # Index per array element:\n    index = cp.arange(1,array.shape[0]+1)\n    # Number of array elements:\n    n = array.shape[0]\n    # Gini coefficient:\n    return float(((cp.sum((2 * index - n  - 1) * array)) / (n * cp.sum(array))))","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:25.876054Z","iopub.execute_input":"2022-02-10T16:29:25.876859Z","iopub.status.idle":"2022-02-10T16:29:25.884716Z","shell.execute_reply.started":"2022-02-10T16:29:25.876822Z","shell.execute_reply":"2022-02-10T16:29:25.883833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def set_distr_percentiles(data, dataset_info, prefix, dtype='int'):\n    if data is None:\n        #If the dataframe with stats is not available, fill columns with None\n        dataset_info[f\"min_{prefix}\"] = None\n        dataset_info[f\"p25_{prefix}\"] = None\n        dataset_info[f\"p50_{prefix}\"] = None\n        dataset_info[f\"p75_{prefix}\"] = None\n        dataset_info[f\"p90_{prefix}\"] = None\n        dataset_info[f\"p95_{prefix}\"] = None\n        dataset_info[f\"p99_{prefix}\"] = None\n        dataset_info[f\"max_{prefix}\"] = None\n    else: \n        dataset_info[f\"min_{prefix}\"] = data['min'].astype(dtype)\n        dataset_info[f\"p25_{prefix}\"] = data['25%'].astype(dtype)\n        dataset_info[f\"p50_{prefix}\"] = data['50%'].astype(dtype)\n        dataset_info[f\"p75_{prefix}\"] = data['75%'].astype(dtype)\n        dataset_info[f\"p90_{prefix}\"] = data['90%'].astype(dtype)\n        dataset_info[f\"p95_{prefix}\"] = data['95%'].astype(dtype)\n        dataset_info[f\"p99_{prefix}\"] = data['99%'].astype(dtype)\n        dataset_info[f\"max_{prefix}\"] = data['max'].astype(dtype)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:25.88587Z","iopub.execute_input":"2022-02-10T16:29:25.886605Z","iopub.status.idle":"2022-02-10T16:29:25.897603Z","shell.execute_reply.started":"2022-02-10T16:29:25.886569Z","shell.execute_reply":"2022-02-10T16:29:25.896869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#The percentiles that will be extracted for all distributions\nPERCENTILES=np.concatenate([np.arange(0.0, 1.1, 0.1), np.array([0.25, 0.75, 0.95, 0.99])])\nPERCENTILES","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:25.898767Z","iopub.execute_input":"2022-02-10T16:29:25.899489Z","iopub.status.idle":"2022-02-10T16:29:25.909015Z","shell.execute_reply.started":"2022-02-10T16:29:25.899452Z","shell.execute_reply":"2022-02-10T16:29:25.908157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### # Interactions per Item distribution","metadata":{}},{"cell_type":"code","source":"%%time\nitems_freq_df = ddf.groupby('ItemId').size().to_frame('freq').compute().sort_values('freq', ascending=False)\nitems_freq_df.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:25.91707Z","iopub.execute_input":"2022-02-10T16:29:25.917523Z","iopub.status.idle":"2022-02-10T16:29:26.567087Z","shell.execute_reply.started":"2022-02-10T16:29:25.917492Z","shell.execute_reply":"2022-02-10T16:29:26.566313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"items_freq_gini_index = gini_index_cupy(items_freq_df['freq'].values.astype('float'))\ndataset_info['items_freq_gini_index'] = items_freq_gini_index","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:26.568598Z","iopub.execute_input":"2022-02-10T16:29:26.569102Z","iopub.status.idle":"2022-02-10T16:29:29.502594Z","shell.execute_reply.started":"2022-02-10T16:29:26.569062Z","shell.execute_reply":"2022-02-10T16:29:29.5017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nitems_freq_cum_perc = (items_freq_df['freq'].cumsum() / items_freq_df['freq'].sum()).to_frame('cum_interactions_by_item_freq')\nitems_freq_cum_perc['dummy'] = 1\nitems_freq_cum_perc['cum_perc_items'] = items_freq_cum_perc['dummy'].cumsum() / items_freq_cum_perc['dummy'].sum()\ndel items_freq_cum_perc['dummy']\nitems_freq_cum_perc","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:29.504176Z","iopub.execute_input":"2022-02-10T16:29:29.504468Z","iopub.status.idle":"2022-02-10T16:29:30.670784Z","shell.execute_reply.started":"2022-02-10T16:29:29.50443Z","shell.execute_reply":"2022-02-10T16:29:30.669837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"items_freq_cum_perc_pdf = items_freq_cum_perc.set_index('cum_perc_items').to_pandas()","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:30.672271Z","iopub.execute_input":"2022-02-10T16:29:30.672524Z","iopub.status.idle":"2022-02-10T16:29:30.681846Z","shell.execute_reply.started":"2022-02-10T16:29:30.672487Z","shell.execute_reply":"2022-02-10T16:29:30.680736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ax = items_freq_cum_perc_pdf.plot.line(figsize=(15,8))\n\nax.set_title('Cumulative distribution of Items Frequency')\nax.set_ylabel('% of interactions')\nax.set_xlabel('% of items')","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:30.683828Z","iopub.execute_input":"2022-02-10T16:29:30.68411Z","iopub.status.idle":"2022-02-10T16:29:31.107704Z","shell.execute_reply.started":"2022-02-10T16:29:30.684075Z","shell.execute_reply":"2022-02-10T16:29:31.10682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Reindexing the distribution by 1% increments\nitems_freq_cum_perc_reindexed_pdf = items_freq_cum_perc_pdf.reindex(np.arange(0.0, 1.01, 0.01), method='pad').fillna(0.)\nitems_freq_cum_perc_reindexed_pdf.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:31.111411Z","iopub.execute_input":"2022-02-10T16:29:31.111662Z","iopub.status.idle":"2022-02-10T16:29:31.13579Z","shell.execute_reply.started":"2022-02-10T16:29:31.111624Z","shell.execute_reply":"2022-02-10T16:29:31.134851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Extracting some percentiles to save\nitems_freq_cum_perc_selected_pdf = items_freq_cum_perc_reindexed_pdf.loc[[0.01, 0.05, 0.10, 0.25, 0.50]]\ndataset_info['top-01%_item_cum_freq'] = items_freq_cum_perc_selected_pdf.loc[0.01][0]\ndataset_info['top-05%_item_cum_freq'] = items_freq_cum_perc_selected_pdf.loc[0.05][0]\ndataset_info['top-10%_item_cum_freq'] = items_freq_cum_perc_selected_pdf.loc[0.10][0]\ndataset_info['top-25%_item_cum_freq'] = items_freq_cum_perc_selected_pdf.loc[0.25][0]\ndataset_info['top-50%_item_cum_freq'] = items_freq_cum_perc_selected_pdf.loc[0.50][0]\nitems_freq_cum_perc_selected_pdf","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:31.137633Z","iopub.execute_input":"2022-02-10T16:29:31.137949Z","iopub.status.idle":"2022-02-10T16:29:31.153245Z","shell.execute_reply.started":"2022-02-10T16:29:31.137891Z","shell.execute_reply":"2022-02-10T16:29:31.15204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"items_freq_percentiles_df = items_freq_df['freq'].describe(percentiles=PERCENTILES)\nset_distr_percentiles(items_freq_percentiles_df, dataset_info, prefix='item_freq')\nitems_freq_percentiles_df","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:31.155595Z","iopub.execute_input":"2022-02-10T16:29:31.156315Z","iopub.status.idle":"2022-02-10T16:29:31.377996Z","shell.execute_reply.started":"2022-02-10T16:29:31.156224Z","shell.execute_reply":"2022-02-10T16:29:31.37724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ax = items_freq_df.groupby('freq').size().to_pandas().hist(bins=100, figsize=(15,8))\nax.set_title('Items Frequency histogram')","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:31.379345Z","iopub.execute_input":"2022-02-10T16:29:31.379771Z","iopub.status.idle":"2022-02-10T16:29:31.758522Z","shell.execute_reply.started":"2022-02-10T16:29:31.379731Z","shell.execute_reply":"2022-02-10T16:29:31.757836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### # Interactions per User distribution","metadata":{}},{"cell_type":"code","source":"%%time\nif HAS_USER:\n    user_freq_series = ddf.groupby('UserId').size().compute()\n    display(user_freq_series.head(10))","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:31.759864Z","iopub.execute_input":"2022-02-10T16:29:31.760149Z","iopub.status.idle":"2022-02-10T16:29:32.733617Z","shell.execute_reply.started":"2022-02-10T16:29:31.760102Z","shell.execute_reply":"2022-02-10T16:29:32.732349Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if HAS_USER:\n    users_freq_gini_index = gini_index_cupy(user_freq_series.values.astype('float'))\nelse:\n    users_freq_gini_index = None\ndataset_info['users_freq_gini_index'] = users_freq_gini_index","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:32.734765Z","iopub.execute_input":"2022-02-10T16:29:32.735099Z","iopub.status.idle":"2022-02-10T16:29:32.751344Z","shell.execute_reply.started":"2022-02-10T16:29:32.735059Z","shell.execute_reply":"2022-02-10T16:29:32.750696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_USER:\n    users_freq_df = ddf.groupby('UserId').size().to_frame('freq').compute().sort_values('freq', ascending=False)\n    display(users_freq_df.head(10))","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:32.752388Z","iopub.execute_input":"2022-02-10T16:29:32.752723Z","iopub.status.idle":"2022-02-10T16:29:33.556475Z","shell.execute_reply.started":"2022-02-10T16:29:32.75266Z","shell.execute_reply":"2022-02-10T16:29:33.555797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_USER:\n    users_freq_cum_perc = (users_freq_df['freq'].cumsum() / users_freq_df['freq'].sum()).to_frame('cum_interactions_by_user_freq')\n    users_freq_cum_perc['dummy'] = 1\n    users_freq_cum_perc['cum_perc_users'] = users_freq_cum_perc['dummy'].cumsum() / users_freq_cum_perc['dummy'].sum()\n    del users_freq_cum_perc['dummy']\n    display(users_freq_cum_perc)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:33.557644Z","iopub.execute_input":"2022-02-10T16:29:33.559159Z","iopub.status.idle":"2022-02-10T16:29:33.616388Z","shell.execute_reply.started":"2022-02-10T16:29:33.559111Z","shell.execute_reply":"2022-02-10T16:29:33.615697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if HAS_USER:\n    users_freq_cum_perc_pdf = users_freq_cum_perc.set_index('cum_perc_users').to_pandas()","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:33.617521Z","iopub.execute_input":"2022-02-10T16:29:33.618172Z","iopub.status.idle":"2022-02-10T16:29:33.642953Z","shell.execute_reply.started":"2022-02-10T16:29:33.618136Z","shell.execute_reply":"2022-02-10T16:29:33.642297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if HAS_USER:\n    ax = users_freq_cum_perc_pdf.plot.line(figsize=(15,8))\n\n    ax.set_title('Cumulative distribution of Users Frequency')\n    ax.set_ylabel('% of interactions')\n    ax.set_xlabel('% of users')","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:33.643985Z","iopub.execute_input":"2022-02-10T16:29:33.644229Z","iopub.status.idle":"2022-02-10T16:29:35.167016Z","shell.execute_reply.started":"2022-02-10T16:29:33.644196Z","shell.execute_reply":"2022-02-10T16:29:35.1663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if HAS_USER:\n    users_freq_cum_perc_reindexed_pdf = users_freq_cum_perc_pdf.reindex(np.arange(0.0, 1.01, 0.01), method='pad').fillna(0.)\n    display(users_freq_cum_perc_reindexed_pdf.head())","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:35.168156Z","iopub.execute_input":"2022-02-10T16:29:35.169055Z","iopub.status.idle":"2022-02-10T16:29:35.235326Z","shell.execute_reply.started":"2022-02-10T16:29:35.169012Z","shell.execute_reply":"2022-02-10T16:29:35.234458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if HAS_USER:\n    users_freq_cum_perc_selected_pdf = users_freq_cum_perc_reindexed_pdf.reindex(np.arange(0.0, 1.01, 0.01), method='pad') \\\n                            .loc[[0.01, 0.05, 0.10, 0.25, 0.50]]\n    dataset_info['top-01%_user_cum_freq'] = users_freq_cum_perc_selected_pdf.loc[0.01][0]\n    dataset_info['top-05%_user_cum_freq'] = users_freq_cum_perc_selected_pdf.loc[0.05][0]\n    dataset_info['top-10%_user_cum_freq'] = users_freq_cum_perc_selected_pdf.loc[0.10][0]\n    dataset_info['top-25%_user_cum_freq'] = users_freq_cum_perc_selected_pdf.loc[0.25][0]\n    dataset_info['top-50%_user_cum_freq'] = users_freq_cum_perc_selected_pdf.loc[0.50][0]\n    users_freq_cum_perc_selected_pdf\nelse:\n    dataset_info['top-01%_user_cum_freq'] = None\n    dataset_info['top-05%_user_cum_freq'] = None\n    dataset_info['top-10%_user_cum_freq'] = None\n    dataset_info['top-25%_user_cum_freq'] = None\n    dataset_info['top-50%_user_cum_freq'] = None","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:35.236928Z","iopub.execute_input":"2022-02-10T16:29:35.237208Z","iopub.status.idle":"2022-02-10T16:29:35.245745Z","shell.execute_reply.started":"2022-02-10T16:29:35.23717Z","shell.execute_reply":"2022-02-10T16:29:35.245083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if HAS_USER:\n    users_freq_percentiles_df = users_freq_df['freq'].describe(percentiles=PERCENTILES)\nelse:    \n    users_freq_percentiles_df = None\n    \nset_distr_percentiles(users_freq_percentiles_df, dataset_info, prefix='user_freq')\ndisplay(users_freq_percentiles_df)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:35.247458Z","iopub.execute_input":"2022-02-10T16:29:35.247939Z","iopub.status.idle":"2022-02-10T16:29:35.333838Z","shell.execute_reply.started":"2022-02-10T16:29:35.247885Z","shell.execute_reply":"2022-02-10T16:29:35.333044Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if HAS_USER:\n    ax = users_freq_df.groupby('freq').size().to_pandas().hist(bins=100, figsize=(15,8))\n    ax.set_title('Users Frequency histogram')","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:35.335301Z","iopub.execute_input":"2022-02-10T16:29:35.335559Z","iopub.status.idle":"2022-02-10T16:29:35.754399Z","shell.execute_reply.started":"2022-02-10T16:29:35.335522Z","shell.execute_reply":"2022-02-10T16:29:35.753698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### # User-Item repeated interaction distribution","metadata":{}},{"cell_type":"code","source":"%%time\nif HAS_USER:\n    items_per_user_df = ddf.groupby('UserId')['ItemId'].size().to_frame('count')\n    display(items_per_user_df)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:35.755804Z","iopub.execute_input":"2022-02-10T16:29:35.756074Z","iopub.status.idle":"2022-02-10T16:29:35.78693Z","shell.execute_reply.started":"2022-02-10T16:29:35.756039Z","shell.execute_reply":"2022-02-10T16:29:35.786147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_USER:\n    unique_items_per_user_df = ddf[['UserId', 'ItemId']].drop_duplicates().groupby('UserId').size().to_frame('nunique')\n    display(unique_items_per_user_df)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:35.788319Z","iopub.execute_input":"2022-02-10T16:29:35.78858Z","iopub.status.idle":"2022-02-10T16:29:35.819514Z","shell.execute_reply.started":"2022-02-10T16:29:35.788544Z","shell.execute_reply":"2022-02-10T16:29:35.818798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_USER:\n    items_per_user_joined_df = items_per_user_df.merge(unique_items_per_user_df)\n    items_per_user_joined_df['user_perc_repeated_interactions'] = (items_per_user_joined_df['count'] - items_per_user_joined_df['nunique']) / items_per_user_joined_df['count']\n    display(items_per_user_joined_df)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:35.820478Z","iopub.execute_input":"2022-02-10T16:29:35.823559Z","iopub.status.idle":"2022-02-10T16:29:36.340971Z","shell.execute_reply.started":"2022-02-10T16:29:35.82353Z","shell.execute_reply":"2022-02-10T16:29:36.340183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_USER:\n    repeat_user_item_interactions_percentiles_df = \\\n            items_per_user_joined_df['user_perc_repeated_interactions'].compute().describe(percentiles=PERCENTILES) #.compute()\nelse:\n    repeat_user_item_interactions_percentiles_df = None\nset_distr_percentiles(repeat_user_item_interactions_percentiles_df, dataset_info, prefix='user_perc_repeated_interactions', dtype='float')\ndisplay(repeat_user_item_interactions_percentiles_df)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:36.342516Z","iopub.execute_input":"2022-02-10T16:29:36.343033Z","iopub.status.idle":"2022-02-10T16:29:56.535421Z","shell.execute_reply.started":"2022-02-10T16:29:36.342992Z","shell.execute_reply":"2022-02-10T16:29:56.534129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Temporal aspects","metadata":{}},{"cell_type":"markdown","source":"### # Interactions per day, week, month","metadata":{}},{"cell_type":"code","source":"%%time\nif HAS_TIMESTAMP:\n    ddf.groupby('DateStr').size().compute().to_pandas().sort_index().plot.line(figsize=(15,8))","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:56.536585Z","iopub.execute_input":"2022-02-10T16:29:56.538059Z","iopub.status.idle":"2022-02-10T16:29:57.33898Z","shell.execute_reply.started":"2022-02-10T16:29:56.538017Z","shell.execute_reply":"2022-02-10T16:29:57.338254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_TIMESTAMP:\n    ddf.groupby('year-week').size().compute().to_pandas().sort_index().plot.line(figsize=(15,8))","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:57.340557Z","iopub.execute_input":"2022-02-10T16:29:57.341332Z","iopub.status.idle":"2022-02-10T16:29:58.14441Z","shell.execute_reply.started":"2022-02-10T16:29:57.341292Z","shell.execute_reply":"2022-02-10T16:29:58.143682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_TIMESTAMP:\n    ddf.groupby('year-month').size().compute().to_pandas().sort_index().plot.line(figsize=(15,8))","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:58.145743Z","iopub.execute_input":"2022-02-10T16:29:58.146401Z","iopub.status.idle":"2022-02-10T16:29:58.932443Z","shell.execute_reply.started":"2022-02-10T16:29:58.146359Z","shell.execute_reply":"2022-02-10T16:29:58.931741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Items lifetime: Items interactions decay over time","metadata":{}},{"cell_type":"code","source":"%%time\nif HAS_TIMESTAMP:\n    # Getting the first time each item was seen in the dataset\n    first_date_item_df = ddf.groupby('ItemId')[['Time', 'DateStr', 'year-week', 'year-month']].min()\n    first_date_item_df = first_date_item_df.compute().rename(\n                              {'Time': 'first_Time',\n                               'DateStr': 'first_DateStr',\n                               'year-week': 'first_year-week',\n                               'year-month': 'first_year-month'}, axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:58.93409Z","iopub.execute_input":"2022-02-10T16:29:58.934578Z","iopub.status.idle":"2022-02-10T16:29:59.598076Z","shell.execute_reply.started":"2022-02-10T16:29:58.934536Z","shell.execute_reply":"2022-02-10T16:29:59.597359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_TIMESTAMP:\n    #Computing for each interaction how many days since the item was first seen\n    item_interactions_time_df = ddf[['ItemId', 'Time']].merge(first_date_item_df, left_on='ItemId', right_index=True)#.compute()\n    item_interactions_time_df['item_elapsed_days_since_first_seen'] = (item_interactions_time_df['Time'] - item_interactions_time_df['first_Time']).dt.days\n    display(item_interactions_time_df.head())","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:59.599191Z","iopub.execute_input":"2022-02-10T16:29:59.59973Z","iopub.status.idle":"2022-02-10T16:29:59.721278Z","shell.execute_reply.started":"2022-02-10T16:29:59.59969Z","shell.execute_reply":"2022-02-10T16:29:59.720574Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_TIMESTAMP:\n    item_elapsed_days_since_available_describe_df = item_interactions_time_df['item_elapsed_days_since_first_seen'] \\\n                                    .compute().describe(percentiles=PERCENTILES)  #.compute()\n    set_distr_percentiles(item_elapsed_days_since_available_describe_df, dataset_info, prefix='item_interactions_by_age_days')\n    display(item_elapsed_days_since_available_describe_df)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:29:59.722359Z","iopub.execute_input":"2022-02-10T16:29:59.723004Z","iopub.status.idle":"2022-02-10T16:30:00.601659Z","shell.execute_reply.started":"2022-02-10T16:29:59.722967Z","shell.execute_reply":"2022-02-10T16:30:00.60101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if HAS_TIMESTAMP:\n    item_interactions_time_df.groupby('item_elapsed_days_since_first_seen').size().compute() \\\n                .to_pandas().sort_index().plot.bar(figsize=(100,15))","metadata":{"execution":{"iopub.status.busy":"2022-02-10T17:27:07.597135Z","iopub.execute_input":"2022-02-10T17:27:07.597389Z","iopub.status.idle":"2022-02-10T17:27:16.625596Z","shell.execute_reply.started":"2022-02-10T17:27:07.59736Z","shell.execute_reply":"2022-02-10T17:27:16.624771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Item Cold-start","metadata":{}},{"cell_type":"markdown","source":"#### How many new items every week / month","metadata":{}},{"cell_type":"code","source":"%%time\nif HAS_TIMESTAMP:\n    new_items_by_day = first_date_item_df.groupby('first_DateStr').size().to_pandas()\n    display(new_items_by_day)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:00.60289Z","iopub.execute_input":"2022-02-10T16:30:00.603177Z","iopub.status.idle":"2022-02-10T16:30:00.62149Z","shell.execute_reply.started":"2022-02-10T16:30:00.603143Z","shell.execute_reply":"2022-02-10T16:30:00.620811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_TIMESTAMP:\n    #Ignoring first day\n    dataset_info['new_items_by_day_p50'] = new_items_by_day[1:].median()\n    print(dataset_info['new_items_by_day_p50'])","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:00.622712Z","iopub.execute_input":"2022-02-10T16:30:00.623113Z","iopub.status.idle":"2022-02-10T16:30:00.62973Z","shell.execute_reply.started":"2022-02-10T16:30:00.623076Z","shell.execute_reply":"2022-02-10T16:30:00.62883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if HAS_TIMESTAMP:\n    display(new_items_by_day[1:].sort_index().plot.bar(figsize=(100,10)))","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:00.631959Z","iopub.execute_input":"2022-02-10T16:30:00.632272Z","iopub.status.idle":"2022-02-10T16:30:11.014591Z","shell.execute_reply.started":"2022-02-10T16:30:00.632239Z","shell.execute_reply":"2022-02-10T16:30:11.01396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_TIMESTAMP:\n    new_items_by_week = first_date_item_df.groupby('first_year-week').size().to_pandas()\n    display(new_items_by_week)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:11.018525Z","iopub.execute_input":"2022-02-10T16:30:11.020454Z","iopub.status.idle":"2022-02-10T16:30:11.043737Z","shell.execute_reply.started":"2022-02-10T16:30:11.020389Z","shell.execute_reply":"2022-02-10T16:30:11.04295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if HAS_TIMESTAMP:\n    dataset_info['new_items_by_week_p50'] = new_items_by_week[1:].median()\n    print(dataset_info['new_items_by_week_p50'])","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:11.045613Z","iopub.execute_input":"2022-02-10T16:30:11.046212Z","iopub.status.idle":"2022-02-10T16:30:11.052023Z","shell.execute_reply.started":"2022-02-10T16:30:11.046173Z","shell.execute_reply":"2022-02-10T16:30:11.051138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if HAS_TIMESTAMP:\n    if len(new_items_by_week[1:]) > 0:\n        new_items_by_week[1:].sort_index().plot.bar(figsize=(20,8))","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:11.053614Z","iopub.execute_input":"2022-02-10T16:30:11.053986Z","iopub.status.idle":"2022-02-10T16:30:12.671293Z","shell.execute_reply.started":"2022-02-10T16:30:11.05395Z","shell.execute_reply":"2022-02-10T16:30:12.670584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_TIMESTAMP:\n    new_items_by_month = first_date_item_df.groupby('first_year-month').size().to_pandas()\n    display(new_items_by_month)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:12.672423Z","iopub.execute_input":"2022-02-10T16:30:12.672813Z","iopub.status.idle":"2022-02-10T16:30:12.695742Z","shell.execute_reply.started":"2022-02-10T16:30:12.672777Z","shell.execute_reply":"2022-02-10T16:30:12.695038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_TIMESTAMP:\n    #Ignoring first month\n    dataset_info['new_items_by_month_p50'] = new_items_by_month[1:].median()\n    print(dataset_info['new_items_by_month_p50'])","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:12.696797Z","iopub.execute_input":"2022-02-10T16:30:12.700264Z","iopub.status.idle":"2022-02-10T16:30:12.708884Z","shell.execute_reply.started":"2022-02-10T16:30:12.700186Z","shell.execute_reply":"2022-02-10T16:30:12.708152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_TIMESTAMP:\n    if len(new_items_by_month[1:]) > 0:\n        new_items_by_month[1:].sort_index().plot.bar(figsize=(15,8))","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:12.710288Z","iopub.execute_input":"2022-02-10T16:30:12.711Z","iopub.status.idle":"2022-02-10T16:30:13.031738Z","shell.execute_reply.started":"2022-02-10T16:30:12.710965Z","shell.execute_reply":"2022-02-10T16:30:13.031104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### % of interactions on new items first seen in the same day, week, month","metadata":{}},{"cell_type":"code","source":"%%time\nif HAS_TIMESTAMP:\n    item_cold_start_df = ddf.merge(first_date_item_df,  left_on='ItemId', right_index=True)[['DateStr', 'first_DateStr',\n                                                                        'year-week', 'first_year-week',\n                                                                        'year-month', 'first_year-month']] #.compute()\n    min_year_week = item_cold_start_df['year-week'].compute().min()\n    #Ignoring first week (where most of the items will occur first)\n    item_cold_start_df = item_cold_start_df[item_cold_start_df['year-week'] != min_year_week]\n    #Checking if the item was created in the same day, week or month of the interaction\n    item_cold_start_df['item_created_same_day'] = (item_cold_start_df['DateStr'] == item_cold_start_df['first_DateStr'])\n    item_cold_start_df['item_created_same_week'] = (item_cold_start_df['year-week'] == item_cold_start_df['first_year-week'])\n    item_cold_start_df['item_created_same_month'] = (item_cold_start_df['year-month'] == item_cold_start_df['first_year-month'])","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:13.032893Z","iopub.execute_input":"2022-02-10T16:30:13.033149Z","iopub.status.idle":"2022-02-10T16:30:14.064768Z","shell.execute_reply.started":"2022-02-10T16:30:13.033114Z","shell.execute_reply":"2022-02-10T16:30:14.064033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_TIMESTAMP:\n    perc_interactions_on_items_created_same_day = item_cold_start_df['item_created_same_day'].mean().compute()\n    dataset_info['perc_interact_items_created_same_day'] = perc_interactions_on_items_created_same_day\n    print(perc_interactions_on_items_created_same_day)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:14.066086Z","iopub.execute_input":"2022-02-10T16:30:14.06649Z","iopub.status.idle":"2022-02-10T16:30:15.310398Z","shell.execute_reply.started":"2022-02-10T16:30:14.066453Z","shell.execute_reply":"2022-02-10T16:30:15.308799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_TIMESTAMP:\n    perc_interactions_on_items_created_same_week = item_cold_start_df['item_created_same_week'].mean().compute()\n    dataset_info['perc_interact_items_created_same_week'] = perc_interactions_on_items_created_same_week\n    print(perc_interactions_on_items_created_same_week)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:15.31165Z","iopub.execute_input":"2022-02-10T16:30:15.311896Z","iopub.status.idle":"2022-02-10T16:30:16.560144Z","shell.execute_reply.started":"2022-02-10T16:30:15.311861Z","shell.execute_reply":"2022-02-10T16:30:16.559426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_TIMESTAMP:\n    perc_interactions_on_items_created_same_month = item_cold_start_df['item_created_same_month'].mean().compute()\n    dataset_info['perc_interact_items_created_same_month'] = perc_interactions_on_items_created_same_month\n    print(perc_interactions_on_items_created_same_month)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:16.561513Z","iopub.execute_input":"2022-02-10T16:30:16.561813Z","iopub.status.idle":"2022-02-10T16:30:17.802465Z","shell.execute_reply.started":"2022-02-10T16:30:16.561778Z","shell.execute_reply":"2022-02-10T16:30:17.801617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Users lifetime: User interactions decay over time","metadata":{}},{"cell_type":"code","source":"%%time\nif HAS_USER and HAS_TIMESTAMP:\n    first_date_user_df = ddf.groupby('UserId')[['Time', 'DateStr', 'year-week', 'year-month']].min()\n    first_date_user_df = first_date_user_df.compute().rename(\n                              {'Time': 'first_Time',\n                               'DateStr': 'first_DateStr',\n                               'year-week': 'first_year-week',\n                               'year-month': 'first_year-month'}, axis=1)\n    display(first_date_user_df)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:17.803708Z","iopub.execute_input":"2022-02-10T16:30:17.804406Z","iopub.status.idle":"2022-02-10T16:30:18.82077Z","shell.execute_reply.started":"2022-02-10T16:30:17.804361Z","shell.execute_reply":"2022-02-10T16:30:18.820099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_USER and  HAS_TIMESTAMP:\n    user_interactions_time_df = ddf[['UserId', 'Time']].merge(first_date_user_df, left_on='UserId', right_index=True) #.compute()\n    user_interactions_time_df['user_elapsed_days_since_first_seen'] = (user_interactions_time_df['Time'] - user_interactions_time_df['first_Time']).dt.days","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:18.822416Z","iopub.execute_input":"2022-02-10T16:30:18.822955Z","iopub.status.idle":"2022-02-10T16:30:18.931192Z","shell.execute_reply.started":"2022-02-10T16:30:18.822916Z","shell.execute_reply":"2022-02-10T16:30:18.930428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_USER and HAS_TIMESTAMP:\n    user_elapsed_days_since_first_seen_describe_df = \\\n            user_interactions_time_df['user_elapsed_days_since_first_seen'].compute().describe(percentiles=PERCENTILES) #.compute()\nelse:\n    user_elapsed_days_since_first_seen_describe_df = None\n    \nset_distr_percentiles(user_elapsed_days_since_first_seen_describe_df, dataset_info, prefix='user_interactions_by_age_days')\ndisplay(user_elapsed_days_since_first_seen_describe_df)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:18.932387Z","iopub.execute_input":"2022-02-10T16:30:18.93316Z","iopub.status.idle":"2022-02-10T16:30:20.321056Z","shell.execute_reply.started":"2022-02-10T16:30:18.933107Z","shell.execute_reply":"2022-02-10T16:30:20.32019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_USER and HAS_TIMESTAMP:\n    user_interactions_time_df.groupby('user_elapsed_days_since_first_seen').size().compute().to_pandas().sort_index().plot.bar(figsize=(100,8))#","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:20.322597Z","iopub.execute_input":"2022-02-10T16:30:20.322879Z","iopub.status.idle":"2022-02-10T16:30:29.602712Z","shell.execute_reply.started":"2022-02-10T16:30:20.322842Z","shell.execute_reply":"2022-02-10T16:30:29.602034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### User Cold-start","metadata":{}},{"cell_type":"markdown","source":"#### How many new users every week / month","metadata":{}},{"cell_type":"code","source":"%%time\nif HAS_USER and HAS_TIMESTAMP:\n    new_users_by_day = first_date_user_df.groupby('first_DateStr').size()\n    display(new_users_by_day)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:29.603782Z","iopub.execute_input":"2022-02-10T16:30:29.604179Z","iopub.status.idle":"2022-02-10T16:30:29.626877Z","shell.execute_reply.started":"2022-02-10T16:30:29.604137Z","shell.execute_reply":"2022-02-10T16:30:29.626089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_USER and HAS_TIMESTAMP:\n    #Ignoring first day\n    dataset_info['new_users_by_day_p50'] = new_users_by_day[1:].median()\n    print(dataset_info['new_users_by_day_p50'])\nelse:\n    dataset_info['new_users_by_day_p50'] = None","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:29.628247Z","iopub.execute_input":"2022-02-10T16:30:29.628496Z","iopub.status.idle":"2022-02-10T16:30:29.636566Z","shell.execute_reply.started":"2022-02-10T16:30:29.628462Z","shell.execute_reply":"2022-02-10T16:30:29.63573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_USER and HAS_TIMESTAMP:\n    new_users_by_week = first_date_user_df.groupby('first_year-week').size().to_pandas()\n    display(new_users_by_week)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:29.638174Z","iopub.execute_input":"2022-02-10T16:30:29.638692Z","iopub.status.idle":"2022-02-10T16:30:29.657458Z","shell.execute_reply.started":"2022-02-10T16:30:29.638654Z","shell.execute_reply":"2022-02-10T16:30:29.65679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_USER and HAS_TIMESTAMP:\n    #Ignoring first week\n    dataset_info['new_users_by_week_p50'] = new_users_by_week[1:].median()\n    print(dataset_info['new_users_by_week_p50'])\nelse:\n    dataset_info['new_users_by_week_p50'] = None","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:29.658575Z","iopub.execute_input":"2022-02-10T16:30:29.659281Z","iopub.status.idle":"2022-02-10T16:30:29.666507Z","shell.execute_reply.started":"2022-02-10T16:30:29.659246Z","shell.execute_reply":"2022-02-10T16:30:29.665739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_USER and HAS_TIMESTAMP:\n    if len(new_users_by_week[1:]) > 0:\n        new_users_by_week[1:].sort_index().plot.bar(figsize=(20,8))","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:29.667805Z","iopub.execute_input":"2022-02-10T16:30:29.668368Z","iopub.status.idle":"2022-02-10T16:30:31.074514Z","shell.execute_reply.started":"2022-02-10T16:30:29.668331Z","shell.execute_reply":"2022-02-10T16:30:31.073769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_USER and HAS_TIMESTAMP:\n    new_users_by_month = first_date_user_df.groupby('first_year-month').size().to_pandas()\n    display(new_users_by_month)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:31.076217Z","iopub.execute_input":"2022-02-10T16:30:31.076702Z","iopub.status.idle":"2022-02-10T16:30:31.09828Z","shell.execute_reply.started":"2022-02-10T16:30:31.076662Z","shell.execute_reply":"2022-02-10T16:30:31.097575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if HAS_USER and HAS_TIMESTAMP:\n    #Ignoring first month\n    dataset_info['new_users_by_month_p50'] = new_users_by_month[1:].median()\n    print(dataset_info['new_users_by_month_p50'])\nelse:\n    dataset_info['new_users_by_month_p50'] = None","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:31.09938Z","iopub.execute_input":"2022-02-10T16:30:31.100069Z","iopub.status.idle":"2022-02-10T16:30:31.106194Z","shell.execute_reply.started":"2022-02-10T16:30:31.100031Z","shell.execute_reply":"2022-02-10T16:30:31.105306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if HAS_USER and HAS_TIMESTAMP:\n    if len(new_users_by_month[1:]) > 0:\n        new_users_by_month[1:].sort_index().plot.bar(figsize=(15,8))","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:31.107732Z","iopub.execute_input":"2022-02-10T16:30:31.108188Z","iopub.status.idle":"2022-02-10T16:30:31.429418Z","shell.execute_reply.started":"2022-02-10T16:30:31.10815Z","shell.execute_reply":"2022-02-10T16:30:31.428745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### % of interactions on new items first seen in the same day, week, month","metadata":{}},{"cell_type":"code","source":"%%time\nif HAS_USER and HAS_TIMESTAMP:\n    user_cold_start_df = ddf.merge(first_date_user_df,  left_on='UserId', right_index=True)[['DateStr', 'first_DateStr',\n                                                                        'year-week', 'first_year-week',\n                                                                        'year-month', 'first_year-month']] #.compute()\n    min_year_week = user_cold_start_df['year-week'].compute().min()\n    #Ignoring first week (where most of the items will occur first)\n    user_cold_start_df = user_cold_start_df[user_cold_start_df['year-week'] != min_year_week]\n    #Checking if the item was created in the same day, week or month of the interaction\n    user_cold_start_df['user_created_same_day'] = (user_cold_start_df['DateStr'] == user_cold_start_df['first_DateStr'])\n    user_cold_start_df['user_created_same_week'] = (user_cold_start_df['year-week'] == user_cold_start_df['first_year-week'])\n    user_cold_start_df['user_created_same_month'] = (user_cold_start_df['year-month'] == user_cold_start_df['first_year-month'])","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:31.430514Z","iopub.execute_input":"2022-02-10T16:30:31.430785Z","iopub.status.idle":"2022-02-10T16:30:32.931144Z","shell.execute_reply.started":"2022-02-10T16:30:31.430748Z","shell.execute_reply":"2022-02-10T16:30:32.930283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_USER and HAS_TIMESTAMP:\n    perc_interactions_by_users_first_seen_same_day = user_cold_start_df['user_created_same_day'].mean().compute()\n    dataset_info['perc_interact_users_created_same_day'] = perc_interactions_by_users_first_seen_same_day\n    print(perc_interactions_by_users_first_seen_same_day)\nelse:\n    dataset_info['perc_interact_users_created_same_day'] = None","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:32.932702Z","iopub.execute_input":"2022-02-10T16:30:32.932974Z","iopub.status.idle":"2022-02-10T16:30:34.60076Z","shell.execute_reply.started":"2022-02-10T16:30:32.932938Z","shell.execute_reply":"2022-02-10T16:30:34.600031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_USER and HAS_TIMESTAMP:\n    perc_interactions_by_users_first_seen_same_week = user_cold_start_df['user_created_same_week'].mean().compute()\n    dataset_info['perc_interact_users_created_same_week'] = perc_interactions_by_users_first_seen_same_week\n    print(perc_interactions_by_users_first_seen_same_week)\nelse:\n    dataset_info['perc_interact_users_created_same_week'] = None","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:34.602051Z","iopub.execute_input":"2022-02-10T16:30:34.602369Z","iopub.status.idle":"2022-02-10T16:30:36.25137Z","shell.execute_reply.started":"2022-02-10T16:30:34.602331Z","shell.execute_reply":"2022-02-10T16:30:36.250599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nif HAS_USER and HAS_TIMESTAMP:\n    perc_interactions_by_users_first_seen_same_month = user_cold_start_df['user_created_same_month'].mean().compute()\n    dataset_info['perc_interact_users_created_same_month'] = perc_interactions_by_users_first_seen_same_month\n    print(perc_interactions_by_users_first_seen_same_month)\nelse:\n    dataset_info['perc_interact_users_created_same_month'] = None","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:30:36.252697Z","iopub.execute_input":"2022-02-10T16:30:36.253063Z","iopub.status.idle":"2022-02-10T16:30:37.892391Z","shell.execute_reply.started":"2022-02-10T16:30:36.253019Z","shell.execute_reply":"2022-02-10T16:30:37.891435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exporting","metadata":{}},{"cell_type":"markdown","source":"In this last section we generate and export two CSV files under `\\kaggle\\working` folder:  \n- `H&B_fashion_dataset_profile.csv` - Main statistics collected for this dataset\n- `H&B_fashion_dataset_user_item_cum_freq_distr.csv\"` - The cumulative distributions of interactions per users and per items","metadata":{"execution":{"iopub.status.busy":"2022-02-10T17:12:13.726423Z","iopub.execute_input":"2022-02-10T17:12:13.726683Z","iopub.status.idle":"2022-02-10T17:12:13.732926Z","shell.execute_reply.started":"2022-02-10T17:12:13.726654Z","shell.execute_reply":"2022-02-10T17:12:13.731859Z"}}},{"cell_type":"markdown","source":"## Exporting dataset info profile","metadata":{}},{"cell_type":"code","source":"dataset_info_df = pd.DataFrame.from_dict(dataset_info, orient='index').T\ndataset_info_df","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:57:11.91863Z","iopub.execute_input":"2022-02-10T16:57:11.919434Z","iopub.status.idle":"2022-02-10T16:57:11.942996Z","shell.execute_reply.started":"2022-02-10T16:57:11.919393Z","shell.execute_reply":"2022-02-10T16:57:11.942282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'Creating output path if it does not exists: {OUTPUT_PATH}')\nos.makedirs(OUTPUT_PATH, exist_ok=True)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:57:13.747738Z","iopub.execute_input":"2022-02-10T16:57:13.748582Z","iopub.status.idle":"2022-02-10T16:57:13.755581Z","shell.execute_reply.started":"2022-02-10T16:57:13.748529Z","shell.execute_reply":"2022-02-10T16:57:13.753838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if SAVE_DATASET_INFO_CSV:\n    output_path = os.path.join(OUTPUT_PATH, f\"{dataset_info['name']}_dataset_profile.csv\")\n    print(f'Saving the dataset info to: {output_path}')\n    dataset_info_df.to_csv(output_path, index=False)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:57:15.232775Z","iopub.execute_input":"2022-02-10T16:57:15.233215Z","iopub.status.idle":"2022-02-10T16:57:15.241868Z","shell.execute_reply.started":"2022-02-10T16:57:15.233182Z","shell.execute_reply":"2022-02-10T16:57:15.240782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Exporting items and users freq. cumulative distributions","metadata":{}},{"cell_type":"code","source":"if HAS_USER:\n    user_item_cum_freq_distr_df = items_freq_cum_perc_reindexed_pdf.merge(users_freq_cum_perc_reindexed_pdf, \n                                                                          left_index=True, right_index=True)\nelse:\n    user_item_cum_freq_distr_df = items_freq_cum_perc_reindexed_pdf\n    user_item_cum_freq_distr_df['cum_interactions_by_user_freq'] = 0.0\nuser_item_cum_freq_distr_df","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:57:20.980978Z","iopub.execute_input":"2022-02-10T16:57:20.981507Z","iopub.status.idle":"2022-02-10T16:57:20.999256Z","shell.execute_reply.started":"2022-02-10T16:57:20.981467Z","shell.execute_reply":"2022-02-10T16:57:20.998411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if SAVE_USERS_ITEMS_CUM_FREQ_DISTR:\n    output_path = os.path.join(OUTPUT_PATH, f\"{dataset_info['name']}_dataset_user_item_cum_freq_distr.csv\")\n    print(f'Saving the users and items cumulative freq. distribution to: {output_path}')\n    user_item_cum_freq_distr_df.to_csv(output_path)","metadata":{"execution":{"iopub.status.busy":"2022-02-10T16:57:50.457275Z","iopub.execute_input":"2022-02-10T16:57:50.457524Z","iopub.status.idle":"2022-02-10T16:57:50.466416Z","shell.execute_reply.started":"2022-02-10T16:57:50.457495Z","shell.execute_reply":"2022-02-10T16:57:50.465574Z"},"trusted":true},"execution_count":null,"outputs":[]}]}