{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## SUPER FAST FEAT GEN (100x)\nThis notebook is a revamped version of [OTTO Generate Features (For Training)](https://www.kaggle.com/code/seholee/otto-generate-features-for-training).\nI mixed polars and pandas to speed up the process 100x times faster. I also shared the datasets generated as **public**.\nThe whole thing runs within 20 minutes on 16 core computer on gcp.\nSeems to run even faster with the 96 core computer with TPU machine.\n\n\nI didn't work on it yet, but you can improve this notebook to generate data on the 30gb ram machine as well. That works by not joining the columns generated to the large dataset, but using a separate result dataset to keep join only the result columns with generated feature. Although you can do so, you probably will need over 100 gb to train the entire thing anyways(since tree models need the entire dataset to be on the memory.).","metadata":{}},{"cell_type":"code","source":"!pip install gcloud polars kaggle pyarrow fastparquet","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:25:37.382096Z","iopub.execute_input":"2023-01-21T11:25:37.382865Z","iopub.status.idle":"2023-01-21T11:25:44.611189Z","shell.execute_reply.started":"2023-01-21T11:25:37.382832Z","shell.execute_reply":"2023-01-21T11:25:44.610134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# !kaggle datasets download radek1/otto-train-and-test-data-for-local-validation","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# !kaggle datasets download radek1/otto-full-optimized-memory-footprint","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# !unzip otto-full-optimized-memory-footprint.zip","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# !mkdir otto-train-and-test-data-for-local-validation\n# !unzip otto-train-and-test-data-for-local-validation.zip otto-train-and-test-data-for-local-validation","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom datetime import datetime\nimport matplotlib.pyplot as plt\n# from google.cloud import storage\nimport polars as pl\nimport gc\n\n# multiprocessing \nimport psutil\nN_CORES = psutil.cpu_count()     # Available CPU cores\nprint(f\"N Cores : {N_CORES}\")\nfrom multiprocessing import Pool","metadata":{"tags":[],"execution":{"iopub.status.busy":"2023-01-21T11:25:59.187189Z","iopub.execute_input":"2023-01-21T11:25:59.188053Z","iopub.status.idle":"2023-01-21T11:25:59.490289Z","shell.execute_reply.started":"2023-01-21T11:25:59.188011Z","shell.execute_reply":"2023-01-21T11:25:59.489420Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## General Approach\n- use test.parquet for session data\n- use train.parquet + test.parquet for aid data\n- merge them for final training data\n- you can split test.parquet in two pieces for training and validation so that you don't have to submit every time.","metadata":{}},{"cell_type":"markdown","source":"## Prepare Session data","metadata":{}},{"cell_type":"code","source":"df = pl.read_parquet('/kaggle/input/otto-train-and-test-data-for-local-validation/test.parquet')","metadata":{"tags":[],"execution":{"iopub.status.busy":"2023-01-21T11:26:12.078441Z","iopub.execute_input":"2023-01-21T11:26:12.078790Z","iopub.status.idle":"2023-01-21T11:26:12.608110Z","shell.execute_reply.started":"2023-01-21T11:26:12.078763Z","shell.execute_reply":"2023-01-21T11:26:12.607294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(datetime.fromtimestamp(df['ts'].min()))\nprint(datetime.fromtimestamp(df['ts'].max()))","metadata":{"tags":[],"execution":{"iopub.status.busy":"2023-01-21T11:26:13.989860Z","iopub.execute_input":"2023-01-21T11:26:13.990419Z","iopub.status.idle":"2023-01-21T11:26:14.002599Z","shell.execute_reply.started":"2023-01-21T11:26:13.990375Z","shell.execute_reply":"2023-01-21T11:26:14.001647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df","metadata":{"tags":[],"execution":{"iopub.status.busy":"2023-01-21T11:26:14.857548Z","iopub.execute_input":"2023-01-21T11:26:14.858318Z","iopub.status.idle":"2023-01-21T11:26:14.870202Z","shell.execute_reply.started":"2023-01-21T11:26:14.858282Z","shell.execute_reply":"2023-01-21T11:26:14.869277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Session Feature - session count","metadata":{}},{"cell_type":"code","source":"counts = df.groupby('session').count()\ncounts.columns = ['session', 'sess_cnt']","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:26:17.200499Z","iopub.execute_input":"2023-01-21T11:26:17.201303Z","iopub.status.idle":"2023-01-21T11:26:17.971608Z","shell.execute_reply.started":"2023-01-21T11:26:17.201271Z","shell.execute_reply":"2023-01-21T11:26:17.970714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df.join(counts, on='session', how='left')","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:26:17.973327Z","iopub.execute_input":"2023-01-21T11:26:17.973649Z","iopub.status.idle":"2023-01-21T11:26:18.272523Z","shell.execute_reply.started":"2023-01-21T11:26:17.973620Z","shell.execute_reply":"2023-01-21T11:26:18.271598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.hist(df['sess_cnt'], bins=200)\nplt.show()","metadata":{"tags":[],"execution":{"iopub.status.busy":"2023-01-21T11:26:18.273744Z","iopub.execute_input":"2023-01-21T11:26:18.274026Z","iopub.status.idle":"2023-01-21T11:26:18.915871Z","shell.execute_reply.started":"2023-01-21T11:26:18.274000Z","shell.execute_reply":"2023-01-21T11:26:18.915109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Session Feature - session average time of the day + std of that","metadata":{}},{"cell_type":"code","source":"%%time\ndts = pd.to_datetime(df['ts'].to_pandas(), unit='s')\n\ndf = df.with_column(pl.from_pandas(dts.dt.weekday.rename('day')))\ndf = df.with_column(pl.from_pandas(dts.dt.hour.rename('hour')))\ndf = df.with_column(pl.from_pandas((dts.dt.hour*100 + dts.dt.minute*100//60).rename('hm')))\ndf","metadata":{"tags":[],"execution":{"iopub.status.busy":"2023-01-21T11:26:19.471736Z","iopub.execute_input":"2023-01-21T11:26:19.472490Z","iopub.status.idle":"2023-01-21T11:26:22.515571Z","shell.execute_reply.started":"2023-01-21T11:26:19.472444Z","shell.execute_reply":"2023-01-21T11:26:22.514697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## mean\nhm_mean = df.groupby('session').mean().select(['session', 'hm'])\nhm_mean.columns = ['session', 'hm_mean']\ndf = df.join(hm_mean, on='session', how='left')","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:26:22.517028Z","iopub.execute_input":"2023-01-21T11:26:22.517338Z","iopub.status.idle":"2023-01-21T11:26:26.048659Z","shell.execute_reply.started":"2023-01-21T11:26:22.517311Z","shell.execute_reply":"2023-01-21T11:26:26.047793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## median\nhm_med = df.groupby('session').median().select(['session', 'hm'])\nhm_med.columns = ['session', 'hm_median']\ndf = df.join(hm_med, on='session', how='left')","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:26:26.049925Z","iopub.execute_input":"2023-01-21T11:26:26.050201Z","iopub.status.idle":"2023-01-21T11:26:27.204249Z","shell.execute_reply.started":"2023-01-21T11:26:26.050177Z","shell.execute_reply":"2023-01-21T11:26:27.203112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n## std\ntemp = df.to_pandas()\ntemp['hm_std'] = temp.groupby('session')['hm'].transform('std')\ntemp['hm_std'] = temp['hm_std'].fillna(0)\ndf = df.with_column(pl.from_pandas(temp['hm_std']))","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:26:27.206474Z","iopub.execute_input":"2023-01-21T11:26:27.206801Z","iopub.status.idle":"2023-01-21T11:26:27.908687Z","shell.execute_reply.started":"2023-01-21T11:26:27.206773Z","shell.execute_reply":"2023-01-21T11:26:27.907559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:26:27.909836Z","iopub.execute_input":"2023-01-21T11:26:27.910124Z","iopub.status.idle":"2023-01-21T11:26:27.922144Z","shell.execute_reply.started":"2023-01-21T11:26:27.910099Z","shell.execute_reply":"2023-01-21T11:26:27.921237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Session Features - weekday ratio + hour ratio","metadata":{}},{"cell_type":"code","source":"%%time\nfor i in range(7):\n    df_temp_cnt = df.filter(pl.col('day') == i).groupby('session').count()\n    df_temp_cnt.columns = ['session', f'day{i}cnt']\n    df = df.join(df_temp_cnt, on='session', how='left')\n    df = df.with_column(pl.col(f'day{i}cnt').fill_null(0))\n    df = df.with_column(df[f'day{i}cnt']/df['sess_cnt'])\n    del df_temp_cnt\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:26:27.923266Z","iopub.execute_input":"2023-01-21T11:26:27.923560Z","iopub.status.idle":"2023-01-21T11:26:32.396328Z","shell.execute_reply.started":"2023-01-21T11:26:27.923534Z","shell.execute_reply":"2023-01-21T11:26:32.395362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:26:32.397642Z","iopub.execute_input":"2023-01-21T11:26:32.398199Z","iopub.status.idle":"2023-01-21T11:26:32.405903Z","shell.execute_reply.started":"2023-01-21T11:26:32.398170Z","shell.execute_reply":"2023-01-21T11:26:32.405121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfor i in range(24):\n    df_temp_cnt = df.filter(pl.col('hour') == i).groupby('session').count()\n    df_temp_cnt.columns = ['session', f'hour{i}cnt']\n    df = df.join(df_temp_cnt, on='session', how='left')\n    df = df.with_column(pl.col(f'hour{i}cnt').fill_null(0))\n    df = df.with_column(df[f'hour{i}cnt']/df['sess_cnt'])\n    del df_temp_cnt\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:26:32.406986Z","iopub.execute_input":"2023-01-21T11:26:32.407287Z","iopub.status.idle":"2023-01-21T11:26:45.533182Z","shell.execute_reply.started":"2023-01-21T11:26:32.407262Z","shell.execute_reply":"2023-01-21T11:26:45.532327Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:26:45.534369Z","iopub.execute_input":"2023-01-21T11:26:45.534695Z","iopub.status.idle":"2023-01-21T11:26:45.545356Z","shell.execute_reply.started":"2023-01-21T11:26:45.534666Z","shell.execute_reply":"2023-01-21T11:26:45.544515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Session Feature - number of group of consecutive events within 1600 seconds (actual session count)","metadata":{}},{"cell_type":"code","source":"df = df.with_column(df['ts'].diff().rename('ts_diff'))\ndf_pd = df.to_pandas()\ndf_pd['index1'] = df_pd.index","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:26:45.548696Z","iopub.execute_input":"2023-01-21T11:26:45.549284Z","iopub.status.idle":"2023-01-21T11:26:45.809010Z","shell.execute_reply.started":"2023-01-21T11:26:45.549256Z","shell.execute_reply":"2023-01-21T11:26:45.807910Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = df_pd.groupby('session').first()\ntemp['flag'] = True\ntemp = temp[['flag', 'index1']]","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:26:45.810210Z","iopub.execute_input":"2023-01-21T11:26:45.810496Z","iopub.status.idle":"2023-01-21T11:26:50.377638Z","shell.execute_reply.started":"2023-01-21T11:26:45.810471Z","shell.execute_reply":"2023-01-21T11:26:50.376517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_pd = df_pd.merge(temp, how='left', left_on='index1', right_on='index1')","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:26:50.378890Z","iopub.execute_input":"2023-01-21T11:26:50.379251Z","iopub.status.idle":"2023-01-21T11:26:55.296850Z","shell.execute_reply.started":"2023-01-21T11:26:50.379221Z","shell.execute_reply":"2023-01-21T11:26:55.295688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_pd.drop(['index1'], inplace=True, axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:26:55.298264Z","iopub.execute_input":"2023-01-21T11:26:55.298634Z","iopub.status.idle":"2023-01-21T11:26:56.503095Z","shell.execute_reply.started":"2023-01-21T11:26:55.298602Z","shell.execute_reply":"2023-01-21T11:26:56.502098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf_pd.loc[df_pd['flag'] == True, 'ts_diff'] = -1\ndf_pd.drop(['flag'], inplace=True, axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:26:56.504381Z","iopub.execute_input":"2023-01-21T11:26:56.504687Z","iopub.status.idle":"2023-01-21T11:26:57.997330Z","shell.execute_reply.started":"2023-01-21T11:26:56.504660Z","shell.execute_reply":"2023-01-21T11:26:57.996371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_pd","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:26:57.998493Z","iopub.execute_input":"2023-01-21T11:26:57.998799Z","iopub.status.idle":"2023-01-21T11:26:59.668624Z","shell.execute_reply.started":"2023-01-21T11:26:57.998773Z","shell.execute_reply":"2023-01-21T11:26:59.667753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get sess_cnt2\ndf = pl.from_pandas(df_pd)\nsess_cnt2 = df.filter(pl.col('ts_diff') > 1600).groupby('session').count()\nsess_cnt2.columns = ['session', 'sess_cnt2']\nsess_cnt2 = sess_cnt2.with_column(sess_cnt2['sess_cnt2'] + 1)  # 원래 세션 갯수는 1개니까 1 더해주기.\ndf = df.join(sess_cnt2, on='session', how='left')\ndf = df.with_column(pl.col('sess_cnt2').fill_null(1))","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:26:59.669818Z","iopub.execute_input":"2023-01-21T11:26:59.670159Z","iopub.status.idle":"2023-01-21T11:27:03.844025Z","shell.execute_reply.started":"2023-01-21T11:26:59.670129Z","shell.execute_reply":"2023-01-21T11:27:03.842908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.hist(df['sess_cnt2'], bins=200)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:27:03.845247Z","iopub.execute_input":"2023-01-21T11:27:03.845580Z","iopub.status.idle":"2023-01-21T11:27:04.380905Z","shell.execute_reply.started":"2023-01-21T11:27:03.845549Z","shell.execute_reply":"2023-01-21T11:27:04.380003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Session Feature - session buys, carts, clicks cnt + ratio","metadata":{}},{"cell_type":"code","source":"cl_cnt = df.filter(pl.col('type')==0).groupby('session').count()\ncl_cnt.columns = ['session', 'cl_cnt']\nca_cnt = df.filter(pl.col('type')==1).groupby('session').count()\nca_cnt.columns = ['session', 'ca_cnt'] \nor_cnt = df.filter(pl.col('type')==2).groupby('session').count()\nor_cnt.columns = ['session', 'or_cnt'] ","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:27:04.382177Z","iopub.execute_input":"2023-01-21T11:27:04.382523Z","iopub.status.idle":"2023-01-21T11:27:05.544448Z","shell.execute_reply.started":"2023-01-21T11:27:04.382493Z","shell.execute_reply":"2023-01-21T11:27:05.543326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df.join(cl_cnt, on='session', how='left')\ndf = df.join(ca_cnt, on='session', how='left')\ndf = df.join(or_cnt, on='session', how='left')","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:27:05.545740Z","iopub.execute_input":"2023-01-21T11:27:05.546108Z","iopub.status.idle":"2023-01-21T11:27:06.188415Z","shell.execute_reply.started":"2023-01-21T11:27:05.546063Z","shell.execute_reply":"2023-01-21T11:27:06.187257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df.with_column(pl.col('cl_cnt').fill_null(0))\ndf = df.with_column(pl.col('ca_cnt').fill_null(0))\ndf = df.with_column(pl.col('or_cnt').fill_null(0))","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:27:06.189758Z","iopub.execute_input":"2023-01-21T11:27:06.190051Z","iopub.status.idle":"2023-01-21T11:27:06.550302Z","shell.execute_reply.started":"2023-01-21T11:27:06.190023Z","shell.execute_reply":"2023-01-21T11:27:06.549377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:27:06.551418Z","iopub.execute_input":"2023-01-21T11:27:06.551699Z","iopub.status.idle":"2023-01-21T11:27:06.578640Z","shell.execute_reply.started":"2023-01-21T11:27:06.551674Z","shell.execute_reply":"2023-01-21T11:27:06.577700Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# click to cart, click to order, cart to order ratio\ndf = df.with_column((df['ca_cnt']/df['cl_cnt']).alias('ca_cl_ratio'))\ndf = df.with_column((df['or_cnt']/df['cl_cnt']).alias('or_cl_ratio'))\ndf = df.with_column((df['or_cnt']/df['ca_cnt']).alias('or_ca_ratio'))","metadata":{"tags":[],"execution":{"iopub.status.busy":"2023-01-21T11:27:06.579748Z","iopub.execute_input":"2023-01-21T11:27:06.580035Z","iopub.status.idle":"2023-01-21T11:27:06.936499Z","shell.execute_reply.started":"2023-01-21T11:27:06.580009Z","shell.execute_reply":"2023-01-21T11:27:06.935309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df","metadata":{"tags":[],"execution":{"iopub.status.busy":"2023-01-21T11:27:06.938518Z","iopub.execute_input":"2023-01-21T11:27:06.938851Z","iopub.status.idle":"2023-01-21T11:27:06.967127Z","shell.execute_reply.started":"2023-01-21T11:27:06.938821Z","shell.execute_reply":"2023-01-21T11:27:06.966060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df.with_column(pl.col('ca_cl_ratio').fill_nan(-1))\ndf = df.with_column(pl.col('or_cl_ratio').fill_nan(-1))\ndf = df.with_column(pl.col('or_ca_ratio').fill_nan(-1))","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:27:06.968276Z","iopub.execute_input":"2023-01-21T11:27:06.968571Z","iopub.status.idle":"2023-01-21T11:27:07.289238Z","shell.execute_reply.started":"2023-01-21T11:27:06.968545Z","shell.execute_reply":"2023-01-21T11:27:07.288152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['or_cl_ratio'].is_infinite().sum()","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:27:07.290401Z","iopub.execute_input":"2023-01-21T11:27:07.290697Z","iopub.status.idle":"2023-01-21T11:27:07.304229Z","shell.execute_reply.started":"2023-01-21T11:27:07.290669Z","shell.execute_reply":"2023-01-21T11:27:07.303296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df.with_column(\n    pl.when(pl.col(['ca_cl_ratio', 'or_cl_ratio', 'or_ca_ratio']).is_infinite())\n    .then(0)\n    .otherwise(pl.col(['ca_cl_ratio', 'or_cl_ratio','or_ca_ratio']))\n    .keep_name()\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:27:07.305318Z","iopub.execute_input":"2023-01-21T11:27:07.305582Z","iopub.status.idle":"2023-01-21T11:27:07.393065Z","shell.execute_reply.started":"2023-01-21T11:27:07.305558Z","shell.execute_reply":"2023-01-21T11:27:07.391827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['or_cl_ratio'].is_infinite().sum()","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:27:07.394515Z","iopub.execute_input":"2023-01-21T11:27:07.394870Z","iopub.status.idle":"2023-01-21T11:27:07.411421Z","shell.execute_reply.started":"2023-01-21T11:27:07.394839Z","shell.execute_reply":"2023-01-21T11:27:07.410055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Session - minimum ts, average ts, maximum ts ( time passed since last visit, most active visit, first visit )","metadata":{}},{"cell_type":"code","source":"ss_ts_max = df.groupby('session').max()[['session', 'ts']]\nss_ts_max.columns = ['session', 'ts_max']\nss_ts_min = df.groupby('session').min()[['session', 'ts']]\nss_ts_min.columns = ['session', 'ts_min']\nss_ts_mean = df.groupby('session').mean()[['session', 'ts']]\nss_ts_mean.columns = ['session', 'ts_mean']","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:27:07.416521Z","iopub.execute_input":"2023-01-21T11:27:07.417279Z","iopub.status.idle":"2023-01-21T11:27:08.870210Z","shell.execute_reply.started":"2023-01-21T11:27:07.417247Z","shell.execute_reply":"2023-01-21T11:27:08.869060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df.join(ss_ts_max, on='session', how='left')\ndf = df.join(ss_ts_min, on='session', how='left')\ndf = df.join(ss_ts_mean, on='session', how='left')\ndf","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:27:08.871687Z","iopub.execute_input":"2023-01-21T11:27:08.872004Z","iopub.status.idle":"2023-01-21T11:27:09.647736Z","shell.execute_reply.started":"2023-01-21T11:27:08.871978Z","shell.execute_reply":"2023-01-21T11:27:09.646758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def check_df_ok(df):\n    for col in df.columns:\n        print(col, end=' ')\n        print(df[col].is_null().sum(), end=' ')\n        try:\n            print(df[col].is_nan().sum(), end=' ')\n        except:\n            print('exc', end=' ')\n        try:\n            print(df[col].is_infinite().sum(), end=' ')\n        except:\n            print('exc', end=' ')        \n        print('')\ncheck_df_ok(df)","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:27:09.648968Z","iopub.execute_input":"2023-01-21T11:27:09.649299Z","iopub.status.idle":"2023-01-21T11:27:10.438299Z","shell.execute_reply.started":"2023-01-21T11:27:09.649265Z","shell.execute_reply":"2023-01-21T11:27:10.437272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Cleanup","metadata":{}},{"cell_type":"code","source":"df = df.groupby('session').first().sort(pl.col('session'))","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:27:10.439424Z","iopub.execute_input":"2023-01-21T11:27:10.439720Z","iopub.status.idle":"2023-01-21T11:27:10.814738Z","shell.execute_reply.started":"2023-01-21T11:27:10.439695Z","shell.execute_reply":"2023-01-21T11:27:10.813731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df.drop(['aid', 'ts', 'type', 'day', 'hour', 'hm', 'ts_diff'])","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:27:10.815965Z","iopub.execute_input":"2023-01-21T11:27:10.816356Z","iopub.status.idle":"2023-01-21T11:27:10.824137Z","shell.execute_reply.started":"2023-01-21T11:27:10.816324Z","shell.execute_reply":"2023-01-21T11:27:10.823329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.write_parquet('sess_feature.parquet')","metadata":{"tags":[],"execution":{"iopub.status.busy":"2023-01-21T11:27:10.825138Z","iopub.execute_input":"2023-01-21T11:27:10.825467Z","iopub.status.idle":"2023-01-21T11:27:12.470253Z","shell.execute_reply.started":"2023-01-21T11:27:10.825442Z","shell.execute_reply":"2023-01-21T11:27:12.469239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Prepare Aid Data","metadata":{}},{"cell_type":"code","source":"df = pl.concat([pl.read_parquet('/kaggle/input/otto-train-and-test-data-for-local-validation/train.parquet'),\n                pl.read_parquet('/kaggle/input/otto-train-and-test-data-for-local-validation/test.parquet')])","metadata":{"tags":[],"execution":{"iopub.status.busy":"2023-01-21T11:27:44.767663Z","iopub.execute_input":"2023-01-21T11:27:44.768091Z","iopub.status.idle":"2023-01-21T11:28:04.865455Z","shell.execute_reply.started":"2023-01-21T11:27:44.768046Z","shell.execute_reply":"2023-01-21T11:28:04.864301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.shape","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:28:04.867100Z","iopub.execute_input":"2023-01-21T11:28:04.867403Z","iopub.status.idle":"2023-01-21T11:28:04.873577Z","shell.execute_reply.started":"2023-01-21T11:28:04.867377Z","shell.execute_reply":"2023-01-21T11:28:04.872808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:28:04.874570Z","iopub.execute_input":"2023-01-21T11:28:04.875186Z","iopub.status.idle":"2023-01-21T11:28:04.888786Z","shell.execute_reply.started":"2023-01-21T11:28:04.875158Z","shell.execute_reply":"2023-01-21T11:28:04.888087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Aid Features - aid count","metadata":{}},{"cell_type":"code","source":"counts = df.groupby('aid').count()\ncounts.columns = ['aid', 'aid_cnt']\ndf = df.join(counts, on='aid', how='left')","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:39:04.693741Z","iopub.execute_input":"2023-01-21T11:39:04.695104Z","iopub.status.idle":"2023-01-21T11:39:13.654763Z","shell.execute_reply.started":"2023-01-21T11:39:04.695050Z","shell.execute_reply":"2023-01-21T11:39:13.653427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.hist(df['aid_cnt'], bins=200)\nplt.show()","metadata":{"tags":[],"execution":{"iopub.status.busy":"2023-01-21T11:39:13.656302Z","iopub.execute_input":"2023-01-21T11:39:13.656588Z","iopub.status.idle":"2023-01-21T11:39:16.944471Z","shell.execute_reply.started":"2023-01-21T11:39:13.656561Z","shell.execute_reply":"2023-01-21T11:39:16.943581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Aid Features - aid average time mean and std","metadata":{}},{"cell_type":"code","source":"%%time\ndts = pd.to_datetime(df['ts'].to_pandas(), unit='s')\n\ndf = df.with_column(pl.from_pandas(dts.dt.weekday.rename('day')))\ndf = df.with_column(pl.from_pandas(dts.dt.hour.rename('hour')))\ndf = df.with_column(pl.from_pandas((dts.dt.hour*100 + dts.dt.minute*100//60).rename('hm')))\ndf","metadata":{"tags":[],"execution":{"iopub.status.busy":"2023-01-21T11:28:57.230752Z","iopub.execute_input":"2023-01-21T11:28:57.231789Z","iopub.status.idle":"2023-01-21T11:30:04.153018Z","shell.execute_reply.started":"2023-01-21T11:28:57.231745Z","shell.execute_reply":"2023-01-21T11:30:04.151850Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n## mean\nhm_mean = df.groupby('aid').mean().select(['aid', 'hm'])\nhm_mean.columns = ['aid', 'aid_hm_mean']\ndf = df.join(hm_mean, on='aid', how='left')","metadata":{"tags":[],"execution":{"iopub.status.busy":"2023-01-21T11:30:04.154844Z","iopub.execute_input":"2023-01-21T11:30:04.155151Z","iopub.status.idle":"2023-01-21T11:30:22.884944Z","shell.execute_reply.started":"2023-01-21T11:30:04.155124Z","shell.execute_reply":"2023-01-21T11:30:22.883795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n## median\nhm_med = df.groupby('aid').median().select(['aid', 'hm'])\nhm_med.columns = ['aid', 'aid_hm_median']\ndf = df.join(hm_med, on='aid', how='left')","metadata":{"tags":[],"execution":{"iopub.status.busy":"2023-01-21T11:30:22.886165Z","iopub.execute_input":"2023-01-21T11:30:22.886449Z","iopub.status.idle":"2023-01-21T11:30:30.663262Z","shell.execute_reply.started":"2023-01-21T11:30:22.886425Z","shell.execute_reply":"2023-01-21T11:30:30.662255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n## std\ntemp = df.to_pandas()\ntemp['aid_hm_std'] = temp.groupby('aid')['hm'].transform('std')\ntemp['aid_hm_std'] = temp['aid_hm_std'].fillna(0)\ndf = df.with_column(pl.from_pandas(temp['aid_hm_std']))","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:30:30.665645Z","iopub.execute_input":"2023-01-21T11:30:30.665925Z","iopub.status.idle":"2023-01-21T11:30:47.302105Z","shell.execute_reply.started":"2023-01-21T11:30:30.665901Z","shell.execute_reply":"2023-01-21T11:30:47.300831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:38:39.858380Z","iopub.execute_input":"2023-01-21T11:38:39.859561Z","iopub.status.idle":"2023-01-21T11:38:39.874964Z","shell.execute_reply.started":"2023-01-21T11:38:39.859502Z","shell.execute_reply":"2023-01-21T11:38:39.874031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Aid Feature - weekday and hour ratio","metadata":{}},{"cell_type":"code","source":"%%time\nfor i in range(7):\n    df_temp_cnt = df.filter(pl.col('day') == i).groupby('aid').count()\n    df_temp_cnt.columns = ['aid', f'aid_day{i}cnt']\n    df = df.join(df_temp_cnt, on='aid', how='left')\n    df = df.with_column(pl.col(f'aid_day{i}cnt').fill_null(0))\n    df = df.with_column(df[f'aid_day{i}cnt']/df['aid_cnt'])\n    del df_temp_cnt\n    gc.collect()","metadata":{"tags":[],"execution":{"iopub.status.busy":"2023-01-21T11:40:26.474246Z","iopub.execute_input":"2023-01-21T11:40:26.475010Z","iopub.status.idle":"2023-01-21T11:41:54.276089Z","shell.execute_reply.started":"2023-01-21T11:40:26.474967Z","shell.execute_reply":"2023-01-21T11:41:54.275059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:41:54.278202Z","iopub.execute_input":"2023-01-21T11:41:54.278554Z","iopub.status.idle":"2023-01-21T11:41:54.288705Z","shell.execute_reply.started":"2023-01-21T11:41:54.278523Z","shell.execute_reply":"2023-01-21T11:41:54.287846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfor i in range(24):\n    df_temp_cnt = df.filter(pl.col('hour') == i).groupby('aid').count()\n    df_temp_cnt.columns = ['aid', f'aid_hour{i}cnt']\n    df = df.join(df_temp_cnt, on='aid', how='left')\n    df = df.with_column(pl.col(f'aid_hour{i}cnt').fill_null(0))\n    df = df.with_column(df[f'aid_hour{i}cnt']/df['aid_cnt'])\n    del df_temp_cnt\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:41:54.289813Z","iopub.execute_input":"2023-01-21T11:41:54.290495Z","iopub.status.idle":"2023-01-21T11:46:33.786074Z","shell.execute_reply.started":"2023-01-21T11:41:54.290464Z","shell.execute_reply":"2023-01-21T11:46:33.785121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:46:33.788314Z","iopub.execute_input":"2023-01-21T11:46:33.788617Z","iopub.status.idle":"2023-01-21T11:46:33.800141Z","shell.execute_reply.started":"2023-01-21T11:46:33.788589Z","shell.execute_reply":"2023-01-21T11:46:33.799405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"check_df_ok(df)","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:32:04.180595Z","iopub.execute_input":"2023-01-21T11:32:04.180881Z","iopub.status.idle":"2023-01-21T11:32:11.085101Z","shell.execute_reply.started":"2023-01-21T11:32:04.180855Z","shell.execute_reply":"2023-01-21T11:32:11.083887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Aid Feature - aid buys, carts, clicks cnt + ratio","metadata":{}},{"cell_type":"code","source":"%%time\ncl_cnt = df.filter(pl.col('type')==0).groupby('aid').count()\ncl_cnt.columns = ['aid', 'aid_cl_cnt']\nca_cnt = df.filter(pl.col('type')==1).groupby('aid').count()\nca_cnt.columns = ['aid', 'aid_ca_cnt'] \nor_cnt = df.filter(pl.col('type')==2).groupby('aid').count()\nor_cnt.columns = ['aid', 'aid_or_cnt'] ","metadata":{"tags":[],"execution":{"iopub.status.busy":"2023-01-21T11:32:11.086248Z","iopub.execute_input":"2023-01-21T11:32:11.086529Z","iopub.status.idle":"2023-01-21T11:32:21.118631Z","shell.execute_reply.started":"2023-01-21T11:32:11.086504Z","shell.execute_reply":"2023-01-21T11:32:21.117664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf = df.join(cl_cnt, on='aid', how='left')\ndf = df.join(ca_cnt, on='aid', how='left')\ndf = df.join(or_cnt, on='aid', how='left')","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:32:21.119842Z","iopub.execute_input":"2023-01-21T11:32:21.120139Z","iopub.status.idle":"2023-01-21T11:32:40.750154Z","shell.execute_reply.started":"2023-01-21T11:32:21.120114Z","shell.execute_reply":"2023-01-21T11:32:40.748892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf = df.with_column(pl.col('aid_cl_cnt').fill_null(0))\ndf = df.with_column(pl.col('aid_ca_cnt').fill_null(0))\ndf = df.with_column(pl.col('aid_or_cnt').fill_null(0))","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:32:40.751377Z","iopub.execute_input":"2023-01-21T11:32:40.751655Z","iopub.status.idle":"2023-01-21T11:32:49.101790Z","shell.execute_reply.started":"2023-01-21T11:32:40.751631Z","shell.execute_reply":"2023-01-21T11:32:49.100422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# click to cart, click to order, cart to order ratio\ndf = df.with_column((df['aid_ca_cnt']/df['aid_cl_cnt']).alias('aid_ca_cl_ratio'))\ndf = df.with_column((df['aid_or_cnt']/df['aid_cl_cnt']).alias('aid_or_cl_ratio'))\ndf = df.with_column((df['aid_or_cnt']/df['aid_ca_cnt']).alias('aid_or_ca_ratio'))","metadata":{"tags":[],"execution":{"iopub.status.busy":"2023-01-21T11:32:49.103165Z","iopub.execute_input":"2023-01-21T11:32:49.103467Z","iopub.status.idle":"2023-01-21T11:32:56.834807Z","shell.execute_reply.started":"2023-01-21T11:32:49.103443Z","shell.execute_reply":"2023-01-21T11:32:56.833776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf = df.with_column(pl.col('aid_ca_cl_ratio').fill_nan(-1))\ndf = df.with_column(pl.col('aid_or_cl_ratio').fill_nan(-1))\ndf = df.with_column(pl.col('aid_or_ca_ratio').fill_nan(-1))","metadata":{"tags":[],"execution":{"iopub.status.busy":"2023-01-21T11:32:56.836079Z","iopub.execute_input":"2023-01-21T11:32:56.836415Z","iopub.status.idle":"2023-01-21T11:33:03.671921Z","shell.execute_reply.started":"2023-01-21T11:32:56.836386Z","shell.execute_reply":"2023-01-21T11:33:03.670792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['aid_or_cl_ratio'].is_infinite().sum()","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:33:03.673189Z","iopub.execute_input":"2023-01-21T11:33:03.673493Z","iopub.status.idle":"2023-01-21T11:33:03.845982Z","shell.execute_reply.started":"2023-01-21T11:33:03.673467Z","shell.execute_reply":"2023-01-21T11:33:03.845003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf = df.with_column(\n    pl.when(pl.col(['aid_ca_cl_ratio',\n                    'aid_or_cl_ratio',\n                    'aid_or_ca_ratio']).is_infinite())\n    .then(0)\n    .otherwise(pl.col(['aid_ca_cl_ratio',\n                       'aid_or_cl_ratio',\n                       'aid_or_ca_ratio']))\n    .keep_name()\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:33:03.847263Z","iopub.execute_input":"2023-01-21T11:33:03.847593Z","iopub.status.idle":"2023-01-21T11:33:05.477498Z","shell.execute_reply.started":"2023-01-21T11:33:03.847565Z","shell.execute_reply":"2023-01-21T11:33:05.476253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['aid_or_cl_ratio'].is_infinite().sum()","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:33:05.478945Z","iopub.execute_input":"2023-01-21T11:33:05.479596Z","iopub.status.idle":"2023-01-21T11:33:05.662856Z","shell.execute_reply.started":"2023-01-21T11:33:05.479563Z","shell.execute_reply":"2023-01-21T11:33:05.661836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Aid Feature - same ts stuff","metadata":{}},{"cell_type":"code","source":"%%time\naid_ts_max = df.groupby('aid').max()[['aid', 'ts']]\naid_ts_max.columns = ['aid', 'aid_ts_max']\naid_ts_min = df.groupby('aid').min()[['aid', 'ts']]\naid_ts_min.columns = ['aid', 'aid_ts_min']\naid_ts_mean = df.groupby('aid').mean()[['aid', 'ts']]\naid_ts_mean.columns = ['aid', 'aid_ts_mean']","metadata":{"tags":[],"execution":{"iopub.status.busy":"2023-01-21T11:48:06.288054Z","iopub.execute_input":"2023-01-21T11:48:06.288560Z","iopub.status.idle":"2023-01-21T11:48:25.962190Z","shell.execute_reply.started":"2023-01-21T11:48:06.288524Z","shell.execute_reply":"2023-01-21T11:48:25.960825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df.join(aid_ts_max, on='aid', how='left')\ndf = df.join(aid_ts_min, on='aid', how='left')\ndf = df.join(aid_ts_mean, on='aid', how='left')\ndf","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:48:25.964008Z","iopub.execute_input":"2023-01-21T11:48:25.964369Z","iopub.status.idle":"2023-01-21T11:48:39.437407Z","shell.execute_reply.started":"2023-01-21T11:48:25.964338Z","shell.execute_reply":"2023-01-21T11:48:39.436409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df.drop(['session', 'ts', 'type', 'day', 'hour', 'hm'])\ndf = df.groupby('aid').first()","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:48:39.438655Z","iopub.execute_input":"2023-01-21T11:48:39.438955Z","iopub.status.idle":"2023-01-21T11:48:40.521339Z","shell.execute_reply.started":"2023-01-21T11:48:39.438929Z","shell.execute_reply":"2023-01-21T11:48:40.520312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.write_parquet('aid_features.parquet')","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:48:40.523474Z","iopub.execute_input":"2023-01-21T11:48:40.523815Z","iopub.status.idle":"2023-01-21T11:48:43.356504Z","shell.execute_reply.started":"2023-01-21T11:48:40.523788Z","shell.execute_reply":"2023-01-21T11:48:43.355470Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df","metadata":{"execution":{"iopub.status.busy":"2023-01-21T11:48:43.358127Z","iopub.execute_input":"2023-01-21T11:48:43.358479Z","iopub.status.idle":"2023-01-21T11:48:43.386423Z","shell.execute_reply.started":"2023-01-21T11:48:43.358451Z","shell.execute_reply":"2023-01-21T11:48:43.385574Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Important Notes\n- There should be High Advantage in doing Normalization, because the distribution of original dataset and the dataset for submission has completely different numbers.","metadata":{}},{"cell_type":"markdown","source":"#### Idea Bank\n0. user이 어떤 아이템을 주로 사는지 cluster 번호로 넣어주어야함.\n1. 쇼핑몰 웹사이트 보면서 패턴 확인하기. 첫 페이지 UI들을 보면 insight들을 얻을 수 있다.\n만약 OTTO에서 제공한 데이터가 어떤 종류의 쇼핑몰인지 알면, 더욱 자세히 알 수 있다. 독일 => 쇼핑몰 => 가전 처럼 국가와 도메인을 좁혀서 보면\n어떤 걸 더 클릭할 가능성이 많은지 볼 수 있음.\n2. Youtube Recommendation System처럼 다른 recommendation system 만드는 내용 나오는 논문이나 자료에서 자기들이 썼다고 하는 feature들을 가져와서 쓰기\n3. Chris Deotte가 말한 features\n\n4. discussion에서 말한 feature\n\n5. word2vec - windowsize? => https://www.kaggle.com/competitions/otto-recommender-system/discussion/370751\n6. 3 types of classification loss => https://www.kaggle.com/competitions/otto-recommender-system/discussion/370502\n7. numba fast pipeline study => https://www.kaggle.com/code/carnozhao/otto-fast-cpu-end-to-end-pipeline\n","metadata":{}},{"cell_type":"markdown","source":"#### Deotte's Features\nYou make features for users and features for items. (Note in this competition, \"session\" actually means \"user\"). If we use a binary loss, then we feed our model user-item pairs and our model predicts a single number probability between 0 and 1 (that the user will order this item in future. Another model or output can predict click. And another cart.)\n\nInstead of inputting the user id and item id into our model, we input all of the user features concatenated with all of the item features for that particular user and particular item.\n\nExample user features\n\n- how many items has user already clicked\n- how many items has user already ordered\n- what is average hour that user clicks\n- what is average hour that user orders\n- how many real sessions does user have (real session define by time gap between activity)\n- what is average number of items in each user real session\n- what is last day of week user made activity (i.e. monday, tuesday)\n- what is first day of week user made activity\n- what is average time between clicks\n\nExample item features\n\n- has this item already been clicked by user\n- has this item already been added to cart by user\n- if already clicked, what is its relative order? 1 means last clicked, 2 means second to last clicked etc\n- has user clicked this item multiple times already? how many\n- how many items (that user has already clicked) have recommended this item with their co-visitation matrix\n- !!!when was date that this item was first seen in train\n- how many times what this item clicked in train\n- what is the average hour of day that this item is clicked\n- what is the average hour of day that this item is ordered\n- how popular is this item on monday (i.e. what percentage of monday clicks are this item)\n- how popular is this item on tuesday\n- what is the most common day of week this item is clicked\n\n- count up all unique items that were clicked immediately before and after. How many unique items have been clicked immediately before and after. (For example, maybe item only has 10 unique items that get clicked before and after. Whereas another item has 1000 unique items clicked before and after)\n\n- what percentage of users click this item more than once\n- has this item ever been bought in train data\n\nAbove are just examples, we can brainstorm 1000s of more features.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}