{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"- I cannot understand algorithm especially the part of `merge` in `In[3]` from [Candidate ReRank Model - [LB 0.575]](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575)\n- So I will check by implementating little by little.","metadata":{}},{"cell_type":"code","source":"VER = 5\n\nimport pandas as pd, numpy as np\nfrom tqdm.notebook import tqdm\nimport os, sys, pickle, glob, gc\nfrom collections import Counter\nimport cudf, itertools\nprint('We will use RAPIDS version',cudf.__version__)","metadata":{"papermill":{"duration":3.036143,"end_time":"2022-11-10T16:03:24.014816","exception":false,"start_time":"2022-11-10T16:03:20.978673","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-01-02T07:12:46.546681Z","iopub.execute_input":"2023-01-02T07:12:46.547087Z","iopub.status.idle":"2023-01-02T07:12:49.284305Z","shell.execute_reply.started":"2023-01-02T07:12:46.547002Z","shell.execute_reply":"2023-01-02T07:12:49.283311Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# CACHE FUNCTIONS\ndef read_file(f):\n    return cudf.DataFrame( data_cache[f] )\ndef read_file_to_cache(f):\n    df = pd.read_parquet(f)\n    df.ts = (df.ts/1000).astype('int32')\n    df['type'] = df['type'].map(type_labels).astype('int8')\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-01-02T07:13:50.283055Z","iopub.execute_input":"2023-01-02T07:13:50.283436Z","iopub.status.idle":"2023-01-02T07:13:50.289341Z","shell.execute_reply.started":"2023-01-02T07:13:50.283402Z","shell.execute_reply":"2023-01-02T07:13:50.288266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# CACHE THE DATA ON CPU BEFORE PROCESSING ON GPU\ndata_cache = {}\ntype_labels = {'clicks':0, 'carts':1, 'orders':2}\nfiles = glob.glob('../input/otto-chunk-data-inparquet-format/*_parquet/*')\nfor f in files: data_cache[f] = read_file_to_cache(f)\n\n# CHUNK PARAMETERS\nREAD_CT = 5\nCHUNK = int( np.ceil( len(files)/6 ))\nprint(f'We will process {len(files)} files, in groups of {READ_CT} and chunks of {CHUNK}.')","metadata":{"papermill":{"duration":0.063943,"end_time":"2022-11-10T16:03:24.091816","exception":false,"start_time":"2022-11-10T16:03:24.027873","status":"completed"},"tags":[],"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-01-02T07:13:51.962018Z","iopub.execute_input":"2023-01-02T07:13:51.962397Z","iopub.status.idle":"2023-01-02T07:14:48.451658Z","shell.execute_reply.started":"2023-01-02T07:13:51.962359Z","shell.execute_reply":"2023-01-02T07:14:48.450456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This section is important to understand algorithms(original `In [3]` part).\nI'll check here.","metadata":{"papermill":{"duration":0.004089,"end_time":"2022-11-10T16:03:24.100502","exception":false,"start_time":"2022-11-10T16:03:24.096413","status":"completed"},"tags":[]}},{"cell_type":"code","source":"%%time\n\nPART = 0\nj = 0\na = j*CHUNK\nb = min( (j+1)*CHUNK, len(files) )\nstep_number = 1\ntail_number = 30\n\nprint(f\"a:{a}, b:{b}\")\nprint(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\nk = a\n\ndfs = [read_file(files[k])]\nfor i in range(1, READ_CT):\n    if k + i < b:\n        dfs.append(read_file(files[k + i]))\ndf = cudf.concat(dfs, ignore_index=True, axis=0)\nif step_number == 1:\n    df = df.loc[df[\"type\"].isin([1, 2])]  # ONLY WANT CARTS AND ORDERS\ndf = df.reset_index(drop=True)\ndf[\"n\"] = df.groupby(\"session\").cumcount()\ndf = df.loc[df.n < tail_number].drop(\"n\", axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-01-02T07:19:20.016418Z","iopub.execute_input":"2023-01-02T07:19:20.017104Z","iopub.status.idle":"2023-01-02T07:19:20.118282Z","shell.execute_reply.started":"2023-01-02T07:19:20.017067Z","shell.execute_reply":"2023-01-02T07:19:20.117305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(df.head())\ndisplay(df.describe())\ndisplay(len(df))","metadata":{"execution":{"iopub.status.busy":"2023-01-02T07:20:04.572804Z","iopub.execute_input":"2023-01-02T07:20:04.573177Z","iopub.status.idle":"2023-01-02T07:20:04.648373Z","shell.execute_reply.started":"2023-01-02T07:20:04.573146Z","shell.execute_reply":"2023-01-02T07:20:04.647286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I thought the length of `df` is the same of that after merge oneself, but it is wrong.","metadata":{}},{"cell_type":"code","source":"df_merged = df.merge(df, on=\"session\")\ndisplay(df_merged.head())\ndisplay(df_merged.describe())\ndisplay(len(df_merged))","metadata":{"execution":{"iopub.status.busy":"2023-01-02T07:20:29.998118Z","iopub.execute_input":"2023-01-02T07:20:29.998812Z","iopub.status.idle":"2023-01-02T07:20:30.125257Z","shell.execute_reply.started":"2023-01-02T07:20:29.998777Z","shell.execute_reply":"2023-01-02T07:20:30.124120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Why did it happen? Just because here:\n\n```python\n>>> import pandas as pd\n>>> df = pd.DataFrame({\"item\": [\"apple\", \"orange\", \"apple\", \"apple\"], \"value\":[20, 10, 30, 20]})\n>>> df\n     item  value\n0   apple     20\n1  orange     10\n2   apple     30\n3   apple     20\n>>> df.merge(df) # I'm not sure why it happends but I don't use this way.\n     item  value\n0   apple     20\n1   apple     20\n2   apple     20\n3   apple     20\n4  orange     10\n5   apple     30\n>>> df.merge(df, on=\"value\") # 20: 2 C 2 + 2 = 4, 10 and 30: unique\n   item_x  value  item_y\n0   apple     20   apple\n1   apple     20   apple\n2   apple     20   apple\n3   apple     20   apple\n4  orange     10  orange\n5   apple     30   apple\n>>> df.merge(df, on=\"item\") # apple: 3 C 2 + 3 = 9, orange: unique\n     item  value_x  value_y\n0   apple       20       20\n1   apple       20       30\n2   apple       20       20\n3   apple       30       20\n4   apple       30       30\n5   apple       30       20\n6   apple       20       20\n7   apple       20       30\n8   apple       20       20\n9  orange       10       10\n```","metadata":{}},{"cell_type":"markdown","source":"This notebook is done.","metadata":{}},{"cell_type":"markdown","source":"Postscript\n\nFor me, there are many things to understand about the specifications of various libraries.\nIn addition to this, it often takes time for me to understand them before participating in the competition in earnest.\n\nSo, I will continue to work steadily and without haste.","metadata":{}},{"cell_type":"markdown","source":"Thank you for reading it :)","metadata":{}}]}