{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Start the Analysis - Basic Details and Prepare\n\nThis notebook shows the initial understanding for the analysis. I just begin the competition. Hopefully learn more from all of you.\n\n# Goal\n\nFor each \\<session_id\\>\\_\\<question\\_number\\>, you are identifying whether you believe the user for this particular session will answer this question correctly","metadata":{}},{"cell_type":"markdown","source":"# Assumptions\n\n* level segments 0-4, 5-12, and 13-22 are each provided in sequence\n* Assume each question's level are different\n* Each session has 18 questions","metadata":{}},{"cell_type":"markdown","source":"# Questions?\n\nSome sub questions may want to analyze:\n* For each group of level, how much precentage do we believe the rate of success? \n* For each question, how much precentage do we believe the rate of success? \n* How do we seperate each question's data?\n* How to analyze the texts?\n* Is there any trends that question level decreases for each group?\n\n","metadata":{}},{"cell_type":"markdown","source":"# Memory Cost\n\n#### Int\n* int8 is a signed integer type that can store values in the range of -128 to 127, \n* int16 can store values in the range of -32768 to 32767\n* int32 can store values in the range of -2147483648 to 2147483647\n* int64 can store values in the range of -9223372036854775808 to 9223372036854775807\n\n#### Uint\n* uint8 can represent integer values between 0 and 255 (2^8 - 1).\n* uint16 can represent integer values between 0 and 65,535 (2^16 - 1).\n* uint32 can represent integer values between 0 and 4,294,967,295 (2^32 - 1).\n* uint64 can represent integer values between 0 and 18,446,744,073,709,551,615 (2^64 - 1).\n\n#### Float\n* float16: 16 bits of storage. It has a range of approximately -65,000 to +65,000, and a precision of approximately 1 part in 65,000. It is often used in machine learning applications to reduce memory usage, since it uses only half the storage of float32.\n\n* float32: 32 bits of storage. It has a range of approximately -3.4 x 10^38 to +3.4 x 10^38, and a precision of approximately 1 part in 10^7.\n\n* float64: 64 bits of storage. It has a range of approximately -1.8 x 10^308 to +1.8 x 10^308, and a precision of approximately 1 part in 10^15. It is used when high precision is required\n\n#### Object\nWhen a column is cast as \"category\" type, Pandas internally maps each unique value to an integer code, and stores the data as an array of integers instead of an array of strings or other object types. This reduces the memory usage, as the integer codes take up less space than the original values.\n\nFor example, if a column has only three unique values, \"A\", \"B\", and \"C\", and those values are repeated many times, it would be more memory-efficient to store the column as a \"category\" type, where \"A\", \"B\", and \"C\" are mapped to the integer codes 0, 1, and 2, respectively.\n\n#### Default for pandas and numpy:\n* int64, uint64, float64 and object","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Reference: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384359\n# I made tiny updates\n# You can Modify it by \n# 1. https://www.kaggle.com/competitions/predict-student-performance-from-game-play/data?select=train.csv\n# 2. click Column, there are detailed summaries for each column\ndtypes={\n    'index': np.uint16,\n    'elapsed_time':np.int32,\n    'event_name':'category',\n    'name':'category',\n    'level':np.uint8,\n    'page': np.float32,\n    'room_coor_x':np.float32,\n    'room_coor_y':np.float32,\n    'screen_coor_x':np.float32,\n    'screen_coor_y':np.float32,\n    'hover_duration':np.float32,\n    'text':'category',\n    'fqid':'category',\n    'room_fqid':'category',\n    'text_fqid':'category',\n    'fullscreen':'category',\n    'hq':'category',\n    'music':'category',\n    'level_group':'category'}\n\ndataset_df = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv', dtype=dtypes)\ndataset_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:10:52.683530Z","iopub.execute_input":"2023-04-18T20:10:52.684063Z","iopub.status.idle":"2023-04-18T20:12:46.082416Z","shell.execute_reply.started":"2023-04-18T20:10:52.684012Z","shell.execute_reply":"2023-04-18T20:12:46.081273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# try the code by run a subset\n# if you want to run on whole dataset then comment this cell\n# first 100 samples\nunique_session_ids = dataset_df['session_id'].unique()[:100]\ndataset_df = dataset_df[dataset_df['session_id'].isin(unique_session_ids)]","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:44:13.929278Z","iopub.execute_input":"2023-04-18T20:44:13.929725Z","iopub.status.idle":"2023-04-18T20:44:13.948309Z","shell.execute_reply.started":"2023-04-18T20:44:13.929688Z","shell.execute_reply":"2023-04-18T20:44:13.946799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset_df.session_id.nunique()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:46:36.345424Z","iopub.execute_input":"2023-04-18T20:46:36.345840Z","iopub.status.idle":"2023-04-18T20:46:36.355120Z","shell.execute_reply.started":"2023-04-18T20:46:36.345802Z","shell.execute_reply":"2023-04-18T20:46:36.354162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Understand the data by one sample\n\nHelpful article to understand the game: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384796","metadata":{}},{"cell_type":"code","source":"df_one_sample = dataset_df[dataset_df['session_id'] == 20090312431273200]\ndf_one_sample.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:12:46.091950Z","iopub.execute_input":"2023-04-18T20:12:46.092557Z","iopub.status.idle":"2023-04-18T20:12:46.136155Z","shell.execute_reply.started":"2023-04-18T20:12:46.092513Z","shell.execute_reply":"2023-04-18T20:12:46.134973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# each column's unique values\ndf_one_sample.nunique()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:12:46.159946Z","iopub.execute_input":"2023-04-18T20:12:46.160546Z","iopub.status.idle":"2023-04-18T20:12:46.172318Z","shell.execute_reply.started":"2023-04-18T20:12:46.160506Z","shell.execute_reply":"2023-04-18T20:12:46.171292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Event_name\n* navigate_click: This could refer to a click on a navigation button, link, or element within the user interface that allows the user to move between different sections or pages of the application.\n\n* person_click: This could refer to a click on a person icon or button, such as a user profile or avatar.\n\n* cutscene_click: A cutscene is a sequence in a video game that is not interactive and interrupts the gameplay. A cutscene_click could refer to a click on a button or element within the cutscene to skip or end the sequence.\n\n* object_click: This could refer to a click on an object or item within the game or application. For example, in a game like Minecraft, this could refer to a click on a block or item in the game world.\n\n* object_hover: This could refer to a mouse hover over an object or item within the game or application, without actually clicking on it.\n\n* map_hover and map_click: These could refer to interactions with a map within the application, such as hovering over or clicking on different locations.\n\n* notification_click: This could refer to a click on a notification or alert within the application.\n\n* observation_click: This is less clear without more context, but it could refer to a click on an observation or data point within the application.\n\n* checkpoint: This could refer to a specific point in the game or application where progress is saved or a milestone is reached.\n\n* notebook_click: Again, without more context it is difficult to say for sure, but this could refer to a click on a notebook or document within the application.","metadata":{}},{"cell_type":"code","source":"df_one_sample.event_name.value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:12:46.173597Z","iopub.execute_input":"2023-04-18T20:12:46.173926Z","iopub.status.idle":"2023-04-18T20:12:46.187952Z","shell.execute_reply.started":"2023-04-18T20:12:46.173895Z","shell.execute_reply":"2023-04-18T20:12:46.186651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Name\n\nIdentifies whether a notebook_click is is opening or closing the notebook.","metadata":{}},{"cell_type":"code","source":"df_one_sample.name.value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:12:46.189717Z","iopub.execute_input":"2023-04-18T20:12:46.190133Z","iopub.status.idle":"2023-04-18T20:12:46.210519Z","shell.execute_reply.started":"2023-04-18T20:12:46.190098Z","shell.execute_reply":"2023-04-18T20:12:46.209161Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Level\n\nWhat level of the game the event occurred in (0 to 22). Each game has 18 levels","metadata":{}},{"cell_type":"code","source":"df_one_sample.level.value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:12:46.211969Z","iopub.execute_input":"2023-04-18T20:12:46.212656Z","iopub.status.idle":"2023-04-18T20:12:46.226302Z","shell.execute_reply.started":"2023-04-18T20:12:46.212611Z","shell.execute_reply":"2023-04-18T20:12:46.225097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Text\n\nWhich show the display information.","metadata":{}},{"cell_type":"code","source":"df_one_sample['text'].dropna()[0:60]","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:12:46.227987Z","iopub.execute_input":"2023-04-18T20:12:46.228703Z","iopub.status.idle":"2023-04-18T20:12:46.243244Z","shell.execute_reply.started":"2023-04-18T20:12:46.228661Z","shell.execute_reply":"2023-04-18T20:12:46.241682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_one_sample['text'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:12:46.247980Z","iopub.execute_input":"2023-04-18T20:12:46.248777Z","iopub.status.idle":"2023-04-18T20:12:46.261334Z","shell.execute_reply.started":"2023-04-18T20:12:46.248726Z","shell.execute_reply":"2023-04-18T20:12:46.260165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Fqid\n\nThe fully qualified ID of the event, including different charactors.","metadata":{}},{"cell_type":"code","source":"df_one_sample.fqid.value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:12:46.263180Z","iopub.execute_input":"2023-04-18T20:12:46.263736Z","iopub.status.idle":"2023-04-18T20:12:46.279935Z","shell.execute_reply.started":"2023-04-18T20:12:46.263693Z","shell.execute_reply":"2023-04-18T20:12:46.278193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Text_fqid \n\nThe fully qualified ID of the event, including different props and special events in the game.\n\n* tunic: This could refer to a specific game or application, possibly a video game.\n\n* wildlife.center: This could refer to a specific location or area within the game or application, possibly related to a wildlife center.\n\n* crane_ranger.crane: This could refer to a specific crane or bird within the game, possibly related to a ranger or wildlife center.\n\n* historicalsociety: This could refer to a specific organization or group within the game or application, possibly related to history or a museum.\n\n* cage.confrontation: This could refer to a specific event or encounter within the game, possibly related to a cage or imprisonment.\n\n* frontdesk.archivist.newspaper: This could refer to a specific location or area within the game, possibly related to a front desk, archivist, or newspaper.\n\n* entry.groupconvo: This could refer to a specific event or encounter within the game, possibly related to a group conversation or meeting.\n\n* wildlife.center.wells.nodeer: This could refer to a specific location or area within the game, possibly related to a wildlife center, wells, or the absence of deer.","metadata":{}},{"cell_type":"code","source":"df_one_sample.text_fqid.value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:12:46.281952Z","iopub.execute_input":"2023-04-18T20:12:46.282663Z","iopub.status.idle":"2023-04-18T20:12:46.294012Z","shell.execute_reply.started":"2023-04-18T20:12:46.282616Z","shell.execute_reply":"2023-04-18T20:12:46.292965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# size of the data\ndataset_df.shape","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:12:46.295530Z","iopub.execute_input":"2023-04-18T20:12:46.296759Z","iopub.status.idle":"2023-04-18T20:12:46.307509Z","shell.execute_reply.started":"2023-04-18T20:12:46.296711Z","shell.execute_reply":"2023-04-18T20:12:46.306269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Read label dataset","metadata":{}},{"cell_type":"code","source":"labels = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train_labels.csv')\nlabels['session_id'] = labels.session_id.apply(lambda x: int(x.split('_')[0]) )\nlabels['q'] = labels.session_id.apply(lambda x: int(x.split('_')[-1][1:]) )\nlabels.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:12:46.309290Z","iopub.execute_input":"2023-04-18T20:12:46.310061Z","iopub.status.idle":"2023-04-18T20:12:47.439320Z","shell.execute_reply.started":"2023-04-18T20:12:46.310008Z","shell.execute_reply":"2023-04-18T20:12:47.438419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Which shows the prediction goal is to identify the user doing each question correctly or not","metadata":{}},{"cell_type":"code","source":"# try the code by run a subset\n# if you want to run on whole dataset then comment this cell\n# first 100 samples\nunique_session_ids = dataset_df['session_id'].unique()[:100]\nlabels = labels[labels['session_id'].isin(unique_session_ids)]","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:46:04.202946Z","iopub.execute_input":"2023-04-18T20:46:04.203404Z","iopub.status.idle":"2023-04-18T20:46:04.297233Z","shell.execute_reply.started":"2023-04-18T20:46:04.203364Z","shell.execute_reply":"2023-04-18T20:46:04.296018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Combine two datasets\n\nCurrently we know each section's information and each section ID's prediction resutls. The prediction result is for each question, but we don't know about which part of the data the user answering the question. Instead we know about the level group for each question. \n\nIf you knows how to get each question's data from train.csv, please show in the comment, thanks. (as an example get all data for 20090313571836404_q1 in train.csv)\n\nCurrently I have categorical data and numeric data, and apply the following steps:\n\n1. Group with session_id and level_group, generate statistics\n1. convert each level group in to the same row\n1. for each level build a model and return the probability of 0 and 1\n1. for each question build a model and return the probability of 0 and 1, this model can add the probability from the last step as a input of the new model\n1. Combination of the probability and return the result","metadata":{}},{"cell_type":"markdown","source":"## Step 1","metadata":{}},{"cell_type":"code","source":"grouped = dataset_df.groupby(['session_id', 'level_group'])\ngrouped.describe()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:12:47.440690Z","iopub.execute_input":"2023-04-18T20:12:47.441903Z","iopub.status.idle":"2023-04-18T20:12:54.058602Z","shell.execute_reply.started":"2023-04-18T20:12:47.441857Z","shell.execute_reply":"2023-04-18T20:12:54.057723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset_df.nunique()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:12:54.060015Z","iopub.execute_input":"2023-04-18T20:12:54.061139Z","iopub.status.idle":"2023-04-18T20:12:54.097180Z","shell.execute_reply.started":"2023-04-18T20:12:54.061097Z","shell.execute_reply":"2023-04-18T20:12:54.096069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset_df.info()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:12:54.098604Z","iopub.execute_input":"2023-04-18T20:12:54.098952Z","iopub.status.idle":"2023-04-18T20:12:54.124622Z","shell.execute_reply.started":"2023-04-18T20:12:54.098913Z","shell.execute_reply":"2023-04-18T20:12:54.123376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_cols = ['event_name', 'name','fqid', 'room_fqid', 'text_fqid','fullscreen','hq','music']\nnum_cols = ['elapsed_time','level','page','room_coor_x', 'room_coor_y', \n        'screen_coor_x', 'screen_coor_y', 'hover_duration']\ndummy_cols = ['event_name', 'name','fqid','fullscreen','hq','music']","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:36:52.478516Z","iopub.execute_input":"2023-04-18T20:36:52.479000Z","iopub.status.idle":"2023-04-18T20:36:52.485612Z","shell.execute_reply.started":"2023-04-18T20:36:52.478954Z","shell.execute_reply":"2023-04-18T20:36:52.484643Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset_df[cat_cols].nunique()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:35:53.802561Z","iopub.execute_input":"2023-04-18T20:35:53.802984Z","iopub.status.idle":"2023-04-18T20:35:53.821896Z","shell.execute_reply.started":"2023-04-18T20:35:53.802947Z","shell.execute_reply":"2023-04-18T20:35:53.820960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Basic statistics and Preprocessing\n\nCategorical data: frequency encoding, unique counts\n\nNumeric data: mean, standard deviation","metadata":{}},{"cell_type":"code","source":"# modify from Reference: https://www.kaggle.com/code/cdeotte/random-forest-baseline-0-664/notebook\nfrom tqdm import tqdm\ndef feature_engineer(df):\n    df = df[['session_id', 'level_group'] + cat_cols + num_cols]\n\n    dfs = []\n    groups = df.groupby(['session_id', 'level_group'])\n\n    # Frequency encoding and dummy variables for categorical variables\n    for c in tqdm(cat_cols):\n        freq = groups[c].agg('count') / groups[c].agg('count').sum()\n        freq.name = c + '_freq'\n        dfs.append(freq)\n\n        # Number of unique values\n        nunique = groups[c].nunique()\n        nunique.name = c + '_nunique'\n        dfs.append(nunique)\n    # Dummy Variable\n    for c in tqdm(dummy_cols):\n        dummies = pd.get_dummies(df[c], prefix=c)\n        dummies_grouped = dummies.groupby([df['session_id'], df['level_group']]).sum()\n        dfs.append(dummies_grouped)\n\n    # Mean and standard deviation for numeric variables\n    for c in tqdm(num_cols):\n        mean = groups[c].mean()\n        mean.name = c + '_mean'\n        dfs.append(mean)\n\n        std = groups[c].std()\n        std.name = c + '_std'\n        dfs.append(std)\n    df = pd.concat(dfs, axis=1)\n    df = df.fillna(-1)\n    df = df.reset_index()\n    df.set_index(['session_id', 'level_group'], inplace=True)  # set index to both session_id and level_group\n    return df\n\ndataset_df_transformed = feature_engineer(dataset_df)\ndataset_df_transformed.head(9)","metadata":{"execution":{"iopub.status.busy":"2023-04-18T20:38:14.193394Z","iopub.execute_input":"2023-04-18T20:38:14.193815Z","iopub.status.idle":"2023-04-18T20:38:14.668945Z","shell.execute_reply.started":"2023-04-18T20:38:14.193778Z","shell.execute_reply":"2023-04-18T20:38:14.667375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thanks for watching, and I'll keep updating it.","metadata":{}}]}