{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"font-family:calibri;font-size:250%;text-align:center;\">📈 EDA Adventure: Navigating the Children Learning Domain</h1>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"table\"></a>\n<h1 style=\"background-color:lightskyblue;font-family:calibri;font-size:250%;text-align:center;border-radius: 25px 25px;\">Table of Contents</h1>\n\n* [1. Introduction](#1)\n\n    * [1.1 Foreword](#1.1)\n    \n    * [1.2 Reference and Inspiration Sources](#1.2)\n\n* [2. Copmetition information](#2)\n\n    * [2.1 What's This Competition All About? 🤔💡](#2.1)\n    \n    * [2.2 A Look at the Available Data 🔍](#2.2)\n    \n    * [2.3 Understanding the Competition Goal 🎯](#2.3)\n    \n    * [2.4 Little bit about evaluation metric 🧐](#2.4)\n    \n    * [2.5 🏆 Efficiency Prize? Let's Learn More!](#2.5)\n    \n    * [2.6 🔎 Why it Matters: Uncovering the Importance of the Competition](#2.6)\n\n* [3. Data Exploration](#3)\n    \n    * [3.1 Load libraries and data](#3.1)\n    \n    * [3.2 Explore train data](#3.2)\n    \n    * [3.3 Explore train labels](#3.3)\n    \n    * [3.4 Explore sample submission](#3.4)\n    \n    * [3.5 Explore test data](#3.5)\n\n* [4. Data Preprocessing](#4)\n\n    * [4.1 Make data smaller](#4.1)\n    \n    * [4.2 Downsampling for EDA](#4.2)\n\n* [5. Exploratory Data Analysis](#5)\n    \n    * [5.1 Number of sessions in dataset](#5.1)\n    \n    * [5.2 Number of events in sessions](#5.2)\n    \n    * [5.3 Types of events in dataset](#5.3)\n    \n    * [5.4 Distribution of rooms](#5.4)\n    \n    * [5.5 How many events for each level group](#5.5)\n    \n    * [5.6 Distribution of the event name by level](#5.6)\n    \n    * [5.7 Number of words in text of events](#5.7)\n    \n    * [5.8 What about music?](#5.8)\n\n* [6. Conclusion](#6)","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">1. Introduction</p>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1.1\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">1.1 Foreword</p>","metadata":{}},{"cell_type":"markdown","source":"🤗 Hey there! Welcome to my first notebook for this competition 🎉 I'm so excited to dive into some exploratory data analysis (EDA) and see what insights we can uncover. My goal is to make this notebook a useful resource for myself and anyone else who is participating in this competition.\n\n💻 So, grab your laptops, some tea (or coffee, if that's your thing), and let's get started!\n\n🧐 If you find my work helpful, don't forget to give it an upvote! I would really appreciate it. And if you have any comments or questions, please don't hesitate to reach out in the comments section.\n\n🚀 Let's get started! 🚀","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1.2\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">1.2 Reference and Inspiration Sources</p>","metadata":{}},{"cell_type":"markdown","source":"- 💬 [ChatGPT](https://chat.openai.com/chat) was a big help there! I'm not the best with emojis, but ChatGPT always knows the right ones 🤗 (even there😉)\n- My fellow Petr was just a bit quicker than me with their [notebook](https://www.kaggle.com/code/asimple/eda-game-play). Make sure to check it out if you haven't already! 🔍\n- Nice [notebook](https://www.kaggle.com/code/mohammad2012191/reduce-memory-usage-2gb-780mb) to reduce the size of training data. Give the upvote to this man😉","metadata":{}},{"cell_type":"markdown","source":"<a id=\"2\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">2. Competition Information</p>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"2.1\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">2.1 What's This Competition All About? 🤔💡</p>","metadata":{}},{"cell_type":"markdown","source":"Get ready for a thrilling adventure in the world of online education! 🎓💻 This competition is all about using time series data from an educational game to predict whether players will answer questions correctly. 🤔 It's like playing detective and using the information given to you to make an educated guess!\n\nYou'll have access to the training data and labels, with 18 questions for each session. The objective is to use the previous information for each session to predict whether the user will answer each question correctly. 🧐\n\nThe data will be split into three question checkpoints (level 4, level 12, and level 22) and served up as Pandas dataframes. All you have to do is use the sample notebook, make your predictions, and voila! You'll have a submission.csv file ready for you to submit. 📈📊\n\nSo put on your thinking cap, get ready for some mind-bending fun, and let's see how well you can do! 🤩💡","metadata":{}},{"cell_type":"markdown","source":"<a id=\"2.2\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">2.2 A Look at the Available Data 🔍</p>","metadata":{}},{"cell_type":"markdown","source":"Get ready to dive into the data! 📊🔍 For this competition, you'll be using the Kaggle Time Series API to analyze time series data generated by an online educational game. The goal is to determine whether players will answer questions correctly at three different checkpoints (level 4, level 12, and level 22).\n\nYou'll have access to the training data and labels which includes information about 18 questions for each session. Some of the columns you'll be working with include session_id, elapsed_time, event_name, level, text, and much more.\n\nSo sharpen your skills, and get ready to turn this data into insights! 💪💻","metadata":{}},{"cell_type":"markdown","source":"<a id=\"2.3\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">2.3 Understanding the Competition Goal 🎯</p>","metadata":{}},{"cell_type":"markdown","source":"Let's talk about the goal of this competition. The objective is to use the time series data generated from an online educational game to determine the players' ability to answer questions correctly. You will have access to the training data and labels for 18 questions per session. Your task is to predict, for each session and question, whether the player will answer the question correctly or not, based on the previous information available in the session. ","metadata":{}},{"cell_type":"markdown","source":"<a id=\"2.4\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">2.4 Little bit about evaluation metric 🧐</p>","metadata":{}},{"cell_type":"markdown","source":"🤔 Let's talk about the evaluation metric for this competition! 💻\n\nThe F1 score is the main evaluation metric used to judge the performance of your model. 🧐 It's a balance between precision and recall, two important measures of a model's accuracy.\n\nPrecision is the number of correct positive predictions divided by the number of positive predictions made by the model. Recall is the number of correct positive predictions divided by the number of positive samples in the data.\n\nF1 score is calculated as 2 * (Precision * Recall) / (Precision + Recall). The higher the F1 score, the better the model is at making correct positive predictions while also minimizing false positive predictions.\n\nSo, aim for the highest F1 score possible to show off the skills of your model! 💪","metadata":{}},{"cell_type":"markdown","source":"<a id=\"2.5\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">2.5 🏆 Efficiency Prize? Let's Learn More!</p>","metadata":{}},{"cell_type":"markdown","source":"🤔 Interested in the Efficiency Prize in our competition? Let's dive in! 💦\n\nThe Efficiency Prize is a special track within our competition that focuses on creating models that are both highly accurate and computationally efficient. We understand that big and powerful models can be tough to use in real-world educational settings with limited computational resources, so the Efficiency Prize rewards submissions that are both accurate and lightweight.\n\nTo be eligible for the Efficiency Prize, a submission must first be selected for the Leaderboard Prize or automatically selected under certain conditions. It must also be ranked higher than the sample_submission.csv benchmark on the private leaderboard and can only be run on CPU, not GPU.\n\nThe Efficiency Prize is awarded based on a unique evaluation metric that balances a submission's predictive performance (F1 score) and runtime. The score is calculated by subtracting a submission's F1 score and runtime from the benchmark, with the goal of minimizing the score.\n\nSo, if you're up for a challenge, give the Efficiency Prize a shot! Show us that you can create an accurate and efficient model that can be used in real-world educational settings. Good luck! 🚀","metadata":{}},{"cell_type":"markdown","source":"<a id=\"2.6\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">2.6 🔎 Why it Matters: Uncovering the Importance of the Competition</p>","metadata":{}},{"cell_type":"markdown","source":"💡 The competition is all about making a difference in the world of education!\n\nAs we all know, the education sector is crucial in shaping the future of our society. However, not all students have access to the same resources and tools, which can limit their learning potential. This competition is aimed at bridging that gap by developing models that are not only accurate, but also efficient and easily accessible for educational organizations with limited computational capabilities.\n\n📚 By participating in this competition, you have the opportunity to play a role in providing equal education opportunities for all. Your models will help create more effective and efficient learning experiences for students around the world!","metadata":{}},{"cell_type":"markdown","source":"Now, we are ready to write some code and play with data😛","metadata":{}},{"cell_type":"markdown","source":"<a id=\"3\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">3. Data Exploration</p>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"3.1\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">3.1 Load libraries and data</p>","metadata":{}},{"cell_type":"markdown","source":"Without importing this plotly version, plots become invisible. Maybe, their new version aims at some invisible-watcher people, but perhaps I will never know...","metadata":{}},{"cell_type":"code","source":"!pip install plotly==5.11.0","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:24:36.170646Z","iopub.execute_input":"2023-02-09T00:24:36.171256Z","iopub.status.idle":"2023-02-09T00:25:32.780778Z","shell.execute_reply.started":"2023-02-09T00:24:36.171211Z","shell.execute_reply":"2023-02-09T00:25:32.779132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from pathlib import Path\n\nimport pandas as pd\nimport numpy as np\nimport plotly.express as px","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:25:32.786508Z","iopub.execute_input":"2023-02-09T00:25:32.786940Z","iopub.status.idle":"2023-02-09T00:25:34.036220Z","shell.execute_reply.started":"2023-02-09T00:25:32.786894Z","shell.execute_reply":"2023-02-09T00:25:34.035148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Keep seed to pleasure The Devastator","metadata":{}},{"cell_type":"code","source":"RANDOM_SEED = 42","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:25:34.037700Z","iopub.execute_input":"2023-02-09T00:25:34.038083Z","iopub.status.idle":"2023-02-09T00:25:34.042834Z","shell.execute_reply.started":"2023-02-09T00:25:34.038049Z","shell.execute_reply":"2023-02-09T00:25:34.041848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_path = Path(\"/kaggle/input/predict-student-performance-from-game-play\")","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:25:34.044527Z","iopub.execute_input":"2023-02-09T00:25:34.044892Z","iopub.status.idle":"2023-02-09T00:25:34.056992Z","shell.execute_reply.started":"2023-02-09T00:25:34.044859Z","shell.execute_reply":"2023-02-09T00:25:34.055832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = pd.read_csv(data_path / \"train.csv\")\ntest_data = pd.read_csv(data_path / \"test.csv\")\ntrain_labels = pd.read_csv(data_path / \"train_labels.csv\")\nsample_submission = pd.read_csv(data_path / \"sample_submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:25:34.059948Z","iopub.execute_input":"2023-02-09T00:25:34.060346Z","iopub.status.idle":"2023-02-09T00:26:40.658744Z","shell.execute_reply.started":"2023-02-09T00:25:34.060310Z","shell.execute_reply":"2023-02-09T00:26:40.657656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3.2\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">3.2 Explore train data</p>","metadata":{}},{"cell_type":"markdown","source":"First things first, let's start with training data","metadata":{}},{"cell_type":"code","source":"train_data","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:26:40.660413Z","iopub.execute_input":"2023-02-09T00:26:40.660753Z","iopub.status.idle":"2023-02-09T00:26:40.715503Z","shell.execute_reply.started":"2023-02-09T00:26:40.660721Z","shell.execute_reply":"2023-02-09T00:26:40.714565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, I observe here next columns:\n  - session_id - the ID of the session the event took place in\n  - index - the index of the event for the session\n  - elapsed_time - how much time has passed (in milliseconds) between the start of the session and when the event was recorded\n  - event_name - the name of the event type\n  - name - the event name (e.g. identifies whether a notebook_click is is opening or closing the notebook)\n  - level - what level of the game the event occurred in (0 to 22)\n  - page - the page number of the event (only for notebook-related events)\n  - room_coor_x - the coordinates of the click in reference to the in-game room (only for click events)\n  - room_coor_y - the coordinates of the click in reference to the in-game room (only for click events)\n  - screen_coor_x - the coordinates of the click in reference to the player’s screen (only for click events)\n  - screen_coor_y - the coordinates of the click in reference to the player’s screen (only for click events)\n  - hover_duration - how long (in milliseconds) the hover happened for (only for hover events)\n  - text - the text the player sees during this event\n  - fqid - the fully qualified ID of the event\n  - room_fqid - the fully qualified ID of the room the event took place in\n  - text_fqid - the fully qualified ID of the text\n  - fullscreen - whether the player is in fullscreen mode\n  - hq - whether the game is in high-quality\n  - music - whether the game music is on or off\n  - level_group - which group of levels - and group of questions - this row belongs to (0-4, 5-12, 13-22)\n  \n(Brazenly stole the description of the columns in the description of the competition)","metadata":{}},{"cell_type":"markdown","source":"So, there we have full information about events: type of event, the sequence number of event and its time, the level at which event occurred, coordinates of the click, and additional information like is music on, is game in fullscreen, is game in high-quality and **text, that player sees during event**.\n\nThis game is separated into levels and rooms. I suppose, that on each level there are some rooms available to a player. And rooms differ in their difficulty.","metadata":{}},{"cell_type":"markdown","source":"Takeaways:\n  - The main component of this dataset is the session, and we have to analyze the session fully, to make reasonable predictions\n  - For each session, we have a number of events (rows in the table). Most of all, all events are clicks of the player. Each event has its type, index in session, emergence time, text seen while this event, room id, coordinates of click, and parameters of the environment at the time of this click (music, fullscreen, quality of picture)\n  - This dataset looks very interesting for investigating the nature of attention and information acquisition depending on different conditions\n  - Most of all, the transformer would be the best architecture for this competition. Because I assume a big role in the text, and also we have many different parameters, that may have their own embeddings. Also, transformers are one of the most wildly-used architectures, that took place in absolutely different tasks.\n  - Predictions must be in *train_labels* table","metadata":{}},{"cell_type":"markdown","source":"<a id=\"3.3\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">3.3 Explore train labels</p>","metadata":{}},{"cell_type":"markdown","source":"Second things second, inspect train labels. It's also important. It's kinda our targets","metadata":{}},{"cell_type":"code","source":"train_labels","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:26:40.717242Z","iopub.execute_input":"2023-02-09T00:26:40.718376Z","iopub.status.idle":"2023-02-09T00:26:40.750734Z","shell.execute_reply.started":"2023-02-09T00:26:40.718335Z","shell.execute_reply":"2023-02-09T00:26:40.749834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, I wasn't wrong. For each question in the session, we have labeled 1, if the player answered the question correctly. And 0 if not.\n\nSo, possibly, we will have to separate events also by levels, to get precise information about the actions of players, related to each question.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"3.4\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">3.4 Explore sample submissions</p>","metadata":{}},{"cell_type":"markdown","source":"Now sample submission. It's how our submission must look like.","metadata":{}},{"cell_type":"code","source":"sample_submission","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:26:40.752124Z","iopub.execute_input":"2023-02-09T00:26:40.752963Z","iopub.status.idle":"2023-02-09T00:26:40.771989Z","shell.execute_reply.started":"2023-02-09T00:26:40.752924Z","shell.execute_reply":"2023-02-09T00:26:40.771073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For each session we have to predict, they will user give the correct answer to questions, based on information of type, given on the train table.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"3.5\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">3.5 Explore test data</p>","metadata":{}},{"cell_type":"markdown","source":"And now, the test data","metadata":{}},{"cell_type":"code","source":"test_data","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:26:40.773310Z","iopub.execute_input":"2023-02-09T00:26:40.774088Z","iopub.status.idle":"2023-02-09T00:26:40.817408Z","shell.execute_reply.started":"2023-02-09T00:26:40.774048Z","shell.execute_reply":"2023-02-09T00:26:40.816227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It looks identical to the train data. But we don't have test_labels table. And consequently, we don't know, did the player give the correct answer or not","metadata":{}},{"cell_type":"markdown","source":"<a id=\"4\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">4. Data Preprocessing</p>","metadata":{}},{"cell_type":"markdown","source":"I'll keep it short, but there are some things that will make work with this data easier","metadata":{}},{"cell_type":"markdown","source":"<a id=\"4.1\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">4.1 Make data smaller</p>","metadata":{}},{"cell_type":"markdown","source":"Train data weights around 2 GB. Not a bit! And there is a way to make it smaller. This function took from this nice notebook: https://www.kaggle.com/code/mohammad2012191/reduce-memory-usage-2gb-780mb","metadata":{}},{"cell_type":"code","source":"def reduce_mem_usage(df):\n    start_mem = df.memory_usage().sum() / 1024**2\n    print(f'Memory usage of dataframe is {start_mem:.2f} MB')\n    for col in df.columns:\n        if df[col].dtype == 'object':\n            df[col] = df[col].astype('category')\n        elif df[col].dtype == 'int':\n            int_types = [np.int8, np.int16, np.int32, np.int64]\n            for int_type in int_types:\n                if df[col].min() >= np.iinfo(int_type).min and df[col].max() <= np.iinfo(int_type).max:\n                    df[col] = df[col].astype(int_type)\n                    break\n        elif df[col].dtype == 'float':\n            float_types = [np.float16, np.float32, np.float64]\n            for float_type in float_types:\n                if df[col].min() >= np.finfo(float_type).min and df[col].max() <= np.finfo(float_type).max:\n                    df[col] = df[col].astype(float_type)\n                    break\n    mem_usage = df.memory_usage().sum() / 1024**2 \n    print(f\"Memory usage became: {mem_usage:.2f} MB\")\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:26:40.818779Z","iopub.execute_input":"2023-02-09T00:26:40.819221Z","iopub.status.idle":"2023-02-09T00:26:40.830645Z","shell.execute_reply.started":"2023-02-09T00:26:40.819185Z","shell.execute_reply":"2023-02-09T00:26:40.829470Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = reduce_mem_usage(train_data)","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:26:40.832108Z","iopub.execute_input":"2023-02-09T00:26:40.832579Z","iopub.status.idle":"2023-02-09T00:26:54.474500Z","shell.execute_reply.started":"2023-02-09T00:26:40.832541Z","shell.execute_reply":"2023-02-09T00:26:54.473364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:26:54.476229Z","iopub.execute_input":"2023-02-09T00:26:54.476572Z","iopub.status.idle":"2023-02-09T00:26:54.521225Z","shell.execute_reply.started":"2023-02-09T00:26:54.476540Z","shell.execute_reply":"2023-02-09T00:26:54.520181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, it looks like before permutation but weighs more than 2 times less. Impressive","metadata":{}},{"cell_type":"markdown","source":"<a id=\"4.2\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">4.2 Downsampling for EDA</p>","metadata":{}},{"cell_type":"markdown","source":"Because this dataset is quite big, and plotly poorly works with tables of this size, I have to do downsampling. I will take 1000 random sessions from the training dataset, and perform analysis on this subset.\n\nI will do something about this problem, to avoid downsampling","metadata":{}},{"cell_type":"code","source":"rng = np.random.default_rng(seed=RANDOM_SEED)\n\neda_sessions = rng.choice(train_data[\"session_id\"].unique(), 1000) \ntrain_for_eda = train_data[train_data[\"session_id\"].isin(eda_sessions)].reset_index(drop=True)\ntrain_for_eda","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:26:54.522628Z","iopub.execute_input":"2023-02-09T00:26:54.522981Z","iopub.status.idle":"2023-02-09T00:26:54.934072Z","shell.execute_reply.started":"2023-02-09T00:26:54.522949Z","shell.execute_reply":"2023-02-09T00:26:54.932795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Still a big number of events","metadata":{}},{"cell_type":"markdown","source":"<a id=\"5\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">5. Exploratory Data Analysis</p>","metadata":{}},{"cell_type":"markdown","source":"Here we go with EDA on the notebook titled EDA. Expected that earlier, didn't you? I'm too","metadata":{}},{"cell_type":"markdown","source":"<a id=\"5.1\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">5.1 Number of sessions in dataset</p>","metadata":{}},{"cell_type":"markdown","source":"First I want to calculate the number of sessions in the whole dataset","metadata":{}},{"cell_type":"code","source":"n_train_sessions = train_data[\"session_id\"].nunique()\nn_test_sessions = test_data[\"session_id\"].nunique()\n\nn_of_sessions_df = pd.DataFrame({\"dataset\": [\"train\", \"test\"], \"n_of_sessions\": [n_train_sessions, n_test_sessions]})\n\nfig = px.bar(n_of_sessions_df, x='dataset', y='n_of_sessions', text_auto=True, title=\"Number of sessions in datasets\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:26:54.938634Z","iopub.execute_input":"2023-02-09T00:26:54.939024Z","iopub.status.idle":"2023-02-09T00:26:55.650079Z","shell.execute_reply.started":"2023-02-09T00:26:54.938989Z","shell.execute_reply":"2023-02-09T00:26:55.649098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It's hard to see, but there are only 3 sessions in the test dataset🤪. So, the public leaderboard may be not very representative.\n\nBut in a train, we have a lot of sessions, and that's good.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"5.2\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">5.2 Number of events in sessions</p>","metadata":{}},{"cell_type":"markdown","source":"Now let's look at number of events in each session","metadata":{}},{"cell_type":"code","source":"events_in_session = pd.DataFrame(train_data.groupby('session_id').size())\nevents_in_session.columns = [\"n_of_events\"]","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:26:55.652928Z","iopub.execute_input":"2023-02-09T00:26:55.653268Z","iopub.status.idle":"2023-02-09T00:26:55.894360Z","shell.execute_reply.started":"2023-02-09T00:26:55.653236Z","shell.execute_reply":"2023-02-09T00:26:55.893205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.histogram(events_in_session, x=\"n_of_events\", title=\"Types of events in dataset\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:26:55.895861Z","iopub.execute_input":"2023-02-09T00:26:55.896184Z","iopub.status.idle":"2023-02-09T00:26:55.968581Z","shell.execute_reply.started":"2023-02-09T00:26:55.896153Z","shell.execute_reply":"2023-02-09T00:26:55.967529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, most sessions have around 1000 events, but there are some unique cases, where sessions have up to 19k events","metadata":{}},{"cell_type":"markdown","source":"<a id=\"5.3\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">5.3 Types of events in dataset</p>","metadata":{}},{"cell_type":"code","source":"fig = px.bar(train_data[\"event_name\"].value_counts(), text_auto=True, title=\"Types of events in dataset\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:26:55.969986Z","iopub.execute_input":"2023-02-09T00:26:55.970544Z","iopub.status.idle":"2023-02-09T00:26:56.110956Z","shell.execute_reply.started":"2023-02-09T00:26:55.970510Z","shell.execute_reply":"2023-02-09T00:26:56.110086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, the prevalence of navigate click and person click events is observed. Other events occur less often","metadata":{}},{"cell_type":"markdown","source":"<a id=\"5.4\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">5.4 Distribution of rooms</p>","metadata":{}},{"cell_type":"code","source":"fig = px.bar(train_data[\"room_fqid\"].value_counts(), text_auto=True, title=\"Number of rooms\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:26:56.112219Z","iopub.execute_input":"2023-02-09T00:26:56.112718Z","iopub.status.idle":"2023-02-09T00:26:56.247912Z","shell.execute_reply.started":"2023-02-09T00:26:56.112685Z","shell.execute_reply":"2023-02-09T00:26:56.246989Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, there are 19 rooms in the training set and they all are in use. Some are more popular, others less. But it still looks like all rooms are important","metadata":{}},{"cell_type":"markdown","source":"<a id=\"5.5\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">5.5 How many events for each level group</p>","metadata":{}},{"cell_type":"markdown","source":"Because we will produce predictions for questions answering, using all data of user groups, it's essential to us to know, how many events will occur in each level group. Let's figure that out","metadata":{}},{"cell_type":"code","source":"events_in_level_group = pd.DataFrame(train_data.groupby(['session_id', 'level_group']).size().reset_index(name='n_of_events'))","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:26:56.249203Z","iopub.execute_input":"2023-02-09T00:26:56.249745Z","iopub.status.idle":"2023-02-09T00:26:56.783585Z","shell.execute_reply.started":"2023-02-09T00:26:56.249709Z","shell.execute_reply":"2023-02-09T00:26:56.782643Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.histogram(events_in_level_group, x=\"n_of_events\", color=\"level_group\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:26:56.784935Z","iopub.execute_input":"2023-02-09T00:26:56.785463Z","iopub.status.idle":"2023-02-09T00:26:56.866608Z","shell.execute_reply.started":"2023-02-09T00:26:56.785428Z","shell.execute_reply":"2023-02-09T00:26:56.865731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we observe, there are much fewer actions needed to get done with the 0-4 level group. Another requires mostly around 500 events to end with this level group. There are some outliers, but there are not so important.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"5.6\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">5.6 Distribution of the event name by level</p>","metadata":{}},{"cell_type":"markdown","source":"Now we can look, at what levels events occur more often","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_for_eda, x='event_name', color='level', title=\"Distribution of the event name by level (undersampled)\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:26:56.867542Z","iopub.execute_input":"2023-02-09T00:26:56.867864Z","iopub.status.idle":"2023-02-09T00:27:00.998519Z","shell.execute_reply.started":"2023-02-09T00:26:56.867825Z","shell.execute_reply":"2023-02-09T00:27:00.997408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"5.7\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">5.7 Number of words in text of events</p>","metadata":{}},{"cell_type":"code","source":"n_of_words = train_for_eda[\"text\"].apply(lambda x: len(x.split()))","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:27:00.999700Z","iopub.execute_input":"2023-02-09T00:27:01.000254Z","iopub.status.idle":"2023-02-09T00:27:01.016382Z","shell.execute_reply.started":"2023-02-09T00:27:01.000216Z","shell.execute_reply":"2023-02-09T00:27:01.015459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.histogram(n_of_words, title=\"Number of words in events (undersampled)\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:27:01.017662Z","iopub.execute_input":"2023-02-09T00:27:01.018198Z","iopub.status.idle":"2023-02-09T00:27:01.417856Z","shell.execute_reply.started":"2023-02-09T00:27:01.018163Z","shell.execute_reply":"2023-02-09T00:27:01.416696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, there are not so many words in the text field. So, we don't need to use a model, capable of working with big texts","metadata":{}},{"cell_type":"markdown","source":"<a id=\"5.8\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">5.8 What about music?</p>","metadata":{}},{"cell_type":"markdown","source":"At last, let's discover, how many events was with music","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_data[\"music\"], title=\"Number of events with music\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:27:01.419151Z","iopub.execute_input":"2023-02-09T00:27:01.419779Z","iopub.status.idle":"2023-02-09T00:27:05.202186Z","shell.execute_reply.started":"2023-02-09T00:27:01.419726Z","shell.execute_reply":"2023-02-09T00:27:05.201181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Wait, is there no music?","metadata":{}},{"cell_type":"code","source":"train_data[\"music\"].count()","metadata":{"execution":{"iopub.status.busy":"2023-02-09T00:29:05.165506Z","iopub.execute_input":"2023-02-09T00:29:05.166082Z","iopub.status.idle":"2023-02-09T00:29:05.201564Z","shell.execute_reply.started":"2023-02-09T00:29:05.166040Z","shell.execute_reply":"2023-02-09T00:29:05.200447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Yes, there is no music information in training set. So, I will not use this feature while modeling, because it seems unuseful","metadata":{}},{"cell_type":"markdown","source":"<a id=\"6\"></a>\n## <p style=\"padding:10px;background-color:lightskyblue;margin:0;color:black;font-family:calibri;font-size:120%;text-align:center;border-radius: 25px 25px;overflow:hidden;font-weight:500\">6. Conclusion</p>","metadata":{}},{"cell_type":"markdown","source":"So, in this notebook, we discovered this dataset and perform some starting EDA on this data. Now we know, that:\n  - The main goal of this competition is to predict, will player answer the question of the level\n  - To make a prediction, we have sessions of the users, and each session has many events in it\n  - There are up to 12k sessions and 13 million events in the training set\n  - There are not so much text in the events, up to 17 words\n  - navigate_click and person_click events are the most often events","metadata":{}},{"cell_type":"markdown","source":"**This notebook is under construction and I'm working on the updates. Stay tuned**\n\n**Upvote if you want to support me and if this notebook was useful for you. It's very important for me**","metadata":{}}]}