{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"text-align: center;\"><b>EDA Simplified:<span style=\"color:\n#6d6d6d;\"> Predicting Student Performance from Game Play</span></b></h1>","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #6d6d6d; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Introduction</h2>\n\nRegarding about the education on students around the world, they learned about everything they were teached based on the subjects: math, science, english, and on top of that, history. Speaking of that subject, there's one \"point-and-click\" game that teaches history about the United States, which is \"Jo Wilder and the Capitol Case\", which shows how a girl uncovers the real stories behind mysterious artifacts from two movements in Wisconsin State history. And with reference to this game, a lot of students ranging from grades 3-6 started playing and learning the facts from that game, leading to that competition created by the Learning Agency Lab. The whole point behind mentioning the \"Jo Wilder\" game is that almost all of the game-based learning platforms insufficiently make use of knowledge tracing to support self-learning students, as they don't know how to retain that information they learned while playing any specific educational game. And as for knowledge tracing, they have been developed as well as studied in learning environments and online tutoring, however there's less focus of knowledge tracing in all of the educational games. Therefore, Field Day Lab, the creator of the Jo Wilder game teamed up with the Learning Agency Lab for researching the trace from the student learning in educational games, which they covered the Jo Wilder game. And with that, let's blast into EDA for tracing the student performances!\n\n<center>\n    <img src=\"https://dpi.wi.gov/sites/default/files/imce/news/dpi-connected/Jo_Wilder_Game.png\" width=500>\n    <figcaption style=\"color: gray;\">A scene taken from the gameplay of Jo Wilder and the Capitol Case. Note that the creator of this game, Field Day Lab, teamed up with the Learning Agency Lab over tracing student learning.</figcaption>\n</center>","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #6d6d6d; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Imports + Setup</h2>\n\nWhen we start our journey of visualizing the data here, we must import our necessary modules as well as creating our dataframes. So with that being said, we import the pandas module as pd along with the numpy module as np first for data science stuff like dataframe creation and linear algebra. Next, for importing the plotting modules, we import the plotly module with the express attribute as px together with the altair module as alt, thus disabling their MaxRowsError by calling the disable_max_rows function from the alt module's data_transformers attribute, so that we can plot down interactive graphs about the data we visualized.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nimport plotly.express as px\nimport altair as alt\nalt.data_transformers.disable_max_rows()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once we imported the modules, let's create our dataframes! We characterize the train_labels_df dataframe, alongside with the train_df dataframe to reading out the csv files to the competition data, but as for reading out the train.csv file from the data, we configure the nrows parameter to 10000 so that we'll not crash our notebook. After we created our two dataframes, we display their first five rows with the head function.","metadata":{}},{"cell_type":"code","source":"train_labels_df = pd.read_csv(\"/kaggle/input/predict-student-performance-from-game-play/train_labels.csv\")\ntrain_df = pd.read_csv(\"/kaggle/input/predict-student-performance-from-game-play/train.csv\", nrows=10000)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels_df.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style=\"background-color: #6d6d6d; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 8px; border-radius: 15px;\">Basic Dataframe Analysis</h3>\n\nNow that we have our dataframes loaded for our analysis, let's visualize the basics of the two dataframes! First, let's find out the number of data in two dataframes with the len function.","metadata":{}},{"cell_type":"code","source":"len(train_labels_df)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(train_df)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As for the train_labels_df dataframe, we found out that there are 212022 data entities in there, while we saw 10000 data entities in the train_df dataframe because we truncated the number of data entities in this dataframe. In other words, the interface of the Jo Wilder game makes the data in the two dataframes complex and huge.","metadata":{}},{"cell_type":"markdown","source":"Now let's find the number of NaN values that is present in the two dataframes! We use each dataframe to find the NaN values with the isna function along with summing them up with the sum function.","metadata":{}},{"cell_type":"code","source":"train_labels_df.isna().sum()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.isna().sum()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"While we saw no NaN values in the train_labels_df dataframe, we noticed a lot of NaN values in the train_df dataframe. Specifically for the NaN values from the train_df dataframe, there are 9732 of them in the page column, 1170 in the room_coor_x, room_coor_y, screen_coor_x, and screen_coor_y columns, 8856 in hover_duration column, 6617 in the text column, 3184 in the fqid column, 6617 in the text_fqid column, and all 10000 columns in the specified columns of fullscreen, hq, and music. Additionally, the number of NaN data present in the train_df dataframe hinted us that most of the NaN values came from the irrelevant data gathered in some interfaces of the Jo Wilder game.","metadata":{}},{"cell_type":"markdown","source":"Finally, let's visualize the number of columns in both dataframes! We use each dataframe out of the two to find the shape of it with the shape attribute, setting the slice index of 1 to pull out the last slice of the shape.","metadata":{}},{"cell_type":"code","source":"train_labels_df.shape[1]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.shape[1]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we ran our two code cells above, we counted two columns in the train_labels_df dataframe, as well as counting the twenty columns in the train_df dataframe. In other words, the two columns specified the labels in the train_labels_df dataframe whereas the twenty columns specified the complex game interactions in the Jo Wilder game.","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #6d6d6d; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">The train_labels_df dataframe Analysis</h2>","metadata":{}},{"cell_type":"markdown","source":"Once we visualized the basic parts of our two dataframes, let's take a quick visualization of the train_labels_df dataframe! First, let's detail the columns in this dataframe.\n* **session_id**: ID specified for each session\n* **correct**: Whether the student answered each task correct.\n\nWithout a doubt, let's visualize the correct column distribution into Altair's histogram! We use the alt module's Chart function to configure our graph, setting the train_labels_df dataframe as our data for the graph, then mark out the bars to our graph with the mark_bar function, as well as encoding the configuration of the bars in the histogram with the encode function, setting the alt module's X function inside of it, which contains the correct column and the bin parameter being set to True for displaying the bins, and the y parameter to the count function in strings.","metadata":{}},{"cell_type":"code","source":"alt.Chart(train_labels_df).mark_bar().encode(\n    alt.X(\"correct\", bin=True),\n    y=\"count()\"\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From our histogram we've created, we found out that most of the students who played the Jo Wilder game got the tasks correct (indicated as 1) than those who got them incorrect (indicated as 0). In other words, the highest data counted is the ones, which there are almost 150,000 data entities counted, and the lowest is the zeroes, in which there are approximatedly 60,150 data entities counted. Specifically, the number of zeroes and ones in the correct column from the train_labels_df dataframe shows that most students understood their knowledge of history once they played as Jo Wilder, trying to solve the clues from the mysterious artifacts she spotted.","metadata":{}},{"cell_type":"markdown","source":"And with just our data distribution based on the correct column, we've completed our short data analysis of the train_labels_df dataframe! Now, let's look forward to propel into our data visualization in the train_df dataframe.","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #6d6d6d; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">The train_df Dataframe Analysis</h2>\n\nAs we head into our analysis in the train_df dataframe, here's what we need to know about the data in each column!\n* **session_id**: ID specified in each session.\n* **index**: the index of each event session.\n* **elapsed_time**: the overall time between the starting point of the session and when the event was recorded.\n* **event_name**: name specified of each event type.\n* **name**: name specified for the event name.\n* **level**: specified for the event occurances in levels.\n* **page**: the page number of the event.\n* **room_coor_x**, **room_coor_y**: the click coordinates specified to the in-game room.\n* **screen_coor_x**, **screen_coor_y**: the click coordinates specified to the player's screen.\n* **hover_duration**: time elapsed for how long the hover happened for.\n* **text**: the text players see during the specified event.\n* **fqid**: the fully qualified ID for the event, hence for its short name.\n* **room_fqid**: the fully qualified ID for each room in the game.\n* **text_fqid**: same for the room_fqid, but it specified the text.\n* **fullscreen**: whether the player is in fullscreen mode. (Considered NULL)\n* **hq**: whether the player plays the game in high quality. (Considered NULL)\n* **music**: whether the music in the game is on or off. (Considered NULL)\n* **level_group**: which group of levels (and questions) this row belongs to, in range.\n\nWith all of that 20 columns being listed from the train_df dataframe, let's begin our data analysis in this section!","metadata":{}},{"cell_type":"markdown","source":"First, let's visualize the elapsed_time data column distribution into our subplot that includes a histogram and a box-plot with Altair! We characterize our first variable one, to use the alt module's Chart function to indicate that we're creating our graph, setting the train_df dataframe as our data for the graph, then we configure our bar chart into our chart configuration with the mark_bar function, as well as arranging the setup of it with the encode function, setting the alt module's X function that contains the elapsed_time column as well as the bin parameter configured to True for specifying the x-axes of the histogram, along with setting the y parameter to the count function encased in strings for counting the values from the specified parameter in the alt module's X function. \n\nAfter that, we define another variable, two, to the same setup as variable one with the alt module's Chart function for configuring another chart, then we generate our boxplot with the mark_boxplot function, setting the extent parameter to min-max for making our box chart extend from the minimum and maximum of the specified data, and then encode the arrangement of the boxplot graph with the encode function, setting the x parameter to the elapsed_time column for configuring the x-axes of the boxplot, thus applying the properties function to configure the properties of our boxplot, setting the height parameter to 300 for configuring the height of our diagram. With our variables one and two characterize, we use the alt module's concat function for placing the two graphs into our subplot, setting the two graph variables one and two into it.","metadata":{}},{"cell_type":"code","source":"one = alt.Chart(train_df).mark_bar().encode(\n    alt.X(\"elapsed_time\", bin=True),\n    y=\"count()\"\n)\n\ntwo = alt.Chart(train_df).mark_boxplot(extent=\"min-max\").encode(\n    x=\"elapsed_time\",\n).properties(height=300)\n\nalt.concat(one, two)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the histogram we created, we found out that there's a skew to the right of the diagram, as we understood that there's a lot of data in the left side of the histogram. Nevertheless, the highest number of data in ranges counted in the histogram is from 0 to 500000, with approximately 2990 entities caluclated, while the lowest is between 4M and 4.5M (M indicates million), as there are almost 250 data entities listed in that range. \n\nOn the other graph, we noticed how the box part in the boxplot diagram was shifted to the left just like the histogram, as we saw how the distance to the maximum data of the elapsed_time column is longer than the distance from the box part to the minimum data of the elapsed_time column. In other words, the median of the elapsed_time column data is 1062161, the first quartile is 459932, the third quatile is 2073021.5, and the interquartile range is 1613089.5. Furthermore, the right-skew display in the histogram and box diagrams based on the elapsed_time column shows us that most players in the Jo Wilder game spend less time playing this game, as they possibly gave up on finding the clues in each level they're on or unable to apply critical thinking and historical inquiry towards learning about the two movements in the Wisconsin state history.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the event_name column data into a pie chart with Plotly! Before we begin this visualization, we create our dataframe, event_name, to count the values from the train_df dataframe's event_name data column with the value_counts function. After that, we characterize our fig variable figure to create our pie chart with the px module's pie function, setting the event_name dataframe as our data for the pie chart along with the names parameter to the event_name dataframe's indexes specified by the index attribute for configuring the labels for the pie chart and the values parameter to the event_name dataframe's values specified by the values attribute for configuring the values to the pie chart. With all of that completed, we use the show function to the fig variable figure to display our pie chart in the notebook outputs.","metadata":{}},{"cell_type":"code","source":"event_names = train_df[\"event_name\"].value_counts()\n\nfig = px.pie(event_names, names=event_names.index, values=event_names.values)\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From our pie chart we compiled, we realized that 41.8 percent of the data in the event_name column is navigate_click, while 20.9 percent of the data is person_click, 9.68 percent is cutscene_click, 7.88 percent is object_click, 6.11 percent is map_hover, 5.33 percent is notebook_click, 2.68 percent is notification_click, 2.14 percent is map_click, 0.84 percent is observation_click, and 0.26 percent is checkpoint. Nevertheless, the highest data counted is navigate_click, with 4177 data entities counted, while checkpoint is counted the least is checkpoint, with just only 26 entities counted. Additionally, the navigate_click data in the event_name column shows us that most students who played the Jo Wilder game mostly used this event for navigating the main character they're playing for obtaining objects for finding clues, or talk to the other characters that are present in the game.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the name column data into the bar chart that is created with Altair! We use the alt module's Chart function to create our graph, setting the train_df dataframe as our data for the graph, then applying the mark_bar function for indicating that we're creating a bar chart as well as configuring the settings of the bar chart with the encode function, setting the x parameter to the name data column for placing the x-axis to the bar chart and the y parameter to the count function encased in strings for counting the values of the name data column to encode it to the y-axes, thus setting the properties function to change our layout of our graph in which we set the width parameter to 500 for widen our graph with the width in size.","metadata":{}},{"cell_type":"code","source":"alt.Chart(train_df).mark_bar().encode(\n    x=\"name\",\n    y=\"count()\"\n).properties(width=500)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the bar chart we created out from Altair, we visualized that the undefined data in the name column is counted the most than the other data inside this column. In other words, the data undefined has approximately 4850 entities, while basic has around 4750 entities, close has estimated 250 entities, open has around 100 entities, prev has around 50 entities, and next has 10-ish entities. In addition, the highest number of data that was marked as \"undefined\" shows us that there are unclear events that cannot be identified while the students navigate Jo Wilder in the game, while the close and open occurred sometimes when students open or close the notebook Jo Wilder recorded it down followed by prev and next data that occurred rarely, since most students didn't complete the stages in the gameplay.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the level distribution with a boxplot and the histogram combined into one diagram with Plotly! We characterize the fig variable figure to the px module's histogram function for creating our histogram, placing the train_df dataframe as the histogram's data, as well as configuring the x parameter to the level column for placing the histogram's x-axis, and the marginal parameter to box for placing the boxplot graph into our histogram diagram. Finally, we apply the show function to the fig variable figure so that our graph is displayed in the notebook output.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_df, x=\"level\", marginal=\"box\")\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see from the graph that helds the histogram and the boxplot all at once, we noticed that the bins in the histogram plot showed a uniform distribution, as there are peaks of high data seen from the left and right of the histogram. Nevertheless, level 18 is counted as the most, with 1452 entities counted while level 12 is counted as the least, as it contained 93 data entities. Meanwhile in the box plot, we noticed that the box part inside the boxplot was skewed to the right, as we spotted that the distance from the data's lower fence is longer than the distance to the upper fence. In other words, the first quartile is 7, the median is 14, the third quartile is 18, and the interquartile range is 11. Specifically, the data distribution based out of the level column from the train_df dataframe hinted us that most players in the Jo Wilder game landed in variate levels, as some were in the beginning, while others landed in the middle or in the end of all of the levels.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the page data column into Altair's histogram and boxplot together into one subplot again! We characterize our first graph variable, one, to the alt module's Chart function for configuring our graph, setting the train_df dataframe as the graph's data, as well as marking the bars to the graph with the mark_bar function for indicating that we're creating a histogram, as well as configuring the characteristics of it with the encode function, setting the x parameter to the alt module's X function for configuring the x-axes of the histogram in which it contains the page column alongside with the bin parameter being set to True for displaying the bins, and the y parameter to the count function that is encased in strings for counting the data in the specified column for the y-axes. \n\nAfter we configured the variable graph figure one, we characterize another graph variable, two, to same thing as we did for the graph variable, one, but we use the mark_boxplot function for configuring our boxplot, setting the extent parameter to min-max for making our boxplot stretch from the minimum of the data to the maximum of the data, then we configure the characteristics of our boxplot with the encode function, setting the x parameter to the page column for configuring the x-axes of the boxplot we made, as well as setting the properties of our graph with the properties function, setting the height parameter to 300 for configuring the height of our graph. Finally, we place our graphs together with the alt module's concat function, placing the graph variables one and two together.","metadata":{}},{"cell_type":"code","source":"one = alt.Chart(train_df).mark_bar().encode(\n    x=alt.X(\"page\", bin=True),\n    y=\"count()\"\n)\n\ntwo = alt.Chart(train_df).mark_boxplot(extent=\"min-max\").encode(\n    x=\"page\",\n).properties(height=300)\n\nalt.concat(one, two)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From what we saw in the histogram and the box-plot charts, we visualized that the histogram graph created a mostly bimodal display based on the page column's data we distributed. In other words, the highest number of data listed in the histogram is between 5 and 6, with almost 88 data entities counted, while the lowest is under the range from 2 to 3, with nearly 18 data entities listed. Meanwhile for the box plox, we found out that the box part is in the middle of the graph, as the distance from the box part to the minimum and maximum of the page column is equally the same. Nevertheless, the first quartile of the boxplot based on the page column data is 1, the median is 3, the third quartile is 5, and the interquartile range is 4. Additionally, the most data under the range between 5 and 6 in the page data pointed out to us that the students that played the Jo Wilder game landed on variate page numbers of the events, as some landed in the lower page numbers, while others landed on the higher page numbers.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the room_coor_x and room_coor_y columns into each two histograms with a box plot separately with Plotly! In each code cell, we characterize the fig variable figure to the px module's histogram function to create our histogram graph, setting the train_df dataframe as the histogram's data input, alongside with the x parameter to the room_coor_x and room_coor_y columns separately, and the marginal parameter to box for placing the boxplot our histogram. Finally, we display off our fig variables separately with the show function that is plugged to the fig graph variable.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_df, x=\"room_coor_x\", marginal=\"box\")\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.histogram(train_df, x=\"room_coor_y\", marginal=\"box\")\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the room_coor_x histogram distribution, we envisaged that the bins somehow showed a right skew, as we see the tall peak of data on the right of our diagram we compiled. Specifically, the highest range of data counted in this histogram is from -50 to 0, with 709 data entities listed, while the lowest range of data counted is between -920 to -900, with only one data entity listed. And for the box-plot graph based on the room_coor_x column on top of the room_coor_x histogram data, we spotted there are outliers on the left of the plot, as we saw how they were clamped up together thus we noticed that the size of the first and third quartiles from the median were merely the same. Funnily enough, the first quartile of the room_coor_x data in the boxplot is -353.5411, the median is -17.3485, the third quartile is 304.2546, and the interquartile range is 657.7957.\n\nOn the other hand, we noticed that the room_coor_y data distribution in another histogram showed mostly a multimodal display on the mostly middle right in the diagram. In other words, the highest data range in the histogram is from -120 to -100, with 405 data entities counted, while the data from -920 to -900 is counted the least. As for the boxplot, we glimpsed that the outliers were clamped up together as well as seeing the mostly symmetrical display of the box part just like the one we saw in the room_coor_x boxplot, but however we noticed that the stuffed outliers were plotted both in the left and right of the diagram. Furthermore, the first quartile of the room_coor_y data is -218.0651, the median is -17.3485, the third quartile is 24, and the interquartile range is 242.0651. And as for the distributions and visualizations we saw for the room_coor_x and room_coor_y, we revealed that there's a lot of variation of the room coordinates in the Jo Wilder game, as there were some rooms that guided the players to the clues of each object, while others guided them to interact with other characters in the game.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the plots from room_coor_x and room_coor_y columns by the name and level columns separately into Altair's scatter plot! We configure our chart with the alt module's Chart function in which we plug-in the train_df dataframe we created, then we characterize the scatterplot into our graph by using the mark_circle function, setting the size parameter to 60 for adjusting the area of circle in the scatterplot, and then we configure our chart's characteristics with the encode function, setting the x parameter to the room_coor_x column for placing the x-axis for the graph, the y parameter to the room_coor_y column for placing the y-axis for the graph, the color parameter to the name and level columns separately for indicating the color legend for the scatterplot, and the tooltop parameter to the list containing the room_coor_x and room_coor_y columns as well as the individual name and level columns for enabling the tooltip of the scatterplot. Thus, on the outside of the chart configuration since we arranged the parameters in the encode function, we plug in the interactive function to make our scatterplot graph be interactive.","metadata":{}},{"cell_type":"code","source":"alt.Chart(train_df).mark_circle(size=60).encode(\n    x=\"room_coor_x\",\n    y=\"room_coor_y\",\n    color=\"name\",\n    tooltip=[\"room_coor_x\", \"room_coor_y\", \"name\"]\n).interactive()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"alt.Chart(train_df).mark_circle(size=60).encode(\n    x=\"room_coor_x\",\n    y=\"room_coor_y\",\n    color=\"level\",\n    tooltip=[\"room_coor_x\", \"room_coor_y\", \"level\"]\n).interactive()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the two scatterplots consisting of the room_coor_x and room_coor_y columns by the name and level columns in color, we noticed that the plots together displayed the outline of the capitol that Jo Wilder was in. Not only that, we saw most of the plots were marked as undefined, while some of the other plots were marked as basic, close, and open. And as for the scatter plot that has the level column color scale, we found out that most plots were light-colored on the top-middle portion of our graph as it was ranged from 0 to 10, while we spotted dark-colored plots in the bottom of the graph as it was ranged from 18 through 20-ish. In other words, most plots that were marked as unknown in the name column color legend indicates that there are unclear events that cannot be identified while the students navigate Jo Wilder in the game as mentioned previously in the name column distribution in a bar chart, while the scattered light and dark colored plots in another scatterplot graph displayed variations of the level data since some of the Jo Wilder users played on the beginning levels, while others played on the middle or end levels.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the screen_coor_x and screen_coor_y data columns into the histogram and box-plot with Plotly! To get us weaving to plot this data out, we generate the one and only fig variable to creating our histogram model with the px module's histogram function, inputting the train_df dataframe as the data for the histogram, alongside with arranging the x parameter to the screen_coor_x and screen_coor_y separately and the marginal parameter to box so that we'll plot the box chart on the top of the histogram. Last but not least, we present out the two fig variable's combined histogram and box-plot chart based on screen_coor_x and screen_coor_y columns separately with the show function.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_df, x=\"screen_coor_x\", marginal=\"box\")\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.histogram(train_df, x=\"screen_coor_y\", marginal=\"box\")\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the very top histogram which is based on the screen_coor_x data, we noticed that there are a lot a peaked data that were close together on the left part of the diagram, indicating this as a right-skew display. Aside from that, the range from 360 to 379 is counted the most with 293 units of data summed together, while the range between 1220 to 1239 is the least counted data, with only 5 entities. And as for the box-plot, we spotted a cluster of outliers on the far right, as we didn't see any outliers on the left of the boxplot. Specifically, the first and third quartiles of the screen_coor_x data specified by the box-plot is 286 and 694, the median is 459, and the interquartile range is 408.\n\nMeanwhile on the bottom histogram which consists over the screen_coor_y data distribution, we found out that the peaks of data were placed in the middle left of the histogram, as it nearly showed a symmetrical display. Funnily enough, the highest number of data in the screen_coor_y data histogram is from 430 to 439, with 311 entities counted, while the ranges 950 to 959 and 1050 to 1059 was counted the least since there's one entity seen on the two data ranges specified. And for the box-plot, we spotted a few clusters of outliers on the left of the boxplot but also we glimpsed a long string of clustered outliers with a break in between. Thus, we noticed that the data distance from first quartile to median is longer than the data distance from median to the third quartile, making the box portion asymmetrical. Other than the boxplot's appearance, the first and third quartiles of this graph is 314 and 500, the median is 415, and the interquartile range is 236. Additionally, the number of data distributed from the screen_coor_x and screen_coor_y columns explains to us that the values in the screen coordinates vary as players use the click coordinates throughout the Jo Wilder game.","metadata":{}},{"cell_type":"markdown","source":"Now let's use Altair for plotting another scatter plot based out of the screen_coor_x and screen_coor_y columns! We use the alt module's Chart function to plot down our chart, setting the train_df dataframe as the data for plotting the graph, then we plug the mark_circle function to create our scatter plot as we set the size parameter to 60 for configuring the scatter plot size, as well as applying the encode function to encode our scatter plot, setting the x parameter to the screen_coor_x column for specifying the graph's x-axis, the y parameter to the screen_coor_y column for specifying the graph's y-axis, the color parameter to the name and level columns separately for specifying the scatter plot color legend, and the tooltip parameter to a list containing the three columns specified for enabling the tooltip, as well as setting the graph interactive with the interactive function.","metadata":{}},{"cell_type":"code","source":"alt.Chart(train_df).mark_circle(size=60).encode(\n    x=\"screen_coor_x\",\n    y=\"screen_coor_y\",\n    color=\"name\",\n    tooltip=[\"screen_coor_x\", \"screen_coor_y\", \"name\"]\n).interactive()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"alt.Chart(train_df).mark_circle(size=60).encode(\n    x=\"screen_coor_x\",\n    y=\"screen_coor_y\",\n    color=\"level\",\n    tooltip=[\"screen_coor_x\", \"screen_coor_y\", \"level\"]\n).interactive()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we plotted two scatter plots based out of the data from the screen_coor_x and screen_coor_y columns by the color indication of name and level columns, we found out that the plots graphed in the diagram resembled an outline based on the Jo Wilder game. In other words, we spotted a lot of plots that labeled basic and undefined thus seeing the tiny clusters of the data plots labeled as prev and close in the first scatter plot consisting on the name data column, whereas we see a lot of varieties of light and dark plots clustered together or scattered around the graph. Additionally, the plots scattered and clustered around the graph based out from plotting down the screen_coor_x and screen_coor_y columns gave us a clue that almost all players who played the Jo Wilder game clicked everywhere around the game screen, as some navigated Jo Wilder to talk to other characters and obtain items, while others open or close their notebook Jo Wilder equips.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the fqid data into a bar chart with Plotly! Before we begin this plotting process, we create the fqid dataframe to count the train_df dataframe's fqid data column with the value_counts function. Following from that, we generate the fig variable to the px module's bar function for creating our bar chart, setting the fqid dataframe as the data for the bar chart, the x and color parameters to the fqid dataframe's indexes with the index attribute and the y parameter to the fqid dataframe's values with the values attribute. With that completed, we use the show function to the fig varaible so that we display the graph on the notebook output cell.","metadata":{}},{"cell_type":"code","source":"fqid = train_df[\"fqid\"].value_counts()\n\nfig = px.bar(fqid, x=fqid.index, y=fqid.values, color=fqid.index)\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From this bar chart we displayed above, we visualized the thin, colorful bars that were high on the left and then getting lower towards the right of the graph. In other words, the data labeled as \"worker\" is counted the highest with 631 data entities, while the data marked as \"block_0\" is counted the lowest with 2 data entities. In specific way of explaining this fqid data distribution, the most data that was marked as \"worker\", \"gramps\", and other 9 data labels hinted that most players who played as Jo Wilder interacted with the workers or her gramps inside the capitol, while the other data that is counted the least shows how players sometimes navigate Jo Wilder to find objects that was scattered around the capitol or interact with her notebook that she equipped.","metadata":{}},{"cell_type":"markdown","source":"Let's now distribute the hover_duration into Altair's histogram and box chart at the same time! First of all, we characterize the fig1 variable to use the alt module's Chart function to create our Altair chart, placing the train_df dataframe as the graph's data, then we mark the bars of the graphs with the mark_bar function, and then we configure the graph's characteristics with the encode function, setting the x-axis configuration with the alt module's X function that contains the hover_duration column as well as the bin parameter set to True for creating a histogram, along with setting the y parameter to the count function encased in strings for counting the values of a specific column.\n\nThenceforth, we define another variable, fig2 to create another chart with the alt module's Chart function in which the train_df dataframe is placed inside of it for loading the data to the graph, then create a box plot with the mark_boxplot function, setting the extent parameter to min-max for extending the boxplot figure based on the data's minimum and maximum points and then encode the graph's characteristics with the encode function, setting the x parameter to the hover_duration column for specifying the x-axes of the plot, as well as configuring the properties of the graph with the properties function, setting the height parameter to 300 for adjusting the height of the graph. Finally, we assemble the two graphs together into a subplot with the alt module's concat function, setting the graph variables, fig1 and fig2 inside.","metadata":{}},{"cell_type":"code","source":"fig1 = alt.Chart(train_df).mark_bar().encode(\n    alt.X(\"hover_duration\", bin=True),\n    y=\"count()\"\n)\n\nfig2 = alt.Chart(train_df).mark_boxplot(extent=\"min-max\").encode(\n    x=\"hover_duration\",\n).properties(height=300)\n\nalt.concat(fig1, fig2)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the data distribution that was based on the hover_duration column, we glimpsed that there's a data peak on the first left of the histogram, as it showed a right-skew distribution. Aside from the graph's display, the highest counted data range is from 0 to 5000, with nearly 1075 data entities while the range from 15000 to 20000 is counted the least with approximately 10 data entities.\n\nOn the other graph, we could see that the whisker plot from the third quartile to the maximum of the hover_duration data column is longer than the box part and the whisker plot from the minimum of the hover_duration to the first quartile combined thus we noticed that the box part is shifted to the left of the graph, indicating this as a right-skew. In other words, the first and third quartiles specified in the box plot is 77.75 and 903.5, the median is 351, and the 825.75. Furthermore, the right-skew in the histogram as well as the box plot showed us that most players in the Jo Wilder Game hovers some items or other interfaces in less than 5 seconds (that's 5000 ms).","metadata":{}},{"cell_type":"markdown","source":"Let's then proceeed to graph the level_group data into Plotly's pie chart! Beforehand, we create the level_group dataframe to count the values of the train_df dataframe's level_group column with the value_counts function. After that, we create the fig variable to the px module's pie function for generating a pie chart, setting the level_group dataframe as the graph's data, along with the names parameter to the indexes of the level_group dataframe with the index attribute for specifying the pie chart's names and the values parameter to the values of the level_group dataframe with the values attribute for specifying the values of the pie chart. With our graph being finalized, we use the show function to the fig variable figure for displaying the graph to the notebook output.","metadata":{}},{"cell_type":"code","source":"level_group = train_df[\"level_group\"].value_counts()\n\nfig = px.pie(level_group, names=level_group.index, values=level_group.values)\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the pie chart we compiled, we found out that 54.9% of the data in the level_group column were marked as \"13-22\" since it has 5494 entities, while 31.1% of the data were marked as \"5-12\" with 3107 data entities and 14% of the data were labeled as \"0-4\". Specifically, most of the data marked as \"13-22\" hinted us that most questions were belonged to the levels from 13 to 22 in the Jo Wilder game, as they challenged players to find complicated clues about a specific case assigned by the other characters in the game.","metadata":{}},{"cell_type":"markdown","source":"Lastly for this section, we're going to visualize the text data column into a wordcloud with style by stylecloud! Beforehand, we install the stylecloud module with pip first and then we import the stylecloud module followed by the Image module from the IPython module's display attribute.","metadata":{}},{"cell_type":"code","source":"!pip3 install stylecloud\n\nimport stylecloud\nfrom IPython.display import Image","metadata":{"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Following from importing the modules, we characterize the concat_text variable to an empty string with a space that is joined with the join function by the i variable that is looped over the train_df dataframe's text data column attribute in which it was converted to str with the astype function that held the str type inside, which it was encased in a square bracket.\n\nThenceforth, we call out the gen_stylecloud function from the stylecloud module for creating a styled wordcloud, setting the text parameter to the concat_text variable for loading the text to the wordcloud, the icon_name parameter to \"fas fa-comment\" encased in strings for shaping the wordcloud into a specific icon, the palette parameter to a string containing \"scientific.diverging.Roma_17\" for setting the color palette of the wordcloud, the background_color parameter to \"black\" for setting the color background to black, the gradient parameter to \"center\" for adjusting the color gradient of the wordcloud, and the size parameter to 1024 for specifying the width and height of the wordcloud graph.","metadata":{}},{"cell_type":"code","source":"concat_text = ' '.join([i for i in train_df.text.astype(str)])\n\nstylecloud.gen_stylecloud(\n    text=concat_text,\n    icon_name=\"fas fa-comment\",\n    palette=\"scientific.diverging.Roma_17\",\n    background_color=\"black\",\n    gradient=\"center\",\n    size=1024\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once we completed creating our wordcloud, the gen_stylecloud function saves the styled wordcloud as an image that is known as \"stylecloud.png\" in the notebook output files of our notebook, so we call the Image function to display the image that was named as \"stylecloud.png\" from the output folder of our notebook, as well as setting the width and height parameters to 1024 for specifying the width and height display of a specific image.","metadata":{}},{"cell_type":"code","source":"Image(filename=\"./stylecloud.png\", width=1024, height=1024)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the wordcloud we generated, we visualized a lot of concatenated words that formed a message bubble icon from the text data column in the train_df dataframe, ranging from big words to little words blended in the specific color gradient we configured. Specifically, one of the common words we glimpsed in the train_df dataframe's text column is \"Basketball jersey\", \"Nan nan\", \"Teddy\", and \"Capitol\". Additionally, the commom words seen in our styled wordcloud graph hinted us that most of the text were from the history references in Wisconsin or from the character's dialogue in the Jo Wilder game, as one of the words, \"basketball jersey\", were used in the beginning of the game based on searching the clues about the misconceptions based out of women's basketball history.","metadata":{}},{"cell_type":"markdown","source":"After a long analysis based on most of the 20 columns, we finally complete our chapter of going through the data of the train_df dataframe despite skipping a few columns because of all the Nan values as well as some complicated data to plot.","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #6d6d6d; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Conclusion</h2>\n\nBased on the data we visualized in the train_df and the train_labels_df dataframes, we inferred that the data used in the competition contained the questions that were present in the Jo Wilder game for predicting whether the player answered it correctly based on the clues Jo Wilder found and recorded around the capitol based on the student learning. And following from predicting the student's performance to all 18 questions correctly, we will make game developers to amplify educational games and further support the teachers around every school as well as foreshadow broader support for game-based learning platforms, which we are going to visualize something outside [Jo Wilder's horizons](https://www.kaggle.com/code/dinowun/eda-simplified-isolated-asl-recognition).","metadata":{}}]}