{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"While working on my notebook that explores text aspects of the dataset, I discovered the concept of undefined values in addition to NaN values. What is the difference between these two types of missing values, and how should they be treated in data analysis? I will attempt to answer these questions, which are essential for proper data interpretation. A brief search revealed that no notebooks specifically address this issue.","metadata":{}},{"cell_type":"code","source":"import pandas as pd","metadata":{"execution":{"iopub.status.busy":"2023-03-31T04:46:24.499656Z","iopub.execute_input":"2023-03-31T04:46:24.500086Z","iopub.status.idle":"2023-03-31T04:46:24.505795Z","shell.execute_reply.started":"2023-03-31T04:46:24.500050Z","shell.execute_reply":"2023-03-31T04:46:24.504580Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Only the \"text\" and \"event_name\" fields are required, and they should be cast to more compact data types.","metadata":{}},{"cell_type":"code","source":"dtypes={'event_name':'category', 'text':'category'}\nuse_columns = ['event_name', 'text']\ndf = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv', dtype=dtypes, usecols=use_columns)","metadata":{"execution":{"iopub.status.busy":"2023-03-31T04:46:24.830643Z","iopub.execute_input":"2023-03-31T04:46:24.831057Z","iopub.status.idle":"2023-03-31T04:46:59.439793Z","shell.execute_reply.started":"2023-03-31T04:46:24.831020Z","shell.execute_reply":"2023-03-31T04:46:59.438457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['text'].value_counts(dropna=False).head()","metadata":{"execution":{"iopub.status.busy":"2023-03-31T04:37:05.769848Z","iopub.status.idle":"2023-03-31T04:37:05.770456Z","shell.execute_reply.started":"2023-03-31T04:37:05.770253Z","shell.execute_reply":"2023-03-31T04:37:05.770275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" These two values are the most common in the dataset, but what distinguishes them?\n\nThe main difference between NaN and undefined values in a text column is that NaN value indicates an event that doesn't have text in general (like navigation_click or notebook_click) or should have text but is missing. On the other hand, undefined values are reserved by developers and indicate intermediate clicks in cutscenes that have text.\n\nTo prove it I found some information on official game GitHub [page](https://github.com/fielddaylab/jo_wilder/tree/5082da4057f30dd0917c97a65f2aa7be13469f79), the picture is below. \n\n<img src=\"https://i.ibb.co/fvm4DWt/9.png\" width=40%>","metadata":{}},{"cell_type":"markdown","source":"We can also explore the dataset in more detail.","metadata":{}},{"cell_type":"code","source":"df[df['text'] == 'undefined']['event_name'].unique()","metadata":{"execution":{"iopub.status.busy":"2023-03-31T04:37:05.771578Z","iopub.status.idle":"2023-03-31T04:37:05.772203Z","shell.execute_reply.started":"2023-03-31T04:37:05.771987Z","shell.execute_reply":"2023-03-31T04:37:05.772009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df[df['text'].isna()]['event_name'].unique()","metadata":{"execution":{"iopub.status.busy":"2023-03-31T04:37:05.773319Z","iopub.status.idle":"2023-03-31T04:37:05.773934Z","shell.execute_reply.started":"2023-03-31T04:37:05.773732Z","shell.execute_reply":"2023-03-31T04:37:05.773755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's examine the sequence of events in which undefined values appear.","metadata":{}},{"cell_type":"code","source":"df.iloc[174:185]","metadata":{"execution":{"iopub.status.busy":"2023-03-31T04:37:12.364560Z","iopub.execute_input":"2023-03-31T04:37:12.365004Z","iopub.status.idle":"2023-03-31T04:37:12.377482Z","shell.execute_reply.started":"2023-03-31T04:37:12.364959Z","shell.execute_reply":"2023-03-31T04:37:12.376463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.iloc[574:585]","metadata":{"execution":{"iopub.status.busy":"2023-03-31T04:37:12.565527Z","iopub.execute_input":"2023-03-31T04:37:12.566994Z","iopub.status.idle":"2023-03-31T04:37:12.578279Z","shell.execute_reply.started":"2023-03-31T04:37:12.566936Z","shell.execute_reply":"2023-03-31T04:37:12.577357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In my opinion, these values occur at the beginning of cutscenes and when the character or camera moves within cutscenes. \n\nAs a result, we should remove these values when investigating the text data.\n\nThank you for reading this article. If you have any questions, suggestions, or ideas, please leave them in the comments section","metadata":{}}]}