{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30775,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Missing Values EDA","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"This notebook is to tell you the importance of missing values in this competition. \n### Read this: Mentioned in Data\nThe competition data is compiled into two sources, parquet files containing the accelerometer (actigraphy) series and csv files containing the remaining tabular data. The majority of measures are missing for most participants. In particular, the target sii is missing for a portion of the participants in the training set. You may wish to apply non-supervised learning techniques to this data. The sii value is present for all instances in the test set","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib\nimport matplotlib.pyplot as plt\nimport os\nimport missingno as msno","metadata":{"execution":{"iopub.status.busy":"2024-09-30T16:58:38.289580Z","iopub.execute_input":"2024-09-30T16:58:38.290443Z","iopub.status.idle":"2024-09-30T16:58:38.298833Z","shell.execute_reply.started":"2024-09-30T16:58:38.290389Z","shell.execute_reply":"2024-09-30T16:58:38.297573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train=pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\nsample_actigraphy=pd.read_parquet('/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet/id=00115b9f/part-0.parquet')","metadata":{"execution":{"iopub.status.busy":"2024-09-30T16:58:38.755292Z","iopub.execute_input":"2024-09-30T16:58:38.756494Z","iopub.status.idle":"2024-09-30T16:58:38.824803Z","shell.execute_reply.started":"2024-09-30T16:58:38.756443Z","shell.execute_reply":"2024-09-30T16:58:38.823251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Let's learn about missing values**\n\nI feel missingno library would be helpful to check only what is needed rather than diving too deep at the initial step. Saves some time and effort","metadata":{}},{"cell_type":"code","source":"def read_ids_from_csv(csv_file_path):\n    df = pd.read_csv(csv_file_path)\n    ids = set(df['id'])\n    return ids\n\ndef get_folders_in_directory(directory_path):\n    folders = set(name[3:] for name in os.listdir(directory_path) if os.path.isdir(os.path.join(directory_path, name)))\n    return folders\n\ntrain_ids=read_ids_from_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\nfolders=get_folders_in_directory('/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet/')\nprint (len(set(train_ids)),len(set(folders)))\nmissing_folders = train_ids - folders\nmissing_percentage = (len(missing_folders) / len(train_ids)) * 100\nprint(f\"Percentage of train IDs missing as folders: {missing_percentage:.2f}%\")\n","metadata":{"execution":{"iopub.status.busy":"2024-09-30T16:58:40.526073Z","iopub.execute_input":"2024-09-30T16:58:40.527126Z","iopub.status.idle":"2024-09-30T16:58:41.278249Z","shell.execute_reply.started":"2024-09-30T16:58:40.527076Z","shell.execute_reply":"2024-09-30T16:58:41.276892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## test\nspecific_id='00115b9f'\nif specific_id in folders:\n    print(f\"The folder '{specific_id}' is present in the directory.\")\nelse:\n    print(f\"The folder '{specific_id}' is NOT present in the directory.\")","metadata":{"execution":{"iopub.status.busy":"2024-09-30T16:58:41.835711Z","iopub.execute_input":"2024-09-30T16:58:41.836166Z","iopub.status.idle":"2024-09-30T16:58:41.843502Z","shell.execute_reply.started":"2024-09-30T16:58:41.836111Z","shell.execute_reply":"2024-09-30T16:58:41.842109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df=pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\nprint(df.shape)\nprint(df.info())\nprint(df.isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2024-09-30T16:58:43.074805Z","iopub.execute_input":"2024-09-30T16:58:43.075305Z","iopub.status.idle":"2024-09-30T16:58:43.164686Z","shell.execute_reply.started":"2024-09-30T16:58:43.075260Z","shell.execute_reply":"2024-09-30T16:58:43.163356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"msno.matrix(df)\nplt.title('Missing Data Matrix')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-30T16:58:44.830874Z","iopub.execute_input":"2024-09-30T16:58:44.831379Z","iopub.status.idle":"2024-09-30T16:58:45.680514Z","shell.execute_reply.started":"2024-09-30T16:58:44.831337Z","shell.execute_reply":"2024-09-30T16:58:45.679218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"msno.bar(df)\nplt.title('Missing Data Bar Chart')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-30T16:58:45.682470Z","iopub.execute_input":"2024-09-30T16:58:45.682865Z","iopub.status.idle":"2024-09-30T16:58:48.972744Z","shell.execute_reply.started":"2024-09-30T16:58:45.682818Z","shell.execute_reply":"2024-09-30T16:58:48.971094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(25, 12))  # Adjust the figsize as needed\n\n# Generate the missing data heatmap\nmsno.heatmap(df, fontsize=12, ax=ax)\n\n# Rotate the x-axis labels\nax.set_xticklabels(ax.get_xticklabels(), rotation=90, ha='center')\n\n# Rotate the y-axis labels if needed\nax.set_yticklabels(ax.get_yticklabels(), rotation=0, ha='right')\n\n# Adjust layout\nplt.tight_layout()\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-30T16:58:48.975549Z","iopub.execute_input":"2024-09-30T16:58:48.976059Z","iopub.status.idle":"2024-09-30T16:58:59.134798Z","shell.execute_reply.started":"2024-09-30T16:58:48.976007Z","shell.execute_reply":"2024-09-30T16:58:59.133250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"msno.dendrogram(df)\nplt.title('Missing Data Dendrogram')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-30T16:58:59.136885Z","iopub.execute_input":"2024-09-30T16:58:59.137441Z","iopub.status.idle":"2024-09-30T16:59:00.576488Z","shell.execute_reply.started":"2024-09-30T16:58:59.137382Z","shell.execute_reply":"2024-09-30T16:59:00.575225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This dendrogram, from a missing values plot, provides a hierarchical clustering of variables based on their pattern of missing values. Here's what it generally indicates:\n\n1) **Clusters of Similar Missingness Patterns**: The variables grouped closely together share similar missing value patterns. This means that when one variable has missing data, the others in the same cluster are likely to also have missing values.\n\n2) **Distances Indicate Missing Value Correlation**: The horizontal axis (distance) shows how correlated the missingness of variables is. Variables with a low distance between them (closer on the horizontal axis) tend to have a similar pattern of missing values.\n\n3) **Hierarchical Structure of Missing Data**: The hierarchical structure formed by the tree (dendrogram) helps identify how variables are related based on missing data. Larger branches indicate groups of variables that collectively share a missingness pattern, while smaller branches show more specific groups or pairs.\n\n4) **Potential Impact on Data Analysis**: Understanding the clusters of missing values can help decide how to handle the missing data—whether to impute, drop, or analyze separately—based on the grouping and correlations in missingness.","metadata":{}},{"cell_type":"markdown","source":"From the above analysis what we learn is that \n1) Demographics insturment does not have any missing values\n\n2) Missing values are clustered to the instruments (PCIAT, FGC ,BIA, Physical)\n\n3) In the correlation plot we also learn that the instruments missing values are highly correlated and is less correlated to other instruments. So we could assume that if a instrument values are missing it does not mean that other instruments would also be empty. ","metadata":{}},{"cell_type":"code","source":"import dask.dataframe as dd\n# dask used as pandas causes OOM error\nfull_antigraphy=dd.read_parquet('/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet')","metadata":{"execution":{"iopub.status.busy":"2024-09-30T16:51:41.643434Z","iopub.execute_input":"2024-09-30T16:51:41.643901Z","iopub.status.idle":"2024-09-30T16:51:45.807936Z","shell.execute_reply.started":"2024-09-30T16:51:41.643860Z","shell.execute_reply":"2024-09-30T16:51:45.806170Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"null_percentage = (full_antigraphy.isnull().sum() / len(full_antigraphy)) * 100\n\n# Compute the result and display it\nnull_percentage_computed = null_percentage.compute()\n\nprint(null_percentage_computed)","metadata":{"execution":{"iopub.status.busy":"2024-09-30T16:53:23.384167Z","iopub.execute_input":"2024-09-30T16:53:23.384713Z","iopub.status.idle":"2024-09-30T16:55:13.384929Z","shell.execute_reply.started":"2024-09-30T16:53:23.384663Z","shell.execute_reply":"2024-09-30T16:55:13.383483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Conclusion\n1) No null present in antigraphy files. These would be really helpful in EDA and some feature engineering. \n\n2) We should go with semi supervised learning approach to fill in the null. \n\nPlease upvote incase you find this useful\n","metadata":{}}]}