{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\n# for dirname, _, filenames in os.walk('/kaggle/input'):\n#     for filename in filenames:\n#         print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-11-15T07:26:44.203888Z","iopub.execute_input":"2024-11-15T07:26:44.204415Z","iopub.status.idle":"2024-11-15T07:26:44.211041Z","shell.execute_reply.started":"2024-11-15T07:26:44.204371Z","shell.execute_reply":"2024-11-15T07:26:44.209578Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Child Mind Institute — Problematic Internet Use: Exploratory Data Analysis\n\n## Objective\nThis notebook provides an exploratory analysis of the dataset from the Child Mind Institute competition on problematic internet use. Our goal is to understand the data structure, identify key patterns, and prepare the dataset for modelng.\n\n## Outline\n1. Import Libraries\n2. Data Loading and Initial Overview\n3. Data Dictionary and Description\n4. Data Engineering (Handling Missing Data\n6. Data Visualization\n7. Insights and Observations","metadata":{}},{"cell_type":"markdown","source":"## Import Libraries","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.express as px\nfrom sklearn.impute import SimpleImputer, KNNImputer\nfrom scipy import stats\n\n# Configure visualization styles\nsns.set_theme(style=\"whitegrid\")\nplt.rcParams['figure.figsize'] = (10, 6)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T07:26:44.213452Z","iopub.execute_input":"2024-11-15T07:26:44.213855Z","iopub.status.idle":"2024-11-15T07:26:44.226787Z","shell.execute_reply.started":"2024-11-15T07:26:44.213818Z","shell.execute_reply":"2024-11-15T07:26:44.225476Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data Loading and Initial Overview","metadata":{}},{"cell_type":"code","source":"# Load data\ntrain_df = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\ntest_df = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv')\nprint(\"train size: {}, test size: {}\".format(train_df.shape, test_df.shape))\n\n# Display initial rows\ntrain_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T07:26:44.228944Z","iopub.execute_input":"2024-11-15T07:26:44.229319Z","iopub.status.idle":"2024-11-15T07:26:44.302264Z","shell.execute_reply.started":"2024-11-15T07:26:44.229283Z","shell.execute_reply":"2024-11-15T07:26:44.301121Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Data overview (train)\ntrain_df.info()\n\n# Missing values visualization\nsns.heatmap(train_df.isnull(), cbar=False, cmap=\"viridis\")\nplt.title(\"Missing Values in Dataset\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T07:26:44.305234Z","iopub.execute_input":"2024-11-15T07:26:44.305714Z","iopub.status.idle":"2024-11-15T07:26:45.670946Z","shell.execute_reply.started":"2024-11-15T07:26:44.305675Z","shell.execute_reply":"2024-11-15T07:26:45.669731Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Data overview (test)\ntest_df.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T07:26:45.672406Z","iopub.execute_input":"2024-11-15T07:26:45.672850Z","iopub.status.idle":"2024-11-15T07:26:45.688658Z","shell.execute_reply.started":"2024-11-15T07:26:45.672804Z","shell.execute_reply":"2024-11-15T07:26:45.687630Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Columns present in train but not in test\ntrain_only_columns = set(train_df.columns) - set(test_df.columns)\nprint(\"Columns only in train:\", train_only_columns)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T07:26:45.689916Z","iopub.execute_input":"2024-11-15T07:26:45.690312Z","iopub.status.idle":"2024-11-15T07:26:45.697777Z","shell.execute_reply.started":"2024-11-15T07:26:45.690266Z","shell.execute_reply":"2024-11-15T07:26:45.696738Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data Dictionary and Feature Description\n\nThe dataset contains 3960 training examples, out of which only 2736 are predictable with the target variable `sii`. Many of the features contain missing values, particularly those related to the child’s fitness status and certain survey data. Below are key points regarding some important features:\n\n### Predictability of Target:\n- **Target Variable**: `sii` (Predictable for 2736 out of 3960 examples).\n\n### Missing Data:\n- A significant portion of the dataset has missing values. Some features, especially those describing the child’s fitness status, have fewer than 50% valid entries. These features may not provide enough reliable information and should be handled carefully during analysis.\n\n### Key Features with Missing Data:\n1. **Fitness-Related Features**:\n   - `Fitness_Endurance-Season`: Only 1308 non-null values (about 33% valid).\n   - `Fitness_Endurance-Max_Stage`: Only 743 non-null values (about 19% valid).\n   - `Fitness_Endurance-Time_Mins`: Only 740 non-null values (about 19% valid).\n   - `Fitness_Endurance-Time_Sec`: Only 740 non-null values (about 19% valid).\n   - **Recommendation**: These columns may need to be dropped or imputed, as they contain substantial missing data.\n\n2. **Physical Measurements**:\n   - `Physical-Waist_Circumference`: Only 898 non-null values (about 23% valid).\n   - **Recommendation**: This column has significant missing data and should be considered for exclusion.\n\n3. **PAQ-A Features**:\n   - `PAQ_A-Season`: Only 475 non-null values (about 12% valid).\n   - `PAQ_A-PAQ_A_Total`: Only 475 non-null values (about 12% valid).\n   - **Recommendation**: Since these columns have less than 40% valid entries, they may not contribute meaningful information to the analysis.\n\n### PCIAT Features (Parent-Child Internet Addiction Test):\n- **PCIAT Columns**: All fields starting with `PCIAT-` (e.g., `PCIAT-PCIAT_01`, `PCIAT-PCIAT_02`, ...) are missing in the test dataset.\n  - **Recommendation**: These columns should be eliminated from the analysis as they are not present in the test data and cannot be used for predictive modeling.\n\n### Column Removal:\n- Columns with low non-null values (below 40%) such as those describing fitness status (`Fitness_Endurance-Season`, `Fitness_Endurance-Max_Stage`, etc.), waist circumference (`Physical-Waist_Circumference`), PAQ-A data (`PAQ_A-Season`, `PAQ_A-PAQ_A_Total`), and PCIAT-related data (since it's missing in the test set) should be considered for removal during preprocessing to avoid skewing the model’s results with unreliable data.","metadata":{}},{"cell_type":"markdown","source":"## Data Engineering (Handling Missing Data)","metadata":{}},{"cell_type":"code","source":"# Filter the training data to keep only rows where 'sii' is not null\ntrain_df = train_df[train_df['sii'].notnull()]\n\n# Store targets\ntarget = train_df['sii']\n\n# Display the number of rows after filtering\nprint(f\"Number of rows with valid 'sii' target: {train_df.shape[0]}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T07:26:45.699185Z","iopub.execute_input":"2024-11-15T07:26:45.700025Z","iopub.status.idle":"2024-11-15T07:26:45.711198Z","shell.execute_reply.started":"2024-11-15T07:26:45.699981Z","shell.execute_reply":"2024-11-15T07:26:45.710005Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Set the threshold for the minimum percentage of non-missing values required\nthreshold = 40\n\n# Calculate the percentage of non-missing values for each column in both train and test\ntrain_non_missing = train_df.notnull().mean() * 100\ntest_non_missing = test_df.notnull().mean() * 100\n\n# Identify columns that meet the threshold in both train and test\ncolumns_to_keep = list(train_non_missing[(train_non_missing >= threshold) & \n                       (test_non_missing >= threshold)].index)\n\n# Filter both dataframes to keep only the columns that meet the threshold\ntrain_filt = train_df[columns_to_keep]\ntest_filt = test_df[columns_to_keep]\n\nprint(\"Remaining columns after filtering:\", train_filt.columns)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T07:26:45.712762Z","iopub.execute_input":"2024-11-15T07:26:45.713148Z","iopub.status.idle":"2024-11-15T07:26:45.732352Z","shell.execute_reply.started":"2024-11-15T07:26:45.713111Z","shell.execute_reply":"2024-11-15T07:26:45.731185Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Data overview (train after filtering)\ntrain_filt.info()\n\n# Missing values visualization\nsns.heatmap(train_filt.isnull(), cbar=False, cmap=\"viridis\")\nplt.title(\"Missing Values in Dataset\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T07:26:45.735261Z","iopub.execute_input":"2024-11-15T07:26:45.735646Z","iopub.status.idle":"2024-11-15T07:26:46.986884Z","shell.execute_reply.started":"2024-11-15T07:26:45.735609Z","shell.execute_reply":"2024-11-15T07:26:46.985718Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Data overview (test after filtering)\ntest_filt.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T07:26:46.988298Z","iopub.execute_input":"2024-11-15T07:26:46.988644Z","iopub.status.idle":"2024-11-15T07:26:47.003255Z","shell.execute_reply.started":"2024-11-15T07:26:46.988609Z","shell.execute_reply":"2024-11-15T07:26:47.002018Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Identify numerical and categorical columns in both train and test DataFrames\nnum_cols_train = train_filt.select_dtypes(include=['int64', 'float64']).columns\ncat_cols_train = train_filt.select_dtypes(include=['object', 'category']).columns\n\nnum_cols_test = test_filt.select_dtypes(include=['int64', 'float64']).columns\ncat_cols_test = test_filt.select_dtypes(include=['object', 'category']).columns\n\n# Print numerical and categorical columns for train\nprint(f\"Train Numerical Columns: {num_cols_train}\\nTrain Categorical Columns: {cat_cols_train}\")\n\n# Impute the training data\n# num_imputer = SimpleImputer(strategy='median')\n# cat_imputer = SimpleImputer(strategy='most_frequent')\n\n# train_filt[num_cols_train] = num_imputer.fit_transform(train_filt[num_cols_train])\n# train_filt[cat_cols_train] = cat_imputer.fit_transform(train_filt[cat_cols_train])\n\n# # Impute the test data with the same transformers\n# test_filt[num_cols_test] = num_imputer.transform(test_filt[num_cols_test])\n# test_filt[cat_cols_test] = cat_imputer.transform(test_filt[cat_cols_test])\n\n# # For more complex patterns, using KNNImputer for numerical columns\n# knn_imputer = KNNImputer(n_neighbors=5)\n\n# # Apply KNNImputer to train and test datasets\n# train_filt[num_cols_train] = knn_imputer.fit_transform(train_filt[num_cols_train])\n# test_filt[num_cols_test] = knn_imputer.transform(test_filt[num_cols_test])\n\n# print(\"Imputation completed for both training and testing data.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T07:26:47.005115Z","iopub.execute_input":"2024-11-15T07:26:47.005523Z","iopub.status.idle":"2024-11-15T07:26:47.016280Z","shell.execute_reply.started":"2024-11-15T07:26:47.005484Z","shell.execute_reply":"2024-11-15T07:26:47.015153Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data Visualization","metadata":{}},{"cell_type":"code","source":"# Distribution of the target variable\nsns.histplot(target, kde=True)\nplt.title(\"Distribution of Target Variable - Severity Impairment Index (sii)\")\nplt.show()\n\n# Distribution of numerical features\nfor col in num_cols_train:\n    plt.figure()\n    sns.histplot(train_filt[col], kde=True)\n    plt.title(f\"Distribution of {col}\")\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T07:26:47.017562Z","iopub.execute_input":"2024-11-15T07:26:47.018065Z","iopub.status.idle":"2024-11-15T07:27:45.159506Z","shell.execute_reply.started":"2024-11-15T07:26:47.018018Z","shell.execute_reply":"2024-11-15T07:27:45.158425Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Target vs. selected features\ninterest_cols = ['Basic_Demos-Age', 'Basic_Demos-Sex', 'SDS-SDS_Total_T', 'PreInt_EduHx-computerinternet_hoursday']\nfor col in interest_cols:\n    plt.figure()\n    sns.boxplot(x=train_filt[col], y=target)\n    plt.title(f\"Severity Impairment Index (sii) vs. {col}\")\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T07:27:45.160889Z","iopub.execute_input":"2024-11-15T07:27:45.161262Z","iopub.status.idle":"2024-11-15T07:27:47.474672Z","shell.execute_reply.started":"2024-11-15T07:27:45.161228Z","shell.execute_reply":"2024-11-15T07:27:47.473534Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Insights and Observations:\n\n1. **Severity Impairment Index (SII) Distribution**:\n   - Approximately **60%** of the training data has a **Severity Impairment Index (SII) of 0**. This suggests that most children in the dataset have no significant impairment or are classified as having low severity of impairment.\n   - **~1%** of the data has an SII of 3, indicating that very few children are experiencing high levels of impairment. This distribution is **highly skewed**, with a large portion of the dataset concentrated around 0, making the dataset **imbalanced**.\n   - The imbalance may present challenges in predictive modeling, as the model may tend to predict the majority class (SII=0) more frequently, leading to potential underperformance in predicting higher SII levels.\n\n2. **Gender Distribution**:\n   - The male-to-female ratio is approximately **1.75:1**, with more males than females in the dataset. This could influence the interpretation of results related to gender-specific factors.\n\n3. **Numeric Columns**:\n   - **Most of the numeric columns are right-skewed**, with a concentration of low values, such as in the **Bio-electric Impedance Analysis (BIA)** index scores. This suggests that the majority of children in the dataset exhibit lower BIA values, with a few outliers exhibiting higher values.\n   - This right-skewness could indicate that the data might need **log transformation** or other methods to normalize the distribution for some of the columns, particularly if they are used in predictive modeling.\n\n4. **Sleep Disturbance**:\n   - The data related to **sleep disturbance** appears to be **right-skewed**. This indicates that while most children report lower levels of sleep disturbance, a smaller group may be experiencing significantly higher levels of disruption. This could suggest that a few children are suffering from more severe sleep disturbances, which could impact their overall health and well-being.\n   - **Possible interpretation**: Children who experience higher levels of sleep disturbance might also have a higher likelihood of being impacted by internet addiction or related issues. This correlation could be important to investigate further.\n\n5. **Internet Use and SII**:\n   - A large portion of the data reflects **low internet use**, which is highly correlated with an SII of 0 (no impairment). This may indicate that children who use the internet less are less likely to exhibit behavioral issues related to internet addiction or technology use.\n   - This could suggest that the internet use behavior is a factor in determining the severity of impairment, where children with higher internet usage might be at a higher risk of developing internet addiction-related issues (hence, a higher SII).\n\n---\n\n### Future Steps:\n\n1. **Handling Imbalanced Data**:\n   - Explore techniques such as **oversampling** or **undersampling** to balance the dataset. Alternatively, consider using **weighted loss functions** in predictive models to account for the class imbalance.\n\n2. **Further Exploration of Sleep Disturbance**:\n   - Investigate if there is a **causal relationship** between sleep disturbance and other factors such as screen time, behavioral issues, or mental health. Use **statistical tests** or **correlation analysis** to confirm these relationships.\n\n3. **Feature Engineering**:\n   - Create new features from existing ones, such as **combined internet use and sleep disturbance** indicators, which may reveal hidden patterns in the data that could improve predictive modeling.\n\n4. **Multivariate Analysis**:\n   - Perform a **multivariate analysis** to explore how combinations of features (e.g., gender, sleep disturbance, and internet use) relate to the Severity Impairment Index (SII). This could uncover more complex interactions that simple univariate analysis might miss.\n\n5. **Addressing Missing Data**:\n   - Ensure that missing values in critical columns, such as fitness-related and behavioral columns, are appropriately handled using imputation or removal. Check if the missingness pattern is **random** or **biased**.\n\n6. **Model Development**:\n   - Test different machine learning models to see how well they handle the class imbalance, including tree-based models like **Random Forest** or **XGBoost**, which can handle imbalances more effectively.\n   - Use **cross-validation** to evaluate the model’s performance and avoid overfitting, particularly with skewed data distributions.","metadata":{}},{"cell_type":"code","source":"# Combine filtered train data and target\ncombined_df = pd.concat([train_filt, target], axis=1)\n\n# Export train_filt (with target 'sii') to CSV\ncombined_df.to_csv('train.csv', index=False)\n\n# Export test_filt to CSV\ntest_filt.to_csv('test.csv', index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T07:27:47.476239Z","iopub.execute_input":"2024-11-15T07:27:47.476694Z","iopub.status.idle":"2024-11-15T07:27:47.578323Z","shell.execute_reply.started":"2024-11-15T07:27:47.476648Z","shell.execute_reply":"2024-11-15T07:27:47.576799Z"}},"outputs":[],"execution_count":null}]}