{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30839,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-01-16T22:18:05.687949Z","iopub.execute_input":"2025-01-16T22:18:05.688265Z","iopub.status.idle":"2025-01-16T22:18:09.751965Z","shell.execute_reply.started":"2025-01-16T22:18:05.688236Z","shell.execute_reply":"2025-01-16T22:18:09.750711Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# Import required libraries\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport warnings\n\n# Suppress warnings\nwarnings.filterwarnings(\"ignore\")\n\n# Reload the train.csv file\ntrain_path = '/kaggle/input/child-mind-institute-problematic-internet-use/train.csv'\ntrain_df = pd.read_csv(train_path)\n\n# Display general information about the dataset to confirm successful loading\ntrain_info = train_df.info()\ntrain_info\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-16T22:18:27.366370Z","iopub.execute_input":"2025-01-16T22:18:27.366777Z","iopub.status.idle":"2025-01-16T22:18:28.403643Z","shell.execute_reply.started":"2025-01-16T22:18:27.366746Z","shell.execute_reply":"2025-01-16T22:18:28.402487Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Step 1: Distribution of the target variable (sii)\nsii_distribution = train_df['sii'].value_counts(normalize=True).sort_index()\n\n# Plot the distribution of 'sii'\nplt.figure(figsize=(8, 5))\nsns.barplot(x=sii_distribution.index, y=sii_distribution.values, palette='viridis')\nplt.title('Distribution of Severity Impairment Index (sii)', fontsize=14)\nplt.xlabel('sii (Severity Level)', fontsize=12)\nplt.ylabel('Proportion of Participants', fontsize=12)\nplt.xticks(ticks=[0, 1, 2, 3], labels=['None', 'Mild', 'Moderate', 'Severe'])\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-16T22:18:32.255149Z","iopub.execute_input":"2025-01-16T22:18:32.255476Z","iopub.status.idle":"2025-01-16T22:18:32.548155Z","shell.execute_reply.started":"2025-01-16T22:18:32.255449Z","shell.execute_reply":"2025-01-16T22:18:32.546927Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# Step 2: Missing data visualization\nmissing_data = train_df.isnull().mean().sort_values(ascending=False)\n\n# Plot the missing data rates\nplt.figure(figsize=(12, 6))\nsns.barplot(x=missing_data.index[:30], y=missing_data.values[:30], palette='coolwarm')\nplt.title('Top 30 Features with Highest Missing Rates', fontsize=14)\nplt.ylabel('Proportion of Missing Values', fontsize=12)\nplt.xticks(rotation=90)\nplt.show()\n\n\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-16T22:18:44.655481Z","iopub.execute_input":"2025-01-16T22:18:44.655830Z","iopub.status.idle":"2025-01-16T22:18:45.104655Z","shell.execute_reply.started":"2025-01-16T22:18:44.655799Z","shell.execute_reply":"2025-01-16T22:18:45.103693Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# Step 3: Correlation with sii\n# Ensure 'id' and 'sii' are dropped only if they exist in the DataFrame\nnumeric_features = train_df.select_dtypes(include=['float64', 'int64']).copy()\ncolumns_to_drop = [col for col in ['id', 'sii'] if col in numeric_features.columns]\nnumeric_features = numeric_features.drop(columns=columns_to_drop, errors='ignore')\n\n# Compute correlations with 'sii' safely\ncorrelations = numeric_features.corrwith(train_df['sii']).sort_values(ascending=False)\n\n# Plot the top 10 correlations with 'sii'\nplt.figure(figsize=(10, 6))\nsns.barplot(x=correlations.index[:10], y=correlations.values[:10], palette='magma')\nplt.title('Top 10 Features Correlated with sii', fontsize=14)\nplt.ylabel('Correlation Coefficient', fontsize=12)\nplt.xticks(rotation=45)\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-16T22:19:00.371907Z","iopub.execute_input":"2025-01-16T22:19:00.372271Z","iopub.status.idle":"2025-01-16T22:19:00.657563Z","shell.execute_reply.started":"2025-01-16T22:19:00.372236Z","shell.execute_reply":"2025-01-16T22:19:00.656475Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# Step 4: Advanced Insight - Feature Interaction Analysis\n# Explore potential interactions between top correlated features\n# Select the top 2 positively and negatively correlated features\npositive_features = correlations.head(2).index.tolist()\nnegative_features = correlations.tail(2).index.tolist()\nselected_features = positive_features + negative_features\n\n# Pairplot to visualize potential interactions\nsns.pairplot(train_df, vars=selected_features, hue='sii', palette='coolwarm', corner=True)\nplt.suptitle('Feature Interactions and sii', y=1.02, fontsize=16)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-16T22:19:10.184598Z","iopub.execute_input":"2025-01-16T22:19:10.184970Z","iopub.status.idle":"2025-01-16T22:19:14.931089Z","shell.execute_reply.started":"2025-01-16T22:19:10.184938Z","shell.execute_reply":"2025-01-16T22:19:14.929648Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Step 5: Advanced Insight - Clustering based on correlated features\nfrom sklearn.cluster import KMeans\nfrom sklearn.preprocessing import StandardScaler\n\n# Scale the data for clustering\nscaler = StandardScaler()\nscaled_features = scaler.fit_transform(numeric_features[selected_features].dropna())\n\n# Apply K-Means clustering\nkmeans = KMeans(n_clusters=4, random_state=42)\nkmeans_clusters = kmeans.fit_predict(scaled_features)\n\n# Add clusters to the dataset\nclustered_data = train_df.dropna(subset=selected_features).copy()\nclustered_data['Cluster'] = kmeans_clusters\n\n# Visualize clusters with a scatter plot\nplt.figure(figsize=(10, 6))\nsns.scatterplot(\n    x=clustered_data[selected_features[0]],\n    y=clustered_data[selected_features[1]],\n    hue=clustered_data['Cluster'],\n    palette='tab10',\n    style=clustered_data['sii']\n)\nplt.title('Clusters based on Key Features', fontsize=14)\nplt.xlabel(selected_features[0], fontsize=12)\nplt.ylabel(selected_features[1], fontsize=12)\nplt.legend(title='Cluster/SII')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-16T22:19:28.954515Z","iopub.execute_input":"2025-01-16T22:19:28.954853Z","iopub.status.idle":"2025-01-16T22:19:30.292399Z","shell.execute_reply.started":"2025-01-16T22:19:28.954823Z","shell.execute_reply":"2025-01-16T22:19:30.291261Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n**Objective:**  \nThis notebook provides an in-depth exploratory data analysis (EDA) and advanced feature insights for the Child Mind Institute's Problematic Internet Use dataset. The aim is to uncover patterns and relationships between physical activity, demographic data, and the severity of internet use impairment (`sii`), enabling better understanding and preparation for predictive modeling.\n\n---\n\n**Structure of the Notebook:**\n\n1. **Dataset Overview and Setup:**\n   - Load and describe the dataset to understand its structure and key attributes.\n   - Suppress unnecessary warnings to enhance readability.\n\n2. **Exploratory Data Analysis (EDA):**\n   - **Target Variable Analysis (`sii`):**\n     - Examine the distribution of severity levels (`None`, `Mild`, `Moderate`, `Severe`) to assess balance and representation.\n   - **Missing Data Exploration:**\n     - Visualize missing data rates across features to prioritize data cleaning or imputation strategies.\n   - **Correlation Analysis:**\n     - Identify numerical features most strongly correlated (positively and negatively) with `sii`.\n\n3. **Advanced Feature Insights:**\n   - **Feature Interaction Analysis:**\n     - Use pairplots to explore potential interactions between the most positively and negatively correlated features with `sii`.\n   - **Clustering Analysis:**\n     - Apply K-Means clustering on key correlated features to identify natural groupings within the data.\n     - Visualize clusters with scatter plots, enhancing interpretability by combining cluster information with severity levels (`sii`).\n\n4. **Key Insights and Visualizations:**\n   - Use bar plots, pairplots, and scatter plots to communicate findings effectively.\n   - Highlight actionable insights, such as clusters or feature relationships that could be critical for building predictive models.\n\n---\n\n**Highlights:**\n- **Pairplot Analysis:** A focused look at feature interactions with the target variable.\n- **Clustering:** Leverage unsupervised learning to reveal natural groupings and their overlap with `sii`.\n- **Correlation Focus:** Pinpoint features with the strongest relationships to `sii` for modeling purposes.\n\n---\n\n**Applications:**\nThis notebook is a strong foundation for:\n- Preparing data for machine learning pipelines.\n- Feature engineering and selection.\n- Gaining domain insights into problematic internet use behaviors.\n\nFeel free to use and adapt this notebook to advance your analysis or improve prediction models. 🚀","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}}]}