{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 📊 Predicting Problematic Internet Usage in Children and Adolescents Based on Physical Activity\n\n## Overview\nIn today’s digital age, excessive internet use among children and adolescents has become a growing concern, impacting mental health, academic performance, and physical well-being. This competition seeks to address this issue by predicting levels of problematic internet usage through an analysis of physical activity and fitness data. By identifying patterns early on, we can take proactive steps to encourage healthier digital habits.\n\n## Objectives\n- **Develop a Predictive Model**: Build and optimize a model to predict the likelihood of problematic internet usage based on physical activity features.\n\n## What Are We Predicting?\nThe target variable in this competition is derived from the **Parent-Child Internet Addiction Test (PCIAT)**—a 20-item scale that measures characteristics and behaviors associated with compulsive internet use, including **compulsivity, escapism, and dependency**. Specifically, we are predicting the **Severity Impairment Index (SII)**, which classifies the level of problematic internet usage into four categories:\n\n- **0 = None** (PCIAT score: 0-30)\n- **1 = Mild** (PCIAT score: 31-49)\n- **2 = Moderate** (PCIAT score: 50-79)\n- **3 = Severe** (PCIAT score: 80-100)\n\n- **Analyze Key Factors**: Explore the relationships between various physical activity metrics (e.g., exercise frequency, fitness levels) and digital behavior.\n- **Enable Early Interventions**: Generate insights to help educators, parents, and healthcare professionals identify children and adolescents at risk for problematic internet use.\n\n## Exploratory Data Analysis (EDA) Goals\n1. **Understand Dataset Structure and Features**: Identify the main characteristics of physical activity and any additional demographic data.\n2. **Data Quality Assessment**: Check for missing values, outliers, or inconsistencies to ensure reliable inputs for modeling.\n3. **Pattern Identification**: Examine trends and correlations between physical activity levels and internet usage, focusing on potential risk factors.\n4. **Hypothesis Formation**: Generate hypotheses on how various physical or lifestyle factors might contribute to or mitigate problematic internet behavior.\n\n---\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport matplotlib.gridspec as gridspec\nimport seaborn as sns\nimport warnings\n%matplotlib inline\n# Set up visual style\nsns.set(style=\"whitegrid\")\nplt.figure(figsize=(10, 6))\nwarnings.filterwarnings('ignore', category=UserWarning)  # Only suppress UserWarnings","metadata":{"execution":{"iopub.status.busy":"2024-11-15T06:07:37.937211Z","iopub.execute_input":"2024-11-15T06:07:37.938121Z","iopub.status.idle":"2024-11-15T06:07:37.952395Z","shell.execute_reply.started":"2024-11-15T06:07:37.938072Z","shell.execute_reply":"2024-11-15T06:07:37.951261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Read input data","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\ntest = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv')\ndata_dict = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/data_dictionary.csv')","metadata":{"execution":{"iopub.status.busy":"2024-11-15T06:07:38.968607Z","iopub.execute_input":"2024-11-15T06:07:38.969059Z","iopub.status.idle":"2024-11-15T06:07:39.027002Z","shell.execute_reply.started":"2024-11-15T06:07:38.969001Z","shell.execute_reply":"2024-11-15T06:07:39.025773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head(5)","metadata":{"execution":{"iopub.status.busy":"2024-11-15T06:07:39.543172Z","iopub.execute_input":"2024-11-15T06:07:39.543638Z","iopub.status.idle":"2024-11-15T06:07:39.623593Z","shell.execute_reply.started":"2024-11-15T06:07:39.543595Z","shell.execute_reply":"2024-11-15T06:07:39.622449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Checking if there are NAs in train data","metadata":{}},{"cell_type":"code","source":"\npd.set_option('display.max_rows',None)\npd.set_option('display.max_columns',None)\ntrain['PCIAT-PCIAT_Total_bin'] =pd.cut(train['PCIAT-PCIAT_Total'], bins=10)\npd.DataFrame(train.isna().sum()*100/train.shape[0]).transpose()","metadata":{"execution":{"iopub.status.busy":"2024-11-15T06:07:41.099275Z","iopub.execute_input":"2024-11-15T06:07:41.099735Z","iopub.status.idle":"2024-11-15T06:07:41.163273Z","shell.execute_reply.started":"2024-11-15T06:07:41.099689Z","shell.execute_reply":"2024-11-15T06:07:41.161952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"warnings.filterwarnings('ignore', category=FutureWarning)  ","metadata":{"execution":{"iopub.status.busy":"2024-11-15T06:07:41.766470Z","iopub.execute_input":"2024-11-15T06:07:41.766893Z","iopub.status.idle":"2024-11-15T06:07:41.772348Z","shell.execute_reply.started":"2024-11-15T06:07:41.766853Z","shell.execute_reply":"2024-11-15T06:07:41.771274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# 1. Age Distribution\n\nfig, axes = plt.subplots(3, 2, figsize=(30, 30))\n\nsns.histplot(train['Basic_Demos-Age'], bins=20, kde=True, color='skyblue',ax=axes[0, 0])\naxes[0, 0].set_title('Age Distribution')\naxes[0, 0].set_xlabel('Age')\naxes[0, 0].set_ylabel('Frequency')\n\n# 2. Boxplot of SII by Age\nsns.boxplot(x='PCIAT-PCIAT_Total_bin', y='Basic_Demos-Age', data=train, palette='coolwarm',ax=axes[0, 1])\naxes[0, 1].set_title('Problematic Internet Usage Severity by Age')\naxes[0, 1].set_xlabel('Severity Level (SII)')\naxes[0, 1].set_ylabel('Age')\n\n\n\n\n# 3. Bar Plot of SII by Gender\nsns.countplot(x='PCIAT-PCIAT_Total_bin', hue='Basic_Demos-Sex', data=train, palette='pastel',ax=axes[1, 0])\naxes[1, 0].set_title('Internet Usage Severity by Gender')\naxes[1, 0].set_xlabel('Severity Level (SII)')\naxes[1, 0].set_ylabel('Count')\n\n# 4. SII by Season\n\nsns.countplot(x='PCIAT-PCIAT_Total_bin', hue='Basic_Demos-Enroll_Season', data=train, palette='muted',ax=axes[1, 1])\naxes[1, 1].set_title('Problematic Internet Usage Severity by Enrollment Season')\naxes[1, 1].set_xlabel('Severity Level (SII)')\naxes[1, 1].set_ylabel('Count')\n\n# 5. Distribution of SII Scores\nsns.histplot(train['PCIAT-PCIAT_Total'], bins=10, kde=True, color='coral',ax=axes[2, 0])\naxes[2, 0].set_title('Distribution of Problematic Internet Usage Severity')\naxes[2, 0].set_xlabel('Severity Level (SII)')\naxes[2, 0].set_ylabel('Frequency')\n\n\n# 6. Genber based SII distribution\nfor group, subset in train.groupby('Basic_Demos-Sex'):\n    sns.kdeplot(subset['PCIAT-PCIAT_Total'], label=group, fill=True,ax=axes[2, 1])\n    axes[2, 1].set_title('KDE of Values by Group')\n    axes[2, 1].set_xlabel('Value')\n    axes[2, 1].set_ylabel('Density')\n\n\n# Remove any empty subplots if necessary\naxes[0, 1].set_xticklabels(axes[0, 1].get_xticklabels(), rotation=45)\naxes[1, 0].legend(title='Gender', labels=['Male', 'Female'])\naxes[1, 1].set_xticklabels(axes[1, 1].get_xticklabels(), rotation=45)\naxes[1, 1].legend(title='Season')\naxes[2, 1].legend(title='Group')\n\n# Adjust layout to prevent overlap\nplt.tight_layout()\n\n# Show the plots\nplt.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2024-11-15T06:07:42.285096Z","iopub.execute_input":"2024-11-15T06:07:42.285561Z","iopub.status.idle":"2024-11-15T06:07:45.780324Z","shell.execute_reply.started":"2024-11-15T06:07:42.285472Z","shell.execute_reply":"2024-11-15T06:07:45.779116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Intital Analysis \n\n* The clinical sample is centered around age group of 8 \n* we can see that SII score gets worse with age\n* There doesnt seem to be much of a difference in SII distrbution between male and female","metadata":{}},{"cell_type":"markdown","source":"### identifying trends based on the SII score\n","metadata":{}},{"cell_type":"code","source":"print(\"\"\"we have around \"\"\"+str(round(100*train[train['PCIAT-PCIAT_Total'].isnull()].shape[0]/train.shape[0],0))+\n      \"\"\" of people whose scores are missing\"\"\")\n","metadata":{"execution":{"iopub.status.busy":"2024-11-15T06:07:45.782520Z","iopub.execute_input":"2024-11-15T06:07:45.782977Z","iopub.status.idle":"2024-11-15T06:07:45.791816Z","shell.execute_reply.started":"2024-11-15T06:07:45.782930Z","shell.execute_reply":"2024-11-15T06:07:45.790433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_dict=pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/data_dictionary.csv')\n","metadata":{"execution":{"iopub.status.busy":"2024-11-15T06:07:45.793056Z","iopub.execute_input":"2024-11-15T06:07:45.793451Z","iopub.status.idle":"2024-11-15T06:07:45.811462Z","shell.execute_reply.started":"2024-11-15T06:07:45.793410Z","shell.execute_reply":"2024-11-15T06:07:45.809999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pcit_questions=data_dict[data_dict['Instrument'].str.contains('Parent-Child Internet Addiction Test')][['Field','Description']]\nfig, axes = plt.subplots(11, 2, figsize=(100, 200))\ni=0\nj=0\nfor s,row in pcit_questions.iterrows():\n    subset=train[[row['Field'],'Basic_Demos-Age']]\n    na_count = subset.groupby('Basic_Demos-Age')[row['Field']].apply(lambda x: x.isna().sum()).reset_index()\n    na_count.columns = ['Basic_Demos-Age', 'na_count']\n\n    \n    sns.barplot(data=na_count, x='Basic_Demos-Age', y='na_count', color='skyblue',ax=axes[i, j])\n    axes[i, j].set_title(row['Description'])\n    axes[i, j].set_xlabel('Age')\n    axes[i, j].set_ylabel('Count of NaNs')\n    j=j+1\n    if j==2:\n        i=i+1\n        j=0\n# Adjust layout to prevent overlap\nplt.tight_layout()\n\n# Show the plots\nplt.show()\n\n    ","metadata":{"execution":{"iopub.status.busy":"2024-11-15T06:08:54.899030Z","iopub.execute_input":"2024-11-15T06:08:54.899454Z","iopub.status.idle":"2024-11-15T06:09:19.340937Z","shell.execute_reply.started":"2024-11-15T06:08:54.899416Z","shell.execute_reply":"2024-11-15T06:09:19.339716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Analysis\n\nThe above plots show that there is no dependency between questions unanswered and age of the sample under consideration as they follow similar distrbution ","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}