{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30822,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:29.787549Z","iopub.execute_input":"2025-01-02T01:04:29.787908Z","iopub.status.idle":"2025-01-02T01:04:34.790008Z","shell.execute_reply.started":"2025-01-02T01:04:29.787880Z","shell.execute_reply":"2025-01-02T01:04:34.788988Z"},"collapsed":true,"jupyter":{"outputs_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Understanding the task\n\nThe aim of this competition is to predict the Severity Impairment Index (sii), which measures the level of problematic internet use among children and adolescents, based on physical activity data and other features. \n\nsii is derived from `PCIAT-PCIAT_Total`, the sum of scores from the Parent-Child Internet Addiction Test (PCIAT: 20 questions, scored 0-5).\n\nTarget Variable (sii) is defined as:\n- 0: None (PCIAT-PCIAT_Total from 0 to 30)\n- 1: Mild (PCIAT-PCIAT_Total from 31 to 49)\n- 2: Moderate (PCIAT-PCIAT_Total from 50 to 79)\n- 3: Severe (PCIAT-PCIAT_Total 80 and more)\n\nThis makes sii an ordinal categorical variable with four levels, where the order of categories is meaningful.\nType of Machine Learning Problem we can use with sii as a target:\n\n1. Ordinal classification (ordinal logistic regression, models with custom ordinal loss functions)\n2. Multiclass classification (treat sii as a nominal categorical variable without considering the order)\n3. Regression (ignore the discrete nature of categories and treat sii as a continuous variable, then round prediction)\n4. Custom (e.g. loss functions that penalize errors based on the distance between categories)\n\nWe can also use `PCIAT-PCIAT_Total` as a continuous target variable, and implement regression on `PCIAT-PCIAT_Total` and then map predictions to sii categories.\n\nFinally, another strategy involves predicting responses to each question of the Parent-Child Internet Addiction Test: i.e. pedict individual question scores as separate targets, sum the predicted scores to get the `PCIAT-PCIAT_Total` and map predictions to the corresponding sii category.\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport matplotlib.gridspec as gridspec\nimport plotly.express as px\nimport seaborn as sns\nimport warnings","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:34.791105Z","iopub.execute_input":"2025-01-02T01:04:34.791623Z","iopub.status.idle":"2025-01-02T01:04:36.011903Z","shell.execute_reply.started":"2025-01-02T01:04:34.791583Z","shell.execute_reply":"2025-01-02T01:04:36.010929Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"warnings.filterwarnings('ignore', category=FutureWarning)\n\nsns.set(style=\"whitegrid\")\n%matplotlib inline","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:36.013413Z","iopub.execute_input":"2025-01-02T01:04:36.013978Z","iopub.status.idle":"2025-01-02T01:04:36.020339Z","shell.execute_reply.started":"2025-01-02T01:04:36.013946Z","shell.execute_reply":"2025-01-02T01:04:36.019248Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Exploratory Data Analysis","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\ntest = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv')\ndata_dict = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/data_dictionary.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:36.021854Z","iopub.execute_input":"2025-01-02T01:04:36.022187Z","iopub.status.idle":"2025-01-02T01:04:36.115869Z","shell.execute_reply.started":"2025-01-02T01:04:36.022161Z","shell.execute_reply":"2025-01-02T01:04:36.114950Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:36.117050Z","iopub.execute_input":"2025-01-02T01:04:36.117429Z","iopub.status.idle":"2025-01-02T01:04:36.156159Z","shell.execute_reply.started":"2025-01-02T01:04:36.117390Z","shell.execute_reply":"2025-01-02T01:04:36.155173Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:36.157042Z","iopub.execute_input":"2025-01-02T01:04:36.157372Z","iopub.status.idle":"2025-01-02T01:04:36.179574Z","shell.execute_reply.started":"2025-01-02T01:04:36.157342Z","shell.execute_reply":"2025-01-02T01:04:36.178654Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Columns not present in test set","metadata":{}},{"cell_type":"code","source":"\nset(train.columns) - set(test.columns)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:36.180677Z","iopub.execute_input":"2025-01-02T01:04:36.181022Z","iopub.status.idle":"2025-01-02T01:04:36.197292Z","shell.execute_reply.started":"2025-01-02T01:04:36.180987Z","shell.execute_reply":"2025-01-02T01:04:36.196387Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"makes sense considering how sii is derived from what range the value of PCIAT-PCIAT_Total falls in, which is the sum of all values of PCIAT-PCIAT_01 to PCIAT-PCIAT_20 of all the nan values, this indicates another problem which is that the values  of sii can be incorrect depending on if all values of PCIAT-PCIAT are present or not","metadata":{}},{"cell_type":"code","source":"train.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:36.200217Z","iopub.execute_input":"2025-01-02T01:04:36.200504Z","iopub.status.idle":"2025-01-02T01:04:36.236944Z","shell.execute_reply.started":"2025-01-02T01:04:36.200480Z","shell.execute_reply":"2025-01-02T01:04:36.236048Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_dict.head(20)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:36.238546Z","iopub.execute_input":"2025-01-02T01:04:36.238840Z","iopub.status.idle":"2025-01-02T01:04:36.251172Z","shell.execute_reply.started":"2025-01-02T01:04:36.238790Z","shell.execute_reply":"2025-01-02T01:04:36.250116Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Missing Values\n\nAll columns have a substantial proportion of missing values, except `id` and the three basic demographic columns for sex, age and season of enrollment. Even the target `sii` has missing values:","metadata":{}},{"cell_type":"code","source":"\n# Percentage of missing values per column\nmissing_percent = (train.isnull().sum() / len(train)) * 100\nmissing_percent.plot(kind=\"barh\", figsize=(10, 20))\nplt.xlabel(\"Percentage of Missing Data\")\nplt.ylabel(\"Features\")\nplt.title(\"Missing Data Percentage per Feature\")\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:36.252072Z","iopub.execute_input":"2025-01-02T01:04:36.252424Z","iopub.status.idle":"2025-01-02T01:04:37.324677Z","shell.execute_reply.started":"2025-01-02T01:04:36.252395Z","shell.execute_reply":"2025-01-02T01:04:37.323658Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# Copy the data\ntrain_df = train.copy()\n\n# Define the mapping for `sii` levels\nsii_map = {\n    0: \"None\",\n    1: \"Mild\",\n    2: \"Moderate\",\n    3: \"Severe\",\n    np.nan: \"NaN\"  # Explicitly map NaN for clarity\n}\n\n# Map the sii levels, leaving NaNs as they are\ntrain_df['sii_label'] = train_df['sii'].map(sii_map)\n\n# Count occurrences for each `sii` level, including NaN values\nsii_counts = train_df['sii_label'].value_counts(dropna=False)\nsii_percentages = (sii_counts / sii_counts.sum()) * 100\n\n# Create DataFrame for plotting\nsii_data = pd.DataFrame({\n    'SII Level': sii_counts.index,\n    'Count': sii_counts.values,\n    'Percentage': sii_percentages.values\n})\n\n# Reorder categories to place NaN last\nsii_data['SII Level'] = pd.Categorical(\n    sii_data['SII Level'],\n    categories=[\"None\", \"Mild\", \"Moderate\", \"Severe\", \"NaN\"],\n    ordered=True\n)\nsii_data = sii_data.sort_values('SII Level')\n\n# Define new color mapping for sii levels\ncolors = {\n    \"None\": \"#6A5ACD\",       # Slate Blue\n    \"Mild\": \"#4682B4\",       # Steel Blue\n    \"Moderate\": \"#00CED1\",   # Dark Turquoise\n    \"Severe\": \"#DC143C\",     # Crimson\n    \"NaN\": \"#808080\"         # Gray for NaN\n}\n\n# Plot the bar chart using Plotly\nfig_bar = px.bar(\n    sii_data,\n    x='SII Level',\n    y='Count',\n    text='Count',\n    title=\"Severity Impairment Index (sii) Distribution\",\n    labels={'Count': 'Number of Participants', 'SII Level': 'SII Severity Level'},\n    hover_data={'Percentage': ':.2f'}\n)\n\n# Apply the new colors to each bar\nfig_bar.update_traces(marker=dict(color=[colors[level] for level in sii_data['SII Level']]))\n\n# Customize the layout\nfig_bar.update_layout(\n    title_font=dict(size=20, color=\"#002366\"),  # Navy Title\n    plot_bgcolor='white',\n    xaxis=dict(showgrid=False, tickfont=dict(size=12)),\n    yaxis=dict(showgrid=True, gridcolor='lightgrey', tickfont=dict(size=12)),\n    margin=dict(t=80, l=50, r=50, b=80)\n)\n\n# Show the bar chart\nfig_bar.show()\n\n# Calculate participants with non-null sii\ntotal_participants = len(train)\nnon_null_sii_count = train['sii'].notnull().sum()\nnull_sii_count = total_participants - non_null_sii_count\n\n# Data for the pie chart\nlabels = ['sii Not Null', 'sii Null']\nsizes = [non_null_sii_count, null_sii_count]\ncolors_pie = ['#4CAF50', '#FF6347']  # Green for not null, red for null\nexplode = (0.1, 0)  # Slightly explode the 'sii Not Null' slice\n\n# Create the pie chart\nplt.figure(figsize=(8, 8))\nplt.pie(\n    sizes,\n    labels=labels,\n    autopct=lambda p: f'{p:.1f}%\\n({int(p*total_participants/100):,})',\n    colors=colors_pie,\n    explode=explode,\n    startangle=140,\n    textprops={'fontsize': 12}\n)\n\n# Title for the pie chart\nplt.title('Participants with sii Availability', fontsize=16)\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:55:50.091711Z","iopub.execute_input":"2025-01-02T01:55:50.092091Z","iopub.status.idle":"2025-01-02T01:55:50.289755Z","shell.execute_reply.started":"2025-01-02T01:55:50.092060Z","shell.execute_reply":"2025-01-02T01:55:50.288549Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# Filter the DataFrame to include only rows where `sii` is not null\nfiltered_train = train[train['sii'].notnull()]\n\n# Calculate the percentage of missing values per column in the filtered DataFrame\nmissing_percent_filtered = (filtered_train.isnull().sum() / len(filtered_train)) * 100\n\n# Plot the missing percentages for the  features\nplt.figure(figsize=(10, 20))\nmissing_percent_filtered.sort_values().plot(kind=\"barh\", color=\"teal\", edgecolor=\"black\")\nplt.xlabel(\"Percentage of Missing Data\")\nplt.ylabel(\"Features\")\nplt.title(\"Missing Data Percentage per Feature (Non-Null `sii`)\")\nplt.grid(axis='x', linestyle='--', alpha=0.7)\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:38.789859Z","iopub.execute_input":"2025-01-02T01:04:38.790229Z","iopub.status.idle":"2025-01-02T01:04:39.825490Z","shell.execute_reply.started":"2025-01-02T01:04:38.790197Z","shell.execute_reply":"2025-01-02T01:04:39.824243Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# List of PCIAT features\npciat_columns = [f\"PCIAT-PCIAT_{i:02}\" for i in range(1, 21)]\n\n# Check for rows with all PCIAT data available\npciat_complete = train[pciat_columns].notnull().all(axis=1)\n\n# Count rows with complete and incomplete PCIAT data\ncomplete_count = pciat_complete.sum()  # Rows with all PCIAT columns filled\nincomplete_count = len(pciat_complete) - complete_count  # Rows with missing PCIAT columns\n\n# Data for the pie chart\nlabels = ['Complete PCIAT Data', 'Incomplete PCIAT Data']\ncounts = [complete_count, incomplete_count]\ncolors = ['#4CAF50', '#FF5722']  # Green and Red\n\n# Plot the pie chart\nplt.figure(figsize=(6, 6))\nplt.pie(\n    counts,\n    labels=labels,\n    autopct=lambda p: f'{p:.1f}%\\n({int(p*sum(counts)/100):,})',\n    startangle=90,\n    colors=colors,\n    textprops={'fontsize': 12}\n)\nplt.title(\"Completeness of PCIAT Data\", fontsize=16)\nplt.show()\nnull_sii = train['sii'].isnull().sum()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:39.826554Z","iopub.execute_input":"2025-01-02T01:04:39.826845Z","iopub.status.idle":"2025-01-02T01:04:39.981335Z","shell.execute_reply.started":"2025-01-02T01:04:39.826820Z","shell.execute_reply":"2025-01-02T01:04:39.980050Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Filter participants with non-null `sii`\nnon_null_sii = train[~train['sii'].isnull()]\n\n# Select the PCIAT columns\npciat_columns = [f\"PCIAT-PCIAT_{i:02}\" for i in range(1, 21)]\n\n# Count participants with all PCIAT data not null\ncomplete_pciat_data = non_null_sii[pciat_columns].notnull().all(axis=1).sum()\n\n# Count total participants with non-null `sii`\ntotal_non_null_sii = len(non_null_sii)\nincomplete_pciat_data = total_non_null_sii - complete_pciat_data\n\n# Data for the pie chart\nlabels = ['Complete PCIAT Data', 'Incomplete PCIAT Data']\ncounts = [complete_pciat_data, incomplete_pciat_data]\ncolors = ['#4CAF50', '#FF5722']  # Green and Red\n\n# Plot the pie chart\nplt.figure(figsize=(6, 6))\nplt.pie(\n    counts,\n    labels=labels,\n    autopct=lambda p: f'{p:.1f}%\\n({int(p*sum(counts)/100):,})',\n    startangle=90,\n    colors=colors,\n    textprops={'fontsize': 12}\n)\nplt.title(\"Completeness of PCIAT Data for Participants with Non-Null 'sii'\", fontsize=14)\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:39.982737Z","iopub.execute_input":"2025-01-02T01:04:39.983143Z","iopub.status.idle":"2025-01-02T01:04:40.155903Z","shell.execute_reply.started":"2025-01-02T01:04:39.983109Z","shell.execute_reply":"2025-01-02T01:04:40.154589Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# Calculate data for outer ring (SII completeness)\ntotal_participants = len(train)\nnull_sii = train['sii'].isnull().sum()\nnon_null_sii = total_participants - null_sii\n\n# Calculate data for inner ring (PCIAT completeness for non-null SII)\nnon_null_sii_df = train[~train['sii'].isnull()]\npciat_columns = [f\"PCIAT-PCIAT_{i:02}\" for i in range(1, 21)]\ncomplete_pciat_data = non_null_sii_df[pciat_columns].notnull().all(axis=1).sum()\nincomplete_pciat_data = len(non_null_sii_df) - complete_pciat_data\n\n# Print summary statistics\nprint(f\"Total participants: {total_participants}\")\nprint(f\"Participants with non-null 'sii': {non_null_sii}\")\nprint(f\"Participants with null 'sii': {null_sii} percentage of total participants : {null_sii/total_participants*100:.2f}%\")\nprint(f\"Participants with complete PCIAT (among non-null sii): {complete_pciat_data}\")\nprint(f\"Participants with incomplete PCIAT (among non-null sii): {incomplete_pciat_data}\")\nprint(f\"Percentage of non-null sii participants with complete PCIAT: {complete_pciat_data/non_null_sii*100:.2f}%\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:40.158653Z","iopub.execute_input":"2025-01-02T01:04:40.159075Z","iopub.status.idle":"2025-01-02T01:04:40.176658Z","shell.execute_reply.started":"2025-01-02T01:04:40.159038Z","shell.execute_reply":"2025-01-02T01:04:40.175312Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"For now, we can conclude that the SII score is sometimes incorrect, exactly for 65 participants. we can choose to remove those in further analysis.","metadata":{}},{"cell_type":"markdown","source":"## Different Features Analysis ","metadata":{}},{"cell_type":"markdown","source":"In this part we will analyse different features and their relationship to the taret variable sii and PCIAT columns, we will only proceed the analysis with participants with complete PCIAT (among non-null sii) for more trustful results\n","metadata":{}},{"cell_type":"code","source":"# List of PCIAT columns\npciat_columns = [f\"PCIAT-PCIAT_{i:02}\" for i in range(1, 21)]\n\n# Filter the DataFrame where 'sii' is not NaN and all PCIAT columns are not NaN\nfiltered_train = train.dropna(subset=['sii'] + pciat_columns, how='any')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:40.178509Z","iopub.execute_input":"2025-01-02T01:04:40.178930Z","iopub.status.idle":"2025-01-02T01:04:40.198090Z","shell.execute_reply.started":"2025-01-02T01:04:40.178894Z","shell.execute_reply":"2025-01-02T01:04:40.196678Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"filtered_train.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:40.199926Z","iopub.execute_input":"2025-01-02T01:04:40.200352Z","iopub.status.idle":"2025-01-02T01:04:40.219082Z","shell.execute_reply.started":"2025-01-02T01:04:40.200317Z","shell.execute_reply":"2025-01-02T01:04:40.218054Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Demographics","metadata":{}},{"cell_type":"code","source":"\n# Define the mapping for `sii` levels\nsii_labels = {0: \"None\", 1: \"Mild\", 2: \"Moderate\", 3: \"Severe\"}\n\n# Map `sii` values to their labels\nfiltered_train = filtered_train.copy()  # Create an explicit copy to avoid warnings\nfiltered_train['sii_label'] = filtered_train['sii'].map(sii_labels)\n\n# Set the order for the SII levels\nsii_order = [\"None\", \"Mild\", \"Moderate\", \"Severe\"]\nfiltered_train['sii_label'] = pd.Categorical(\n    filtered_train['sii_label'], \n    categories=sii_order, \n    ordered=True\n)\n\n# Calculate percentages for each `sii` level by sex\nsii_distribution = (\n    filtered_train\n    .groupby(['Basic_Demos-Sex', 'sii_label'])['id']\n    .count()\n    .reset_index()\n)\nsii_distribution['percentage'] = (\n    sii_distribution.groupby('Basic_Demos-Sex')['id']\n    .transform(lambda x: (x / x.sum()) * 100)\n)\n\n# Rename columns for clarity\nsii_distribution.rename(columns={'Basic_Demos-Sex': 'Sex', 'sii_label': 'SII Level'}, inplace=True)\n\n# Map Sex to readable labels\nsex_map = {0: 'Boys', 1: 'Girls'}  # Adjust as per your dataset\nsii_distribution['Sex'] = sii_distribution['Sex'].map(sex_map)\n\n# Plot\nplt.figure(figsize=(10, 6))\nsns.barplot(\n    data=sii_distribution,\n    x='SII Level',\n    y='percentage',\n    hue='Sex',\n    palette=['#1E90FF', '#FF69B4'],  # Blue for boys, Pink for girls\n    order=sii_order,  # Maintain order\n)\n\n# Customize the plot\nplt.title('SII Distribution by Sex (in Percentages)', fontsize=16)\nplt.xlabel('SII Severity Level', fontsize=14)\nplt.ylabel('Percentage (%)', fontsize=14)\nplt.legend(title='Sex', fontsize=12, title_fontsize=12)\nplt.grid(axis='y', linestyle='--', alpha=0.7)\n\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T02:43:45.185907Z","iopub.execute_input":"2025-01-02T02:43:45.186388Z","iopub.status.idle":"2025-01-02T02:43:45.551997Z","shell.execute_reply.started":"2025-01-02T02:43:45.186347Z","shell.execute_reply":"2025-01-02T02:43:45.550837Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- while both genders follow a similar pattern of decreasing percentages as severity increases, girls are more likely to have no impairment, while boys show slightly higher percentages in the moderate category.","metadata":{}},{"cell_type":"code","source":"# Ensure sii_label is categorized and ordered using .loc to avoid warnings\nfiltered_train.loc[:, 'sii_label'] = pd.Categorical(\n    filtered_train['sii_label'], \n    categories=[\"None\", \"Mild\", \"Moderate\", \"Severe\"], \n    ordered=True\n)\n\n# Plot\nplt.figure(figsize=(12, 6))\n\n# Create violin plot with box plot inside\nsns.violinplot(\n    data=filtered_train, \n    x='sii_label',\n    y='Basic_Demos-Age',\n    hue='Basic_Demos-Sex',\n    split=True,\n    inner='box',  # Shows box plot inside violin\n    palette=['#1E90FF', '#FF69B4'],  # Blue for boys, Pink for girls\n    order=[\"None\", \"Mild\", \"Moderate\", \"Severe\"]  # Ensure the correct order\n\n)\n\n# Customize the plot\nplt.title('Age Distribution by SII Severity Level and Sex', fontsize=16)\nplt.xlabel('SII Severity Level', fontsize=14)\nplt.ylabel('Age', fontsize=14)\n\n# Manually create a legend to fix the issue\nhandles = [\n    plt.Line2D([0], [0], color='#1E90FF', lw=4, label='Boys'),\n    plt.Line2D([0], [0], color='#FF69B4', lw=4, label='Girls')\n]\nplt.legend(handles=handles, title='Sex', fontsize=12, title_fontsize=12, loc='upper right')\n\n# Add grid for better readability\nplt.grid(axis='y', linestyle='--', alpha=0.7)\n\nplt.show()\n\n# Calculate and print some summary statistics\nprint(\"\\nAge Summary Statistics by SII Level:\")\nsummary_stats = filtered_train.groupby(['sii_label', 'Basic_Demos-Sex'])['Basic_Demos-Age'].agg(['mean', 'std', 'count']).round(2)\nprint(summary_stats)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T02:02:25.732725Z","iopub.execute_input":"2025-01-02T02:02:25.733111Z","iopub.status.idle":"2025-01-02T02:02:26.152772Z","shell.execute_reply.started":"2025-01-02T02:02:25.733083Z","shell.execute_reply":"2025-01-02T02:02:26.151713Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- Older children tend to have more severe impairment, and the progression from \"None\" → \"Mild\" → \"Moderate\" → \"Severe\" shows a gradual increase in mean age, suggesting a potential time-dependent progression of impairment severity.","metadata":{}},{"cell_type":"markdown","source":"### Internet Use","metadata":{}},{"cell_type":"code","source":"# Ensure you're working on a copy of the DataFrame to avoid SettingWithCopyWarning\nfiltered_train = filtered_train.copy()\n\n# Define the mapping for internet use categories\nparam_map = {0: '< 1h/day', 1: '~ 1h/day', 2: '~ 2hs/day', 3: '> 3hs/day'}\n\n# Map the raw hour data to categories and handle missing values safely\nfiltered_train.loc[:, 'internet_use_encoded'] = filtered_train['PreInt_EduHx-computerinternet_hoursday'].map(param_map).fillna('Missing')\n\n# Step 1: Group data and count occurrences\ninternet_table = filtered_train.groupby(['sii', 'internet_use_encoded'])['id'].count().unstack(fill_value=0)\n\n# Step 2: Calculate percentages\ninternet_table_percentage = internet_table.div(internet_table.sum(axis=1), axis=0) * 100\n\n# Step 3: Combine counts and percentages\nformatted_table = internet_table.astype(str) + \" (\" + internet_table_percentage.round(1).astype(str) + \"%)\"\n\n# Step 4: Dynamically adjust the column labels\nexpected_columns = ['Missing'] + list(param_map.values())\nformatted_table = formatted_table.reindex(columns=expected_columns, fill_value=\"0 (0.0%)\")\n\n# Update the index labels for SII levels\nformatted_table.index = ['0 (None)', '1 (Mild)', '2 (Moderate)', '3 (Severe)']\n\n# Display the formatted table\nformatted_table\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T02:04:31.857934Z","iopub.execute_input":"2025-01-02T02:04:31.858286Z","iopub.status.idle":"2025-01-02T02:04:31.880647Z","shell.execute_reply.started":"2025-01-02T02:04:31.858260Z","shell.execute_reply":"2025-01-02T02:04:31.879508Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create figure with two subplots\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(15, 6))\n\n# Plot 1: Box plot\nsns.boxplot(data=filtered_train, \n            x='sii_label',\n            y='PreInt_EduHx-computerinternet_hoursday',\n            palette='Set2',\n            order=[\"None\", \"Mild\", \"Moderate\", \"Severe\"],\n            ax=ax1)\n\nax1.set_title('Internet Usage Hours by SII Level', fontsize=14)\nax1.set_xlabel('SII Severity Level', fontsize=12)\nax1.set_ylabel('Hours per Day', fontsize=12)\nax1.grid(axis='y', linestyle='--', alpha=0.7)\n\n# Plot 2: Violin plot to show distribution\nsns.violinplot(data=filtered_train,\n               x='sii_label',\n               y='PreInt_EduHx-computerinternet_hoursday',\n               palette='Set2',\n               inner='box',\n               order=[\"None\", \"Mild\", \"Moderate\", \"Severe\"],\n               ax=ax2)\n\nax2.set_title('Internet Usage Hours Distribution by SII Level', fontsize=14)\nax2.set_xlabel('SII Severity Level', fontsize=12)\nax2.set_ylabel('Hours per Day', fontsize=12)\nax2.grid(axis='y', linestyle='--', alpha=0.7)\n\nplt.tight_layout()\nplt.show()\n\n# Calculate and print summary statistics\nprint(\"\\nInternet Usage Summary Statistics by SII Level:\")\nsummary_stats = filtered_train.groupby('sii_label')['PreInt_EduHx-computerinternet_hoursday'].agg([\n    'count',\n    'mean',\n    'median',\n    'std'\n]).round(2)\nprint(summary_stats)\n\n# Calculate correlation\ncorrelation = filtered_train['PreInt_EduHx-computerinternet_hoursday'].corr(filtered_train['sii'])\nprint(f\"\\nCorrelation between Internet Usage Hours and SII: {correlation:.3f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T02:07:50.164485Z","iopub.execute_input":"2025-01-02T02:07:50.164857Z","iopub.status.idle":"2025-01-02T02:07:50.821331Z","shell.execute_reply.started":"2025-01-02T02:07:50.164818Z","shell.execute_reply":"2025-01-02T02:07:50.820373Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Distribution Patterns (from violin plot):\n\n\n- Each severity level shows a bimodal distribution (two peaks)\n- The spread (width) of the distributions is fairly similar across levels\n- The overall shape suggests there might be distinct usage patterns within each severity level\n\n\nMedian Values (from box plot):\n\n\n- \"None\" group has median at 0 hours\n- \"Mild\" and above show higher medians\n- There's a clear upward trend in median usage as severity increases\n  \nThere's a moderate positive correlation (0.333) between SII severity and internet usage, this could suggest either that:\n\n- Higher internet usage might contribute to social impairment, or\n- People with higher social impairment tend to spend more time on the internet\n\nCreating an interaction feature between internet use and age could potentially be useful for modeling.\n","metadata":{}},{"cell_type":"code","source":"groups = data_dict.groupby('Instrument')['Field'].apply(list).to_dict()\n\nfor instrument, features in groups.items():\n    print(f\"{instrument}: {features}\\n\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:41.781636Z","iopub.execute_input":"2025-01-02T01:04:41.781990Z","iopub.status.idle":"2025-01-02T01:04:41.804036Z","shell.execute_reply.started":"2025-01-02T01:04:41.781951Z","shell.execute_reply":"2025-01-02T01:04:41.802808Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# Remove the PCIAT group from the dictionary\nfiltered_groups = {k: v for k, v in groups.items() if k != \"Parent-Child Internet Addiction Test\"}\n\n# Convert the filtered groups dictionary to a dataframe\ngroup_data = []\nfor instrument, features in filtered_groups.items():\n    for feature in features:\n        group_data.append({'Instrument': instrument, 'Feature': feature})\n\ngroup_df = pd.DataFrame(group_data)\n\n# Create the sunburst plot\nfig = px.sunburst(\n    group_df,\n    path=['Instrument', 'Feature'],  # Define hierarchy\n    title='Feature Group Hierarchy (Excluding PCIAT)',\n    color='Instrument',\n    color_discrete_sequence=px.colors.qualitative.Set3,\n)\n\n# Adjust layout for a larger chart\nfig.update_layout(\n    margin=dict(t=40, l=0, r=0, b=0),\n    width=1000,  # Adjust the width\n    height=800   # Adjust the height\n)\n\nfig.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:41.806121Z","iopub.execute_input":"2025-01-02T01:04:41.806443Z","iopub.status.idle":"2025-01-02T01:04:41.923689Z","shell.execute_reply.started":"2025-01-02T01:04:41.806414Z","shell.execute_reply":"2025-01-02T01:04:41.922640Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Children's Global assessement scale ","metadata":{}},{"cell_type":"code","source":"# Drop null values for CGAS score in the entire dataset\ncgas_full_stats = train['CGAS-CGAS_Score'].dropna().describe(percentiles=[0.25, 0.5, 0.75]).round(2)\n\n# Create a structured table\ncgas_full_stats_table = pd.DataFrame({\n    \"Statistic\": [\"Count\", \"Mean\", \"Standard Deviation\", \"Minimum\", \"25th Percentile (Q1)\", \"Median (Q2)\", \"75th Percentile (Q3)\", \"Maximum\"],\n    \"Value\": [\n        cgas_full_stats['count'],\n        cgas_full_stats['mean'],\n        cgas_full_stats['std'],\n        cgas_full_stats['min'],\n        cgas_full_stats['25%'],\n        cgas_full_stats['50%'],\n        cgas_full_stats['75%'],\n        cgas_full_stats['max']\n    ]\n})\n\n# Display the table\ncgas_full_stats_table\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:41.924734Z","iopub.execute_input":"2025-01-02T01:04:41.925118Z","iopub.status.idle":"2025-01-02T01:04:41.940684Z","shell.execute_reply.started":"2025-01-02T01:04:41.925082Z","shell.execute_reply":"2025-01-02T01:04:41.939650Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"--> Contains outliers : maximum value 999, CGAS score can't be over 100 : https://www.corc.uk.net/outcome-experience-measures/childrens-global-assessment-scale-cgas/","metadata":{}},{"cell_type":"code","source":"train[train['CGAS-CGAS_Score'] > 100]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:41.941596Z","iopub.execute_input":"2025-01-02T01:04:41.941920Z","iopub.status.idle":"2025-01-02T01:04:41.961856Z","shell.execute_reply.started":"2025-01-02T01:04:41.941894Z","shell.execute_reply":"2025-01-02T01:04:41.960939Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# Remove null values and outliers\ncgas_cleaned = train['CGAS-CGAS_Score'].dropna()\ncgas_cleaned = cgas_cleaned[cgas_cleaned <= 100]  # Exclude outliers above 100\n\n# Plot the distribution\nplt.figure(figsize=(10, 6))\nsns.histplot(cgas_cleaned, bins=20, kde=True, color=\"#1E90FF\")\n\n# Customize the plot\nplt.title(\"Distribution of CGAS-CGAS_Score\", fontsize=16)\nplt.xlabel(\"CGAS Score\", fontsize=14)\nplt.ylabel(\"Frequency\", fontsize=14)\nplt.grid(axis=\"y\", linestyle=\"--\", alpha=0.7)\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:41.962769Z","iopub.execute_input":"2025-01-02T01:04:41.963064Z","iopub.status.idle":"2025-01-02T01:04:42.387554Z","shell.execute_reply.started":"2025-01-02T01:04:41.963040Z","shell.execute_reply":"2025-01-02T01:04:42.386234Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Filtered dataset\ncgas_filtered = filtered_train[['CGAS-CGAS_Score', 'sii_label']].dropna()\n\n# Step 1: Boxplot\nplt.figure(figsize=(10, 6))\nsns.boxplot(\n    data=cgas_filtered,\n    x='sii_label',\n    y='CGAS-CGAS_Score',\n    palette='Set2',\n    order=[\"None\", \"Mild\", \"Moderate\", \"Severe\"]\n)\n\n# Customize the plot\nplt.title(\"CGAS Scores by SII Severity Level\", fontsize=16)\nplt.xlabel(\"SII Severity Level\", fontsize=14)\nplt.ylabel(\"CGAS Score\", fontsize=14)\nplt.grid(axis=\"y\", linestyle=\"--\", alpha=0.7)\n\nplt.show()\n\n# Step 2: Summary Statistics\nsummary_stats = cgas_filtered.groupby('sii_label')['CGAS-CGAS_Score'].agg(['mean', 'median', 'std', 'count']).round(2)\nprint(\"\\nSummary Statistics:\")\nprint(summary_stats)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:42.388542Z","iopub.execute_input":"2025-01-02T01:04:42.388857Z","iopub.status.idle":"2025-01-02T01:04:42.638300Z","shell.execute_reply.started":"2025-01-02T01:04:42.388820Z","shell.execute_reply":"2025-01-02T01:04:42.637470Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Slight decrease of mean CGAS score as we the sii increases.","metadata":{}},{"cell_type":"markdown","source":"### Physical Measures","metadata":{}},{"cell_type":"code","source":"\n# List of Physical Measures columns with units\nphysical_columns = [\n    ('Physical-BMI', 'Body Mass Index (kg/m²)'),\n    ('Physical-Height', 'Height (in)'),\n    ('Physical-Weight', 'Weight (lbs)'),\n    ('Physical-Waist_Circumference', 'Waist Circumference (in)'),\n    ('Physical-Diastolic_BP', 'Diastolic BP (mmHg)'),\n    ('Physical-HeartRate', 'Heart Rate (beats/min)'),\n    ('Physical-Systolic_BP', 'Systolic BP (mmHg)')\n]\n\n# Create subplots for distributions (4 plots in the first row, 3 in the second row)\nplt.figure(figsize=(18, 10))\nfor i, (col, label) in enumerate(physical_columns, start=1):\n    plt.subplot(2, 4, i)  # 2 rows, 4 columns layout (adjust for 7 plots)\n    sns.histplot(data=filtered_train, x=col, kde=True, bins=30, color='skyblue', edgecolor='black')\n    plt.title(f\"Distribution of {label}\", fontsize=14)\n    plt.xlabel(label, fontsize=12)\n    plt.ylabel(\"Frequency\", fontsize=12)\n    plt.grid(axis='y', linestyle='--', alpha=0.7)\n\n# Adjust layout to prevent overlapping\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:42.639198Z","iopub.execute_input":"2025-01-02T01:04:42.639451Z","iopub.status.idle":"2025-01-02T01:04:45.418097Z","shell.execute_reply.started":"2025-01-02T01:04:42.639430Z","shell.execute_reply":"2025-01-02T01:04:45.416942Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Convert units for Physical-Weight (lbs to kg) and Physical-Height (inches to cm) using .loc to avoid the warning\nfiltered_train.loc[:, 'Physical-Weight'] = filtered_train['Physical-Weight'] * 0.453592\nfiltered_train.loc[:, 'Physical-Height'] = filtered_train['Physical-Height'] * 2.54\nfiltered_train.loc[:, 'Physical-Waist_Circumference'] = filtered_train['Physical-Waist_Circumference'] * 2.54\n\n\n# Drop NaN values for each column independently and calculate stats\nstats_table = pd.DataFrame()\n\nfor col in ['Physical-BMI', 'Physical-Height', 'Physical-Weight', 'Physical-Waist_Circumference', \n            'Physical-Diastolic_BP', 'Physical-HeartRate', 'Physical-Systolic_BP']:\n    stats_table[col] = filtered_train[col].dropna().describe()\n\n# Display the stats table\nstats_table.T","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:45.419090Z","iopub.execute_input":"2025-01-02T01:04:45.419457Z","iopub.status.idle":"2025-01-02T01:04:45.457522Z","shell.execute_reply.started":"2025-01-02T01:04:45.419425Z","shell.execute_reply":"2025-01-02T01:04:45.456474Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"lots of these values are biologically impossible (0 BMI, 0 weight) or out of possible range for a child ( weight 142 kg, waist circumference 127 cm)\nsystolic BP cannot be lower than diastolic BP","metadata":{}},{"cell_type":"code","source":"train[train['Physical-Systolic_BP'] <= train['Physical-Diastolic_BP']][[\n    'Physical-Diastolic_BP', 'Physical-Systolic_BP',\n    'Physical-HeartRate'\n]]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:45.458548Z","iopub.execute_input":"2025-01-02T01:04:45.458932Z","iopub.status.idle":"2025-01-02T01:04:45.471817Z","shell.execute_reply.started":"2025-01-02T01:04:45.458899Z","shell.execute_reply":"2025-01-02T01:04:45.470932Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"error_ranges = {\n    'Physical-BMI': (10, 50),        # Extreme but possible cases exist\n    'Physical-Height': (100, 220),   # From small 5-year-old to very tall adult\n    'Physical-Weight': (15, 200),    # From small 5-year-old to severe obesity\n    'Physical-Waist_Circumference': (40, 150),  # Extreme but possible range\n    'Physical-Diastolic_BP': (40, 120),  # Including very low/high but possible\n    'Physical-HeartRate': (40, 200),     # Including athletic low and exercise high\n    'Physical-Systolic_BP': (70, 180)    # Including very low/high but possible\n}\n\n# Create a function to identify errors\ndef identify_measurement_errors(df, ranges):\n    # Dictionary to store erroneous rows for each measurement\n    errors_by_measurement = {}\n    \n    # For each measurement type\n    for measurement, (min_val, max_val) in ranges.items():\n        # Get rows where values are outside the range or null\n        errors = df[\n            (df[measurement] < min_val) | \n            (df[measurement] > max_val) |\n            (df[measurement].isna())\n        ]\n        \n        # Store in dictionary\n        errors_by_measurement[measurement] = errors\n        \n        # Print summary\n        print(f\"\\n{measurement}:\")\n        print(f\"Total errors: {len(errors)}\")\n        if len(errors) > 0:\n            print(f\"Below {min_val}: {len(df[df[measurement] < min_val])}\")\n            print(f\"Above {max_val}: {len(df[df[measurement] > max_val])}\")\n            print(f\"Null values: {df[measurement].isna().sum()}\")\n    \n    # Get rows with any measurement error\n    any_error = pd.DataFrame()\n    for measurement in ranges.keys():\n        mask = (\n            (df[measurement] < ranges[measurement][0]) |\n            (df[measurement] > ranges[measurement][1]) |\n            (df[measurement].isna())\n        )\n        if any_error.empty:\n            any_error = df[mask]\n        else:\n            any_error = pd.concat([any_error, df[mask]]).drop_duplicates()\n    \n    return errors_by_measurement, any_error\n\n# Run the analysis\nerrors_by_measurement, any_error = identify_measurement_errors(filtered_train, error_ranges)\n\n# Print overall summary\nprint(\"\\nOverall Summary:\")\nprint(f\"Total rows in dataset: {len(filtered_train)}\")\nprint(f\"Rows with at least one measurement error: {len(any_error)}\")\nprint(f\"Percentage of rows with errors: {(len(any_error)/len(filtered_train))*100:.2f}%\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:45.472994Z","iopub.execute_input":"2025-01-02T01:04:45.473298Z","iopub.status.idle":"2025-01-02T01:04:45.619722Z","shell.execute_reply.started":"2025-01-02T01:04:45.473258Z","shell.execute_reply":"2025-01-02T01:04:45.618779Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Bio-electric Impedance Analysis","metadata":{}},{"cell_type":"code","source":"data_dict[data_dict['Instrument']=='Bio-electric Impedance Analysis']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T02:57:02.121082Z","iopub.execute_input":"2025-01-02T02:57:02.121546Z","iopub.status.idle":"2025-01-02T02:57:02.140047Z","shell.execute_reply.started":"2025-01-02T02:57:02.121499Z","shell.execute_reply":"2025-01-02T02:57:02.139099Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n\nbia_columns = [col for col in train.columns if col.startswith(\"BIA-\")]\n\n# Create a new DataFrame with only the BIA columns\nbia_df = train[bia_columns]\n\n# Calculate statistics for the BIA data columns\nbia_stats = bia_df.describe(include='all').T\nbia_stats[\"missing_values\"] = bia_df.isnull().sum()\nbia_stats[\"missing_percentage\"] = (bia_stats[\"missing_values\"] / len(bia_df)) * 100\nbia_stats = bia_stats[[\"count\", \"mean\", \"std\", \"min\", \"25%\", \"50%\", \"75%\", \"max\", \"missing_values\", \"missing_percentage\"]]\n\n# Display the table\nbia_stats","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T03:02:23.135791Z","iopub.execute_input":"2025-01-02T03:02:23.136177Z","iopub.status.idle":"2025-01-02T03:02:23.189177Z","shell.execute_reply.started":"2025-01-02T03:02:23.136148Z","shell.execute_reply":"2025-01-02T03:02:23.188202Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### FitnessGram Vitals and Treadmill","metadata":{}},{"cell_type":"code","source":"data_dict.head(20)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:45.646211Z","iopub.execute_input":"2025-01-02T01:04:45.646490Z","iopub.status.idle":"2025-01-02T01:04:45.667817Z","shell.execute_reply.started":"2025-01-02T01:04:45.646465Z","shell.execute_reply":"2025-01-02T01:04:45.666790Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Columns to analyze\nfitnessgram_columns = [\n    'Fitness_Endurance-Max_Stage',\n    'Fitness_Endurance-Time_Mins',\n    'Fitness_Endurance-Time_Sec',\n]\n\n# Set up the figure\nfig, axes = plt.subplots(1, 3, figsize=(20, 5))\n\n# Flatten axes for easy iteration\naxes = axes.flatten()\n\n# Plot the distributions\nfor i, col in enumerate(fitnessgram_columns):\n    if col == 'Fitness_Endurance-Season':  # Handle categorical data for Season\n        sns.countplot(x=col, data=filtered_train, ax=axes[i], palette=\"Set2\")\n        axes[i].set_title(f'Distribution of {col}')\n        axes[i].set_xlabel('Season')\n        axes[i].set_ylabel('Count')\n    else:  # Handle numerical data\n        sns.histplot(filtered_train[col].dropna(), kde=True, ax=axes[i], color='skyblue', bins=30)\n        axes[i].set_title(f'Distribution of {col}')\n        axes[i].set_xlabel(col)\n        axes[i].set_ylabel('Frequency')\n\n# Adjust layout\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:45.668816Z","iopub.execute_input":"2025-01-02T01:04:45.669108Z","iopub.status.idle":"2025-01-02T01:04:47.095180Z","shell.execute_reply.started":"2025-01-02T01:04:45.669083Z","shell.execute_reply":"2025-01-02T01:04:47.093936Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- On average, participants reached stage 5 in the endurance test.\n- Some participants failed to complete the first stage (min = 0), or these are errors in data again.","metadata":{}},{"cell_type":"markdown","source":"### FitnessGram child","metadata":{}},{"cell_type":"code","source":"data_dict[data_dict['Instrument'] == 'FitnessGram Child']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:47.096155Z","iopub.execute_input":"2025-01-02T01:04:47.096434Z","iopub.status.idle":"2025-01-02T01:04:47.109315Z","shell.execute_reply.started":"2025-01-02T01:04:47.096411Z","shell.execute_reply":"2025-01-02T01:04:47.108154Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# Columns to analyze (non-categorical)\nfitnessgram_child_columns = {\n    'FGC-FGC_CU': 'Curl Up Total',\n    'FGC-FGC_GSND': 'Grip Strength Total (Non-Dominant)',\n    'FGC-FGC_GSD': 'Grip Strength Total (Dominant)',\n    'FGC-FGC_PU': 'Push-Up Total',\n    'FGC-FGC_SRL': 'Sit & Reach Total (Left Side)',\n    'FGC-FGC_SRR': 'Sit & Reach Total (Right Side)',\n    'FGC-FGC_TL': 'Trunk Lift Total'\n}\n\n# Set up the figure\nfig, axes = plt.subplots(3, 3, figsize=(18, 15))\n\n# Flatten axes for easy iteration\naxes = axes.flatten()\n\n# Plot the distributions\nfor i, (col, label) in enumerate(fitnessgram_child_columns.items()):\n    if i < len(axes):  # Ensure we don't exceed available axes\n        sns.histplot(filtered_train[col].dropna(), kde=True, ax=axes[i], color='skyblue', bins=30)\n        axes[i].set_title(f'Distribution of {label}')\n        axes[i].set_xlabel(label)\n        axes[i].set_ylabel('Frequency')\n\n# Remove unused subplots\nfor j in range(len(fitnessgram_child_columns), len(axes)):\n    fig.delaxes(axes[j])\n\n# Adjust layout\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:47.110395Z","iopub.execute_input":"2025-01-02T01:04:47.110721Z","iopub.status.idle":"2025-01-02T01:04:50.144068Z","shell.execute_reply.started":"2025-01-02T01:04:47.110694Z","shell.execute_reply":"2025-01-02T01:04:50.142955Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Columns to analyze (non-categorical)\nfitnessgram_child_columns = [\n    'FGC-FGC_CU',  # Curl Up Total\n    'FGC-FGC_GSND',  # Grip Strength Total (Non-Dominant)\n    'FGC-FGC_GSD',  # Grip Strength Total (Dominant)\n    'FGC-FGC_PU',  # Push-Up Total\n    'FGC-FGC_SRL',  # Sit & Reach Total (Left Side)\n    'FGC-FGC_SRR',  # Sit & Reach Total (Right Side)\n    'FGC-FGC_TL'  # Trunk Lift Total\n]\n\n# Calculate stats\nstats = filtered_train[fitnessgram_child_columns].describe().T\n\n# Add missing values count\nstats['missing_values'] = filtered_train[fitnessgram_child_columns].isna().sum()\n\n# Rename columns for clarity\nstats = stats.rename(\n    columns={\n        'count': 'Valid Count',\n        'mean': 'Mean',\n        'std': 'Std Dev',\n        'min': 'Min',\n        '25%': '25th Percentile',\n        '50%': 'Median',\n        '75%': '75th Percentile',\n        'max': 'Max'\n    }\n)\n\n# Format the stats table\nstats = stats[['Valid Count', 'missing_values', 'Mean', 'Std Dev', 'Min', '25th Percentile', 'Median', '75th Percentile', 'Max']]\n\n# Display the table\nstats\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:50.144979Z","iopub.execute_input":"2025-01-02T01:04:50.145281Z","iopub.status.idle":"2025-01-02T01:04:50.183218Z","shell.execute_reply.started":"2025-01-02T01:04:50.145253Z","shell.execute_reply":"2025-01-02T01:04:50.182282Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Sleep disturbance Scale ","metadata":{}},{"cell_type":"code","source":"# Columns of interest\nsleep_columns = ['SDS-SDS_Total_Raw', 'SDS-SDS_Total_T']\n\n# Initialize a dictionary to store stats\nsleep_stats = {}\n\nfor col in sleep_columns:\n    stats = filtered_train[col].dropna().describe()  # Stats for non-missing values\n    missing_count = filtered_train[col].isna().sum()  # Count of missing values\n    stats['missing'] = missing_count  # Add missing count to stats\n    sleep_stats[col] = stats\n\n# Convert the stats dictionary to a DataFrame\nsleep_stats_df = pd.DataFrame(sleep_stats).T\nsleep_stats_df.columns = ['count', 'mean', 'std', 'min', '25%', '50%', '75%', 'max', 'missing']\n\nsleep_stats_df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:50.184173Z","iopub.execute_input":"2025-01-02T01:04:50.184505Z","iopub.status.idle":"2025-01-02T01:04:50.209406Z","shell.execute_reply.started":"2025-01-02T01:04:50.184468Z","shell.execute_reply":"2025-01-02T01:04:50.208225Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Columns of interest\nsleep_columns = ['SDS-SDS_Total_Raw', 'SDS-SDS_Total_T']\n\n# Create subplots for distributions\nplt.figure(figsize=(12, 5))\nfor i, col in enumerate(sleep_columns, 1):\n    plt.subplot(1, 2, i)\n    sns.histplot(filtered_train[col].dropna(), kde=True, bins=30, color='skyblue')\n    plt.title(f'Distribution of {col}', fontsize=14)\n    plt.xlabel(col, fontsize=12)\n    plt.ylabel('Frequency', fontsize=12)\nplt.tight_layout()\nplt.show()\n\n# Create subplots for boxplots against 'sii_severity_level'\nplt.figure(figsize=(12, 5))\nfor i, col in enumerate(sleep_columns, 1):\n    plt.subplot(1, 2, i)\n    sns.boxplot(\n        data=filtered_train, \n        x='sii', \n        y=col, \n        palette='viridis'\n    )\n    plt.title(f'{col} by SII Severity Level', fontsize=14)\n    plt.xlabel('SII Severity Level', fontsize=12)\n    plt.ylabel(col, fontsize=12)\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:50.210740Z","iopub.execute_input":"2025-01-02T01:04:50.211091Z","iopub.status.idle":"2025-01-02T01:04:51.755205Z","shell.execute_reply.started":"2025-01-02T01:04:50.211062Z","shell.execute_reply":"2025-01-02T01:04:51.754245Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_dict[data_dict['Instrument']=='Physical Activity Questionnaire (Adolescents)']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:51.756086Z","iopub.execute_input":"2025-01-02T01:04:51.756350Z","iopub.status.idle":"2025-01-02T01:04:51.767239Z","shell.execute_reply.started":"2025-01-02T01:04:51.756328Z","shell.execute_reply":"2025-01-02T01:04:51.766272Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_dict[data_dict['Instrument']=='Physical Activity Questionnaire (Children)']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:51.768367Z","iopub.execute_input":"2025-01-02T01:04:51.768730Z","iopub.status.idle":"2025-01-02T01:04:51.788172Z","shell.execute_reply.started":"2025-01-02T01:04:51.768696Z","shell.execute_reply":"2025-01-02T01:04:51.787138Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Columns for analysis\nactivity_columns = ['PAQ_A-PAQ_A_Total', 'PAQ_C-PAQ_C_Total']\n\n# Prepare data for stats table\nactivity_stats = train[activity_columns].describe().T\nactivity_stats['missing'] = train[activity_columns].isnull().sum()\nactivity_stats.rename(columns={'count': 'non-missing'}, inplace=True)\n\n# Plot the distributions\nfig, axes = plt.subplots(1, 2, figsize=(14, 6))\nfor ax, col in zip(axes, activity_columns):\n    sns.histplot(train[col].dropna(), kde=True, ax=ax, bins=30, color='teal')\n    ax.set_title(f\"Distribution of {col}\")\n    ax.set_xlabel('Activity Summary Score')\n    ax.set_ylabel('Frequency')\n\nplt.tight_layout()\nplt.show()\n\nactivity_stats","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:51.789109Z","iopub.execute_input":"2025-01-02T01:04:51.789393Z","iopub.status.idle":"2025-01-02T01:04:52.839793Z","shell.execute_reply.started":"2025-01-02T01:04:51.789368Z","shell.execute_reply":"2025-01-02T01:04:52.838653Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Check the age range for rows where PAQ_A-PAQ_A_Total is not null\nage_range_paq_a = train[train['PAQ_A-PAQ_A_Total'].notnull()]['Basic_Demos-Age'].agg(['min', 'max'])\nprint(\"Age range for PAQ_A-PAQ_A_Total not null:\")\nprint(age_range_paq_a)\n\n# Check the age range for rows where PAQ_C-PAQ_C_Total is not null\nage_range_paq_c = train[train['PAQ_C-PAQ_C_Total'].notnull()]['Basic_Demos-Age'].agg(['min', 'max'])\nprint(\"\\nAge range for PAQ_C-PAQ_C_Total not null:\")\nprint(age_range_paq_c)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:52.840727Z","iopub.execute_input":"2025-01-02T01:04:52.841006Z","iopub.status.idle":"2025-01-02T01:04:52.853152Z","shell.execute_reply.started":"2025-01-02T01:04:52.840982Z","shell.execute_reply":"2025-01-02T01:04:52.852136Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Actigraphy Data (Time Series)","metadata":{}},{"cell_type":"markdown","source":"During their participation in the HBN study, some participants were given an accelerometer to wear for up to 30 days continually while at home and going about their regular daily lives.","metadata":{}},{"cell_type":"code","source":"path = '/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet/id=00115b9f/part-0.parquet'\nseries_train = pd.read_parquet(path)\nseries_train.head(20)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:52.854175Z","iopub.execute_input":"2025-01-02T01:04:52.854541Z","iopub.status.idle":"2025-01-02T01:04:53.080713Z","shell.execute_reply.started":"2025-01-02T01:04:52.854495Z","shell.execute_reply":"2025-01-02T01:04:53.079712Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# Select important columns for time series plotting\nimportant_columns = ['X', 'Y', 'Z', 'enmo', 'anglez', 'light']\ntime_steps = series_train['step']\ncolors = ['b', 'g', 'r', 'c', 'm', 'y']  # Assign a unique color for each column\n\n# Create subplots for each column\nn_cols = 1\nn_rows = len(important_columns)\n\nfig, axes = plt.subplots(n_rows, n_cols, figsize=(12, n_rows * 3))\n\nfor i, column in enumerate(important_columns):\n    ax = axes[i] if n_rows > 1 else axes\n    ax.plot(time_steps, series_train[column], label=column, color=colors[i], alpha=0.8)\n    ax.set_title(f\"Time Series of {column}\")\n    ax.set_xlabel(\"Time Step\")\n    ax.set_ylabel(column)\n    ax.legend()\n    ax.grid(True)\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:53.081713Z","iopub.execute_input":"2025-01-02T01:04:53.082105Z","iopub.status.idle":"2025-01-02T01:04:55.393891Z","shell.execute_reply.started":"2025-01-02T01:04:53.082068Z","shell.execute_reply":"2025-01-02T01:04:55.392772Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train[train['id'] == '00115b9f']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:55.394896Z","iopub.execute_input":"2025-01-02T01:04:55.395200Z","iopub.status.idle":"2025-01-02T01:04:55.416154Z","shell.execute_reply.started":"2025-01-02T01:04:55.395173Z","shell.execute_reply":"2025-01-02T01:04:55.415144Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"makes sense,but we should take a look at a high sii case and compare","metadata":{}},{"cell_type":"code","source":"\n# Paths\nseries_folder = '/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet/'\n\n# Load the tabular data\ntrain_tabular = train.copy()\n# Iterate over the parquet files\nfor root, _, files in os.walk(series_folder):\n    for file in files:\n        if file.endswith('.parquet'):\n            # Extract ID from the file path\n            file_path = os.path.join(root, file)\n            id_part = file_path.split('id=')[-1].split('/')[0]\n            \n            # Check the ID in the tabular data\n            sii_value = train_tabular.loc[train_tabular['id'] == id_part, 'sii'].values\n            if sii_value.size > 0:  # Ensure ID exists in the tabular data\n                sii_value = sii_value[0]\n                if sii_value in [2]:\n                    print(f\"Found ID with SII {sii_value}: {id_part}\")\n                    print(f\"Time series file: {file_path}\")\n                    break\n    else:\n        # Continue if no break occurred\n        continue\n    # Break outer loop if a match is found\n    break","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:55.417489Z","iopub.execute_input":"2025-01-02T01:04:55.417895Z","iopub.status.idle":"2025-01-02T01:04:55.449398Z","shell.execute_reply.started":"2025-01-02T01:04:55.417854Z","shell.execute_reply":"2025-01-02T01:04:55.448536Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train[train['id'] == '7b8842c3']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:55.450255Z","iopub.execute_input":"2025-01-02T01:04:55.450598Z","iopub.status.idle":"2025-01-02T01:04:55.471003Z","shell.execute_reply.started":"2025-01-02T01:04:55.450564Z","shell.execute_reply":"2025-01-02T01:04:55.469815Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"path = '/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet/id=7b8842c3/part-0.parquet'\nseries_train = pd.read_parquet(path)\nseries_train.head(20)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:55.471933Z","iopub.execute_input":"2025-01-02T01:04:55.472273Z","iopub.status.idle":"2025-01-02T01:04:55.516946Z","shell.execute_reply.started":"2025-01-02T01:04:55.472240Z","shell.execute_reply":"2025-01-02T01:04:55.515876Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# Select important columns for time series plotting\nimportant_columns = ['X', 'Y', 'Z', 'enmo', 'anglez', 'light']\ntime_steps = series_train['step']\ncolors = ['b', 'g', 'r', 'c', 'm', 'y']  # Assign a unique color for each column\n\n# Create subplots for each column\nn_cols = 1\nn_rows = len(important_columns)\n\nfig, axes = plt.subplots(n_rows, n_cols, figsize=(12, n_rows * 3))\n\nfor i, column in enumerate(important_columns):\n    ax = axes[i] if n_rows > 1 else axes\n    ax.plot(time_steps, series_train[column], label=column, color=colors[i], alpha=0.8)\n    ax.set_title(f\"Time Series of {column}\")\n    ax.set_xlabel(\"Time Step\")\n    ax.set_ylabel(column)\n    ax.legend()\n    ax.grid(True)\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-02T01:04:55.517916Z","iopub.execute_input":"2025-01-02T01:04:55.518216Z","iopub.status.idle":"2025-01-02T01:04:57.700569Z","shell.execute_reply.started":"2025-01-02T01:04:55.518189Z","shell.execute_reply":"2025-01-02T01:04:57.699399Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## General Insights \n- In 1,224 records in the train data the sii target and all the PCIAT columns are missing - presumably because not available.\n- Overall there are > 100,000 missing values in the train data.\n- Only 2,736 records have a target, the rest are missing.\n- 996 of the young people also have sensor data from a worn device which measures gross motor activity.\n- There are large numbers of missing values remaining in the dataset, with 46 columns missing values and 8 columns missing more than half of their values.\nFor example, most of the waist circumference feature values are missing.","metadata":{}},{"cell_type":"markdown","source":"# Modelling Synthesis","metadata":{}},{"cell_type":"markdown","source":"## First Attempt: Basic Data Cleaning and Encoding\n\n**Data Cleaning:**\n\n* Dropped columns with more than 50% missing values.\n* Imputed remaining missing values using the most frequent strategy.\n* Encoded categorical variables using ordinal and one-hot encoding (get_dummies).\n\n**Models Trained:**\n\n* Random Forest Classifier with stratified K-fold cross-validation and grid search for hyperparameter tuning.\n* XGBoost Classifier.\n\n**Random Forest Classifier Results**\n\n| Metric | Train Data | Test Data |\n|---|---|---|\n| Best CV QWK | 0.3023 | - |\n| Training QWK | 0.9895 | - |\n| Private Score QWK | - | 0.10 |\n| Public Score QWK | - | 0.21 |\n\n**Random Forest with Regularization**\n\n| Metric | Train Data | Test Data |\n|---|---|---|\n| Mean Stratified CV QWK | 0.3960 | - |\n| Training QWK | 0.7571 | - |\n| Private Score QWK | - | 0.42 |\n| Public Score QWK | - | 0.40 |\n\n**XGBoost Classifier Results**\n\n| Metric | Train Data | Test Data |\n|---|---|---|\n| Mean Stratified CV QWK | 0.3208 | - |\n| Mean Training QWK | 0.9894 | - |\n| Private Score QWK | - | 0.24 |\n| Public Score QWK | - | 0.32 |\n\n**XGBoost with Regularization**\n\n| Metric | Train Data | Test Data |\n|---|---|---|\n| Mean Stratified CV QWK | 0.3621 | - |\n| Mean Training QWK | 0.6631 | - |\n| Private Score QWK | - | 0.23 |\n| Public Score QWK | - | 0.32 |\n\n**Observations**\n\n* The Random Forest model exhibited significant overfitting in its initial run, as indicated by a high training QWK and poor test performance.\n* Introducing regularization improved the Random Forest's generalization capability, resulting in the best test performance (Private Score: 0.42, Public Score: 0.40).\n* The XGBoost classifier performed slightly worse than Random Forest, even after introducing regularization.\n* Overall, Random Forest with regularization achieved the best results among all models tested.\n","metadata":{}},{"cell_type":"markdown","source":"## Second Attempt: Regression Approach\n\n### Problem Reformulation\n\nIn this attempt, the problem was approached as a regression task:\n\n* **Objective:** Predict the PCIAT_Total value and map the predicted value to the corresponding SII severity level based on predefined intervals.\n\n### Enhanced Data Cleaning\n\n* Replaced implausible values in various columns with NaN to ensure data consistency. \n* Assigned age groups instead of using raw age values for better generalization (e.g., 18-25, 26-35, etc.) to account for potential non-linear relationships between age and PCIAT_Total.\n* Combined multiple BMI-related columns into a single feature.\n* Retained similar imputation and encoding strategies as in the first attempt:\n    * Imputation: Most frequent strategy.\n    * Encoding: Ordinal encoding for ordinal variables and one-hot encoding for nominal variables.\n\n### Modeling Approach\n\n* Trained three variations of a Random Forest Regressor with different hyperparameter configurations.\n* Used stratified K-fold cross-validation to evaluate model performance.\n\n## Results\n\nUnfortunately, none of the models yielded satisfactory results:\n\n* The predicted PCIAT_Total values showed poor alignment with actual values, as evidenced by high Mean Absolute Error (MAE) and Root Mean Squared Error (RMSE). \n* The mapping back to the SII severity levels resulted in significant inaccuracies, making this approach unfeasible.","metadata":{}},{"cell_type":"markdown","source":"## What Next :","metadata":{}},{"cell_type":"markdown","source":"# Future Directions\n\n* **Incorporate Time Series Data:** Leverage the rich time series data available for feature engineering. Analyzing patterns, trends, and variability in this data could uncover valuable features that improve model performance.\n\n* **Experiment with Regression Models:** Applying regression models to the first data processing pipeline might provide a new perspective and potentially yield better results than the classification approach.\n\n* **Neural Network Models:** Developing and fine-tuning a neural network model could be beneficial, especially for capturing complex patterns in the dataset that traditional models may overlook.\n\n* **Explore Advanced Data Processing Techniques:** Investigating alternative imputation strategies, scaling methods, and encoding approaches might improve the quality of input data and enhance model robustness.","metadata":{}}]}