{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30804,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 1. Project Overview","metadata":{}},{"cell_type":"markdown","source":"This project aim to predict the Severity Impairment Index (sii) - measures the level of problematic internet use among children and adolescents, based on physical activity data and other features.\n\nTarget Variable (sii) is defined as:\n\n- 0: None (PCIAT-PCIAT_Total from 0 to 30)\n- 1: Mild (PCIAT-PCIAT_Total from 31 to 49)\n- 2: Moderate (PCIAT-PCIAT_Total from 50 to 79)\n- 3: Severe (PCIAT-PCIAT_Total 80 and more\n\nNote:\n- sii is derived from PCIAT-PCIAT_Total, the sum of scores from the Parent-Child Internet Addiction Test (PCIAT: 20 questions, scored 0-5).\n- The test dataset doesn't have any PCIAT columns (otherwise predictions would be trivial).\n\n**Insight:**\n\n- We should focus on predicting the target from all other features except the PCIAT results.\n\n- We know the target only for two thirds of the samples. The samples without target can perhaps be used for semi-supervised learning.\n\n- We can directly predict sii (this is the value we have to submit), or we can predict PCIAT-PCIAT_Total and then transform this prediction to a sii prediction for submission. As PCIAT-PCIAT_Total is more granular and informative than sii, training to predict PCIAT-PCIAT_Total has the potential to produce a better model.\n","metadata":{}},{"cell_type":"markdown","source":"# 2. Data overview","metadata":{}},{"cell_type":"markdown","source":"## 2.1. The tabular data","metadata":{}},{"cell_type":"markdown","source":"I. Demographic\n\n1. Demographics - Information about age and sex of participants.\n\nII. Physical Health Measures\n\n2. Bio-electric Impedance Analysis - Measure of key body composition elements, including BMI, fat, muscle, and water content.\n3. FitnessGram Vitals and Treadmill - Measurements of cardiovascular fitness assessed using the NHANES treadmill protocol.\n4. FitnessGram Child - Health related physical fitness assessment measuring five different parameters including aerobic capacity, muscular strength, muscular endurance, flexibility, and body composition.\n5. Physical Measures - Collection of blood pressure, heart rate, height, weight and waist, and hip measurements.\n6. Actigraphy - Objective measure of ecological physical activity through a research-grade biotracker.\n\nIV. Internet Use and Addiction Indicators:\n\n7. Internet Use - Number of hours of using computer/internet per day.\n8. Parent-Child Internet Addiction Test - 20-item scale that measures characteristics and behaviors associated with compulsive use of the Internet including compulsivity, escapism, and dependency.\n9. Physical Activity Questionnaire - Information about children's participation in vigorous activities over the last 7 days.\n\nV. Psychosocial and Emotional Impact:\n\n10. Children's Global Assessment Scale - Numeric scale used by mental health clinicians to rate the general functioning of youths under the age of 18.\n11. Sleep Disturbance Scale - Scale to categorize sleep disorders in children.\n\n","metadata":{}},{"cell_type":"code","source":"from matplotlib.lines import Line2D\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport pandas as pd\n\npath = '/kaggle/input/child-mind-institute-problematic-internet-use/train.csv'\ntrain = pd.read_csv(path)\npath = '/kaggle/input/child-mind-institute-problematic-internet-use/test.csv'\ntest = pd.read_csv(path)\n\npath= '/kaggle/input/child-mind-institute-problematic-internet-use/data_dictionary.csv'\ndata_dict = pd.read_csv(path)\nimport warnings\nwarnings.filterwarnings('ignore')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:43.383808Z","iopub.execute_input":"2024-12-22T04:00:43.384662Z","iopub.status.idle":"2024-12-22T04:00:43.45143Z","shell.execute_reply.started":"2024-12-22T04:00:43.384617Z","shell.execute_reply":"2024-12-22T04:00:43.450146Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"groups = data_dict.groupby('Instrument')['Field'].apply(list).to_dict()\n\nfor instrument, features in groups.items():\n    print(f\"{instrument}: {features}\\n\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:43.453741Z","iopub.execute_input":"2024-12-22T04:00:43.454287Z","iopub.status.idle":"2024-12-22T04:00:43.464672Z","shell.execute_reply.started":"2024-12-22T04:00:43.454235Z","shell.execute_reply":"2024-12-22T04:00:43.463151Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:43.466622Z","iopub.execute_input":"2024-12-22T04:00:43.467174Z","iopub.status.idle":"2024-12-22T04:00:43.502448Z","shell.execute_reply.started":"2024-12-22T04:00:43.467115Z","shell.execute_reply":"2024-12-22T04:00:43.500952Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:43.505165Z","iopub.execute_input":"2024-12-22T04:00:43.505537Z","iopub.status.idle":"2024-12-22T04:00:43.525465Z","shell.execute_reply.started":"2024-12-22T04:00:43.505498Z","shell.execute_reply":"2024-12-22T04:00:43.524048Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:43.526856Z","iopub.execute_input":"2024-12-22T04:00:43.527344Z","iopub.status.idle":"2024-12-22T04:00:43.552298Z","shell.execute_reply.started":"2024-12-22T04:00:43.527292Z","shell.execute_reply":"2024-12-22T04:00:43.551112Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_dict.head(80)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:43.555205Z","iopub.execute_input":"2024-12-22T04:00:43.555587Z","iopub.status.idle":"2024-12-22T04:00:43.577562Z","shell.execute_reply.started":"2024-12-22T04:00:43.555548Z","shell.execute_reply":"2024-12-22T04:00:43.576358Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:10:26.151665Z","iopub.execute_input":"2024-12-22T04:10:26.152112Z","iopub.status.idle":"2024-12-22T04:10:26.159463Z","shell.execute_reply.started":"2024-12-22T04:10:26.152073Z","shell.execute_reply":"2024-12-22T04:10:26.157819Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train1=train.copy()\ntrain1.drop('id', axis=1 , inplace = True)\ntrain1.duplicated().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:11:01.940083Z","iopub.execute_input":"2024-12-22T04:11:01.940618Z","iopub.status.idle":"2024-12-22T04:11:01.976198Z","shell.execute_reply.started":"2024-12-22T04:11:01.940564Z","shell.execute_reply":"2024-12-22T04:11:01.974874Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train1.drop_duplicates(inplace=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:11:22.918677Z","iopub.execute_input":"2024-12-22T04:11:22.919139Z","iopub.status.idle":"2024-12-22T04:11:22.942617Z","shell.execute_reply.started":"2024-12-22T04:11:22.9191Z","shell.execute_reply":"2024-12-22T04:11:22.941525Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train1.describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:12:14.416322Z","iopub.execute_input":"2024-12-22T04:12:14.416816Z","iopub.status.idle":"2024-12-22T04:12:14.591393Z","shell.execute_reply.started":"2024-12-22T04:12:14.416769Z","shell.execute_reply":"2024-12-22T04:12:14.590094Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.2. Time series data","metadata":{}},{"cell_type":"markdown","source":"### Actigraphy Files and Field Descriptions\nDuring their participation in the HBN study, some participants were given an accelerometer to wear for up to 30 days continually while at home and going about their regular daily lives.\n\n- **series_{train|test}.parquet/id={id}** - Series to be used as training data, partitioned by id. Each series is a continuous recording of accelerometer data for a single subject spanning many days.\n\n- **id** - The patient identifier corresponding to the id field in train/test.csv.\n- **step** - An integer timestep for each observation within a series.\n- **X, Y, Z** - Measure of acceleration, in g, experienced by the wrist-worn watch along each standard axis.\n- **enmo** - As calculated and described by the wristpy package, ENMO is the Euclidean Norm Minus One of all accelerometer signals (along each of the x-, y-, and z-axis, measured in g-force) with negative values rounded to zero. Zero values are indicative of periods of no motion. While no standard measure of acceleration exists in this space, this is one of the several commonly computed features.\n- **anglez** - As calculated and described by the wristpy package, Angle-Z is a metric derived from individual accelerometer components and refers to the angle of the arm relative to the horizontal plane.\n- **non-wear_flag** - A flag (0: watch is being worn, 1: the watch is not worn) to help determine periods when the watch has been removed, based on the GGIR definition, which uses the standard deviation and range of the accelerometer data.\n- **light** - Measure of ambient light in lux. See ​​here for details.\n- **battery_voltage** - A measure of the battery voltage in mV.\n- **time_of_day** - Time of day representing the start of a 5s window that the data has been sampled over, with format **%H:%M:%S.%9f.**\n- **weekday** - The day of the week, coded as an integer with 1 being Monday and 7 being Sunday.\n- **quarter** - The quarter of the year, an integer from 1 to 4.\n- **relative_date_PCIAT** - The number of days (integer) since the PCIAT test was administered (negative days indicate that the actigraphy data has been collected before the test was administered).","metadata":{}},{"cell_type":"code","source":"trainSerise1 = pd.read_parquet(\"/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet/id=00115b9f/part-0.parquet\")\ntrainSerise1","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:43.579989Z","iopub.execute_input":"2024-12-22T04:00:43.580917Z","iopub.status.idle":"2024-12-22T04:00:43.623268Z","shell.execute_reply.started":"2024-12-22T04:00:43.580858Z","shell.execute_reply":"2024-12-22T04:00:43.622114Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"trainSerise1.drop(\"step\", axis=1, inplace=True)\ntrainSerise1","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:43.624653Z","iopub.execute_input":"2024-12-22T04:00:43.625033Z","iopub.status.idle":"2024-12-22T04:00:43.65084Z","shell.execute_reply.started":"2024-12-22T04:00:43.624985Z","shell.execute_reply":"2024-12-22T04:00:43.649631Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sensorFeature =trainSerise1.columns.to_list()\nsensorFeature , len(sensorFeature)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:43.653748Z","iopub.execute_input":"2024-12-22T04:00:43.654726Z","iopub.status.idle":"2024-12-22T04:00:43.662451Z","shell.execute_reply.started":"2024-12-22T04:00:43.654669Z","shell.execute_reply":"2024-12-22T04:00:43.661186Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"statFeat = trainSerise1.describe().index.to_list() \nstatFeat, len(statFeat)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:43.663695Z","iopub.execute_input":"2024-12-22T04:00:43.66408Z","iopub.status.idle":"2024-12-22T04:00:43.723803Z","shell.execute_reply.started":"2024-12-22T04:00:43.664041Z","shell.execute_reply":"2024-12-22T04:00:43.722634Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 3.Data Cleaning and Preparation","metadata":{}},{"cell_type":"markdown","source":"## 3.1. Missing value treatment","metadata":{}},{"cell_type":"markdown","source":"All columns have a substantial proportion of missing values, except id (not surprisingly) and the three basic demographic columns for sex, age and season of enrollment. Even the target sii has missing values:","metadata":{}},{"cell_type":"code","source":"train_cols = set(train.columns)\ntest_cols = set(test.columns)\ncolumns_not_in_test = sorted(list(train_cols - test_cols))\nprint('Columns missing in test:')\nprint(columns_not_in_test)\ndata_dict[data_dict['Field'].isin(columns_not_in_test)]\n\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:43.725315Z","iopub.execute_input":"2024-12-22T04:00:43.725714Z","iopub.status.idle":"2024-12-22T04:00:43.743705Z","shell.execute_reply.started":"2024-12-22T04:00:43.725677Z","shell.execute_reply":"2024-12-22T04:00:43.742499Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport numpy as np\nfrom matplotlib.ticker import PercentFormatter\n\nmissing_count = train.isnull().sum().sort_values(ascending=False)\nmissing_ratio = missing_count / len(train)\n\nplt.figure(figsize=(6, 15))\nplt.title('Missing values over the whole training dataset')\nplt.barh(np.arange(len(missing_count)), missing_ratio, color='coral', label='missing')\nplt.barh(np.arange(len(missing_count)), \n         1 - missing_ratio,\n         left=missing_ratio,\n         color='darkseagreen', label='available')\nplt.yticks(np.arange(len(missing_count)), missing_count.index)\nplt.gca().xaxis.set_major_formatter(PercentFormatter(xmax=1, decimals=0))\nplt.xlim(0, 1)\nplt.legend()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:43.744997Z","iopub.execute_input":"2024-12-22T04:00:43.745457Z","iopub.status.idle":"2024-12-22T04:00:44.844605Z","shell.execute_reply.started":"2024-12-22T04:00:43.745406Z","shell.execute_reply":"2024-12-22T04:00:44.843042Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Lọc các mẫu có giá trị 'sii' không rỗng\nsupervised_usable = train[train['sii'].notnull()]\n\n# Đếm số lượng giá trị bị thiếu\nmissing_count = supervised_usable.isnull().sum().sort_values(ascending=False)\nmissing_ratio = missing_count / len(supervised_usable)\n\nplt.figure(figsize=(6, 15))\nplt.title(f'Missing values over the {len(supervised_usable)} samples which have a target')\nplt.barh(np.arange(len(missing_count)), \n         missing_ratio, \n         color='coral', \n         label='missing')\nplt.barh(np.arange(len(missing_count)), \n         1 - missing_ratio,\n         left=missing_ratio,\n         color='darkseagreen', \n         label='available')\nplt.yticks(np.arange(len(missing_count)), missing_count.index)\nplt.gca().xaxis.set_major_formatter(PercentFormatter(xmax=1, decimals=0))\nplt.xlim(0, 1)\nplt.legend()\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:44.848307Z","iopub.execute_input":"2024-12-22T04:00:44.848813Z","iopub.status.idle":"2024-12-22T04:00:46.145035Z","shell.execute_reply.started":"2024-12-22T04:00:44.848758Z","shell.execute_reply":"2024-12-22T04:00:46.143599Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The Parent-Child Internet Addiction Test questions can be ignored by a respondent (missing values in the PCIAT-PCIAT_01 to PCIAT-PCIAT_20 columns), but the SII score is still derived from the the sum of the non-NA values, leading to potentially invalid SII values (unless, of course, some answers were cut out after the data has been collected, just to give us a bit more of a challenge","metadata":{}},{"cell_type":"code","source":"col = ['CGAS-Season', 'Physical-Season', 'Fitness_Endurance-Season', \n       'FGC-Season', 'BIA-Season', 'PAQ_A-Season', 'PAQ_C-Season',\n       'PCIAT-Season', 'SDS-Season', 'PreInt_EduHx-Season']\ncol.extend( \n        ['Physical-Waist_Circumference','Fitness_Endurance-Max_Stage',\n         'Fitness_Endurance-Time_Mins','Fitness_Endurance-Time_Sec','FGC-FGC_GSND',\n         'FGC-FGC_GSND_Zone','FGC-FGC_GSD','FGC-FGC_GSD_Zone','BIA-BIA_Activity_Level_num',\n         'BIA-BIA_BMC','BIA-BIA_BMI','BIA-BIA_BMR','BIA-BIA_DEE','BIA-BIA_ECW','BIA-BIA_FFM',\n         'BIA-BIA_FFMI','BIA-BIA_Fat','BIA-BIA_Frame_num','BIA-BIA_ICW','BIA-BIA_LDM','BIA-BIA_LST',\n         'BIA-BIA_SMM','BIA-BIA_TBW','PAQ_A-PAQ_A_Total','PAQ_C-PAQ_C_Total'])\nlen(col)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.146403Z","iopub.execute_input":"2024-12-22T04:00:46.146746Z","iopub.status.idle":"2024-12-22T04:00:46.155583Z","shell.execute_reply.started":"2024-12-22T04:00:46.146711Z","shell.execute_reply":"2024-12-22T04:00:46.154427Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train1.drop(col , axis =1 ,inplace=True)\ntrain1.shape\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:13:10.169053Z","iopub.execute_input":"2024-12-22T04:13:10.169475Z","iopub.status.idle":"2024-12-22T04:13:10.179736Z","shell.execute_reply.started":"2024-12-22T04:13:10.169439Z","shell.execute_reply":"2024-12-22T04:13:10.178508Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_cleaned= train1.copy()\ndata_cleaned.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:13:20.128866Z","iopub.execute_input":"2024-12-22T04:13:20.129296Z","iopub.status.idle":"2024-12-22T04:13:20.138324Z","shell.execute_reply.started":"2024-12-22T04:13:20.129259Z","shell.execute_reply":"2024-12-22T04:13:20.137067Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Filling Missing Values in first 46 columns with Knn imputer¶\n","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder\nlabel_encoder = LabelEncoder()\ndata_cleaned['Basic_Demos-Enroll_Season'] = label_encoder.fit_transform(data_cleaned['Basic_Demos-Enroll_Season'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:13:23.691067Z","iopub.execute_input":"2024-12-22T04:13:23.691435Z","iopub.status.idle":"2024-12-22T04:13:23.698929Z","shell.execute_reply.started":"2024-12-22T04:13:23.691404Z","shell.execute_reply":"2024-12-22T04:13:23.697713Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.impute import KNNImputer\n\nimputer = KNNImputer(n_neighbors=5)\ndata_cleaned.iloc[:, :45] = imputer.fit_transform(data_cleaned.iloc[:, :45])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:13:29.708211Z","iopub.execute_input":"2024-12-22T04:13:29.708592Z","iopub.status.idle":"2024-12-22T04:13:33.973639Z","shell.execute_reply.started":"2024-12-22T04:13:29.708558Z","shell.execute_reply":"2024-12-22T04:13:33.972406Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder\nlabel_encoder = LabelEncoder()\ndata_cleaned['Severity'] = label_encoder.fit_transform(data_cleaned['Severity'])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.cluster import KMeans\nimport numpy as np\nimport pandas as pd\n\ndef impute_with_kmeans(df, categorical_columns, n_clusters=4):\n    # Fill missing values temporarily with the mode (or any placeholder)\n    df_temp = df.copy()\n    for col in categorical_columns:\n        df_temp[col].fillna(df_temp[col].mode()[0], inplace=True)\n\n    # Perform KMeans clustering after filling missing values\n    kmeans = KMeans(n_clusters=n_clusters, random_state=0)\n    cluster_labels = kmeans.fit_predict(df_temp)\n\n    # Impute missing values within each cluster\n    for col in categorical_columns:\n        for cluster in np.unique(cluster_labels):\n            mask = (cluster_labels == cluster) & df[col].isna()\n            most_frequent = df.loc[cluster_labels == cluster, col].mode()[0]\n            df.loc[mask, col] = most_frequent\n\n    return df\ndata_cleaned = impute_with_kmeans(data_cleaned, ['sii'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:15:31.973524Z","iopub.execute_input":"2024-12-22T04:15:31.973908Z","iopub.status.idle":"2024-12-22T04:15:33.328666Z","shell.execute_reply.started":"2024-12-22T04:15:31.973875Z","shell.execute_reply":"2024-12-22T04:15:33.327778Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_cleaned['sii'].value_counts()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:15:36.609559Z","iopub.execute_input":"2024-12-22T04:15:36.609942Z","iopub.status.idle":"2024-12-22T04:15:36.62148Z","shell.execute_reply.started":"2024-12-22T04:15:36.609908Z","shell.execute_reply":"2024-12-22T04:15:36.62021Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_cleaned.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:15:53.185996Z","iopub.execute_input":"2024-12-22T04:15:53.186415Z","iopub.status.idle":"2024-12-22T04:15:53.216193Z","shell.execute_reply.started":"2024-12-22T04:15:53.186378Z","shell.execute_reply":"2024-12-22T04:15:53.214632Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def detect_outliers_iqr(df):\n    outlier_indices = []\n\n    for column in df.select_dtypes(include=['float64', 'int64']).columns:\n        Q1 = df[column].quantile(0.25)\n        Q3 = df[column].quantile(0.75)\n        IQR = Q3 - Q1\n\n        lower_bound = Q1 - 1.5 * IQR\n        upper_bound = Q3 + 1.5 * IQR\n\n        outliers = df[(df[column] < lower_bound) | (df[column] > upper_bound)]\n        outlier_indices.extend(outliers.index)\n\n        print(f'Outliers in {column}:', outliers.shape[0])\n\n    return outlier_indices","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:16:41.925066Z","iopub.execute_input":"2024-12-22T04:16:41.925507Z","iopub.status.idle":"2024-12-22T04:16:41.933282Z","shell.execute_reply.started":"2024-12-22T04:16:41.92547Z","shell.execute_reply":"2024-12-22T04:16:41.931925Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"outlier_indices = detect_outliers_iqr(train1)\nprint(\"Total outliers detected:\", len(set(outlier_indices)))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:16:54.900836Z","iopub.execute_input":"2024-12-22T04:16:54.901281Z","iopub.status.idle":"2024-12-22T04:16:55.008556Z","shell.execute_reply.started":"2024-12-22T04:16:54.901241Z","shell.execute_reply":"2024-12-22T04:16:55.007369Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"outlier_indices = detect_outliers_iqr(data_cleaned)\nprint(\"Total outliers detected:\", len(set(outlier_indices)))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:17:17.940146Z","iopub.execute_input":"2024-12-22T04:17:17.940528Z","iopub.status.idle":"2024-12-22T04:17:18.04554Z","shell.execute_reply.started":"2024-12-22T04:17:17.940495Z","shell.execute_reply":"2024-12-22T04:17:18.044128Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Số lượng cột cần vẽ\nnum_columns = len(data_cleaned.columns)\n\n# Số box plot mỗi hàng\nplots_per_row = 3  \n\n# Tính số hàng cần thiết\nnum_rows = (num_columns + plots_per_row - 1) // plots_per_row\n\n# Tạo các subplots\nfig, axes = plt.subplots(num_rows, plots_per_row, figsize=(plots_per_row * 6, num_rows * 3))\naxes = axes.flatten()  # Làm phẳng mảng các trục để dễ dàng truy cập\n\n# Vẽ từng box plot\nfor i, column in enumerate(data_cleaned.columns):\n    sns.boxplot(data=data_cleaned[[column]], orient='h', palette=\"Set2\", ax=axes[i])\n    axes[i].set_title(f'Box Plot for {column}', fontsize=12)\n    axes[i].set_xlabel('Values', fontsize=10)\n    axes[i].set_ylabel(column, fontsize=10)\n\n# Ẩn các subplot thừa\nfor j in range(i + 1, len(axes)):\n    fig.delaxes(axes[j])\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:17:22.285993Z","iopub.execute_input":"2024-12-22T04:17:22.286429Z","iopub.status.idle":"2024-12-22T04:17:28.664059Z","shell.execute_reply.started":"2024-12-22T04:17:22.286391Z","shell.execute_reply":"2024-12-22T04:17:28.662481Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 3.1.1 Remove all the columns where the predicatble value is null\n","metadata":{}},{"cell_type":"code","source":"df_train.dropna(subset=['sii'],inplace=True)\nprint('length of df', len(df_train))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.300172Z","iopub.status.idle":"2024-12-22T04:00:46.300614Z","shell.execute_reply.started":"2024-12-22T04:00:46.300406Z","shell.execute_reply":"2024-12-22T04:00:46.300429Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder\ncategorical_cols = df_train.select_dtypes(include=['object']).columns\nother_cols = df_train.select_dtypes(exclude=['object']).columns","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.302388Z","iopub.status.idle":"2024-12-22T04:00:46.302986Z","shell.execute_reply.started":"2024-12-22T04:00:46.302673Z","shell.execute_reply":"2024-12-22T04:00:46.302702Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(df_train.columns)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.304601Z","iopub.status.idle":"2024-12-22T04:00:46.305214Z","shell.execute_reply.started":"2024-12-22T04:00:46.304887Z","shell.execute_reply":"2024-12-22T04:00:46.304917Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_train[categorical_cols].isnull().sum().sort_values(ascending=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.307526Z","iopub.status.idle":"2024-12-22T04:00:46.308498Z","shell.execute_reply.started":"2024-12-22T04:00:46.307884Z","shell.execute_reply":"2024-12-22T04:00:46.307922Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(df_train[other_cols].isnull().sum().sort_values(ascending=False).to_string())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.309883Z","iopub.status.idle":"2024-12-22T04:00:46.310375Z","shell.execute_reply.started":"2024-12-22T04:00:46.310161Z","shell.execute_reply":"2024-12-22T04:00:46.310192Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 3.1.2. Remove the columns that have high null values(>1200)\n","metadata":{}},{"cell_type":"markdown","source":"### 3.1.3. Replace the null values with unknown and then perform label encoding for the null values between 17 and 12001","metadata":{}},{"cell_type":"markdown","source":"## 3.2. Outlier Detection & Treatment","metadata":{}},{"cell_type":"markdown","source":"# 4. Exploratory Data Analysis ","metadata":{}},{"cell_type":"markdown","source":"## 4.1. Univariate analysis","metadata":{}},{"cell_type":"code","source":"df_train_copy = train[[\n    'Basic_Demos-Age', 'Basic_Demos-Sex',\n    'SDS-SDS_Total_Raw', \n    'PreInt_EduHx-computerinternet_hoursday', \n    'PAQ_C-PAQ_C_Total', 'PAQ_A-PAQ_A_Total',\n    'BIA-BIA_SMM', 'BIA-BIA_BMI',\n    'Physical-Height', 'Physical-Weight', 'Physical-BMI',\n    'CGAS-CGAS_Score', \n    'FGC-FGC_CU',\n    'sii'\n]].copy()\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.312084Z","iopub.status.idle":"2024-12-22T04:00:46.312696Z","shell.execute_reply.started":"2024-12-22T04:00:46.312389Z","shell.execute_reply":"2024-12-22T04:00:46.312421Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_train_copy.describe().round(3).T","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.314292Z","iopub.status.idle":"2024-12-22T04:00:46.314744Z","shell.execute_reply.started":"2024-12-22T04:00:46.314523Z","shell.execute_reply":"2024-12-22T04:00:46.314545Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4.2. Bivariate Analysis","metadata":{}},{"cell_type":"markdown","source":"## 4.3. Multivariate Analysis","metadata":{}},{"cell_type":"code","source":"# correlation matrix\nplt.figure(figsize = (8, 4), facecolor = \"white\")\n\n# plotting\nsns.heatmap(\n    data = df_train.corr(numeric_only = True),\n    cmap = \"vlag\",\n    vmin = -1, vmax = 1,\n    linecolor = \"white\", linewidth = 0.5,\n    annot = True,\n    fmt = \".2f\"\n)\n\nplt.title('Correlation Heatmap')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.315939Z","iopub.status.idle":"2024-12-22T04:00:46.316389Z","shell.execute_reply.started":"2024-12-22T04:00:46.316197Z","shell.execute_reply":"2024-12-22T04:00:46.316224Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- AnalysisBIA-BIA_BMI và Physical-BMI: Có mối tương quan cực kỳ mạnh, gần với 1. Điều này phản ánh thực tế rằng chỉ số khối cơ thể (BMI) thường được tính dựa trên các thông số liên quan đến cơ thể như tỷ lệ mỡ.\n- CSGAS-CSGAS_Score và FGC-FGC_CU: Mối tương quan này cũng rất cao, có thể do các biến này đo lường những khía cạnh tương tự hoặc có nguồn dữ liệu liên quan.\n\n- Basic_Demos-Age và PAQ_A_PAQ_Total: Tương quan âm rõ rệt, cho thấy rằng tuổi tác càng lớn thì tổng điểm PAQ (Physical Activity Questionnaire) càng giảm. Điều này phù hợp với thực tế rằng người lớn tuổi thường có mức độ hoạt động thể chất thấp hơn.\n\n- Physical-Weight và PAQ_A_PAQ_Total: Tương quan âm này cũng có thể ám chỉ rằng người có cân nặng cao hơn thường ít hoạt động thể chất.\nCác mối tương quan yếu (màu nhạt):\n\n- Một số biến có tương quan rất yếu hoặc gần bằng 0, như giữa Basic_Demos-Sex và các biến khác. Điều này cho thấy giới tính không ảnh hưởng đáng kể đến các biến này.","metadata":{}},{"cell_type":"markdown","source":"# 5. Features EAD by Groups","metadata":{}},{"cell_type":"code","source":"def calculate_stats(data, columns):\n    if isinstance(columns, str):\n        columns = [columns]\n\n    stats = []\n    for col in columns:\n        if data[col].dtype in ['object', 'category']:\n            counts = data[col].value_counts(dropna=False, sort=False)\n            percents = data[col].value_counts(normalize=True, dropna=False, sort=False) * 100\n            formatted = counts.astype(str) + ' (' + percents.round(2).astype(str) + '%)'\n            stats_col = pd.DataFrame({'count (%)': formatted})\n            stats.append(stats_col)\n        else:\n            stats_col = data[col].describe().to_frame().transpose()\n            stats_col['missing'] = data[col].isnull().sum()\n            stats_col.index.name = col\n            stats.append(stats_col)\n\n    return pd.concat(stats, axis=0)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.319359Z","iopub.status.idle":"2024-12-22T04:00:46.319753Z","shell.execute_reply.started":"2024-12-22T04:00:46.319573Z","shell.execute_reply":"2024-12-22T04:00:46.319597Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_dict = data_dict[data_dict['Instrument'] != 'Parent-Child Internet Addiction Test']\ncontinuous_cols = data_dict[data_dict['Type'].str.contains(\n    'float|int', case=False\n)]['Field'].tolist()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.321763Z","iopub.status.idle":"2024-12-22T04:00:46.32223Z","shell.execute_reply.started":"2024-12-22T04:00:46.322022Z","shell.execute_reply":"2024-12-22T04:00:46.322051Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# I.Demographics\n## 1.Demographics","metadata":{}},{"cell_type":"code","source":"groups.get('Demographics', [])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.32344Z","iopub.status.idle":"2024-12-22T04:00:46.323836Z","shell.execute_reply.started":"2024-12-22T04:00:46.323648Z","shell.execute_reply":"2024-12-22T04:00:46.323667Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import seaborn as sns\nfrom matplotlib.lines import Line2D\n\nfig, axes = plt.subplots(1, 2, figsize=(12, 5))\n\n# Season of Enrollment\nseason_counts = train['Basic_Demos-Enroll_Season'].value_counts(dropna=False)\n\naxes[0].pie(\n    season_counts, labels=season_counts.index,\n    autopct='%1.1f%%', startangle=90,\n    colors=sns.color_palette(\"Set3\")\n)\naxes[0].set_title('Season of Enrollment')\naxes[0].axis('equal')\n\n# Age Distribution by Sex\nsns.histplot(\n    data=train, x='Basic_Demos-Age',\n    hue='Basic_Demos-Sex', multiple='dodge',\n    palette=\"Set2\", bins=20, ax=axes[1]\n)\naxes[1].set_title('Age Distribution by Sex')\naxes[1].set_xlabel('Age')\naxes[1].set_ylabel('Count')\n\nplt.tight_layout()\nplt.show()\nprint('0=Male, 1=Female')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.325171Z","iopub.status.idle":"2024-12-22T04:00:46.325565Z","shell.execute_reply.started":"2024-12-22T04:00:46.325384Z","shell.execute_reply":"2024-12-22T04:00:46.325402Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"calculate_stats(train, 'Basic_Demos-Age')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.326775Z","iopub.status.idle":"2024-12-22T04:00:46.327177Z","shell.execute_reply.started":"2024-12-22T04:00:46.326978Z","shell.execute_reply":"2024-12-22T04:00:46.327009Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"# II. Physical Health Measures\n\n","metadata":{}},{"cell_type":"markdown","source":"## 2. Bio-electric Impedance Analysis","metadata":{}},{"cell_type":"code","source":"data_dict[data_dict['Instrument'] == 'Bio-electric Impedance Analysis']\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.328755Z","iopub.status.idle":"2024-12-22T04:00:46.329263Z","shell.execute_reply.started":"2024-12-22T04:00:46.329071Z","shell.execute_reply":"2024-12-22T04:00:46.329092Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"bia_data_dict = data_dict[data_dict['Instrument'] == 'Bio-electric Impedance Analysis']\ncategorical_columns = bia_data_dict[bia_data_dict['Type'] == 'categorical int']['Field'].tolist()\ncontinuous_columns = bia_data_dict[bia_data_dict['Type'] == 'float']['Field'].tolist()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.330438Z","iopub.status.idle":"2024-12-22T04:00:46.331013Z","shell.execute_reply.started":"2024-12-22T04:00:46.330714Z","shell.execute_reply":"2024-12-22T04:00:46.330742Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"calculate_stats(train, continuous_columns)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.332294Z","iopub.status.idle":"2024-12-22T04:00:46.332682Z","shell.execute_reply.started":"2024-12-22T04:00:46.332502Z","shell.execute_reply":"2024-12-22T04:00:46.332522Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\ncontinuous_columns = [\n    'BIA-BIA_BMC', 'BIA-BIA_BMI', 'BIA-BIA_BMR', 'BIA-BIA_DEE',\n    'BIA-BIA_ECW', 'BIA-BIA_FFM', 'BIA-BIA_FFMI', 'BIA-BIA_FMI',\n    'BIA-BIA_Fat', 'BIA-BIA_ICW', 'BIA-BIA_LDM', 'BIA-BIA_LST',\n    'BIA-BIA_SMM', 'BIA-BIA_TBW'\n]\n\n# Select only the continuous columns and the sii target\ndata_continuous = train[continuous_columns + ['sii']]\n\n# Remove rows where sii is NaN, but keep NaNs in continuous variables\ndata_continuous = data_continuous.dropna(subset=['sii'])\n\n# Melt the data to long format for Seaborn\ndata_long = pd.melt(data_continuous, id_vars='sii', var_name=\"variable\", value_name=\"value\")\n\n# Set up the FacetGrid for KDE plots with increased figure size\ng = sns.FacetGrid(data_long, col=\"variable\", col_wrap=4, height=3.5, aspect=1.2, sharex=False, sharey=False)\n\n# Map KDE plots onto the grid, using hue for sii target\ng.map_dataframe(sns.kdeplot, x=\"value\", hue=\"sii\", fill=True, common_norm=False, palette=\"Set2\", alpha=0.4, linewidth=1.5)\n\n# Add title, adjust layout\ng.fig.suptitle(\"Distribution of Continuous Fitness and Health Metrics by SII Target\", y=1.05, fontsize=18, color=\"#004080\")\ng.set_titles(\"{col_name}\")\ng.set_axis_labels(\"Value\", \"Density\")\n\n# Create custom legend\n# Define the labels and colors (matching `palette=\"Set2\"` used above)\nsii_labels = sorted(data_continuous['sii'].dropna().unique())\ncolors = sns.color_palette(\"Set2\", len(sii_labels))\nlegend_elements = [Line2D([0], [0], color=color, lw=4, label=f'SII {label}') for color, label in zip(colors, sii_labels)]\n\n# Position the custom legend outside of the grid\ng.fig.legend(handles=legend_elements, loc=\"upper center\", ncol=4, bbox_to_anchor=(0.5, 1.15), frameon=False)\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.333744Z","iopub.status.idle":"2024-12-22T04:00:46.334197Z","shell.execute_reply.started":"2024-12-22T04:00:46.333977Z","shell.execute_reply":"2024-12-22T04:00:46.334008Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Some variables, such as Fat Mass Index and Body Fat Percentage, show implausible negative values, and almost all - extreme high values, indicating potential data quality issues","metadata":{}},{"cell_type":"markdown","source":"## 3.FitnessGram Vitals and Treadmill","metadata":{}},{"cell_type":"code","source":"groups.get('FitnessGram Vitals and Treadmill', [])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.336477Z","iopub.status.idle":"2024-12-22T04:00:46.336885Z","shell.execute_reply.started":"2024-12-22T04:00:46.336695Z","shell.execute_reply":"2024-12-22T04:00:46.336715Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fgc_data_dict = data_dict[data_dict['Instrument'] == 'FitnessGram Child']\n\nfgc_columns = []\n\nfor index, row in fgc_data_dict.iterrows():\n    if '_Zone' not in row['Field']:\n        measure_field = row['Field']\n        measure_desc = row['Description']\n        \n        zone_field = measure_field + '_Zone'\n        zone_row = fgc_data_dict[fgc_data_dict['Field'] == zone_field]\n        \n        if not zone_row.empty:\n            zone_desc = zone_row['Description'].values[0]\n            fgc_columns.append((measure_field, zone_field, measure_desc, zone_desc))\n            \nfig, axes = plt.subplots(2, 4, figsize=(24, 10))\n\nfor idx, (measure, zone, measure_desc, zone_desc) in enumerate(fgc_columns):\n    row = idx // 4\n    col = idx % 4\n    \n    sns.histplot(\n        data=train, x=measure,\n        hue=zone, bins=20, palette='Set2',\n        ax=axes[row, col], kde=True\n    )\n    axes[row, col].set_title(f'{measure_desc}')\n\nseason_counts = train['FGC-Season'].value_counts(normalize=True)\naxes[1, 3].pie(\n    season_counts, labels=season_counts.index,\n    autopct='%1.1f%%', startangle=90,\n    colors=sns.color_palette(\"Set3\")\n)\naxes[1, 3].set_title('Season of participation')\naxes[1, 3].axis('equal') \n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.338376Z","iopub.status.idle":"2024-12-22T04:00:46.338758Z","shell.execute_reply.started":"2024-12-22T04:00:46.338582Z","shell.execute_reply":"2024-12-22T04:00:46.338602Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4.FitnessGram Child","metadata":{}},{"cell_type":"code","source":"groups.get('FitnessGram Child', [])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.340054Z","iopub.status.idle":"2024-12-22T04:00:46.340461Z","shell.execute_reply.started":"2024-12-22T04:00:46.340268Z","shell.execute_reply":"2024-12-22T04:00:46.340288Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5. Physical measures","metadata":{}},{"cell_type":"code","source":"groups.get('Physical Measures', [])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.342825Z","iopub.status.idle":"2024-12-22T04:00:46.343291Z","shell.execute_reply.started":"2024-12-22T04:00:46.343073Z","shell.execute_reply":"2024-12-22T04:00:46.343099Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_dict[data_dict['Instrument'] == 'Physical Measures']\ntrain[(train['Physical-Weight'] >= 150) & (train['Physical-Weight'] <= 170)].count()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.344672Z","iopub.status.idle":"2024-12-22T04:00:46.345121Z","shell.execute_reply.started":"2024-12-22T04:00:46.344885Z","shell.execute_reply":"2024-12-22T04:00:46.344904Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"features_physical = groups.get('Physical Measures', [])\ncols = [col for col in features_physical if col in continuous_cols]\n\nplt.figure(figsize=(24, 10))\nn_cols = 4\nn_rows = len(cols) // n_cols + 1\n\nfor i, col in enumerate(cols):\n    plt.subplot(n_rows, n_cols, i + 1)\n    train[col].hist(bins=20)\n    plt.title(col)\n\nplt.subplot(n_rows, n_cols, len(cols) + 1)\nseason_counts = train['Physical-Season'].value_counts(dropna=False)\nplt.pie(\n    season_counts,\n    labels=season_counts.index,\n    autopct='%1.1f%%',\n    startangle=90,\n    colors=sns.color_palette(\"Set3\")\n)\nplt.title('Physical-Season')\n\nplt.suptitle('Histograms for Physical Measures and Physical-Season Pie Chart', y=1.05)\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.34675Z","iopub.status.idle":"2024-12-22T04:00:46.347202Z","shell.execute_reply.started":"2024-12-22T04:00:46.346997Z","shell.execute_reply":"2024-12-22T04:00:46.347024Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\ncontinuous_columns = [\n    'Basic_Demos-Age', 'Physical-BMI', 'Physical-Height', 'Physical-Weight',   \n]\n\n# Select only the continuous columns and the sii target\ndata_continuous = train[continuous_columns + ['sii']]\n\n# Remove rows where sii is NaN, but keep NaNs in continuous variables\ndata_continuous = data_continuous.dropna(subset=['sii'])\n\n# Melt the data to long format for Seaborn\ndata_long = pd.melt(data_continuous, id_vars='sii', var_name=\"variable\", value_name=\"value\")\n\n# Set up the FacetGrid for KDE plots with increased figure size\ng = sns.FacetGrid(data_long, col=\"variable\", col_wrap=4, height=3.5, aspect=1.2, sharex=False, sharey=False)\n\n# Map KDE plots onto the grid, using hue for sii target\ng.map_dataframe(sns.kdeplot, x=\"value\", hue=\"sii\", fill=True, common_norm=False, palette=\"Set2\", alpha=0.4, linewidth=1.5)\n\n# Add title, adjust layout\ng.fig.suptitle(\"Distribution of Continuous Fitness and Health Metrics by SII Target\", y=1.05, fontsize=18, color=\"#004080\")\ng.set_titles(\"{col_name}\")\ng.set_axis_labels(\"Value\", \"Density\")\n\n# Create custom legend\n# Define the labels and colors (matching `palette=\"Set2\"` used above)\nsii_labels = sorted(data_continuous['sii'].dropna().unique())\ncolors = sns.color_palette(\"Set2\", len(sii_labels))\nlegend_elements = [Line2D([0], [0], color=color, lw=4, label=f'SII {label}') for color, label in zip(colors, sii_labels)]\n\n# Position the custom legend outside of the grid\ng.fig.legend(handles=legend_elements, loc=\"upper center\", ncol=4, bbox_to_anchor=(0.5, 1.15), frameon=False)\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.348901Z","iopub.status.idle":"2024-12-22T04:00:46.349485Z","shell.execute_reply.started":"2024-12-22T04:00:46.349206Z","shell.execute_reply":"2024-12-22T04:00:46.349238Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"1. Biểu đồ 1 - Physical-Weight:\nPhân bố của cân nặng: Biểu đồ này thể hiện phân bố mật độ của cân nặng trong dữ liệu. Các đường cong cho thấy mật độ của giá trị cân nặng trong các nhóm sii. Bạn có thể nhận thấy nếu một nhóm có sự tập trung nhiều hơn ở một phạm vi cân nặng nhất định (ví dụ: nhóm sii=1 có cân nặng thường từ 50kg đến 70kg, nhóm sii=2 từ 70kg đến 90kg).\nĐặc điểm của phân phối: Nếu đường cong có một đỉnh rõ ràng (một hoặc nhiều đỉnh), điều này có thể cho thấy sự tập trung của một nhóm ở một mức cân nặng nhất định. Nếu các đường cong của các nhóm có sự chồng lấn, điều này cho thấy các nhóm này có phân phối cân nặng tương tự nhau.\n2. Biểu đồ 2 - Physical-BMI:\nPhân bố của BMI (Body Mass Index): Biểu đồ này thể hiện phân bố mật độ của chỉ số BMI, một chỉ số đánh giá mức độ béo phì hoặc tình trạng cơ thể (dựa trên cân nặng và chiều cao). Những nhóm có chỉ số BMI trong một phạm vi nhất định sẽ có các đường cong chồng lên nhau tại các giá trị BMI nhất định.\nÝ nghĩa: Nếu một nhóm có BMI cao hơn hoặc thấp hơn so với các nhóm khác, điều này có thể phản ánh sự khác biệt về sức khỏe hoặc thể trạng giữa các nhóm đó.\n3. Biểu đồ 3 - Physical-Height:\nPhân bố chiều cao: Biểu đồ này thể hiện phân bố chiều cao của các cá nhân trong từng nhóm sii. Các đường cong biểu thị sự phân bố chiều cao, với các nhóm sii có thể có sự tập trung chiều cao ở các mức khác nhau.\nSo sánh các nhóm: Nếu nhóm sii=1 có phần lớn người có chiều cao thấp hơn, trong khi nhóm sii=2 có người cao hơn, thì bạn có thể kết luận rằng chiều cao có thể ảnh hưởng đến nhóm sii này.\n4. Biểu đồ 4 - Physical-BMR (Basal Metabolic Rate):\nPhân bố BMR: Biểu đồ này thể hiện phân bố của chỉ số BMR (tốc độ trao đổi chất cơ bản). Chỉ số BMR cho biết năng lượng cơ thể cần để duy trì các chức năng cơ bản khi nghỉ ngơi. Nếu BMR của các nhóm có sự khác biệt rõ rệt, điều này có thể chỉ ra sự khác biệt trong mức độ trao đổi chất giữa các nhóm sii.\nĐặc điểm phân phối: Các nhóm có BMR thấp hơn có thể là nhóm có ít hoạt động thể chất hơn hoặc có mức độ trao đổi chất thấp hơn.\n","metadata":{}},{"cell_type":"markdown","source":"# IV. Internet Use and Addiction Indicators","metadata":{}},{"cell_type":"markdown","source":"## 7.Internet Use \n","metadata":{}},{"cell_type":"code","source":"groups.get('Internet Use', [])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.35104Z","iopub.status.idle":"2024-12-22T04:00:46.351586Z","shell.execute_reply.started":"2024-12-22T04:00:46.351313Z","shell.execute_reply":"2024-12-22T04:00:46.351341Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_dict[data_dict['Instrument'] == 'Internet Use']\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.353404Z","iopub.status.idle":"2024-12-22T04:00:46.353781Z","shell.execute_reply.started":"2024-12-22T04:00:46.353597Z","shell.execute_reply":"2024-12-22T04:00:46.353615Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"calculate_stats(train, 'PreInt_EduHx-computerinternet_hoursday')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.355693Z","iopub.status.idle":"2024-12-22T04:00:46.35614Z","shell.execute_reply.started":"2024-12-22T04:00:46.355906Z","shell.execute_reply":"2024-12-22T04:00:46.355925Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data = train[train['PreInt_EduHx-computerinternet_hoursday'].notna()]\nage_range = data['Basic_Demos-Age']\nprint(\n    f\"Age range for participants with measured PreInt_EduHx-computerinternet_hoursday data:\"\n    f\" {age_range.min()} - {age_range.max()} years\"\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.357446Z","iopub.status.idle":"2024-12-22T04:00:46.357798Z","shell.execute_reply.started":"2024-12-22T04:00:46.357628Z","shell.execute_reply":"2024-12-22T04:00:46.357646Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train['PreInt_EduHx-computerinternet_hoursday'].unique()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.359053Z","iopub.status.idle":"2024-12-22T04:00:46.359452Z","shell.execute_reply.started":"2024-12-22T04:00:46.359263Z","shell.execute_reply":"2024-12-22T04:00:46.359283Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"param_map = {0: '< 1h/day', 1: '~ 1h/day', 2: '~ 2hs/day', 3: '> 3hs/day'}\ntrain['internet_use_encoded'] = train[\n    'PreInt_EduHx-computerinternet_hoursday'\n].map(param_map).fillna('Missing')\n\nparam_ord = ['Missing', '< 1h/day', '~ 1h/day', '~ 2hs/day', '> 3hs/day']\ntrain['internet_use_encoded'] = pd.Categorical(\n    train['internet_use_encoded'], categories=param_ord,\n    ordered=True\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.361032Z","iopub.status.idle":"2024-12-22T04:00:46.361439Z","shell.execute_reply.started":"2024-12-22T04:00:46.361257Z","shell.execute_reply":"2024-12-22T04:00:46.361276Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"calculate_stats(train, 'PreInt_EduHx-Season')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.362529Z","iopub.status.idle":"2024-12-22T04:00:46.362899Z","shell.execute_reply.started":"2024-12-22T04:00:46.36272Z","shell.execute_reply":"2024-12-22T04:00:46.362737Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 3, figsize=(18, 5))\n\n# Hours of Internet Use\nax1 = sns.countplot(x='internet_use_encoded', data=train, palette=\"Set3\", ax=axes[0])\naxes[0].set_title('Distribution of Hours of Internet Use')\naxes[0].set_xlabel('Hours per Day Group')\naxes[0].set_ylabel('Count')\n\ntotal = len(train['internet_use_encoded'])\nfor p in ax1.patches:\n    count = int(p.get_height())\n    percentage = '{:.1f}%'.format(100 * count / total)\n    ax1.annotate(f'{count} ({percentage})', (p.get_x() + p.get_width() / 2., p.get_height()), \n                 ha='center', va='baseline', fontsize=10, color='black', xytext=(0, 5), \n                 textcoords='offset points')\n\n# Hours of Internet Use by Age\nsns.boxplot(y=train['Basic_Demos-Age'], x=train['internet_use_encoded'], ax=axes[1], palette=\"Set3\")\naxes[1].set_title('Hours of Internet Use by Age')\naxes[1].set_ylabel('Age')\naxes[1].set_xlabel('Hours per Day Group')\n\n# # Hours of Internet Use (numeric) by Age Group\n# sns.boxplot(y='PreInt_EduHx-computerinternet_hoursday', x='Age Group', data=train, ax=axes[2], palette=\"Set3\")\n# axes[2].set_title('Internet Hours by Age Group')\n# axes[2].set_ylabel('Hours per Day (Numeric)')\n# axes[2].set_xlabel('Age Group')\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.364382Z","iopub.status.idle":"2024-12-22T04:00:46.364743Z","shell.execute_reply.started":"2024-12-22T04:00:46.364576Z","shell.execute_reply":"2024-12-22T04:00:46.364593Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 8. Parent-Child Internet Addiction Test ","metadata":{}},{"cell_type":"markdown","source":"## 9.Physical Activity Questionaire","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"groups.get('Physical Activity Questionnaire (Adolescents)', [])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.366215Z","iopub.status.idle":"2024-12-22T04:00:46.366768Z","shell.execute_reply.started":"2024-12-22T04:00:46.366487Z","shell.execute_reply":"2024-12-22T04:00:46.366517Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data = train[train['PAQ_A-PAQ_A_Total'].notnull()]\nage_range = data['Basic_Demos-Age']\nprint(\n    f\"Age range for Adolescents (with PAQ_A_Total data):\"\n    f\" {age_range.min()} - {age_range.max()} years\"\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.367997Z","iopub.status.idle":"2024-12-22T04:00:46.368374Z","shell.execute_reply.started":"2024-12-22T04:00:46.3682Z","shell.execute_reply":"2024-12-22T04:00:46.368218Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(18, 5))\n\n# PAQ_A-Season\nplt.subplot(1, 3, 1)\ntrain['PAQ_A-Season'].value_counts(normalize=True).plot.pie(\n    autopct='%1.1f%%', colors=plt.cm.Set3.colors\n)\nplt.title('PAQ_A-Season (Adolescents)')\n\n# PAQ_A-PAQ_A_Total\nplt.subplot(1, 3, 2)\nsns.histplot(train['PAQ_A-PAQ_A_Total'], bins=20, kde=True)\nplt.title('PAQ_A-PAQ_A_Total (Adolescents)')\n\n# PAQ_A_Total by Season\nplt.subplot(1, 3, 3)\nsns.violinplot(x='PAQ_A-Season', y='PAQ_A-PAQ_A_Total', data=train, palette=\"Set3\")\nplt.title('PAQ_A_Total by Season (Adolescents)')\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.370576Z","iopub.status.idle":"2024-12-22T04:00:46.371175Z","shell.execute_reply.started":"2024-12-22T04:00:46.370855Z","shell.execute_reply":"2024-12-22T04:00:46.370886Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"calculate_stats(train, ['PAQ_A-PAQ_A_Total'])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.372761Z","iopub.status.idle":"2024-12-22T04:00:46.373321Z","shell.execute_reply.started":"2024-12-22T04:00:46.373038Z","shell.execute_reply":"2024-12-22T04:00:46.373074Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(18, 5))\n\n# PAQ_C-Season\nplt.subplot(1, 3, 1)\ntrain['PAQ_C-Season'].value_counts(normalize=True).plot.pie(\n    autopct='%1.1f%%', colors=plt.cm.Set3.colors\n)\nplt.title('PAQ_C-Season (Children)')\n\n# PAQ_C-PAQ_C_Total\nplt.subplot(1, 3, 2)\nsns.histplot(train['PAQ_C-PAQ_C_Total'], bins=20, kde=True)\nplt.title('PAQ_C-PAQ_C_Total (Children)')\n\n# PAQ_C_Total by Season\nplt.subplot(1, 3, 3)\nsns.violinplot(x='PAQ_C-Season', y='PAQ_C-PAQ_C_Total', data=train, palette=\"Set3\")\nplt.title('PAQ_C_Total by Season (Children)')\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.374841Z","iopub.status.idle":"2024-12-22T04:00:46.375385Z","shell.execute_reply.started":"2024-12-22T04:00:46.375117Z","shell.execute_reply":"2024-12-22T04:00:46.375146Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Note:\nThe division into adolescents and children seems to be incorrect (participants with data in the children columns (PAQ_C_Total) are 7 - 17 years old - overlapping with those with non-missing data in the adolescents columns - 13 - 18 years old).\nPhysical activity levels are fairly stable over the seasons, with only minor variations, although are slightly lower in the fall and winter for adolescents and children, respectively.\nThere are many missing values for these features","metadata":{}},{"cell_type":"markdown","source":"## 10. Children's Global Assessment Scale\n","metadata":{}},{"cell_type":"code","source":"groups.get(\"Children's Global Assessment Scale\", [])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.376499Z","iopub.status.idle":"2024-12-22T04:00:46.377039Z","shell.execute_reply.started":"2024-12-22T04:00:46.376749Z","shell.execute_reply":"2024-12-22T04:00:46.376777Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data = train[train['CGAS-CGAS_Score'].notnull()]\nage_range = data['Basic_Demos-Age']\nprint(\n    f\"Age range for participants with CGAS-CGAS_Score data:\"\n    f\" {age_range.min()} - {age_range.max()} years\"\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.37874Z","iopub.status.idle":"2024-12-22T04:00:46.379295Z","shell.execute_reply.started":"2024-12-22T04:00:46.37901Z","shell.execute_reply":"2024-12-22T04:00:46.379039Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"calculate_stats(train, 'CGAS-CGAS_Score')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.381578Z","iopub.status.idle":"2024-12-22T04:00:46.382137Z","shell.execute_reply.started":"2024-12-22T04:00:46.381833Z","shell.execute_reply":"2024-12-22T04:00:46.381862Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.loc[train['CGAS-CGAS_Score'] == 999, 'CGAS-CGAS_Score'] = np.nan\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.384163Z","iopub.status.idle":"2024-12-22T04:00:46.384564Z","shell.execute_reply.started":"2024-12-22T04:00:46.384369Z","shell.execute_reply":"2024-12-22T04:00:46.384388Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(12, 5))\n\n# CGAS-Season\nplt.subplot(1, 2, 1)\ncgas_season_counts = train['CGAS-Season'].value_counts(normalize=True)\nplt.pie(\n    cgas_season_counts, \n    labels=cgas_season_counts.index, \n    autopct='%1.1f%%', \n    startangle=90, \n    colors=sns.color_palette(\"Set3\")\n)\nplt.title('CGAS-Season')\nplt.axis('equal')\n\n# CGAS-CGAS_Score without outliers (score == 999)\nplt.subplot(1, 2, 2)\nsns.histplot(\n    train['CGAS-CGAS_Score'].dropna(),\n    bins=20, kde=True\n)\nplt.title('CGAS-CGAS_Score (Without Outlier)')\nplt.xlabel('CGAS Score')\nplt.ylabel('Count')\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.386604Z","iopub.status.idle":"2024-12-22T04:00:46.387126Z","shell.execute_reply.started":"2024-12-22T04:00:46.386869Z","shell.execute_reply":"2024-12-22T04:00:46.386889Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"calculate_stats(train, 'CGAS-CGAS_Score')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.38824Z","iopub.status.idle":"2024-12-22T04:00:46.388623Z","shell.execute_reply.started":"2024-12-22T04:00:46.388436Z","shell.execute_reply":"2024-12-22T04:00:46.388454Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"CGAS Interpretation (Reference)\nCGAS is a rating of general functioning for children and young people aged 4-16 years old. The CGAS asks the clinician to rate the child from 1 to 100 based on their lowest level of functioning, regardless of treatment or prognosis, over a specified time period.\n\nSince the CGAS is a measure of general functioning, and the SII reflects the severity of the impact of Internet use on that functioning, I expect this feature, along with Internet use, to be the most important in predicting the SII.\n\nLet's bin the CGAS-CGAS_Score column based on the established score categories and draw counts:","metadata":{}},{"cell_type":"code","source":"bins = np.arange(0, 101, 10)\nlabels = [\n    \"1-10: Needs constant supervision (24 hour care)\",\n    \"11-20: Needs considerable supervision\",\n    \"21-30: Unable to function in almost all areas\",\n    \"31-40: Major impairment in functioning in several areas\",\n    \"41-50: Moderate degree of interference in functioning\",\n    \"51-60: Variable functioning with sporadic difficulties\",\n    \"61-70: Some difficulty in a single area\",\n    \"71-80: No more than slight impairment in functioning\",\n    \"81-90: Good functioning in all areas\",\n    \"91-100: Superior functioning\"\n]\n\ntrain['CGAS_Score_Bin'] = pd.cut(\n    train['CGAS-CGAS_Score'], bins=bins, labels=labels\n)\n\ncounts = train['CGAS_Score_Bin'].value_counts().reindex(labels)\nprop = (counts / counts.sum() * 100).round(1)\ncount_prop_labels = counts.astype(str) + \" (\" + prop.astype(str) + \"%)\"\n\nplt.figure(figsize=(18, 6))\nbars = plt.barh(labels, counts)\nplt.xlabel('Count')\nplt.title('CGAS Score Distribution')\n\nfor bar, label in zip(bars, count_prop_labels):\n    plt.text(\n        bar.get_width(), bar.get_y() + bar.get_height() / 2, label, va='center'\n    )\n\nplt.gca().invert_yaxis()\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.389737Z","iopub.status.idle":"2024-12-22T04:00:46.390338Z","shell.execute_reply.started":"2024-12-22T04:00:46.39002Z","shell.execute_reply":"2024-12-22T04:00:46.390062Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Note:\nThe majority of individuals have CGAS scores between 51-80 (79.7%), i.e. sporadic difficulties to only slight impairments\nTwo participants have extreme difficulty in functioning","metadata":{}},{"cell_type":"code","source":"# Check column names\nprint(\"Columns in the DataFrame:\", train.columns)\n\n# Verify existence of required columns\nrequired_columns = ['CGAS_Score_Bin', 'complete_resp_total']\nmissing_columns = [col for col in required_columns if col not in train.columns]\n\nif missing_columns:\n    print(\"The following required columns are missing:\", missing_columns)\nelse:\n    # Proceed with the filtering\n    train_filt = train.dropna(subset=required_columns)\n    train_filt.loc[:, 'CGAS_Score_Bin'] = train_filt['CGAS_Score_Bin'].cat.remove_unused_categories()\n    train_filt.loc[:, 'sii'] = train_filt['sii'].cat.remove_unused_categories()\n    print(len(train_filt))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.39228Z","iopub.status.idle":"2024-12-22T04:00:46.39282Z","shell.execute_reply.started":"2024-12-22T04:00:46.392546Z","shell.execute_reply":"2024-12-22T04:00:46.392575Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Ensure 'complete_resp_total' exists and is numeric\nif 'complete_resp_total' in train_filt.columns:\n    train_filt['complete_resp_total'] = pd.to_numeric(train_filt['complete_resp_total'], errors='coerce')\nelse:\n    print(\"'complete_resp_total' column is missing from the dataset.\")\n\n# Ensure 'CGAS_Score_Bin' is categorical\nif 'CGAS_Score_Bin' in train_filt.columns:\n    train_filt['CGAS_Score_Bin'] = train_filt['CGAS_Score_Bin'].astype('category')\n    range_labels = [label.split(\":\")[0] for label in train_filt['CGAS_Score_Bin'].cat.categories]\nelse:\n    print(\"'CGAS_Score_Bin' column is missing from the dataset.\")\n\n# Regenerate the plots\nfig, axes = plt.subplots(1, 2, figsize=(16, 5))\n\n# CGAS-CGAS_Score vs sii\nsns.boxplot(\n    data=train_filt,\n    x='sii', y='CGAS-CGAS_Score',\n    palette='Set3', ax=axes[0]\n)\naxes[0].set_xlabel('SII Score')\naxes[0].set_ylabel('CGAS Score')\naxes[0].set_title('Distribution of CGAS Scores by SII')\n\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.394153Z","iopub.status.idle":"2024-12-22T04:00:46.394711Z","shell.execute_reply.started":"2024-12-22T04:00:46.394413Z","shell.execute_reply":"2024-12-22T04:00:46.394443Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# V.Psychosocial and Emotional Impact\n","metadata":{}},{"cell_type":"markdown","source":"## 11. Sleep Disturbance Scale","metadata":{}},{"cell_type":"code","source":"groups.get('Sleep Disturbance Scale', [])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.396073Z","iopub.status.idle":"2024-12-22T04:00:46.396612Z","shell.execute_reply.started":"2024-12-22T04:00:46.396342Z","shell.execute_reply":"2024-12-22T04:00:46.396371Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(18, 5))\n\n# SDS-Season (Pie Chart)\nplt.subplot(1, 3, 1)\nsds_season_counts = train['SDS-Season'].value_counts(normalize=True)\nplt.pie(\n    sds_season_counts, \n    labels=sds_season_counts.index, \n    autopct='%1.1f%%', \n    startangle=90, \n    colors=sns.color_palette(\"Set3\")\n)\nplt.title('SDS-Season')\n\n# SDS-SDS_Total_Raw\nplt.subplot(1, 3, 2)\nsns.histplot(train['SDS-SDS_Total_Raw'].dropna(), bins=20, kde=True)\nplt.title('SDS-SDS_Total_Raw')\nplt.xlabel('Value')\n\n# SDS-SDS_Total_T\nplt.subplot(1, 3, 3)\nsns.histplot(train['SDS-SDS_Total_T'].dropna(), bins=20, kde=True)\nplt.title('SDS-SDS_Total_T')\nplt.xlabel('Value')\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.397769Z","iopub.status.idle":"2024-12-22T04:00:46.398328Z","shell.execute_reply.started":"2024-12-22T04:00:46.398056Z","shell.execute_reply":"2024-12-22T04:00:46.398086Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"calculate_stats(train, ['SDS-SDS_Total_Raw', 'SDS-SDS_Total_T'])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.401233Z","iopub.status.idle":"2024-12-22T04:00:46.401648Z","shell.execute_reply.started":"2024-12-22T04:00:46.401459Z","shell.execute_reply":"2024-12-22T04:00:46.40148Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Note: \n- Seasonal variation is notable, with the largest segment contributing 25.3% and the smallest contributing 7.6%. This suggests that the frequency or severity of sleep disturbances varies across seasons, with significant differences in occurrence.\n- SDS-SDS_Total_Raw have a right-skewed distribution. The majority of values fall between 30-60 points, with a peak around 40-50. This indicates that most children experience mild to moderate sleep disturbances.","metadata":{}},{"cell_type":"markdown","source":"# 2.Processing data","metadata":{}},{"cell_type":"code","source":"!pip install --upgrade category_encoders","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.402866Z","iopub.status.idle":"2024-12-22T04:00:46.403308Z","shell.execute_reply.started":"2024-12-22T04:00:46.403124Z","shell.execute_reply":"2024-12-22T04:00:46.403146Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"numerical_cols = train.select_dtypes(exclude='object').columns\ncategorical_cols = train.select_dtypes(include='object').columns\ntrain = train.dropna(subset=['sii'])\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.405171Z","iopub.status.idle":"2024-12-22T04:00:46.405565Z","shell.execute_reply.started":"2024-12-22T04:00:46.405382Z","shell.execute_reply":"2024-12-22T04:00:46.405402Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train[categorical_cols].head(10)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.407264Z","iopub.status.idle":"2024-12-22T04:00:46.407642Z","shell.execute_reply.started":"2024-12-22T04:00:46.407467Z","shell.execute_reply":"2024-12-22T04:00:46.407487Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import category_encoders as ce\n\nencoder = ce.TargetEncoder(cols=categorical_cols, smoothing=1.0)\ntrain[categorical_cols] = encoder.fit_transform(train[categorical_cols], train['sii'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.408692Z","iopub.status.idle":"2024-12-22T04:00:46.409133Z","shell.execute_reply.started":"2024-12-22T04:00:46.408904Z","shell.execute_reply":"2024-12-22T04:00:46.408924Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train[categorical_cols].head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T04:00:46.411435Z","iopub.status.idle":"2024-12-22T04:00:46.411817Z","shell.execute_reply.started":"2024-12-22T04:00:46.411637Z","shell.execute_reply":"2024-12-22T04:00:46.411655Z"}},"outputs":[],"execution_count":null}]}