{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### EDA for Child Mind Institute - Problematic Internet Use Competition","metadata":{}},{"cell_type":"markdown","source":"<div style=\"border: 2px solid #c9c9c9; padding: 15px; border-radius: 5px; background-color: #f7f7f7;\">\n<h3>Summary</h3>\n<ul>\n  <li>We have a lot of missing data in the dataset. About 30% of the target variable, <code>sii</code>, is missing.</li>\n  <li>The target variable <code>sii</code> seems to increase with age, peaking around 17, and then going down in an inverted U-shape.</li>\n  <li>There are some clearly wrong values in the physical characteristic features.</li>\n  <li>As the stage of hypertension increases, so do the values of the target variable <code>sii</code>, BMI, weight, and abdominal circumference.</li>\n</ul>","metadata":{}},{"cell_type":"code","source":"import os\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport warnings\n\nwarnings.filterwarnings('ignore')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:17.616463Z","iopub.execute_input":"2024-11-29T08:36:17.617197Z","iopub.status.idle":"2024-11-29T08:36:17.624657Z","shell.execute_reply.started":"2024-11-29T08:36:17.617113Z","shell.execute_reply":"2024-11-29T08:36:17.623435Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"input_dir = '/kaggle/input/child-mind-institute-problematic-internet-use'\ndf_train = pd.read_csv(os.path.join(input_dir, 'train.csv'))\ndf_test = pd.read_csv(os.path.join(input_dir, 'test.csv'))\n\nprint(f'train: {df_train.shape}')\nprint(f'test: {df_test.shape}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:17.626932Z","iopub.execute_input":"2024-11-29T08:36:17.627323Z","iopub.status.idle":"2024-11-29T08:36:17.703135Z","shell.execute_reply.started":"2024-11-29T08:36:17.627283Z","shell.execute_reply":"2024-11-29T08:36:17.701773Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There is a difference in the columns included in the training and test datasets. The test dataset is missing 23 columns, including the target variable `sii`.","metadata":{}},{"cell_type":"code","source":"[col for col in df_train.columns if col not in df_test.columns]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:17.704934Z","iopub.execute_input":"2024-11-29T08:36:17.705452Z","iopub.status.idle":"2024-11-29T08:36:17.715017Z","shell.execute_reply.started":"2024-11-29T08:36:17.705399Z","shell.execute_reply":"2024-11-29T08:36:17.713601Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### Missing values\nThe training data has many missing values; only four columns are complete. The target variable `sii` is also missing about 30% of its values.","metadata":{}},{"cell_type":"code","source":"missing_ratios = df_train.isna().mean() * 100\nmissing_ratios = missing_ratios.sort_values(ascending=False)\nfilled_ratios = 100 - missing_ratios\n\nplt.figure(figsize=(8, 15))\nbar_width = 0.8\n\nplt.barh(\n    missing_ratios.index, missing_ratios.values,\n    color='tomato', edgecolor='white', height=bar_width, label='Missing (%)'\n)\n\nplt.barh(\n    missing_ratios.index, filled_ratios.values, left=missing_ratios.values,\n    color='springgreen', edgecolor='white', height=bar_width, label='Filled (%)'\n)\n\nplt.xlabel('Percentage (%)')\nplt.ylabel('Columns')\nplt.title('Percentage of Missing and Filled Values by Column')\nplt.legend(loc='lower right')\nplt.grid(axis='x', linestyle='--', alpha=0.7)\n\nplt.yticks(fontsize=10, rotation=0)\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:17.718377Z","iopub.execute_input":"2024-11-29T08:36:17.718860Z","iopub.status.idle":"2024-11-29T08:36:18.959969Z","shell.execute_reply.started":"2024-11-29T08:36:17.718806Z","shell.execute_reply":"2024-11-29T08:36:18.958589Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"How we deal with the missing data in those columns, and especially the 30% missing in `sii`, is going to affect our final prediction accuracy.","metadata":{}},{"cell_type":"markdown","source":"#### Target Variable\nIn our target variable `sii`, there's a really high concentration of zeros, and then fewer and fewer people as the 'sii' value goes up.","metadata":{}},{"cell_type":"code","source":"data = df_train['sii']\nvalue_counts = data.value_counts(dropna=True).sort_index()\nbins = list(value_counts.index.astype(int).astype(str))\ncounts = list(value_counts.values)\n\nfig, ax = plt.subplots(figsize=(8, 6))\nbars = ax.bar(bins, counts, color='skyblue', edgecolor='black')\n\nfor bar, cnt in zip(bars, counts):\n    ax.text(bar.get_x() + bar.get_width() / 2, bar.get_height(), f'{cnt}', ha='center', va='bottom', fontsize=10)\n\nax.grid(axis='y', linestyle='--', alpha=0.7)\nax.set_title('sii distribution')\nax.set_xlabel('sii')\nax.set_ylabel('Freq')\nfig.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:18.961378Z","iopub.execute_input":"2024-11-29T08:36:18.961711Z","iopub.status.idle":"2024-11-29T08:36:19.232135Z","shell.execute_reply.started":"2024-11-29T08:36:18.961663Z","shell.execute_reply":"2024-11-29T08:36:19.230923Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There's a slight difference in the proportions between genders, but it's probably not significant given our data size.","metadata":{}},{"cell_type":"code","source":"genders = df_train['Basic_Demos-Sex'].unique()\ngender_map = {0: 'male', 1: 'female'}\nfig, axes = plt.subplots(1, 2, figsize=(10, 6))\n\nfor ax, gender in zip(axes, genders):\n    data = df_train.loc[df_train['Basic_Demos-Sex'] == gender]\n    value_counts = data['sii'].value_counts(dropna=True).sort_index()\n    labels = value_counts.index.astype(int).astype(str).tolist()\n\n    ax.pie(\n        value_counts.values, labels=labels, autopct='%1.1f%%',\n        colors=plt.cm.tab20.colors, startangle=90, counterclock=False\n    )\n\n    ax.set_title(gender_map[gender])\n\nplt.suptitle('Distribution of sii by Sex')\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:19.233723Z","iopub.execute_input":"2024-11-29T08:36:19.234143Z","iopub.status.idle":"2024-11-29T08:36:19.504394Z","shell.execute_reply.started":"2024-11-29T08:36:19.234073Z","shell.execute_reply":"2024-11-29T08:36:19.503258Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"As age increases, the average value of the target variable also tends to go up.There's not much data for people older than 19, so we can't be totally sure, but it seems to peak around 17 and then go down in an inverted U-shape after that.","metadata":{}},{"cell_type":"code","source":"average_values = (\n    df_train.groupby('Basic_Demos-Age')['sii'].mean()\n)\n\nfig, ax = plt.subplots()\nax.plot(\n    average_values.index, average_values.values,\n    marker='o', label='Average sii'\n)\nax.grid(True, linestyle='--', alpha=0.7)\nax.set_xticks(range(int(average_values.index.min()), int(average_values.index.max()) + 1))\nax.set_xlabel('Basic_Demos-Age')\nax.set_ylabel('Average sii')\nax.set_title('Average sii by Age')\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:19.505720Z","iopub.execute_input":"2024-11-29T08:36:19.506201Z","iopub.status.idle":"2024-11-29T08:36:19.844508Z","shell.execute_reply.started":"2024-11-29T08:36:19.506131Z","shell.execute_reply":"2024-11-29T08:36:19.843073Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"gender_map = {0: 'male', 1: 'female'}\npivot_df = df_train.pivot_table(\n    index='Basic_Demos-Age', columns='Basic_Demos-Sex', aggfunc='size', fill_value=0\n)\n\nwidth = 0.3\nfig, ax = plt.subplots(figsize=(8, 5))\nax.bar(pivot_df.index - width/2, pivot_df[0], width=width, label=gender_map[0])\nax.bar(pivot_df.index + width/2, pivot_df[1], width=width, label=gender_map[1])\n\nax.set_xticks(range(int(average_values.index.min()), int(average_values.index.max()) + 1))\nax.grid(axis='y', linestyle='--', alpha=0.7)\nax.set_xlabel('Basic_Demos-Age')\nax.set_ylabel('Freq')\nax.set_title('Age distribution by Sex')\nax.legend();","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:19.846508Z","iopub.execute_input":"2024-11-29T08:36:19.846854Z","iopub.status.idle":"2024-11-29T08:36:20.411287Z","shell.execute_reply.started":"2024-11-29T08:36:19.846816Z","shell.execute_reply":"2024-11-29T08:36:20.410215Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- We've got more male data than female data across the board, and the difference is really big between 5 and 12 years old.\n- We saw in that pie chart that males had about 9% more zeros for the target variable, but that's probably just because we have more young guys in our data.\n- So, I think we can say that gender doesn't have much impact on `sii`, but age definitely does.s.","metadata":{}},{"cell_type":"markdown","source":"#### Parent-Child Internet Addiction Test (PCIAT)\n- The target variable `sii` is determined by a threshold based on the total score of the PCIAT questions.","metadata":{}},{"cell_type":"code","source":"(\n    df_train\n    .groupby('sii')['PCIAT-PCIAT_Total']\n    .agg(['min', 'max'])\n    .rename(columns={'min': 'PCIAT-PCIAT_Total min', 'max': 'PCIAT-PCIAT_Total max'})\n    .reset_index()\n    .astype(int)\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:20.414217Z","iopub.execute_input":"2024-11-29T08:36:20.414556Z","iopub.status.idle":"2024-11-29T08:36:20.429197Z","shell.execute_reply.started":"2024-11-29T08:36:20.414520Z","shell.execute_reply":"2024-11-29T08:36:20.427915Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"border: 2px solid #c9c9c9; padding: 15px; border-radius: 5px; background-color: #f7f7f7;\">\nThe PCIAT questions appear to be designed for parents or guardians to answer.<br>\nThe relevance of certain PCIAT questions varies depending on the child's age. Younger children (e.g., 5-year-olds) are unlikely to be performing chores or using email, while internet time limits are less likely to be imposed on young adults (e.g., 22-year-olds).<br>\nResponses to the PCIAT questions are likely influenced by subjective parental opinions, including family circumstances and expectations for their children. While `sii` peaks around age 17, this may be a result of increased parent-child conflict during adolescence, rather than a characteristic specific to that age.","metadata":{}},{"cell_type":"markdown","source":"#### physical characteristics\nSome physical characteristic features have impossible values, like a weight of zero.","metadata":{}},{"cell_type":"code","source":"physical_features = [\n    'Physical-BMI', 'Physical-Height', 'Physical-Weight', 'Physical-Waist_Circumference',\n    'Physical-Diastolic_BP', 'Physical-Systolic_BP', 'Physical-HeartRate'\n]\ndf_physical = df_train[physical_features]\npd.DataFrame({\n    'min': df_physical.min(),\n    'max': df_physical.max().round(1)\n})","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:20.430759Z","iopub.execute_input":"2024-11-29T08:36:20.431265Z","iopub.status.idle":"2024-11-29T08:36:20.450936Z","shell.execute_reply.started":"2024-11-29T08:36:20.431213Z","shell.execute_reply":"2024-11-29T08:36:20.449694Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We'll convert height and waist circumference to centimeters and weight to kilograms, then look at some specific records.","metadata":{}},{"cell_type":"code","source":"tmp_features = ['Basic_Demos-Age', 'Basic_Demos-Sex', 'sii'] + physical_features\ndf_physical = df_train[tmp_features]\ndf_physical['Physical-Height_cm'] = round(df_physical['Physical-Height'] * 2.54, 1)\ndf_physical['Physical-Waist_Circumference_cm'] = round(df_physical['Physical-Waist_Circumference'] * 2.54, 1)\ndf_physical['Physical-Weight_kg'] = round(df_physical['Physical-Weight'] * 0.453, 1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:20.452921Z","iopub.execute_input":"2024-11-29T08:36:20.453393Z","iopub.status.idle":"2024-11-29T08:36:20.464675Z","shell.execute_reply.started":"2024-11-29T08:36:20.453339Z","shell.execute_reply":"2024-11-29T08:36:20.463252Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"(\n    df_physical\n    .loc[~df_physical['Physical-BMI'].isna(), ['Basic_Demos-Age', 'Basic_Demos-Sex', 'Physical-BMI', 'Physical-Height_cm', 'Physical-Weight_kg', 'Physical-Waist_Circumference_cm']]\n    .sort_values('Physical-BMI', ascending=False)\n    .head(10)\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:20.466014Z","iopub.execute_input":"2024-11-29T08:36:20.466558Z","iopub.status.idle":"2024-11-29T08:36:20.496394Z","shell.execute_reply.started":"2024-11-29T08:36:20.466464Z","shell.execute_reply":"2024-11-29T08:36:20.495036Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We found an 8-year-old with both height and waist circumference at 83cm and a BMI of 59. Is this a genuine measurement or a data entry error?  Similarly, we have a 7-year-old recorded as weighing over 100kg. Verifying the accuracy of these records is proving to be very difficult.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots()\nax.scatter(df_physical['Physical-Height_cm'], df_physical['Physical-Weight_kg'])\nax.set_xlabel('Physical-Height_cm')\nax.set_ylabel('Physical-Weight_kg')\nax.set_title('Height vs. Weight')\nax.grid(True, linestyle='--', alpha=0.7)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:20.497769Z","iopub.execute_input":"2024-11-29T08:36:20.498231Z","iopub.status.idle":"2024-11-29T08:36:20.775144Z","shell.execute_reply.started":"2024-11-29T08:36:20.498177Z","shell.execute_reply":"2024-11-29T08:36:20.773942Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We'll calculate relative ratios using US average height and weight data for different age groups (from the CDC's National Health Statistics Reports: https://www.cdc.gov/nchs/data/nhsr/nhsr160-508.pdf). Since we don't know the exact year of our data, we'll use the rounded 2015-2018 values for height and weight in centimeters and kilograms, assuming the same values for ages 19 and up.","metadata":{}},{"cell_type":"code","source":"age = range(5, 23)\nmale_ave_height = [113, 119, 126, 132, 136, 142, 150, 154, 163, 170, 172, 174, 175, 175, 175, 175, 175, 175]\nmale_ave_weight = [21, 24, 28, 32, 35, 41, 47, 49, 60, 65, 73, 71, 77, 75, 80, 80, 80, 80]\nmale_ave_waist_circumference = [55, 57, 60, 63, 67, 70, 73, 73, 80, 80, 85, 82, 86, 85, 90, 90, 90, 90]\nfemale_ave_height = [112, 119, 124, 130, 137, 143, 150, 155, 159, 162, 161, 162, 163, 162, 162, 162, 162, 162]\nfemale_ave_weight = [21, 24, 27, 31, 35, 41, 48, 53, 57, 62, 62, 65, 68, 69, 71, 71, 71, 71]\nfemale_ave_waist_circumference = [56, 57, 61, 63, 65, 70, 73, 77, 79, 81, 80, 82, 84, 86, 88, 88, 88, 88]\n\ndf_average = pd.DataFrame({\n    'age': age,\n    'male_height': male_ave_height,\n    'male_weight': male_ave_weight,\n    'male_waist_circumference': male_ave_waist_circumference,\n    'female_height': female_ave_height,\n    'female_weight': female_ave_weight,\n    'female_waist_circumference': female_ave_waist_circumference\n})","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:20.776579Z","iopub.execute_input":"2024-11-29T08:36:20.776889Z","iopub.status.idle":"2024-11-29T08:36:20.786874Z","shell.execute_reply.started":"2024-11-29T08:36:20.776856Z","shell.execute_reply":"2024-11-29T08:36:20.785421Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_physical['average_height'] = np.nan\ndf_physical['average_weight'] = np.nan\ndf_physical['average_waist_circumference'] = np.nan\n\nfor index, row in df_physical.iterrows():\n    age = row['Basic_Demos-Age']\n    sex = row['Basic_Demos-Sex']\n    if sex == 0:\n        df_physical.loc[index, 'average_height'] = df_average[df_average['age'] == age]['male_height'].values[0]\n        df_physical.loc[index, 'average_weight'] = df_average[df_average['age'] == age]['male_weight'].values[0]\n        df_physical.loc[index, 'average_waist_circumference'] = df_average[df_average['age'] == age]['male_waist_circumference'].values[0]\n    elif sex == 1:\n        df_physical.loc[index, 'average_height'] = df_average[df_average['age'] == age]['female_height'].values[0]\n        df_physical.loc[index, 'average_weight'] = df_average[df_average['age'] == age]['female_weight'].values[0]\n        df_physical.loc[index, 'average_waist_circumference'] = df_average[df_average['age'] == age]['female_waist_circumference'].values[0]\n\ndf_physical['height_ratio'] = df_physical['Physical-Height_cm'] / df_physical['average_height']\ndf_physical['weight_ratio'] = df_physical['Physical-Weight_kg'] / df_physical['average_weight']\ndf_physical['waist_circumference_ratio'] = df_physical['Physical-Waist_Circumference_cm'] / df_physical['average_waist_circumference']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:20.788737Z","iopub.execute_input":"2024-11-29T08:36:20.789600Z","iopub.status.idle":"2024-11-29T08:36:27.382062Z","shell.execute_reply.started":"2024-11-29T08:36:20.789542Z","shell.execute_reply":"2024-11-29T08:36:27.380839Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_tmp = df_physical[~df_physical['sii'].isna()]\nfeatures = ['height_ratio', 'weight_ratio', 'waist_circumference_ratio']\nfig, axes = plt.subplots(1, 3, figsize=(12, 6))\nfor feature, ax in zip(features, axes):\n    sns.stripplot(data=df_tmp, x='sii', y=feature, ax=ax)\n    ax.grid(axis='y', linestyle='--', alpha=0.7)\n    ax.set_title(f'sii vs. {feature}')\n\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:27.383358Z","iopub.execute_input":"2024-11-29T08:36:27.383682Z","iopub.status.idle":"2024-11-29T08:36:28.251288Z","shell.execute_reply.started":"2024-11-29T08:36:27.383646Z","shell.execute_reply":"2024-11-29T08:36:28.250086Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There doesn't seem to be any correlation between the relative physical characteristics and the target variable `sii`. Also, the sample size for `sii` equal to 3 is really small.","metadata":{}},{"cell_type":"markdown","source":"#### Blood Pressure & Heart Rate\nThere are inconsistencies in the blood pressure data. Some records show diastolic pressure exceeding systolic pressure, which is physiologically impossible.","metadata":{}},{"cell_type":"code","source":"bp_hr_cols = ['Physical-Diastolic_BP', 'Physical-Systolic_BP', 'Physical-HeartRate']\ndf_physical[df_physical['Physical-Diastolic_BP'] >= df_physical['Physical-Systolic_BP']][bp_hr_cols]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:28.252549Z","iopub.execute_input":"2024-11-29T08:36:28.252876Z","iopub.status.idle":"2024-11-29T08:36:28.267520Z","shell.execute_reply.started":"2024-11-29T08:36:28.252840Z","shell.execute_reply":"2024-11-29T08:36:28.266232Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We'll categorize blood pressure according to AHA/ACC guidelines (2017): \n-  Normal: <120/<80 mmH\n- Elevated: 120-129/<80 mmHg\n- Stage 1 Hypertension: 130-139 or 80-89 mmHg\n- Stage 2 Hypertension: ≥140 or ≥90 mmHg\n- Hypertensive Crisis: ≥180 or ≥120 mmHg mmHg","metadata":{}},{"cell_type":"code","source":"def classify_blood_pressure(row):\n    systolic = row['Physical-Systolic_BP']\n    diastolic = row['Physical-Diastolic_BP']\n\n    if systolic < 120 and diastolic < 80:\n        return 0\n    elif 120 <= systolic <= 129 and diastolic < 80:\n      return 1\n    elif 130 <= systolic <= 139 or 80 <= diastolic <= 89:\n        return 2\n    elif 140 <= systolic or 90 <= diastolic:\n        return 3\n    elif systolic >= 180 or diastolic >= 120:\n        return 4\n    else:\n        return 0\n\ndf_physical['blood_pressure_level'] = df_physical.apply(classify_blood_pressure, axis=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:28.269232Z","iopub.execute_input":"2024-11-29T08:36:28.269671Z","iopub.status.idle":"2024-11-29T08:36:28.322617Z","shell.execute_reply.started":"2024-11-29T08:36:28.269620Z","shell.execute_reply":"2024-11-29T08:36:28.321500Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(df_physical['blood_pressure_level'].value_counts())\n(\n    df_physical[['Basic_Demos-Age', 'sii', 'Physical-BMI', 'Physical-Height_cm', 'Physical-Weight_kg', 'Physical-Waist_Circumference_cm', 'Physical-HeartRate', 'blood_pressure_level']]\n    .groupby('blood_pressure_level').agg({'mean'})\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:28.324200Z","iopub.execute_input":"2024-11-29T08:36:28.324544Z","iopub.status.idle":"2024-11-29T08:36:28.352795Z","shell.execute_reply.started":"2024-11-29T08:36:28.324509Z","shell.execute_reply":"2024-11-29T08:36:28.351659Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In this dataset, it appears that no one is diagnosed with hypertensive emergency. Also, looking at age and height, the group with a blood_pressure_level of 0 seems to have a higher proportion of young people. As the blood pressure level increases, the target variable sii, weight, BMI, and abdominal circumference also seem to increase (0 < 1 < 2, 3). We could say that people who are highly dependent on the internet tend to lack exercise habits, which may lead to lifestyle-related diseases such as obesity and hypertension.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots()\nax.hist(df_physical['Physical-HeartRate'], edgecolor='white', bins=50)\nax.set_title('Heart Rate Distribution')\nax.grid(axis='y', linestyle='--', alpha=0.7)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:28.354403Z","iopub.execute_input":"2024-11-29T08:36:28.354785Z","iopub.status.idle":"2024-11-29T08:36:28.656474Z","shell.execute_reply.started":"2024-11-29T08:36:28.354745Z","shell.execute_reply":"2024-11-29T08:36:28.655362Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Since the vast majority of heart rate measurements fall within the normal resting range (60-100 bpm), we can assume that these measurements were taken at rest.","metadata":{}},{"cell_type":"markdown","source":"#### Fitness","metadata":{}},{"cell_type":"code","source":"fit_columns = [\n    'Fitness_Endurance-Season', 'Fitness_Endurance-Max_Stage', 'Fitness_Endurance-Time_Mins', 'Fitness_Endurance-Time_Sec',\n    'FGC-Season', 'FGC-FGC_CU', 'FGC-FGC_CU_Zone', 'FGC-FGC_GSND', 'FGC-FGC_GSND_Zone', 'FGC-FGC_GSD', 'FGC-FGC_GSD_Zone',\n    'FGC-FGC_PU', 'FGC-FGC_PU_Zone', 'FGC-FGC_SRL', 'FGC-FGC_SRL_Zone', 'FGC-FGC_SRR', 'FGC-FGC_SRR_Zone', 'FGC-FGC_TL', 'FGC-FGC_TL_Zone'\n]\ndf_fit = pd.concat([df_physical, df_train[fit_columns]], axis=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:28.658058Z","iopub.execute_input":"2024-11-29T08:36:28.658527Z","iopub.status.idle":"2024-11-29T08:36:28.668783Z","shell.execute_reply.started":"2024-11-29T08:36:28.658474Z","shell.execute_reply":"2024-11-29T08:36:28.667536Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_fit[fit_columns].isna().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:28.670286Z","iopub.execute_input":"2024-11-29T08:36:28.670674Z","iopub.status.idle":"2024-11-29T08:36:28.686681Z","shell.execute_reply.started":"2024-11-29T08:36:28.670638Z","shell.execute_reply":"2024-11-29T08:36:28.685569Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The fitness-related columns have varying numbers of missing values, and it looks like some records are missing data even though the individuals should have participated in the tests.","metadata":{}},{"cell_type":"code","source":"value_counts = df_fit['Fitness_Endurance-Season'].value_counts()\ntmp = df_fit[~df_fit['Fitness_Endurance-Season'].isna()]['Basic_Demos-Age']\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(10, 6))\nax1.pie(value_counts, labels=value_counts.index, startangle=90, autopct='%1.1f%%')\nax1.set_title('Fitness Endurance Distribution by Season')\nax2.hist(tmp, edgecolor='white', bins=12)\nax2.set_xticks(range(int(tmp.min()), int(tmp.max()) + 1))\nax2.set_xlabel('Age')\nax2.set_ylabel('Count')\nax2.grid(axis='y', linestyle='--', alpha=0.7);\nax2.set_title('Age Distribution');","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:28.692178Z","iopub.execute_input":"2024-11-29T08:36:28.692531Z","iopub.status.idle":"2024-11-29T08:36:29.116570Z","shell.execute_reply.started":"2024-11-29T08:36:28.692496Z","shell.execute_reply":"2024-11-29T08:36:29.115454Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There doesn't seem to be much difference in participation across seasons. The majority of participants are between the ages of 5 and 12, with a significant drop-off after age 13, and a maximum age of 16.","metadata":{}},{"cell_type":"code","source":"(\n    df_fit\n    .groupby('Fitness_Endurance-Max_Stage')[['Fitness_Endurance-Time_Mins', 'Fitness_Endurance-Time_Sec']]\n    .agg({'min', 'max'})\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:29.118222Z","iopub.execute_input":"2024-11-29T08:36:29.118673Z","iopub.status.idle":"2024-11-29T08:36:29.143460Z","shell.execute_reply.started":"2024-11-29T08:36:29.118621Z","shell.execute_reply":"2024-11-29T08:36:29.142250Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"For the fitness endurance test, I expected that a longer time `Fitness_Endurance-Time_Mins` would correspond to a higher max stage reached `Fitness_Endurance-Max_Stage` , but there are some instances where this relationship is reversed. I wonder if there's a different threshold depending on age. Also, it's concerning that the stage number jumps suddenly from 12 to 26. Are these values accurate?","metadata":{}},{"cell_type":"code","source":"fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(10, 6))\nsns.boxplot(data=df_fit, x='Basic_Demos-Age', y='Fitness_Endurance-Max_Stage', ax=ax1)\nax1.set_xlim(-0.5, 8.5)\nax1.grid(True, linestyle='--', alpha=0.7)\nax1.set_title('Average Max_Stage Distribution by Age')\nax1.set_xlabel('Age')\nax1.set_ylabel('Average Max_Stage')\nsns.boxplot(data=df_fit, x='sii', y='Fitness_Endurance-Max_Stage', ax=ax2)\nax2.grid(True, linestyle='--', alpha=0.7)\nax2.set_title('Average Max_Stage Distribution by sii')\nax2.set_xlabel('sii')\nax2.set_ylabel('Average Max_Stage');","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:29.144947Z","iopub.execute_input":"2024-11-29T08:36:29.145337Z","iopub.status.idle":"2024-11-29T08:36:29.667684Z","shell.execute_reply.started":"2024-11-29T08:36:29.145299Z","shell.execute_reply":"2024-11-29T08:36:29.666483Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"matrix = pd.crosstab(df_fit['Fitness_Endurance-Max_Stage'], df_fit['sii'])\nmatrix_percentage = matrix / matrix.sum().sum() * 100\nmatrix_combined = matrix.astype(str) + \" (\" + round(matrix_percentage, 1).astype(str) + \"%)\"\nmatrix_combined","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-29T08:36:29.669306Z","iopub.execute_input":"2024-11-29T08:36:29.669775Z","iopub.status.idle":"2024-11-29T08:36:29.695519Z","shell.execute_reply.started":"2024-11-29T08:36:29.669718Z","shell.execute_reply":"2024-11-29T08:36:29.694345Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There were only two data points where `Fitness_Endurance-Max_Stage` was 13 or higher. There are so many missing values and so few data points in the Fitness_Endurance data, it's probably going to be difficult to find any trends or correlations with age or the target variable `sii` .","metadata":{}}]}