{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30775,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-09-30T22:01:15.242664Z","iopub.execute_input":"2024-09-30T22:01:15.243196Z","iopub.status.idle":"2024-09-30T22:01:18.688663Z","shell.execute_reply.started":"2024-09-30T22:01:15.243149Z","shell.execute_reply":"2024-09-30T22:01:18.687126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\ntrain_df = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv') \ntest_df = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv')","metadata":{"execution":{"iopub.status.busy":"2024-09-30T22:12:54.520599Z","iopub.execute_input":"2024-09-30T22:12:54.521020Z","iopub.status.idle":"2024-09-30T22:12:54.537861Z","shell.execute_reply.started":"2024-09-30T22:12:54.520980Z","shell.execute_reply":"2024-09-30T22:12:54.536565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Train data')\nprint(train_df.describe())","metadata":{"execution":{"iopub.status.busy":"2024-09-30T22:31:24.294305Z","iopub.execute_input":"2024-09-30T22:31:24.294765Z","iopub.status.idle":"2024-09-30T22:31:24.396889Z","shell.execute_reply.started":"2024-09-30T22:31:24.294725Z","shell.execute_reply":"2024-09-30T22:31:24.395737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. Physical Measurements\n* BMI: The average BMI is approximately 19.84, which is within a healthy range for children, but the standard deviation of 4.93 indicates variability in physical health.\n* Height and Weight: The measurements for height and weight suggest a range of physical development. The maximum weight of 121.6 kg indicates outliers or potential concerns regarding obesity, especially in a young population.\n2. Blood Pressure and Heart Rate\n* Blood Pressure: The average systolic blood pressure is 117.5 mmHg, while diastolic is 70.5 mmHg. While most children are within normal limits, the higher maximums indicate potential outliers that could warrant further investigation.\n* Heart Rate: The average heart rate is 81.67 bpm, which is in the normal range for children. However, variability of 9.31 standard deviation indicates different levels of physical fitness or stress among the children.\n3. Body Composition\n* Measurements such as BIA (Bioelectrical Impedance Analysis) values can provide insight into body composition. Variability in these metrics (e.g., muscle mass vs. fat mass) could reveal underlying health issues or developmental concerns.\n4. Physical Activity Questionnaire\n* The standard deviations across these metrics indicate significant variability, particularly in BIA_LST and BIA_SMM (14.49 and 11.94 standard deviations respectively). The metrics of Lean Soft Tissue and Skeletal Muscle Mass may indicate outliers on fitness level.\n5. Internet Usage\n* The average internet usage of approximately 1.44 hours per day is relatively moderate. The data indicates that some children do not use the internet at all, while others may be at risk of excessive use, with a maximum recorded usage of 3 hours.","metadata":{}},{"cell_type":"code","source":"# Visualization of BMI\nplt.figure(figsize=(10, 6))\nsns.histplot(train_df['Physical-BMI'], bins=15, kde=True, color='blue')\nplt.title('Distribution of BMI')\nplt.xlabel('BMI')\nplt.ylabel('Frequency')\nplt.axvline(train_df['Physical-BMI'].mean(), color='red', linestyle='dashed', linewidth=1, label='Mean BMI')\nplt.legend()\nplt.show()\n\n# Height vs. Weight Scatter Plot\nplt.figure(figsize=(10, 6))\nsns.scatterplot(x='Physical-Height', y='Physical-Weight', data=train_df, alpha=0.5)\nplt.title('Height vs. Weight')\nplt.xlabel('Height (cm)')\nplt.ylabel('Weight (kg)')\nplt.axhline(y=train_df['Physical-Weight'].mean(), color='red', linestyle='dashed', linewidth=1, label='Mean Weight')\nplt.axvline(x=train_df['Physical-Height'].mean(), color='green', linestyle='dashed', linewidth=1, label='Mean Height')\nplt.legend()\nplt.show()\n\n# Blood Pressure Boxplot\nplt.figure(figsize=(10, 6))\nsns.boxplot(data=train_df[['Physical-Systolic_BP', 'Physical-Diastolic_BP']])\nplt.title('Blood Pressure Levels')\nplt.ylabel('Blood Pressure (mmHg)')\nplt.xticks(ticks=[0, 1], labels=['Systolic', 'Diastolic'])\nplt.show()\n\n# Heart Rate Distribution\nplt.figure(figsize=(10, 6))\nsns.histplot(train_df['Physical-HeartRate'], bins=15, kde=True, color='orange')\nplt.title('Distribution of Heart Rate')\nplt.xlabel('Heart Rate (bpm)')\nplt.ylabel('Frequency')\nplt.axvline(train_df['Physical-HeartRate'].mean(), color='red', linestyle='dashed', linewidth=1, label='Mean Heart Rate')\nplt.legend()\nplt.show()\n\n# Internet Usage Histogram\nplt.figure(figsize=(10, 6))\nsns.histplot(train_df['PreInt_EduHx-computerinternet_hoursday'], bins=10, kde=True, color='purple')\nplt.title('Distribution of Internet Usage')\nplt.xlabel('Hours per Day')\nplt.ylabel('Frequency')\nplt.axvline(train_df['PreInt_EduHx-computerinternet_hoursday'].mean(), color='red', linestyle='dashed', linewidth=1, label='Mean Internet Usage')\nplt.legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-30T22:52:36.121257Z","iopub.execute_input":"2024-09-30T22:52:36.121688Z","iopub.status.idle":"2024-09-30T22:52:37.791879Z","shell.execute_reply.started":"2024-09-30T22:52:36.121647Z","shell.execute_reply":"2024-09-30T22:52:37.790790Z"},"trusted":true},"execution_count":null,"outputs":[]}]}