{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Data Overview","metadata":{}},{"cell_type":"markdown","source":"1. **Demographics**\n- **Basic_Demos-Enroll_Season:** Season of enrollment\n- **Basic_Demos-Age :** Age of participant\n- **Basic_Demos-Sex:** Sex of participant{\"0=Male, 1=Female\"}\n\n2. **Children's Global Assessment Scale**\n- **CGAS-Season**: Season of participation\n- **CGAS-CGAS_Score:** Children's Global Assessment Scale Score\n\n3. **Physical Measures**\n\n- **Physical-Season**: Season of participation\n- **Physical-BMI:** Body Mass Index (kg/m^2)\n- **Physical-Height**: Height (in)\n- **Physical-Weight**: Weight (lbs)\n- **Physical-Waist_Circumference**: Waist circumference (in)\n- **Physical-Diastolic_BP**: Diastolic BP (mmHg)\n- **Physical-HeartRate**: Heart rate (beats/min)\n- **Physical-Systolic_BP**: Systolic BP (mmHg)\n\n\n 4. **FitnessGram Vitals and Treadmill**\n\n- **Fitness_Endurance-Season**: Season of participation\n- **Fitness_Endurance-Max_Stage**: Maximum stage reached\n- **Fitness_Endurance-Time_Mins** :Exact time completed : Minutes\n- **Fitness_Endurance-Time_Sec**: Exact time completed: Seconds\n- **FitnessGram Child** \n- **FGC-Season** : Season of participation\n- **FGC-FGC_CU** Curl up total\n- **FGC-FGC_CU_Zone:** Curl up fitness zone {\"0,1\" -> \"0=Needs Improvement, 1=Healthy Fitness Zone\"}\n- **FGC-FGC_GSND:** Grip Strength total (non-dominant)\n- **FGC-FGC_GSND_Zone**: Grip Strength fitness zone (non-dominant) {\"1,2,3\"\t-> \"1=Weak, 2=Normal, 3=Strong\"}\n- **FGC-FGC_GSD:** Grip Strength total (dominant)\n- **FGC-FGC_GSD_Zone:** Grip Strength fitness zone (dominant) {\"1,2,3\" -> 1=Weak, 2=Normal, 3=Strong\"} \n- **FGC-FGC_PU:** Push-up total\n- **FGC-FGC_PU_Zone:** Push-up fitness zone {\"0,1\" -> \"0=Needs Improvement, 1=Healthy Fitness Zone\"} \n\n- **FGC-FGC_SRL:** Sit & Reach total (left side)\n- **FGC-FGC_SRL_Zone:** Sit & Reach fitness zone (left side) {\"0,1\"\t\"0=Needs Improvement, 1=Healthy Fitness Zone\"} \n- **FGC-FGC_SRR:** Sit & Reach total (right side)\n- **FGC-FGC_SRR_Zone:** Sit & Reach fitness zone (right side) {\"0,1\"\t\"0=Needs Improvement, 1=Healthy Fitness Zone\"} \n- **FGC-FGC_TL:** Trunk lift total\n- **FGC-FGC_TL_Zone:** Trunk lift fitness zone {\"0,1\"\t\"0=Needs Improvement, 1=Healthy Fitness Zone\"}\n\n\n\n5. **Bio-electric Impedance Analysis**\n\n- **BIA-Season**: Season of participation\n- **BIA-BIA_Activity_Level_num:** Activity Level  \"1,2,3,4,5\"\t\"1=Very Light, 2=Light, 3=Moderate, 4=Heavy, 5=Exceptional\"\n- **BIA-BIA_BMC:** Bone Mineral Content\n- **BIA-BIA_BMI:** Body Mass Index\n- **BIA-BIA_BMR:** Basal Metabolic Rate\n- **BIA-BIA_DEE:** Daily Energy Expenditure\n- **BIA-BIA_ECW:** Extracellular Water\n- **BIA-BIA_FFM:** Fat Free Mass\n- **BIA-BIA_FFMI:** Fat Free Mass Index\n- **BIA-BIA_FMI:** Fat Mass Index\n- **BIA-BIA_Fat:** Body Fat Percentage\n- **BIA-BIA_Frame_num:** Body Frame {\"1,2,3\"\t\"1=Small, 2=Medium, 3=Large\"}\n- **BIA-BIA_ICW:** Intracellular Water\n- **BIA-BIA_LDM:** Lean Dry Mass\n- **BIA-BIA_LST:** Lean Soft Tissue\n- **BIA-BIA_SMM:** Skeletal Muscle Mass\n- **BIA-BIA_TBW:** Total Body Water\n\n\n\n6. **Physical Activity Questionnaire**\n- **PAQ_A-Season**: Season of participation\tstr\t\"Spring, Summer, Fall, Winter\"\n- **PAQ_A-PAQ_A_Total:** Activity Summary Score (Adolescents)\n- **PAQ_C-PAQ_C_Total:** Activity Summary Score (Children)\n\n\n7. **Parent-Child Internet Addiction Test**\n\n- **PCIAT-Season:** Season of participation\n- **PCIAT-PCIAT_01:** How often does your child disobey time limits you set for online use?  *{ \"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"}*\n- **PCIAT-PCIAT_02:** How often does your child neglect household chores to spend more time online? *{ \"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"}*\n- **PCIAT-PCIAT_03** : How often does your child prefer to spend time online rather than with the rest of your family? {*\"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"}*\n- **PCIAT-PCIAT_04** :How often does your child form new relationships with fellow online users? *{\"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"}*\n- **PCIAT-PCIAT_05:** How often do you complain about the amount of time your child spends online? *{\"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"}*\n- **PCIAT-PCIAT_06:** How often do your child's grades suffer because of the amount of time he or she spends online? *{ \"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"}*\n\n- **PCIAT-PCIAT_07:** How often does your child check his or her e-mail before doing something else? *{\"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"}*\n\n- **PCIAT-PCIAT_08:** How often does your child seem withdrawn from others since discovering the Internet? *{\"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"}*\n\n- *PCIAT-PCIAT_09\tHow often does your child become defensive or secretive when asked what he or she does online?\tcategorical int\t\"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"}*\n\n- **PCIAT-PCIAT_10:** How often have you caught your child sneaking online against your wishes?\t*\"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"*\n\n- **PCIAT-PCIAT_11:** How often does your child spend time along in his or her room playing on the computer? *\"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"*\n\n- **PCIAT-PCIAT_12:** \"How often does your child receive strange phone calls from new \"\"online\"\" friends?\"  *\"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"*\n\n- **PCIAT-PCIAT_13:** \"How often does your child snap, yell, or act annoyed if bothered while online?\"  *\"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"* \n\n- **PCIAT-PCIAT_14:** How often does your child seem more tired and fatigued than he or she did before the Internet came along?\t*\"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"*\n\n- **PCIAT-PCIAT_15:** How often does your child seem preoccupied with being back online when off-line? *\"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"*\n\n- **PCIAT-PCIAT_16:** How often does your child throw tantrums with your interference about how long he or she spends online?\t*\"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"*\n\n- **PCIAT-PCIAT_17:** How often does your child choose to spend time online rather than doing once enjoyed hobbies and/or outside interests? *\"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"*\n\n- **PCIAT-PCIAT_18:** How often does your child become angry or belligerent when your place time limits on how much time he or shes is allowed to spend online? *\"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"*\n\n- **PCIAT-PCIAT_19:** How often does your child choose to spend more time online than going out with friends? *\"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"* \n\n- **PCIAT-PCIAT_20:** \"How often does your child feel depressed, moody, or nervous when off-line which seems to go away once back online?\" *\"0,1,2,3,4,5\"\t\"0=Does Not Apply, 1=Rarely, 2=Occasionally, 3=Frequently, 4=Often, 5=Always\"*\n\n- **PCIAT-PCIAT_Total:** Total Score *Severity Impairment Index: 0-30=None; 31-49=Mild; 50-79=Moderate; 80-100=Severe*\n\n\n\n8. **Sleep Disturbance Scale**\n- **SDS-Season:** Season of participation \n- **SDS-SDS_Total_Raw:** Total Raw Score\n- **SDS-SDS_Total_T:** Total T-Score\n\n9. **Internet Use**\n- **PreInt_EduHx-Season:** Season of participation\n- **PreInt_EduHx-computerinternet_hoursday:** Hours of using computer/internet *\"0,1,2,3\"\t\"0=Less than 1h/day, 1=Around 1h/day, 2=Around 2hs/day, 3=More than 3hs/day\"*\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"# Data Exploration","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport warnings\nimport math\n\nfrom plotly.subplots import make_subplots\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.ensemble import AdaBoostClassifier\nfrom sklearn.ensemble import GradientBoostingClassifier\nimport xgboost as xgb\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import classification_report\nfrom sklearn.preprocessing import  LabelEncoder ,StandardScaler, OneHotEncoder,RobustScaler\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.metrics import confusion_matrix , accuracy_score , recall_score , f1_score , classification_report\nfrom collections import Counter\nfrom sklearn.svm import SVC\n\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:22.861735Z","iopub.execute_input":"2024-10-11T08:02:22.862172Z","iopub.status.idle":"2024-10-11T08:02:23.737442Z","shell.execute_reply.started":"2024-10-11T08:02:22.862131Z","shell.execute_reply":"2024-10-11T08:02:23.736211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data= pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/train.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:23.739464Z","iopub.execute_input":"2024-10-11T08:02:23.739866Z","iopub.status.idle":"2024-10-11T08:02:23.818148Z","shell.execute_reply.started":"2024-10-11T08:02:23.739824Z","shell.execute_reply":"2024-10-11T08:02:23.817137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.set_option('display.max_columns', None)\npd.set_option('display.max_rows', None)","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:23.819554Z","iopub.execute_input":"2024-10-11T08:02:23.820044Z","iopub.status.idle":"2024-10-11T08:02:23.826806Z","shell.execute_reply.started":"2024-10-11T08:02:23.819991Z","shell.execute_reply":"2024-10-11T08:02:23.825661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.shape","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:23.829291Z","iopub.execute_input":"2024-10-11T08:02:23.829796Z","iopub.status.idle":"2024-10-11T08:02:23.842712Z","shell.execute_reply.started":"2024-10-11T08:02:23.829740Z","shell.execute_reply":"2024-10-11T08:02:23.841459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.head()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:23.979287Z","iopub.execute_input":"2024-10-11T08:02:23.979769Z","iopub.status.idle":"2024-10-11T08:02:24.084535Z","shell.execute_reply.started":"2024-10-11T08:02:23.979722Z","shell.execute_reply":"2024-10-11T08:02:24.083435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.info()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:24.086409Z","iopub.execute_input":"2024-10-11T08:02:24.086775Z","iopub.status.idle":"2024-10-11T08:02:24.130800Z","shell.execute_reply.started":"2024-10-11T08:02:24.086734Z","shell.execute_reply":"2024-10-11T08:02:24.129415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = data.drop('id', axis=1)\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:24.239817Z","iopub.execute_input":"2024-10-11T08:02:24.240262Z","iopub.status.idle":"2024-10-11T08:02:24.250593Z","shell.execute_reply.started":"2024-10-11T08:02:24.240220Z","shell.execute_reply":"2024-10-11T08:02:24.249357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data['sii'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:24.252655Z","iopub.execute_input":"2024-10-11T08:02:24.253050Z","iopub.status.idle":"2024-10-11T08:02:24.269971Z","shell.execute_reply.started":"2024-10-11T08:02:24.253007Z","shell.execute_reply":"2024-10-11T08:02:24.268710Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.duplicated().sum()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:24.966885Z","iopub.execute_input":"2024-10-11T08:02:24.967338Z","iopub.status.idle":"2024-10-11T08:02:25.004094Z","shell.execute_reply.started":"2024-10-11T08:02:24.967294Z","shell.execute_reply":"2024-10-11T08:02:25.002359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:25.006213Z","iopub.execute_input":"2024-10-11T08:02:25.006603Z","iopub.status.idle":"2024-10-11T08:02:25.025645Z","shell.execute_reply.started":"2024-10-11T08:02:25.006561Z","shell.execute_reply":"2024-10-11T08:02:25.024027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Cleaning","metadata":{}},{"cell_type":"code","source":"data_clean= data.copy()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:25.535402Z","iopub.execute_input":"2024-10-11T08:02:25.536376Z","iopub.status.idle":"2024-10-11T08:02:25.543397Z","shell.execute_reply.started":"2024-10-11T08:02:25.536316Z","shell.execute_reply":"2024-10-11T08:02:25.541890Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_clean.duplicated().sum()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:25.545832Z","iopub.execute_input":"2024-10-11T08:02:25.546336Z","iopub.status.idle":"2024-10-11T08:02:25.585887Z","shell.execute_reply.started":"2024-10-11T08:02:25.546275Z","shell.execute_reply":"2024-10-11T08:02:25.584540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_clean.drop_duplicates(inplace=True)","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:25.588191Z","iopub.execute_input":"2024-10-11T08:02:25.589546Z","iopub.status.idle":"2024-10-11T08:02:25.627778Z","shell.execute_reply.started":"2024-10-11T08:02:25.589486Z","shell.execute_reply":"2024-10-11T08:02:25.626442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_clean.duplicated().sum()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:25.629110Z","iopub.execute_input":"2024-10-11T08:02:25.629464Z","iopub.status.idle":"2024-10-11T08:02:25.665837Z","shell.execute_reply.started":"2024-10-11T08:02:25.629426Z","shell.execute_reply":"2024-10-11T08:02:25.664502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_cols = [col for col in data.columns if data[col].dtype != 'O']\ncat_cols = [col for col in data.columns if col not in num_cols]\nprint(f'Numerical columns: {num_cols}')\nprint(\"-------------------\")\nprint(f'Categorical columns: {cat_cols}')\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:25.669205Z","iopub.execute_input":"2024-10-11T08:02:25.669802Z","iopub.status.idle":"2024-10-11T08:02:25.683310Z","shell.execute_reply.started":"2024-10-11T08:02:25.669699Z","shell.execute_reply":"2024-10-11T08:02:25.681853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_clean = data_clean.drop(['CGAS-Season', 'Physical-Season', 'Fitness_Endurance-Season', 'FGC-Season', 'BIA-Season', 'PAQ_A-Season', 'PAQ_C-Season', 'PCIAT-Season', 'SDS-Season', 'PreInt_EduHx-Season'], axis=1)\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:25.684921Z","iopub.execute_input":"2024-10-11T08:02:25.685413Z","iopub.status.idle":"2024-10-11T08:02:25.696035Z","shell.execute_reply.started":"2024-10-11T08:02:25.685356Z","shell.execute_reply":"2024-10-11T08:02:25.694769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_clean.head()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:25.697692Z","iopub.execute_input":"2024-10-11T08:02:25.698180Z","iopub.status.idle":"2024-10-11T08:02:25.796991Z","shell.execute_reply.started":"2024-10-11T08:02:25.698125Z","shell.execute_reply":"2024-10-11T08:02:25.795711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_clean.shape","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:25.798662Z","iopub.execute_input":"2024-10-11T08:02:25.799143Z","iopub.status.idle":"2024-10-11T08:02:25.807352Z","shell.execute_reply.started":"2024-10-11T08:02:25.799088Z","shell.execute_reply":"2024-10-11T08:02:25.806037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# One-hot encode the categorical variables\ndata_clean_encoded = pd.get_dummies(data_clean)\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:25.810573Z","iopub.execute_input":"2024-10-11T08:02:25.811110Z","iopub.status.idle":"2024-10-11T08:02:25.823119Z","shell.execute_reply.started":"2024-10-11T08:02:25.811050Z","shell.execute_reply":"2024-10-11T08:02:25.821867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_clean_encoded = data_clean_encoded.drop(['Physical-Waist_Circumference','Fitness_Endurance-Max_Stage','Fitness_Endurance-Time_Mins','Fitness_Endurance-Time_Sec','FGC-FGC_GSND','FGC-FGC_GSND_Zone','FGC-FGC_GSD','FGC-FGC_GSD_Zone','BIA-BIA_Activity_Level_num','BIA-BIA_BMC','BIA-BIA_BMI','BIA-BIA_BMR','BIA-BIA_DEE','BIA-BIA_ECW','BIA-BIA_FFM','BIA-BIA_FFMI','BIA-BIA_Fat','BIA-BIA_Frame_num','BIA-BIA_ICW','BIA-BIA_LDM','BIA-BIA_LST','BIA-BIA_SMM','BIA-BIA_TBW','PAQ_A-PAQ_A_Total','PAQ_C-PAQ_C_Total'], axis=1)","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:25.824926Z","iopub.execute_input":"2024-10-11T08:02:25.825512Z","iopub.status.idle":"2024-10-11T08:02:25.835961Z","shell.execute_reply.started":"2024-10-11T08:02:25.825436Z","shell.execute_reply":"2024-10-11T08:02:25.834678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(data['sii'].dtype)\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:27.269577Z","iopub.execute_input":"2024-10-11T08:02:27.270687Z","iopub.status.idle":"2024-10-11T08:02:27.277323Z","shell.execute_reply.started":"2024-10-11T08:02:27.270594Z","shell.execute_reply":"2024-10-11T08:02:27.276056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_clean_encoded['sii'] = data_clean_encoded['sii'].fillna(4)\nprint(data_clean_encoded['sii'].unique())\ndata_clean_encoded['sii'].replace(4, np.nan, inplace=True)\nprint(data_clean_encoded['sii'].unique())\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:27.279501Z","iopub.execute_input":"2024-10-11T08:02:27.279975Z","iopub.status.idle":"2024-10-11T08:02:27.295868Z","shell.execute_reply.started":"2024-10-11T08:02:27.279929Z","shell.execute_reply":"2024-10-11T08:02:27.294375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.impute import KNNImputer\n\n# Initialize the KNN Imputer\nimputer = KNNImputer(n_neighbors=5)\n\n# Drop the target variable to perform imputation\nfeatures = data_clean_encoded.drop('sii', axis=1)\n\n# Fit the imputer and transform the features\ndata_imputed2 = imputer.fit_transform(features)\n\n# Convert the imputed data back to a DataFrame\ndata_imputed2 = pd.DataFrame(data_imputed2, columns=features.columns)\n\n# If you need to keep the 'sii' column, concatenate it back\ndata_imputed2['sii'] = data_clean_encoded['sii'].reset_index(drop=True)\n\n# Display the resulting DataFrame\ndata_imputed2.head()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:27.306886Z","iopub.execute_input":"2024-10-11T08:02:27.308091Z","iopub.status.idle":"2024-10-11T08:02:31.889094Z","shell.execute_reply.started":"2024-10-11T08:02:27.308026Z","shell.execute_reply":"2024-10-11T08:02:31.888035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.cluster import KMeans\nimport numpy as np\nimport pandas as pd\n\ndef impute_with_kmeans(df, categorical_columns, n_clusters=5):\n    # Fill missing values temporarily with the mode (or any placeholder)\n    df_temp = df.copy()\n    for col in categorical_columns:\n        df_temp[col].fillna(df_temp[col].mode()[0], inplace=True)\n\n    # Perform KMeans clustering after filling missing values\n    kmeans = KMeans(n_clusters=n_clusters, random_state=0)\n    cluster_labels = kmeans.fit_predict(df_temp)\n\n    # Impute missing values within each cluster\n    for col in categorical_columns:\n        for cluster in np.unique(cluster_labels):\n            mask = (cluster_labels == cluster) & df[col].isna()\n            most_frequent = df.loc[cluster_labels == cluster, col].mode()[0]\n            df.loc[mask, col] = most_frequent\n\n    return df\n\n# Apply KMeans-based imputation\ndata_imputed = impute_with_kmeans(data_imputed2, ['sii'])\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:31.891058Z","iopub.execute_input":"2024-10-11T08:02:31.891446Z","iopub.status.idle":"2024-10-11T08:02:33.545190Z","shell.execute_reply.started":"2024-10-11T08:02:31.891405Z","shell.execute_reply":"2024-10-11T08:02:33.544072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_imputed['sii'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:33.546665Z","iopub.execute_input":"2024-10-11T08:02:33.547134Z","iopub.status.idle":"2024-10-11T08:02:33.556638Z","shell.execute_reply.started":"2024-10-11T08:02:33.547079Z","shell.execute_reply":"2024-10-11T08:02:33.555582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#from sklearn.impute import KNNImputer\n\n#imputer = KNNImputer(n_neighbors=5)\n\n#data_imputed = imputer.fit_transform(data_clean_encoded)\n\n#data_imputed = pd.DataFrame(data_imputed, columns=data_clean_encoded.columns)\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:33.559976Z","iopub.execute_input":"2024-10-11T08:02:33.560810Z","iopub.status.idle":"2024-10-11T08:02:33.566218Z","shell.execute_reply.started":"2024-10-11T08:02:33.560764Z","shell.execute_reply":"2024-10-11T08:02:33.565002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_imputed.shape","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:33.567567Z","iopub.execute_input":"2024-10-11T08:02:33.567959Z","iopub.status.idle":"2024-10-11T08:02:33.581580Z","shell.execute_reply.started":"2024-10-11T08:02:33.567918Z","shell.execute_reply":"2024-10-11T08:02:33.580422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_cols = [col for col in data_imputed.columns if data_imputed[col].dtype != 'O']\ncat_cols = [col for col in data_imputed.columns if col not in num_cols]\nprint(f'Numerical columns: {num_cols}')\nprint(\"-------------------------------------------\")\nprint(f'Categorical columns: {cat_cols}')","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:33.583349Z","iopub.execute_input":"2024-10-11T08:02:33.583871Z","iopub.status.idle":"2024-10-11T08:02:33.595793Z","shell.execute_reply.started":"2024-10-11T08:02:33.583814Z","shell.execute_reply":"2024-10-11T08:02:33.594748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_imputed.head()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:33.597151Z","iopub.execute_input":"2024-10-11T08:02:33.597600Z","iopub.status.idle":"2024-10-11T08:02:33.668550Z","shell.execute_reply.started":"2024-10-11T08:02:33.597547Z","shell.execute_reply":"2024-10-11T08:02:33.667434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\ndata_imputed.hist(figsize=(70, 40), bins=30)  # Adjust the figsize and number of bins as needed\nplt.tight_layout()  # Adjust the layout\nplt.show()  # Display the histograms\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:33.670273Z","iopub.execute_input":"2024-10-11T08:02:33.670830Z","iopub.status.idle":"2024-10-11T08:02:49.511781Z","shell.execute_reply.started":"2024-10-11T08:02:33.670776Z","shell.execute_reply":"2024-10-11T08:02:49.510199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nfor column in data_imputed:\n \n \n        plt.figure(figsize=(8, 6))\n       # ax = sns.countplot(x=column, data=data_imputed)\n        sns.histplot(data_imputed[column], kde=True)  # Use kde for smooth curve\n\n        plt.title(f'Count of {column}')\n        plt.xlabel(column)\n        plt.ylabel('Count')\n        \n       \n        plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:02:49.513598Z","iopub.execute_input":"2024-10-11T08:02:49.514033Z","iopub.status.idle":"2024-10-11T08:03:12.562178Z","shell.execute_reply.started":"2024-10-11T08:02:49.513991Z","shell.execute_reply":"2024-10-11T08:03:12.561130Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Handling Outliers","metadata":{}},{"cell_type":"code","source":"def detect_outliers_iqr(df):\n    outlier_indices = []\n\n    for column in df.select_dtypes(include=['float64', 'int64']).columns:\n        Q1 = df[column].quantile(0.25)\n        Q3 = df[column].quantile(0.75)\n        IQR = Q3 - Q1\n\n        lower_bound = Q1 - 1.5 * IQR\n        upper_bound = Q3 + 1.5 * IQR\n\n        outliers = df[(df[column] < lower_bound) | (df[column] > upper_bound)]\n        outlier_indices.extend(outliers.index)\n\n        print(f'Outliers in {column}:', outliers.shape[0])\n\n    return outlier_indices","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:12.566519Z","iopub.execute_input":"2024-10-11T08:03:12.566957Z","iopub.status.idle":"2024-10-11T08:03:12.575306Z","shell.execute_reply.started":"2024-10-11T08:03:12.566914Z","shell.execute_reply":"2024-10-11T08:03:12.573901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Step 1: Define a function to detect outliers using IQR\ndef drop_outliers(df, columns):\n    for column in columns:\n        Q1 = df[column].quantile(0.25)  # First quartile (25th percentile)\n        Q3 = df[column].quantile(0.75)  # Third quartile (75th percentile)\n        IQR = Q3 - Q1  # Interquartile range\n        \n        # Define the outlier bounds\n        lower_bound = Q1 - 1.5 * IQR\n        upper_bound = Q3 + 1.5 * IQR\n        \n        # Filter the dataset to keep only rows within bounds for the current column\n        df = df[(df[column] >= lower_bound) & (df[column] <= upper_bound)]\n    \n    return df\n\n\n\n\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:12.576782Z","iopub.execute_input":"2024-10-11T08:03:12.577155Z","iopub.status.idle":"2024-10-11T08:03:12.592249Z","shell.execute_reply.started":"2024-10-11T08:03:12.577115Z","shell.execute_reply":"2024-10-11T08:03:12.590884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"outlier_indices = detect_outliers_iqr(data_imputed)\nprint(\"Total outliers detected:\", len(set(outlier_indices)))","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:12.593509Z","iopub.execute_input":"2024-10-11T08:03:12.593917Z","iopub.status.idle":"2024-10-11T08:03:12.713746Z","shell.execute_reply.started":"2024-10-11T08:03:12.593876Z","shell.execute_reply":"2024-10-11T08:03:12.712581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nfor column in data_imputed.columns:\n    plt.figure(figsize=(8, 3))  \n    sns.boxplot(data=data_imputed[[column]], orient='h', palette=\"Set2\")\n    \n    plt.title(f'Box Plot for {column} with Outliers', fontsize=16)\n    plt.ylabel(column, fontsize=14)\n    plt.xlabel('Values', fontsize=14)\n    \n    plt.tight_layout()\n    plt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:12.715284Z","iopub.execute_input":"2024-10-11T08:03:12.715741Z","iopub.status.idle":"2024-10-11T08:03:22.102757Z","shell.execute_reply.started":"2024-10-11T08:03:12.715698Z","shell.execute_reply":"2024-10-11T08:03:22.101424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_imputed.shape\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:22.104554Z","iopub.execute_input":"2024-10-11T08:03:22.106041Z","iopub.status.idle":"2024-10-11T08:03:22.114028Z","shell.execute_reply.started":"2024-10-11T08:03:22.105983Z","shell.execute_reply":"2024-10-11T08:03:22.112837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_imputed.info()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:22.115931Z","iopub.execute_input":"2024-10-11T08:03:22.117027Z","iopub.status.idle":"2024-10-11T08:03:22.135054Z","shell.execute_reply.started":"2024-10-11T08:03:22.116971Z","shell.execute_reply":"2024-10-11T08:03:22.133927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"columns_to_check = ['Basic_Demos-Age', 'Physical-Weight', 'Physical-HeartRate', 'Physical-BMI' ,'CGAS-CGAS_Score','Physical-Diastolic_BP' ,'Physical-Systolic_BP' ,'FGC-FGC_CU','FGC-FGC_CU_Zone', \n 'FGC-FGC_PU',            \n 'FGC-FGC_PU_Zone',                       \n 'FGC-FGC_SRL',                            \n 'FGC-FGC_SRL_Zone'  ,                    \n 'FGC-FGC_SRR' ,                          \n 'FGC-FGC_SRR_Zone'  ,                    \n 'FGC-FGC_TL'    ,                       \n 'FGC-FGC_TL_Zone',                     \n 'BIA-BIA_FMI' ]\n\ndata_imputed= drop_outliers(data_imputed, columns_to_check)\n\nprint(data_imputed.shape)\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:22.136500Z","iopub.execute_input":"2024-10-11T08:03:22.136905Z","iopub.status.idle":"2024-10-11T08:03:22.190012Z","shell.execute_reply.started":"2024-10-11T08:03:22.136864Z","shell.execute_reply":"2024-10-11T08:03:22.188832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualization ","metadata":{}},{"cell_type":"code","source":"data_imputed.info()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:22.191661Z","iopub.execute_input":"2024-10-11T08:03:22.192149Z","iopub.status.idle":"2024-10-11T08:03:22.209107Z","shell.execute_reply.started":"2024-10-11T08:03:22.192093Z","shell.execute_reply":"2024-10-11T08:03:22.207721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_imputed.head()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:22.210756Z","iopub.execute_input":"2024-10-11T08:03:22.211164Z","iopub.status.idle":"2024-10-11T08:03:22.276501Z","shell.execute_reply.started":"2024-10-11T08:03:22.211121Z","shell.execute_reply":"2024-10-11T08:03:22.275447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_imputed.describe()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:22.277946Z","iopub.execute_input":"2024-10-11T08:03:22.278411Z","iopub.status.idle":"2024-10-11T08:03:22.436389Z","shell.execute_reply.started":"2024-10-11T08:03:22.278356Z","shell.execute_reply":"2024-10-11T08:03:22.434819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_imputed.shape","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:22.437988Z","iopub.execute_input":"2024-10-11T08:03:22.438369Z","iopub.status.idle":"2024-10-11T08:03:22.446014Z","shell.execute_reply.started":"2024-10-11T08:03:22.438327Z","shell.execute_reply":"2024-10-11T08:03:22.444763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"correlation_matrix = data_imputed.corr()\n\nplt.figure(figsize=(65, 20))\n\nsns.heatmap(correlation_matrix, annot=True, cmap='coolwarm', fmt='.2f', vmin=-1, vmax=1, square=True, cbar_kws={\"shrink\": .8})\n\nplt.title('Correlation Matrix of Features')\nplt.xlabel('Features')\nplt.ylabel('Features')\nplt.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:22.447759Z","iopub.execute_input":"2024-10-11T08:03:22.448244Z","iopub.status.idle":"2024-10-11T08:03:29.008572Z","shell.execute_reply.started":"2024-10-11T08:03:22.448185Z","shell.execute_reply":"2024-10-11T08:03:29.006902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Behavoiural Question ","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Assuming 'data_imputed' is your DataFrame and contains the relevant column\n# Count the frequency of responses for PCIAT-PCIAT_01\nresponse_counts = data['PCIAT-PCIAT_17'].value_counts().sort_index()\n\n# Define the labels corresponding to the values\nlabels = {\n    0: 'Does Not Apply',\n    1: 'Rarely',\n    2: 'Occasionally',\n    3: 'Frequently',\n    4: 'Often',\n    5: 'Always'\n}\n\n# Create a list of formatted labels for the pie chart\nformatted_labels = [f\"{key} = {value}\" for key, value in labels.items()]\n\n# Create a pie chart\nplt.figure(figsize=(10, 8))\nplt.pie(response_counts, labels=formatted_labels, autopct='%1.1f%%', startangle=140, colors=sns.color_palette(\"pastel\"))\nplt.title('How often does your child choose to spend time online rather than doing once enjoyed hobbies and/or outside interests?')\nplt.axis('equal')  # Equal aspect ratio ensures that pie is drawn as a circle.\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:29.010310Z","iopub.execute_input":"2024-10-11T08:03:29.010777Z","iopub.status.idle":"2024-10-11T08:03:29.308322Z","shell.execute_reply.started":"2024-10-11T08:03:29.010727Z","shell.execute_reply":"2024-10-11T08:03:29.306450Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Assuming 'data_imputed' is your DataFrame and contains the relevant column\n# Count the frequency of responses for PCIAT-PCIAT_01\nresponse_counts = data['PCIAT-PCIAT_03'].value_counts().sort_index()\n\n# Define the labels corresponding to the values\nlabels = {\n    0: 'Does Not Apply',\n    1: 'Rarely',\n    2: 'Occasionally',\n    3: 'Frequently',\n    4: 'Often',\n    5: 'Always'\n}\n\n# Create a list of formatted labels for the pie chart\nformatted_labels = [f\"{key} = {value}\" for key, value in labels.items()]\n\n# Create a pie chart\nplt.figure(figsize=(10, 8))\nplt.pie(response_counts, labels=formatted_labels, autopct='%1.1f%%', startangle=140, colors=sns.color_palette(\"pastel\"))\nplt.title('How often does your child prefer to spend time online rather than with the rest of your family?')\nplt.axis('equal')  # Equal aspect ratio ensures that pie is drawn as a circle.\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:29.310711Z","iopub.execute_input":"2024-10-11T08:03:29.311501Z","iopub.status.idle":"2024-10-11T08:03:29.572492Z","shell.execute_reply.started":"2024-10-11T08:03:29.311442Z","shell.execute_reply":"2024-10-11T08:03:29.571303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Assuming 'data_imputed' is your DataFrame and contains the relevant column\n# Count the frequency of responses for PCIAT-PCIAT_01\nresponse_counts = data['PCIAT-PCIAT_02'].value_counts().sort_index()\n\n# Define the labels corresponding to the values\nlabels = {\n    0: 'Does Not Apply',\n    1: 'Rarely',\n    2: 'Occasionally',\n    3: 'Frequently',\n    4: 'Often',\n    5: 'Always'\n}\n\n# Create a list of formatted labels for the pie chart\nformatted_labels = [f\"{key} = {value}\" for key, value in labels.items()]\n\n# Create a pie chart\nplt.figure(figsize=(10, 8))\nplt.pie(response_counts, labels=formatted_labels, autopct='%1.1f%%', startangle=140, colors=sns.color_palette(\"pastel\"))\nplt.title('How often does your child neglect household chores to spend more time online')\nplt.axis('equal')  # Equal aspect ratio ensures that pie is drawn as a circle.\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:29.574042Z","iopub.execute_input":"2024-10-11T08:03:29.574415Z","iopub.status.idle":"2024-10-11T08:03:29.858090Z","shell.execute_reply.started":"2024-10-11T08:03:29.574375Z","shell.execute_reply":"2024-10-11T08:03:29.856666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Assuming 'data_imputed' is your DataFrame and contains the relevant column\n# Count the frequency of responses for PCIAT-PCIAT_01\nresponse_counts = data['PCIAT-PCIAT_18'].value_counts().sort_index()\n\n# Define the labels corresponding to the values\nlabels = {\n    0: 'Does Not Apply',\n    1: 'Rarely',\n    2: 'Occasionally',\n    3: 'Frequently',\n    4: 'Often',\n    5: 'Always'\n}\n\n# Create a list of formatted labels for the pie chart\nformatted_labels = [f\"{key} = {value}\" for key, value in labels.items()]\n\n# Create a pie chart\nplt.figure(figsize=(10, 8))\nplt.pie(response_counts, labels=formatted_labels, autopct='%1.1f%%', startangle=140, colors=sns.color_palette(\"pastel\"))\nplt.title('How often does your child become angry or belligerent when your place time limits on how much time he or shes is allowed to spend online')\nplt.axis('equal')  # Equal aspect ratio ensures that pie is drawn as a circle.\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:29.859964Z","iopub.execute_input":"2024-10-11T08:03:29.861034Z","iopub.status.idle":"2024-10-11T08:03:30.163022Z","shell.execute_reply.started":"2024-10-11T08:03:29.860974Z","shell.execute_reply":"2024-10-11T08:03:30.161827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Assuming 'data_imputed' is your DataFrame and contains the relevant column\n# Count the frequency of responses for PCIAT-PCIAT_01\nresponse_counts = data['PCIAT-PCIAT_01'].value_counts().sort_index()\n\n# Define the labels corresponding to the values\nlabels = {\n    0: 'Does Not Apply',\n    1: 'Rarely',\n    2: 'Occasionally',\n    3: 'Frequently',\n    4: 'Often',\n    5: 'Always'\n}\n\n# Create a list of formatted labels for the pie chart\nformatted_labels = [f\"{key} = {value}\" for key, value in labels.items()]\n\n# Create a pie chart\nplt.figure(figsize=(10, 8))\nplt.pie(response_counts, labels=formatted_labels, autopct='%1.1f%%', startangle=140, colors=sns.color_palette(\"pastel\"))\nplt.title('How often does your child disobey time limits you set for online use')\nplt.axis('equal')  # Equal aspect ratio ensures that pie is drawn as a circle.\nplt.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:30.164686Z","iopub.execute_input":"2024-10-11T08:03:30.165169Z","iopub.status.idle":"2024-10-11T08:03:30.418452Z","shell.execute_reply.started":"2024-10-11T08:03:30.165114Z","shell.execute_reply":"2024-10-11T08:03:30.417164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Box plot for PCIAT-PCIAT_01\nplt.figure(figsize=(12, 6))\nsns.boxplot(x='PCIAT-PCIAT_01', y='Basic_Demos-Age', data=data, palette='pastel')\nplt.title('Age Distribution for Q1: How often does your child disobey time limits you set for online use')\nplt.xlabel('Response')\nplt.ylabel('Age')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:30.420608Z","iopub.execute_input":"2024-10-11T08:03:30.421142Z","iopub.status.idle":"2024-10-11T08:03:30.797677Z","shell.execute_reply.started":"2024-10-11T08:03:30.421086Z","shell.execute_reply":"2024-10-11T08:03:30.796467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\n\nplt.figure(figsize=(12, 6))\nsns.boxplot(x='PCIAT-PCIAT_06', y='Basic_Demos-Age', data=data, palette='pastel')\nplt.title('Age Distribution for Q: How often do your childs grades suffer because of the amount of time he or she spends online?')\nplt.xlabel('Response')\nplt.ylabel('Age')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:30.807547Z","iopub.execute_input":"2024-10-11T08:03:30.808325Z","iopub.status.idle":"2024-10-11T08:03:31.192502Z","shell.execute_reply.started":"2024-10-11T08:03:30.808280Z","shell.execute_reply":"2024-10-11T08:03:31.191365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\n# List of questions to analyze\nquestions = [\n    'PCIAT-PCIAT_01', 'PCIAT-PCIAT_02', 'PCIAT-PCIAT_03', \n    'PCIAT-PCIAT_04', 'PCIAT-PCIAT_05', 'PCIAT-PCIAT_06',\n    'PCIAT-PCIAT_07', 'PCIAT-PCIAT_08', 'PCIAT-PCIAT_09',\n    'PCIAT-PCIAT_10', 'PCIAT-PCIAT_11', 'PCIAT-PCIAT_12',\n    'PCIAT-PCIAT_13', 'PCIAT-PCIAT_14', 'PCIAT-PCIAT_15',\n    'PCIAT-PCIAT_16', 'PCIAT-PCIAT_17', 'PCIAT-PCIAT_18',\n    'PCIAT-PCIAT_19', 'PCIAT-PCIAT_20'\n]\n\nplt.figure(figsize=(15, 40))  # Adjust the size for better visibility\n\n# Create a box plot for each question\nfor i, question in enumerate(questions):\n    plt.subplot(len(questions), 1, i + 1)  # Create subplots\n    sns.boxplot(x=question, y='Basic_Demos-Age', data=data, palette='pastel')\n    plt.title(f'Age Distribution for {question}')\n    plt.xlabel('Response')\n    plt.ylabel('Age')\n\nplt.tight_layout()  # Adjust layout to prevent overlap\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:31.194069Z","iopub.execute_input":"2024-10-11T08:03:31.194446Z","iopub.status.idle":"2024-10-11T08:03:37.115459Z","shell.execute_reply.started":"2024-10-11T08:03:31.194399Z","shell.execute_reply":"2024-10-11T08:03:37.114147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n\nmean_values = data.groupby('PreInt_EduHx-computerinternet_hoursday')['PAQ_C-PAQ_C_Total'].mean().reset_index()\n\n# Create a line plot\nplt.figure(figsize=(10, 8))\nsns.lineplot(x='PreInt_EduHx-computerinternet_hoursday', y='PAQ_C-PAQ_C_Total', data=mean_values, marker='o')\n\n# Customize the plot\nplt.title('Average Fitness Endurance Max Stage vs Total Internet Usage')\nplt.xlabel('Total Internet Usage (hours/day)')\nplt.ylabel('Average Fitness Endurance Max Stage')\nplt.grid(True)\n\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:37.117080Z","iopub.execute_input":"2024-10-11T08:03:37.117515Z","iopub.status.idle":"2024-10-11T08:03:37.497143Z","shell.execute_reply.started":"2024-10-11T08:03:37.117468Z","shell.execute_reply":"2024-10-11T08:03:37.495925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## sleep distrubtion","metadata":{}},{"cell_type":"code","source":"# Calculate the mean sleep time for each internet usage category\nmean_sleep_time = data_imputed.groupby('PreInt_EduHx-computerinternet_hoursday')['SDS-SDS_Total_T'].mean()\n\n# Create a line plot\nplt.figure(figsize=(10, 8))\nplt.plot(mean_sleep_time.index, mean_sleep_time.values, marker='o', linestyle='-', color='blue')\n\n# Customize the plot\nplt.xlabel('Internet Usage Time (hours/day)')\nplt.ylabel('Average Sleep Time (hours)')\nplt.title('Average Sleep Time vs Internet Usage Time')\n\n# Set custom y-ticks with different intervals\ncustom_ticks = [ 45, 50, 55, 60, 65, 70,75]  # Define your custom tick values\nplt.yticks(custom_ticks)\n\n# Add grid lines for better visibility\nplt.grid(True)\n\n# Show the plot\nplt.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:37.498727Z","iopub.execute_input":"2024-10-11T08:03:37.499108Z","iopub.status.idle":"2024-10-11T08:03:37.821781Z","shell.execute_reply.started":"2024-10-11T08:03:37.499068Z","shell.execute_reply":"2024-10-11T08:03:37.820700Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10, 8))\nsns.histplot(data_imputed['sii'], bins=30, kde=True)  \nplt.title('Distribution of SII Feature')\nplt.xlabel('SII')\nplt.ylabel('Frequency')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:37.823444Z","iopub.execute_input":"2024-10-11T08:03:37.823959Z","iopub.status.idle":"2024-10-11T08:03:38.262375Z","shell.execute_reply.started":"2024-10-11T08:03:37.823903Z","shell.execute_reply":"2024-10-11T08:03:38.261091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- *0 for None, 1 for Mild, 2 for Moderate, and 3 for Severe , 0 for unknown*","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10, 8))\nsns.histplot(data_imputed['PCIAT-PCIAT_Total'], bins=30, kde=True)  \nplt.title('Distribution of internet activity')\nplt.xlabel('internet activity')\nplt.ylabel('Frequency')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:38.264110Z","iopub.execute_input":"2024-10-11T08:03:38.264535Z","iopub.status.idle":"2024-10-11T08:03:38.674167Z","shell.execute_reply.started":"2024-10-11T08:03:38.264490Z","shell.execute_reply":"2024-10-11T08:03:38.673065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10, 8))\nsns.histplot(data['PreInt_EduHx-computerinternet_hoursday'], bins=30, kde=True)  \nplt.title('Distribution of computerinternet_hoursday')\nplt.xlabel('computerinternet_hoursday')\nplt.ylabel('Frequency')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:38.675706Z","iopub.execute_input":"2024-10-11T08:03:38.676184Z","iopub.status.idle":"2024-10-11T08:03:39.113473Z","shell.execute_reply.started":"2024-10-11T08:03:38.676131Z","shell.execute_reply":"2024-10-11T08:03:39.112348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Season distribution ","metadata":{}},{"cell_type":"code","source":"# Step 1: Calculate the total count for each season\nseason_counts = {\n    'Fall': data_imputed['Basic_Demos-Enroll_Season_Fall'].sum(),\n    'Spring': data_imputed['Basic_Demos-Enroll_Season_Spring'].sum(),\n    'Summer': data_imputed['Basic_Demos-Enroll_Season_Summer'].sum(),\n    'Winter': data_imputed['Basic_Demos-Enroll_Season_Winter'].sum()\n}\n\n# Step 2: Convert the dictionary to a DataFrame for easier plotting\nseason_counts_df = pd.DataFrame(list(season_counts.items()), columns=['Season', 'Count'])\n\n# Step 3: Create the bar plot\nplt.figure(figsize=(8, 5))\nsns.barplot(data=season_counts_df, x='Season', y='Count', palette='pastel')\n\n# Add labels and title\nplt.xlabel('Season', fontsize=12)\nplt.ylabel('Count', fontsize=12)\nplt.title('Enrollment Counts by Season', fontsize=14)\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:39.115020Z","iopub.execute_input":"2024-10-11T08:03:39.115404Z","iopub.status.idle":"2024-10-11T08:03:39.460631Z","shell.execute_reply.started":"2024-10-11T08:03:39.115364Z","shell.execute_reply":"2024-10-11T08:03:39.459372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## PCIAT_Total VS SII","metadata":{}},{"cell_type":"code","source":"correlation = data_imputed['PCIAT-PCIAT_Total'].corr(data_imputed['sii'])\nprint(f'Correlation between PCIAT_Total and sii: {correlation}')\n\nplt.figure(figsize=(10, 6))\nsns.scatterplot(x='PCIAT-PCIAT_Total', y='sii', data=data_imputed)\n\nplt.title(f'Scatter plot of PCIAT_Total vs. sii\\nCorrelation: {correlation:.2f}', fontsize=16)\nplt.xlabel('PCIAT_Total (Internet Use)')\nplt.ylabel('SII (Self-injurious Ideation)')\n\n# Display the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:39.462168Z","iopub.execute_input":"2024-10-11T08:03:39.462552Z","iopub.status.idle":"2024-10-11T08:03:39.798952Z","shell.execute_reply.started":"2024-10-11T08:03:39.462511Z","shell.execute_reply":"2024-10-11T08:03:39.797809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10, 6))\nsns.boxplot(x='sii', y='PCIAT-PCIAT_Total', data=data_imputed)\n\n# Add a title\nplt.title('Box plot of PCIAT_Total by SII Levels', fontsize=16)\nplt.xlabel('SII (Severity impairment index)')\nplt.ylabel('PCIAT_Total (Internet Use)')\n\n# Display the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:39.800817Z","iopub.execute_input":"2024-10-11T08:03:39.801300Z","iopub.status.idle":"2024-10-11T08:03:40.115247Z","shell.execute_reply.started":"2024-10-11T08:03:39.801246Z","shell.execute_reply":"2024-10-11T08:03:40.114070Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Pysical Measures and SII","metadata":{}},{"cell_type":"markdown","source":"## What is the average BMI, heart rate, or systolic/diastolic blood pressure?","metadata":{}},{"cell_type":"code","source":"averages = {\n    'BMI': data_imputed['Physical-BMI'].mean(),\n    'Heart Rate': data_imputed['Physical-HeartRate'].mean(),\n    'Systolic BP': data_imputed['Physical-Systolic_BP'].mean(),\n    'Diastolic BP': data_imputed['Physical-Diastolic_BP'].mean()\n}\n\n# Convert to DataFrame for easier plotting\naverages_df = pd.DataFrame(list(averages.items()), columns=['Feature', 'Average'])\n\n# Step 2: Create a bar plot\nplt.figure(figsize=(10, 6))\nsns.barplot(x='Feature', y='Average', data=averages_df, palette='coolwarm')\nplt.title('Average Values of BMI, Heart Rate, and Blood Pressure')\nplt.ylabel('Average')\nplt.xlabel('Features')\nplt.ylim(0, averages_df['Average'].max() + 10)  # Adjusting y-axis for better visibility\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:40.116877Z","iopub.execute_input":"2024-10-11T08:03:40.117360Z","iopub.status.idle":"2024-10-11T08:03:40.414739Z","shell.execute_reply.started":"2024-10-11T08:03:40.117302Z","shell.execute_reply":"2024-10-11T08:03:40.413425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10, 8))\nsns.histplot(data_imputed['CGAS-CGAS_Score'], bins=30, kde=True)  \nplt.title('Distribution ofChildrens Global Assessment Scale Score')\nplt.xlabel('Assessment score')\nplt.ylabel('Frequency')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:40.416268Z","iopub.execute_input":"2024-10-11T08:03:40.416736Z","iopub.status.idle":"2024-10-11T08:03:40.838198Z","shell.execute_reply.started":"2024-10-11T08:03:40.416693Z","shell.execute_reply":"2024-10-11T08:03:40.837063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Assuming 'data' is your DataFrame\nmean_values = data.groupby('PreInt_EduHx-computerinternet_hoursday')['CGAS-CGAS_Score'].mean().reset_index()\n\n# Create a bar plot\nplt.figure(figsize=(10, 8))\nsns.barplot(x='PreInt_EduHx-computerinternet_hoursday', y='CGAS-CGAS_Score', data=mean_values, palette='viridis')\n\n# Customize the plot\nplt.title('Average Fitness Endurance Max Stage vs Total Internet Usage')\nplt.xlabel('Total Internet Usage (hours/day)')\nplt.ylabel('Average Fitness Endurance Max Stage')\nplt.ylim(0, 100)  # Set y-axis limits from 0 to 100\nplt.grid(axis='y')\n\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:40.839431Z","iopub.execute_input":"2024-10-11T08:03:40.839811Z","iopub.status.idle":"2024-10-11T08:03:41.148789Z","shell.execute_reply.started":"2024-10-11T08:03:40.839770Z","shell.execute_reply":"2024-10-11T08:03:41.147604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Assuming 'data' is your DataFrame\nmean_values = data.groupby('sii')['CGAS-CGAS_Score'].mean().reset_index()\n\n# Create a bar plot\nplt.figure(figsize=(10, 8))\nsns.barplot(x='sii', y='CGAS-CGAS_Score', data=mean_values, palette='viridis')\n\n# Customize the plot\nplt.title('Average Fitness Endurance Max Stage vs Total Internet Usage')\nplt.xlabel('Total Internet Usage (hours/day)')\nplt.ylabel('Average Fitness Endurance Max Stage')\nplt.ylim(0, 100)  # Set y-axis limits from 0 to 100\nplt.grid(axis='y')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:41.150271Z","iopub.execute_input":"2024-10-11T08:03:41.150674Z","iopub.status.idle":"2024-10-11T08:03:41.458827Z","shell.execute_reply.started":"2024-10-11T08:03:41.150632Z","shell.execute_reply":"2024-10-11T08:03:41.457696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Assuming 'data' is your DataFrame\nmean_values = data.groupby('PCIAT-PCIAT_Total')['CGAS-CGAS_Score'].mean().reset_index()\n\n# Create a bar plot\nplt.figure(figsize=(10, 8))\nsns.barplot(x='PCIAT-PCIAT_Total', y='CGAS-CGAS_Score', data=mean_values, palette='viridis')\n\n# Customize the plot\nplt.title('Average Fitness Endurance Max Stage vs Total Internet Usage')\nplt.xlabel('Total Internet Usage (hours/day)')\nplt.ylabel('Average Fitness Endurance Max Stage')\nplt.ylim(0, 100)  # Set y-axis limits from 0 to 100\nplt.grid(axis='y')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:41.460405Z","iopub.execute_input":"2024-10-11T08:03:41.460900Z","iopub.status.idle":"2024-10-11T08:03:42.588645Z","shell.execute_reply.started":"2024-10-11T08:03:41.460845Z","shell.execute_reply":"2024-10-11T08:03:42.587428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Assuming 'data' is your DataFrame\nmean_values = data.groupby('sii')['CGAS-CGAS_Score'].mean().reset_index()\n\n# Create a bar plot\nplt.figure(figsize=(10, 8))\nsns.barplot(x='sii', y='CGAS-CGAS_Score', data=mean_values, palette='viridis')\n\n# Customize the plot\nplt.title('Average Fitness Endurance Max Stage vs Total Internet Usage')\nplt.xlabel('Total Internet Usage (hours/day)')\nplt.ylabel('Average Fitness Endurance Max Stage')\nplt.ylim(0, 100)  # Set y-axis limits from 0 to 100\nplt.grid(axis='y')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:42.590043Z","iopub.execute_input":"2024-10-11T08:03:42.590398Z","iopub.status.idle":"2024-10-11T08:03:42.906975Z","shell.execute_reply.started":"2024-10-11T08:03:42.590360Z","shell.execute_reply":"2024-10-11T08:03:42.905951Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot 1: BMI vs SII Severity\nplt.figure(figsize=(8, 6))\nsns.boxplot(data=data_imputed, x='sii', y='Physical-BMI', palette='pastel')\nplt.title('BMI vs SII Severity')\nplt.xlabel('SII Severity')\nplt.ylabel('BMI')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:42.908328Z","iopub.execute_input":"2024-10-11T08:03:42.908740Z","iopub.status.idle":"2024-10-11T08:03:43.148321Z","shell.execute_reply.started":"2024-10-11T08:03:42.908700Z","shell.execute_reply":"2024-10-11T08:03:43.147295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(8, 6))\nsns.boxplot(data=data_imputed, x='sii', y='Physical-Systolic_BP', palette='pastel')\nplt.title('Systolic Blood Pressure vs SII Severity')\nplt.xlabel('SII Severity')\nplt.ylabel('Systolic BP')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:43.149492Z","iopub.execute_input":"2024-10-11T08:03:43.149857Z","iopub.status.idle":"2024-10-11T08:03:43.457390Z","shell.execute_reply.started":"2024-10-11T08:03:43.149818Z","shell.execute_reply":"2024-10-11T08:03:43.456121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(8, 6))\nsns.boxplot(data=data_imputed, x='sii', y='Physical-Diastolic_BP', palette='pastel')\nplt.title('Diastolic Blood Pressure vs SII Severity')\nplt.xlabel('SII Severity')\nplt.ylabel('Diastolic BP')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:43.459152Z","iopub.execute_input":"2024-10-11T08:03:43.459947Z","iopub.status.idle":"2024-10-11T08:03:43.699893Z","shell.execute_reply.started":"2024-10-11T08:03:43.459891Z","shell.execute_reply":"2024-10-11T08:03:43.698564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(8, 6))\nsns.boxplot(data=data_imputed, x='sii', y='Physical-HeartRate', palette='pastel')\nplt.title('Heart Rate vs SII Severity')\nplt.xlabel('SII Severity')\nplt.ylabel('Heart Rate')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:43.701438Z","iopub.execute_input":"2024-10-11T08:03:43.701956Z","iopub.status.idle":"2024-10-11T08:03:43.964374Z","shell.execute_reply.started":"2024-10-11T08:03:43.701895Z","shell.execute_reply":"2024-10-11T08:03:43.963404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Demographic Influence","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10, 8))\nsns.histplot(data_imputed['Basic_Demos-Age'], bins=30, kde=True)  \nplt.title('Distribution of Age Feature')\nplt.xlabel('AGE')\nplt.ylabel('Frequency')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:43.965936Z","iopub.execute_input":"2024-10-11T08:03:43.966397Z","iopub.status.idle":"2024-10-11T08:03:44.667466Z","shell.execute_reply.started":"2024-10-11T08:03:43.966342Z","shell.execute_reply":"2024-10-11T08:03:44.666387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#import pandas as pd\n#import numpy as np\n#from sklearn.ensemble import RandomForestClassifier\n#from sklearn.feature_selection import RFE\n\n#X = data_imputed.drop('sii', axis=1)  # replace 'target_column' with the name of your target\n#y = data_imputed['sii']\n# Define the model and RFE\n#model = RandomForestClassifier()\n#rfe = RFE(model, n_features_to_select=10)  # You can adjust the number of features to select\n#rfe.fit(X, y)\n#selected_features = rfe.support_  # Returns a boolean array of selected features\n#ranking = rfe.ranking_  # Returns the ranking of all features\n\n# Print out the names of selected features\n#print(X.columns[selected_features])","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:44.669324Z","iopub.execute_input":"2024-10-11T08:03:44.669957Z","iopub.status.idle":"2024-10-11T08:03:44.676079Z","shell.execute_reply.started":"2024-10-11T08:03:44.669897Z","shell.execute_reply":"2024-10-11T08:03:44.674961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Assuming 'data' is your DataFrame with the required columns\n\n# Step 1: Males vs. Females\nplt.figure(figsize=(14, 6))\n\n# Boxplot for BMI\nplt.subplot(1, 2, 1)\nsns.boxplot(x='Basic_Demos-Sex', y='Physical-BMI', data=data_imputed, palette='pastel')\nplt.title('BMI by Gender')\nplt.ylabel('BMI')\n\n# Boxplot for Heart Rate\nplt.subplot(1, 2, 2)\nsns.boxplot(x='Basic_Demos-Sex', y='Physical-HeartRate', data=data_imputed, palette='pastel')\nplt.title('Heart Rate by Gender')\nplt.ylabel('Heart Rate')\n\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:44.677645Z","iopub.execute_input":"2024-10-11T08:03:44.678101Z","iopub.status.idle":"2024-10-11T08:03:45.234207Z","shell.execute_reply.started":"2024-10-11T08:03:44.678056Z","shell.execute_reply":"2024-10-11T08:03:45.233041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Count the number of occurrences for each sex\nGende_counts = data_imputed['Basic_Demos-Sex'].value_counts()\n\n# Create the pie chart\nplt.figure(figsize=(5, 5))\nplt.pie(Gende_counts, labels=Gende_counts.index, autopct='%1.1f%%', startangle=90, colors=sns.color_palette('pastel'))\nplt.title('Distribution of Sex in the Dataset')\nplt.axis('equal')  # Equal aspect ratio ensures that pie chart is circular\nplt.show()\n# 0-> male  1_> female","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:45.235498Z","iopub.execute_input":"2024-10-11T08:03:45.235867Z","iopub.status.idle":"2024-10-11T08:03:45.385763Z","shell.execute_reply.started":"2024-10-11T08:03:45.235827Z","shell.execute_reply":"2024-10-11T08:03:45.383857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10, 6))\nsns.countplot(data=data_imputed, x='Basic_Demos-Sex', hue='sii', palette='pastel')\n\nplt.xlabel('Sex', fontsize=12)\nplt.ylabel('Count', fontsize=12)\nplt.title('Count of SII Categories by Gender', fontsize=14)\nplt.legend(title='SII', loc='upper right', labels=['None (0)', 'Mild (1)', 'Moderate (2)', 'Severe (3)' ,'unknown (4)'])\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:45.388173Z","iopub.execute_input":"2024-10-11T08:03:45.389384Z","iopub.status.idle":"2024-10-11T08:03:45.723474Z","shell.execute_reply.started":"2024-10-11T08:03:45.389306Z","shell.execute_reply":"2024-10-11T08:03:45.722152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(25, 8))\nsns.countplot(data=data_imputed, x='Basic_Demos-Age', hue='sii', palette='pastel')\n\nplt.xlabel('Sex', fontsize=12)\nplt.ylabel('Count', fontsize=12)\nplt.title('Count of SII Categories by Age', fontsize=14)\nplt.legend(title='SII', loc='upper right', labels=['None (0)', 'Mild (1)', 'Moderate (2)', 'Severe (3)' ,'unknown (4)'])\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:45.725211Z","iopub.execute_input":"2024-10-11T08:03:45.725737Z","iopub.status.idle":"2024-10-11T08:03:46.339319Z","shell.execute_reply.started":"2024-10-11T08:03:45.725682Z","shell.execute_reply":"2024-10-11T08:03:46.338209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- **Sever counts increase with age**\n- **most childern from 6-10 is none**","metadata":{}},{"cell_type":"code","source":"def categorize_severity(score):\n    if score <= 30:\n        return 'None'\n    elif score <= 49:\n        return 'Mild'\n    elif score <= 79:\n        return 'Moderate'\n    else:\n        return 'Severe'\n\ndata['Severity'] = data['PCIAT-PCIAT_Total'].apply(categorize_severity)\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:46.340966Z","iopub.execute_input":"2024-10-11T08:03:46.341336Z","iopub.status.idle":"2024-10-11T08:03:46.350656Z","shell.execute_reply.started":"2024-10-11T08:03:46.341295Z","shell.execute_reply":"2024-10-11T08:03:46.349594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\nplt.figure(figsize=(12, 6))\nsns.boxplot(x='Severity', y='Basic_Demos-Age', data=data, palette='pastel')\nplt.title('Age Distribution Across Severity Impairment Index')\nplt.xlabel('Severity Impairment Index')\nplt.ylabel('Age')\nplt.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:46.352052Z","iopub.execute_input":"2024-10-11T08:03:46.352390Z","iopub.status.idle":"2024-10-11T08:03:46.638053Z","shell.execute_reply.started":"2024-10-11T08:03:46.352353Z","shell.execute_reply":"2024-10-11T08:03:46.636937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Calculate the average age for each severity category\naverage_age = data.groupby('Severity')['Basic_Demos-Age'].mean().reset_index()\n\n# Create a line plot\nplt.figure(figsize=(12, 6))\nsns.lineplot(x='Severity', y='Basic_Demos-Age', data=average_age, marker='o')\nplt.title('Average Age Across Severity Impairment Index')\nplt.xlabel('Severity Impairment Index')\nplt.ylabel('Average Age')\nplt.xticks(rotation=45)  # Rotate x labels for better visibility\nplt.grid(True)  # Add grid for better readability\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:46.639675Z","iopub.execute_input":"2024-10-11T08:03:46.640038Z","iopub.status.idle":"2024-10-11T08:03:46.911231Z","shell.execute_reply.started":"2024-10-11T08:03:46.639998Z","shell.execute_reply":"2024-10-11T08:03:46.909947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\n# Assuming data_imputed is already defined and loaded\n\n# Step 1: Convert 'PreInt_EduHx-computerinternet_hoursday' to integer\nif 'PreInt_EduHx-computerinternet_hoursday' in data_imputed.columns:\n    # Convert the column to integer values, handling NaN by filling with 0\n    data_imputed['PreInt_EduHx-computerinternet_hoursday'] = data_imputed['PreInt_EduHx-computerinternet_hoursday'].fillna(0).astype(int)\nelse:\n    print(\"Column 'PreInt_EduHx-computerinternet_hoursday' does not exist.\")\n\n# Step 2: Group by gender and count the number of occurrences for each internet usage category\ninternet_usage = data_imputed.groupby(['Basic_Demos-Sex', 'PreInt_EduHx-computerinternet_hoursday']).size().unstack()\n\n# Step 3: Create a pie chart for each gender and display them side by side\nfig, axes = plt.subplots(1, len(internet_usage.index), figsize=(16, 8))  # Create subplots for each gender\n\nfor ax, gender in zip(axes, internet_usage.index):\n    ax.pie(internet_usage.loc[gender], \n           labels=internet_usage.columns, \n           autopct='%1.1f%%', \n           startangle=90, \n           colors=sns.color_palette('pastel'))\n    ax.set_title(f'Internet Usage by {gender}')\n    ax.axis('equal')  # Equal aspect ratio ensures that pie chart is circular\n\nplt.tight_layout()  # Adjust layout to prevent overlap\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:46.912744Z","iopub.execute_input":"2024-10-11T08:03:46.913145Z","iopub.status.idle":"2024-10-11T08:03:47.281008Z","shell.execute_reply.started":"2024-10-11T08:03:46.913102Z","shell.execute_reply":"2024-10-11T08:03:47.279872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- **female usage more than 3 h is more than males**\n- **0=Less than 1h/day, 1=Around 1h/day, 2=Around 2hs/day, 3=More than 3hs/day**","metadata":{}},{"cell_type":"code","source":"\ninternet_usage = data_imputed.groupby(['Basic_Demos-Age', 'PreInt_EduHx-computerinternet_hoursday']).size().unstack(fill_value=0)\n\ninternet_usage = internet_usage.reset_index()\n\ninternet_usage_melted = internet_usage.melt(id_vars='Basic_Demos-Age', \n                                              var_name='Internet Usage', \n                                              value_name='Count')\nplt.figure(figsize=(12, 6))\npalette = sns.color_palette('pastel') \nsns.barplot(data=internet_usage_melted, \n            x='Basic_Demos-Age', \n            y='Count', \n            hue='Internet Usage', \n            palette=palette)\n\nplt.xlabel('Age Group', fontsize=12)\nplt.ylabel('Count', fontsize=12)\nplt.title('Internet Usage Distribution by Age Group', fontsize=14)\nplt.xticks(rotation=45)\n\nhandles = []\nfor i, label in enumerate(['0=Less than 1h/day', '1=Around 1h/day', '2=Around 2hs/day', '3=More than 3hs/day']):\n    handles.append(plt.Line2D([0], [0], color=palette[i], lw=4))  # Create a line for each color\n\nplt.legend(handles, \n           ['0=Less than 1h/day', '1=Around 1h/day', '2=Around 2hs/day', '3=More than 3hs/day'], \n           title='Daily Internet Usage')\n\nplt.tight_layout() \nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:47.282447Z","iopub.execute_input":"2024-10-11T08:03:47.282827Z","iopub.status.idle":"2024-10-11T08:03:48.118073Z","shell.execute_reply.started":"2024-10-11T08:03:47.282787Z","shell.execute_reply":"2024-10-11T08:03:48.116780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- *Internet usage increase with age*","metadata":{}},{"cell_type":"code","source":"\ncorrelation_matrix = data_imputed.corr()\n\nsii_correlation = correlation_matrix['sii'].sort_values(ascending=False)\n\ntop_10_features = sii_correlation.index[1:30]\ntop_10_correlations = sii_correlation[1:30]  \n\nplt.figure(figsize=(10, 6))\nsns.barplot(x=top_10_correlations.values, y=top_10_features, palette='coolwarm')\n\nplt.title('Top 10 Correlated Features with Attrition', fontsize=16)\nplt.xlabel('Correlation Coefficient', fontsize=14)\nplt.ylabel('Features', fontsize=14)\n\nplt.axvline(0, color='gray', linestyle='--')  \nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:48.120292Z","iopub.execute_input":"2024-10-11T08:03:48.120795Z","iopub.status.idle":"2024-10-11T08:03:48.742337Z","shell.execute_reply.started":"2024-10-11T08:03:48.120740Z","shell.execute_reply":"2024-10-11T08:03:48.741136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature Engineering","metadata":{}},{"cell_type":"code","source":"data_imputed.info()","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:48.743821Z","iopub.execute_input":"2024-10-11T08:03:48.744223Z","iopub.status.idle":"2024-10-11T08:03:48.761452Z","shell.execute_reply.started":"2024-10-11T08:03:48.744182Z","shell.execute_reply":"2024-10-11T08:03:48.760155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Blood pressure ratio","metadata":{}},{"cell_type":"code","source":"data_imputed['BP_Ratio'] = data_imputed['Physical-Systolic_BP'] / data_imputed['Physical-Diastolic_BP']","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:48.763644Z","iopub.execute_input":"2024-10-11T08:03:48.764214Z","iopub.status.idle":"2024-10-11T08:03:48.775787Z","shell.execute_reply.started":"2024-10-11T08:03:48.764154Z","shell.execute_reply":"2024-10-11T08:03:48.774580Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## BMI To Heart rate Ratio","metadata":{}},{"cell_type":"code","source":"data_imputed['HeartRate_BMI'] = data_imputed['Physical-HeartRate'] * data_imputed['Physical-BMI']","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:48.777330Z","iopub.execute_input":"2024-10-11T08:03:48.777841Z","iopub.status.idle":"2024-10-11T08:03:48.792014Z","shell.execute_reply.started":"2024-10-11T08:03:48.777782Z","shell.execute_reply":"2024-10-11T08:03:48.790725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_imputed['HRmax'] = 220 - data_imputed['Basic_Demos-Age']  # Estimate HRmax based on age","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:48.793394Z","iopub.execute_input":"2024-10-11T08:03:48.793890Z","iopub.status.idle":"2024-10-11T08:03:48.805240Z","shell.execute_reply.started":"2024-10-11T08:03:48.793836Z","shell.execute_reply":"2024-10-11T08:03:48.803730Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#data_imputed['BodyStrength&Flexibility'] = data_imputed['FGC-FGC_CU'] + data_imputed['FGC-FGC_PU'] + data_imputed['FGC-FGC_SRL'] + data_imputed['FGC-FGC_SRR'] + data_imputed['FGC-FGC_TL']\ndata_imputed['BodyStrength&Flexibility_class'] = data_imputed['FGC-FGC_CU_Zone'] + data_imputed['FGC-FGC_PU_Zone'] + data_imputed['FGC-FGC_SRL_Zone'] + data_imputed['FGC-FGC_SRR_Zone'] + data_imputed['FGC-FGC_TL_Zone']\n\n#data_imputed = data_imputed.drop(['FGC-FGC_CU','FGC-FGC_PU','FGC-FGC_SRL','FGC-FGC_SRR','FGC-FGC_TL'], axis=1)\n\ndata_imputed = data_imputed.drop(['FGC-FGC_CU_Zone','FGC-FGC_PU_Zone','FGC-FGC_SRL_Zone','FGC-FGC_SRR_Zone','FGC-FGC_TL_Zone'], axis=1)                          \n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:48.806660Z","iopub.execute_input":"2024-10-11T08:03:48.807053Z","iopub.status.idle":"2024-10-11T08:03:48.825893Z","shell.execute_reply.started":"2024-10-11T08:03:48.807013Z","shell.execute_reply":"2024-10-11T08:03:48.824712Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_imputed.shape","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:48.827546Z","iopub.execute_input":"2024-10-11T08:03:48.828082Z","iopub.status.idle":"2024-10-11T08:03:48.841895Z","shell.execute_reply.started":"2024-10-11T08:03:48.828023Z","shell.execute_reply":"2024-10-11T08:03:48.840795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Transformation and scaling","metadata":{}},{"cell_type":"code","source":"x = data_imputed.drop('sii', axis=1)\ny = data_imputed['sii']\nx_train , x_test , y_train ,y_test =train_test_split(x,y , random_state=1 ,test_size=0.3)","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:48.843348Z","iopub.execute_input":"2024-10-11T08:03:48.843775Z","iopub.status.idle":"2024-10-11T08:03:48.856480Z","shell.execute_reply.started":"2024-10-11T08:03:48.843734Z","shell.execute_reply":"2024-10-11T08:03:48.855326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scaler = RobustScaler()\n\n#scaler=StandardScaler()\nx_train=pd.DataFrame(scaler.fit_transform(x_train), columns=x_train.columns)\nx_test=pd.DataFrame(scaler.fit_transform(x_test), columns=x_train.columns)\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:48.858023Z","iopub.execute_input":"2024-10-11T08:03:48.859232Z","iopub.status.idle":"2024-10-11T08:03:48.903966Z","shell.execute_reply.started":"2024-10-11T08:03:48.859184Z","shell.execute_reply":"2024-10-11T08:03:48.902668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Modeling ","metadata":{}},{"cell_type":"code","source":"result1, result2, result3 = [], [], [] \n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:48.905361Z","iopub.execute_input":"2024-10-11T08:03:48.905755Z","iopub.status.idle":"2024-10-11T08:03:48.910963Z","shell.execute_reply.started":"2024-10-11T08:03:48.905715Z","shell.execute_reply":"2024-10-11T08:03:48.909773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def modeling(model):\n    model.fit(x_train, y_train)\n    \n    train_pred = model.predict(x_train)\n    test_pred = model.predict(x_test)\n    \n    # Calculate metrics with specified average\n    train_accuracy = accuracy_score(y_train, train_pred) * 100\n    test_accuracy = accuracy_score(y_test, test_pred) * 100\n    test_recall = recall_score(y_test, test_pred, average='weighted') * 100\n    test_f1_score = f1_score(y_test, test_pred, average='weighted') * 100\n    \n    # Append results\n    result1.append(test_accuracy)\n    result2.append(test_recall)\n    result3.append(test_f1_score)\n    \n    # Classification report\n    report = classification_report(y_test, test_pred)\n    print(report)\n    print(f'Training Accuracy: {train_accuracy}')\n    print(f'Test Accuracy: {test_accuracy}, Test Recall: {test_recall}, Test F1: {test_f1_score}')\n    \n    # Confusion matrix\n    cm = confusion_matrix(y_test, test_pred)\n    sns.heatmap(cm, annot=True, fmt='0.2f', cmap='YlGnBu', linewidths=1)\n    plt.xlabel('Predicted')\n    plt.ylabel('Actual')\n    plt.title('Confusion Matrix')\n    plt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:48.912468Z","iopub.execute_input":"2024-10-11T08:03:48.913184Z","iopub.status.idle":"2024-10-11T08:03:48.926406Z","shell.execute_reply.started":"2024-10-11T08:03:48.913138Z","shell.execute_reply":"2024-10-11T08:03:48.925125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"svm_model = SVC()\nmodeling(svm_model)","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:48.928051Z","iopub.execute_input":"2024-10-11T08:03:48.928899Z","iopub.status.idle":"2024-10-11T08:03:49.554450Z","shell.execute_reply.started":"2024-10-11T08:03:48.928830Z","shell.execute_reply":"2024-10-11T08:03:49.553209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"logist = LogisticRegression()\nmodeling(logist)","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:49.555847Z","iopub.execute_input":"2024-10-11T08:03:49.556199Z","iopub.status.idle":"2024-10-11T08:03:50.366912Z","shell.execute_reply.started":"2024-10-11T08:03:49.556161Z","shell.execute_reply":"2024-10-11T08:03:50.365741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"decision_tree = DecisionTreeClassifier()\nmodeling(decision_tree)","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:50.368523Z","iopub.execute_input":"2024-10-11T08:03:50.370003Z","iopub.status.idle":"2024-10-11T08:03:50.776476Z","shell.execute_reply.started":"2024-10-11T08:03:50.369954Z","shell.execute_reply":"2024-10-11T08:03:50.775246Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#pip install lazypredict\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:50.777825Z","iopub.execute_input":"2024-10-11T08:03:50.778196Z","iopub.status.idle":"2024-10-11T08:03:50.783103Z","shell.execute_reply.started":"2024-10-11T08:03:50.778158Z","shell.execute_reply":"2024-10-11T08:03:50.781950Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#from lazypredict.Supervised import LazyClassifier\n#from sklearn.model_selection import train_test_split\n#from sklearn.datasets import load_iris\n#import pandas as pd\n\n\n#X = data_imputed.drop('sii', axis=1)\n#y = data_imputed['sii']\n#X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)\n\n#clf = LazyClassifier(verbose=0, ignore_warnings=True, custom_metric=None)\n\n#models, predictions = clf.fit(X_train, X_test, y_train, y_test)\n\n#print(models)\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:50.784740Z","iopub.execute_input":"2024-10-11T08:03:50.785273Z","iopub.status.idle":"2024-10-11T08:03:50.795291Z","shell.execute_reply.started":"2024-10-11T08:03:50.785215Z","shell.execute_reply":"2024-10-11T08:03:50.794208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#from lazypredict.Supervised import LazyClassifier\n#from sklearn.model_selection import train_test_split\n#from sklearn.ensemble import RandomForestClassifier\n#from sklearn.feature_selection import RFE\n#import pandas as pd\n\n# Assuming data_imputed is already loaded and processed\n\n# Step 1: Define X and y\n#X = data_imputed.drop('sii', axis=1)  # Independent features\n#y = data_imputed['sii']               # Target feature\n\n# Step 2: Perform feature selection using RFE\n#model = RandomForestClassifier()      # You can use any classifier here, RandomForest works well\n#rfe = RFE(model, n_features_to_select=10)  # Adjust the number of features to select\n#rfe.fit(X, y)\n\n# Step 3: Select features based on RFE\n#X_selected = X.loc[:, rfe.support_]   # Keep only the selected features\n\n# Step 4: Split the dataset with selected features\n#X_train, X_test, y_train, y_test = train_test_split(X_selected, y, test_size=0.2, random_state=42)\n\n# Step 5: Use LazyClassifier with the selected features\n#clf = LazyClassifier(verbose=0, ignore_warnings=True, custom_metric=None)\n#models, predictions = clf.fit(X_train, X_test, y_train, y_test)\n\n# Output the results\n#print(models)\n","metadata":{"execution":{"iopub.status.busy":"2024-10-11T08:03:50.796944Z","iopub.execute_input":"2024-10-11T08:03:50.797364Z","iopub.status.idle":"2024-10-11T08:03:50.808089Z","shell.execute_reply.started":"2024-10-11T08:03:50.797288Z","shell.execute_reply":"2024-10-11T08:03:50.806946Z"},"trusted":true},"execution_count":null,"outputs":[]}]}