{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30775,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# EASY solution: 1 catboost, no tracker data used","metadata":{}},{"cell_type":"markdown","source":"# <p style=\"padding:15px; background-color:chocolate; font-family:arial; font-weight:bold; color:white; font-size:100%; letter-spacing: 2px; text-align:left; border-radius: 10px 10px\">1/ Introduction</p>","metadata":{}},{"cell_type":"markdown","source":" <div style=\"border-radius:12px; border:chocolate solid; padding: 15px; background-color: #f6f5f5; font-size:120%; text-align:left\">\n\n- The main aim of the competition is to use our training data to predict **sii** or **Severity Impairment Index**, which is a standard measure of Problematic Internet Use (PIU).\n- The training data comprises 3,960 records of children and young people with 81 columns (not including the ID column).\n- Of particular importance in the data are results of the **Parent-Child Internet Addiction Test (PCIAT)**.\n- The target is actually derived from the field PCIAT-PCIAT_Total (scored out of 100).\n- We can therefore choose to predict the PCIAT Total and convert this to sii (making this a regression problem) or stick with sii (making this a classification problem).\n- The test data is really just formatted sample data. The actual test data of about 3,800 instances is hidden.\n- In the sample data none of the 22 PCIAT fields are available (in addition to the target feature). Hence the sample data format has 58 columns compared to 81 in the train data.\n- In 1,224 records in the train data the sii target and all the PCIAT columns are missing - presumably because not available.\n- Overall there are > 1,000,000 missing values in the train data.\n- Only 2,736 records have a target, the rest are missing.\n- 996 of the young people also have sensor data from a worn device which measures gross motor activity.","metadata":{}},{"cell_type":"markdown","source":"# <p style=\"padding:15px; background-color:chocolate; font-family:arial; font-weight:bold; color:white; font-size:100%; letter-spacing: 2px; text-align:left; border-radius: 10px 10px\">2/ Imports</p>","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\nfrom sklearn.model_selection import RandomizedSearchCV\nfrom scipy.stats import uniform, randint\nfrom tqdm import tqdm\nfrom scipy.optimize import minimize\n\nimport numpy as np, pandas as pd, os\nfrom sklearn.model_selection import cross_val_score, StratifiedKFold\nimport xgboost as xgb\nfrom lightgbm import LGBMRegressor\nfrom sklearn.ensemble import VotingRegressor\nfrom catboost import CatBoostRegressor\nimport plotly.express as px, seaborn as sns, matplotlib.pyplot as plt\nsns.set_style('darkgrid')\nfrom sklearn.metrics import make_scorer, cohen_kappa_score\nimport eli5\nfrom eli5.sklearn import PermutationImportance\nimport warnings\nwarnings.simplefilter('ignore')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2025-01-07T14:08:23.797387Z","iopub.execute_input":"2025-01-07T14:08:23.797727Z","iopub.status.idle":"2025-01-07T14:08:41.993254Z","shell.execute_reply.started":"2025-01-07T14:08:23.797685Z","shell.execute_reply":"2025-01-07T14:08:41.992070Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <p style=\"padding:15px; background-color:chocolate; font-family:arial; font-weight:bold; color:white; font-size:100%; letter-spacing: 2px; text-align:left; border-radius: 10px 10px\">3/ Data</p>","metadata":{}},{"cell_type":"code","source":"path = '../input/child-mind-institute-problematic-internet-use/'\n\ntrain = pd.read_csv(path + 'train.csv', index_col = 'id')\ntest = pd.read_csv(path + 'test.csv', index_col = 'id')\n\nprint(\"The train data has the shape: \",train.shape)\nprint(\"The test data has the shape: \",test.shape)\nprint(\"\")\nprint(\"Total number of missing training values: \", train.isna().sum().sum())","metadata":{"execution":{"iopub.status.busy":"2025-01-07T14:08:41.995336Z","iopub.execute_input":"2025-01-07T14:08:41.996028Z","iopub.status.idle":"2025-01-07T14:08:42.101331Z","shell.execute_reply.started":"2025-01-07T14:08:41.995991Z","shell.execute_reply":"2025-01-07T14:08:42.100260Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <p style=\"padding:15px; background-color:chocolate; font-family:arial; font-weight:bold; color:white; font-size:100%; letter-spacing: 2px; text-align:left; border-radius: 10px 10px\">4/ Predictive Features</p>","metadata":{}},{"cell_type":"markdown","source":" <div style=\"border-radius:12px; border:chocolate solid; padding: 15px; background-color: #f6f5f5; font-size:120%; text-align:left\">\n    \n* **Demographics** - Information about age and sex of participants.\n* **Internet Use** - Number of hours of using computer/internet per day.\n* **Children's Global Assessment Scale** - Numeric scale used by mental health clinicians to rate the general functioning of youths under the age of 18.\n* **Physical Measures** - Collection of blood pressure, heart rate, height, weight and waist, and hip measurements.\n* **FitnessGram Vitals and Treadmill** - Measurements of cardiovascular fitness assessed using the NHANES treadmill protocol.\n* **FitnessGram Child** - Health related physical fitness assessment measuring five different parameters including aerobic capacity, muscular strength, muscular endurance, flexibility, and body composition.\n* **Bio-electric Impedance Analysis** - Measure of key body composition elements, including BMI, fat, muscle, and water content.\n* **Physical Activity Questionnaire** - Information about children's participation in vigorous activities over the last 7 days.\n* **Sleep Disturbance Scale** - Scale to categorize sleep disorders in children.\n* **Actigraphy** - Objective measure of ecological physical activity through a research-grade biotracker. Many values seem to relate to a period *after* the PCIAT test was carried out. See discussion [here](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/538082#3017157).\n* **Season** - for each set of measurements there is a 'season' feature which gives the season of the year when the measurements were carried out. These are the only predictive categorical features in the dataset and can be easily preprocessed.","metadata":{}},{"cell_type":"code","source":"train_cat_columns = train.select_dtypes(exclude = 'number').columns\n\nfor season in train_cat_columns:\n    train[season] = train[season].fillna(0)\n    train[season] = train[season].replace({'Spring':1, 'Summer':2, 'Fall':3, 'Winter':4})","metadata":{"execution":{"iopub.status.busy":"2025-01-07T14:08:42.105813Z","iopub.execute_input":"2025-01-07T14:08:42.106118Z","iopub.status.idle":"2025-01-07T14:08:42.149878Z","shell.execute_reply.started":"2025-01-07T14:08:42.106088Z","shell.execute_reply":"2025-01-07T14:08:42.148913Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_cat_columns = test.select_dtypes(exclude = 'number').columns\n\nfor season in test_cat_columns:\n    test[season] = test[season].fillna(0)\n    test[season] = test[season].replace({'Spring':1, 'Summer':2, 'Fall':3, 'Winter':4})","metadata":{"execution":{"iopub.status.busy":"2025-01-07T14:08:42.152933Z","iopub.execute_input":"2025-01-07T14:08:42.153260Z","iopub.status.idle":"2025-01-07T14:08:42.171536Z","shell.execute_reply.started":"2025-01-07T14:08:42.153230Z","shell.execute_reply":"2025-01-07T14:08:42.170442Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <p style=\"padding:15px; background-color:chocolate; font-family:arial; font-weight:bold; color:white; font-size:100%; letter-spacing: 2px; text-align:left; border-radius: 10px 10px\">5/ PCIAT Features</p>","metadata":{}},{"cell_type":"markdown","source":" <div style=\"border-radius:12px; border:chocolate solid; padding: 15px; background-color: #f6f5f5; font-size:120%; text-align:left\">\n    \n* As mentioned there are 22 PCIAT features. These comprise answers to 20 questions (each marked out of 5), the total score and 'season' when the test was carried out.\n* The sii target is derived from the total PCIAT score:\n    - 0-30 gives sii = 0\n    - 31-49 gives sii = 1\n    - 50-79 gives sii = 2\n    - 80-100 gives sii = 3. \n* We show this by simply counting the values. The same information is confirmed [here](https://digitalwellnesslab.org/wp-content/uploads/Scoring-Overview.pdf).\n* We drop all the PCIAT features from the dataset except the PCIAT Total feature which can be used as a regression target.\n* The PCIAT Total visualisation box plot shows us that many of the top scores look like outliers - yet this is our most important category.","metadata":{}},{"cell_type":"code","source":"PCIAT_cols = [val for val in train.columns[train.columns.str.contains('PCIAT')]]\nprint('Number of PCIAT features = ' , len(PCIAT_cols))","metadata":{"execution":{"iopub.status.busy":"2025-01-07T14:08:42.172855Z","iopub.execute_input":"2025-01-07T14:08:42.173162Z","iopub.status.idle":"2025-01-07T14:08:42.190683Z","shell.execute_reply.started":"2025-01-07T14:08:42.173133Z","shell.execute_reply":"2025-01-07T14:08:42.189592Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.sii.value_counts()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-07T14:08:42.191873Z","iopub.execute_input":"2025-01-07T14:08:42.192193Z","iopub.status.idle":"2025-01-07T14:08:42.212761Z","shell.execute_reply.started":"2025-01-07T14:08:42.192149Z","shell.execute_reply":"2025-01-07T14:08:42.211603Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"PCIAT_cols.remove('PCIAT-PCIAT_Total')\ntrain = train.drop(columns = PCIAT_cols)","metadata":{"execution":{"iopub.status.busy":"2025-01-07T14:08:42.214218Z","iopub.execute_input":"2025-01-07T14:08:42.215078Z","iopub.status.idle":"2025-01-07T14:08:42.224865Z","shell.execute_reply.started":"2025-01-07T14:08:42.215026Z","shell.execute_reply":"2025-01-07T14:08:42.223767Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <p style=\"padding:15px; background-color:chocolate; font-family:arial; font-weight:bold; color:white; font-size:100%; letter-spacing: 2px; text-align:left; border-radius: 10px 10px\">6/ Severity Impairment Index </p>","metadata":{}},{"cell_type":"markdown","source":" <div style=\"border-radius:12px; border:chocolate solid; padding: 15px; background-color: #f6f5f5; font-size:120%; text-align:left\">\n\n* When first alerted to this competition by email there was a reference to excessive internet usage amongst children and young people as being the key problem to be assessed.\n* One of the puzzling things about the data is that, even for the 34 'severe' cases of PIU where sii =3, we can see that 5 participants assessed as severe are hardly using the internet at all.\n* How can they have scored so highly on the PCIAT questionnaire? There is a helpful discussion [here](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535525#3003303).","metadata":{}},{"cell_type":"code","source":"sns.countplot(train, x = 'sii').set_title('Count of sii')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-07T14:08:42.225970Z","iopub.execute_input":"2025-01-07T14:08:42.226254Z","iopub.status.idle":"2025-01-07T14:08:42.526321Z","shell.execute_reply.started":"2025-01-07T14:08:42.226228Z","shell.execute_reply":"2025-01-07T14:08:42.525176Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train = train.dropna(subset='sii')","metadata":{"execution":{"iopub.status.busy":"2025-01-07T14:08:42.527931Z","iopub.execute_input":"2025-01-07T14:08:42.528268Z","iopub.status.idle":"2025-01-07T14:08:42.537416Z","shell.execute_reply.started":"2025-01-07T14:08:42.528236Z","shell.execute_reply":"2025-01-07T14:08:42.536486Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <p style=\"padding:15px; background-color:chocolate; font-family:arial; font-weight:bold; color:white; font-size:100%; letter-spacing: 2px; text-align:left; border-radius: 10px 10px\">7/ Correlations </p>","metadata":{"execution":{"iopub.status.busy":"2024-09-28T11:36:45.957167Z","iopub.execute_input":"2024-09-28T11:36:45.958067Z","iopub.status.idle":"2024-09-28T11:36:45.963155Z","shell.execute_reply.started":"2024-09-28T11:36:45.958022Z","shell.execute_reply":"2024-09-28T11:36:45.961779Z"}}},{"cell_type":"markdown","source":" <div style=\"border-radius:12px; border:chocolate solid; padding: 15px; background-color: #f6f5f5; font-size:120%; text-align:left\">\n    \n* With large numbers of features to choose from I decided to do some feature selection and check it's impact on the model.\n* Here I select the features with the strongest correlation with the PCIAT total and drop the weaker ones.","metadata":{}},{"cell_type":"code","source":"corr = pd.DataFrame(train.corr()['PCIAT-PCIAT_Total'].sort_values(ascending = False))\ncorr.style.background_gradient(cmap='YlOrRd')","metadata":{"execution":{"iopub.status.busy":"2025-01-07T14:08:42.538789Z","iopub.execute_input":"2025-01-07T14:08:42.539225Z","iopub.status.idle":"2025-01-07T14:08:42.629935Z","shell.execute_reply.started":"2025-01-07T14:08:42.539170Z","shell.execute_reply":"2025-01-07T14:08:42.628909Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"selection = corr[(corr['PCIAT-PCIAT_Total']>.06) | (corr['PCIAT-PCIAT_Total']<-.01)]\nselection = [val for val in selection.index]\nselection.remove('PCIAT-PCIAT_Total')\nselection.remove('sii')","metadata":{"execution":{"iopub.status.busy":"2025-01-07T14:08:42.631254Z","iopub.execute_input":"2025-01-07T14:08:42.631668Z","iopub.status.idle":"2025-01-07T14:08:42.638645Z","shell.execute_reply.started":"2025-01-07T14:08:42.631635Z","shell.execute_reply":"2025-01-07T14:08:42.637536Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"selection","metadata":{"execution":{"iopub.status.busy":"2025-01-07T14:08:42.639873Z","iopub.execute_input":"2025-01-07T14:08:42.640234Z","iopub.status.idle":"2025-01-07T14:08:42.653475Z","shell.execute_reply.started":"2025-01-07T14:08:42.640202Z","shell.execute_reply":"2025-01-07T14:08:42.652242Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <p style=\"padding:15px; background-color:chocolate; font-family:arial; font-weight:bold; color:white; font-size:100%; letter-spacing: 2px; text-align:left; border-radius: 10px 10px\">8/ Missing Values</p>","metadata":{}},{"cell_type":"markdown","source":" <div style=\"border-radius:12px; border:chocolate solid; padding: 15px; background-color: #f6f5f5; font-size:120%; text-align:left\">\n\n* There are large numbers of missing values remaining in the dataset\n* For example, most of the waist circumference feature values are missing.\n* Let's drop columns where there are more than half values missing.\n* Though we have not included the actigraphy data in this notebook analysis, these records are only available for just over a third of respondents.","metadata":{}},{"cell_type":"code","source":"null = train.isna().sum().sort_values(ascending = False).head(46)\nnull = pd.DataFrame(null)\nnull = null.rename(columns= {0:'Missing'})\nnull.style.background_gradient(cmap='YlOrRd')","metadata":{"execution":{"iopub.status.busy":"2025-01-07T14:08:42.657628Z","iopub.execute_input":"2025-01-07T14:08:42.658099Z","iopub.status.idle":"2025-01-07T14:08:42.679098Z","shell.execute_reply.started":"2025-01-07T14:08:42.658058Z","shell.execute_reply":"2025-01-07T14:08:42.677528Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"half_missing = [val for val in train.columns[train.isnull().sum()>len(train)/2]]\nhalf_missing","metadata":{"execution":{"iopub.status.busy":"2025-01-07T14:08:42.680286Z","iopub.execute_input":"2025-01-07T14:08:42.680651Z","iopub.status.idle":"2025-01-07T14:08:42.690435Z","shell.execute_reply.started":"2025-01-07T14:08:42.680618Z","shell.execute_reply":"2025-01-07T14:08:42.689167Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"selection = [i for i in selection if i not in half_missing]","metadata":{"execution":{"iopub.status.busy":"2025-01-07T14:08:42.691956Z","iopub.execute_input":"2025-01-07T14:08:42.692427Z","iopub.status.idle":"2025-01-07T14:08:42.702351Z","shell.execute_reply.started":"2025-01-07T14:08:42.692379Z","shell.execute_reply":"2025-01-07T14:08:42.701143Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <p style=\"padding:15px; background-color:chocolate; font-family:arial; font-weight:bold; color:white; font-size:100%; letter-spacing: 2px; text-align:left; border-radius: 10px 10px\">9/ Selected Features</p>","metadata":{"execution":{"iopub.status.busy":"2024-10-05T13:56:59.196471Z","iopub.execute_input":"2024-10-05T13:56:59.196979Z","iopub.status.idle":"2024-10-05T13:56:59.20348Z","shell.execute_reply.started":"2024-10-05T13:56:59.196936Z","shell.execute_reply":"2024-10-05T13:56:59.202078Z"}}},{"cell_type":"markdown","source":" <div style=\"border-radius:12px; border:chocolate solid; padding: 15px; background-color: #f6f5f5; font-size:120%; text-align:left\">\n\n* The idea is to create a robust model that focuses on key signals in the data and reduces some of the excess noise from large numbers of features.\n* Some of the min and max values appear to be impossible (such as a minimum weight of zero) or very unlikely.","metadata":{"execution":{"iopub.status.busy":"2024-10-05T13:59:54.552101Z","iopub.execute_input":"2024-10-05T13:59:54.552525Z","iopub.status.idle":"2024-10-05T13:59:54.560829Z","shell.execute_reply.started":"2024-10-05T13:59:54.552484Z","shell.execute_reply":"2024-10-05T13:59:54.559313Z"}}},{"cell_type":"code","source":"describe = train[selection].describe().T\ndescribe = describe[['min','max']].sort_index()\ndescribe.style.background_gradient(cmap='YlOrRd')","metadata":{"execution":{"iopub.status.busy":"2025-01-07T14:08:42.703664Z","iopub.execute_input":"2025-01-07T14:08:42.704099Z","iopub.status.idle":"2025-01-07T14:08:43.032794Z","shell.execute_reply.started":"2025-01-07T14:08:42.704054Z","shell.execute_reply":"2025-01-07T14:08:43.031710Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <p style=\"padding:15px; background-color:chocolate; font-family:arial; font-weight:bold; color:white; font-size:100%; letter-spacing: 2px; text-align:left; border-radius: 10px 10px\">11/ Regression Model</p>","metadata":{}},{"cell_type":"markdown","source":" <div style=\"border-radius:12px; border:chocolate solid; padding: 15px; background-color: #f6f5f5; font-size:120%; text-align:left\">\n\n* In this section I train 3 regression model as a benchmark and use PCIAT-PCIAT_Total as the target.\n* We tweak the quadratic kappa function to convert PCIAT total scores to sii categories, which gives a better cross-validation result.","metadata":{"execution":{"iopub.status.busy":"2024-10-01T12:35:10.605389Z","iopub.execute_input":"2024-10-01T12:35:10.605904Z","iopub.status.idle":"2024-10-01T12:35:10.613456Z","shell.execute_reply.started":"2024-10-01T12:35:10.605858Z","shell.execute_reply":"2024-10-01T12:35:10.612056Z"}}},{"cell_type":"code","source":"X = train[selection]\ntest = test[selection]\ny = train['PCIAT-PCIAT_Total']","metadata":{"execution":{"iopub.status.busy":"2025-01-07T14:08:43.034052Z","iopub.execute_input":"2025-01-07T14:08:43.034377Z","iopub.status.idle":"2025-01-07T14:08:43.041782Z","shell.execute_reply.started":"2025-01-07T14:08:43.034312Z","shell.execute_reply":"2025-01-07T14:08:43.040895Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def convert(scores):\n    scores = np.array(scores)\n    bins = np.zeros_like(scores)\n    bins[scores <= 30] = 0\n    bins[(scores > 30) & (scores < 50)] = 1\n    bins[(scores >= 50) & (scores < 80)] = 2\n    bins[scores >= 80] = 3\n    return bins\ndef convert2(scores):\n    scores = np.array(scores)*1.252\n    bins = np.zeros_like(scores)\n    bins[scores <= 32] = 0\n    bins[(scores > 32) & (scores < 44)] = 1\n    bins[(scores >= 44) & (scores < 82)] = 2\n    bins[scores >= 82] = 3\n    return bins\ndef quadratic_kappa(y_true, y_pred):\n    y_true_cat = convert(y_true)\n    y_pred_cat = convert(y_pred)\n    return cohen_kappa_score(y_true_cat, y_pred_cat, weights='quadratic')\n\nkappa_scorer = make_scorer(quadratic_kappa, greater_is_better=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-07T14:08:43.042981Z","iopub.execute_input":"2025-01-07T14:08:43.043341Z","iopub.status.idle":"2025-01-07T14:08:43.054336Z","shell.execute_reply.started":"2025-01-07T14:08:43.043299Z","shell.execute_reply":"2025-01-07T14:08:43.053197Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#params_xgb\nparams_xgb = {\n    'max_depth': 3,\n    'n_estimators': 100,\n    'learning_rate': 0.064,\n    'subsample': 0.7,\n    'reg_alpha': 1,\n    'reg_lambda': 5,\n    'random_seed': 42,\n    'colsample_bytree': 0.8\n}\n\n#params_catboost\nparams_catboost = {\n    'max_depth': 5,\n    'n_estimators': 100,\n    'learning_rate': 0.054,\n    'subsample': 1.0,\n    'random_seed': 42,\n    'colsample_bylevel': 1.0\n}\n\n#params_light\nparams_light = {\n    'learning_rate': 0.046,\n    'max_depth': 5,\n    'num_leaves': 480,\n    'n_estimators': 100,\n    'min_data_in_leaf': 13,\n    'feature_fraction': 0.898,\n    'bagging_fraction': 0.784,\n    'bagging_freq': 4,\n    'lambda_l1': 10,\n    'lambda_l2': 0.01\n}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-07T14:08:43.055548Z","iopub.execute_input":"2025-01-07T14:08:43.055920Z","iopub.status.idle":"2025-01-07T14:08:43.067403Z","shell.execute_reply.started":"2025-01-07T14:08:43.055879Z","shell.execute_reply":"2025-01-07T14:08:43.066306Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"skf = StratifiedKFold(n_splits=10)\nxgb_model = xgb.XGBRegressor(**params_xgb)\nlight_model = LGBMRegressor(**params_light, random_state=42, verbose=-1)\ncatboost_model = CatBoostRegressor(**params_catboost, verbose=0)\n# Kết hợp mô hình bằng VotingRegressor\nvoting_regressor = VotingRegressor(\n    estimators=[\n        ('xgb', xgb_model),\n        ('light', light_model),\n        ('catboost', catboost_model)\n    ],\n    weights=[6.0, 6.0, 5.0]  # Trọng số cho XGBoost và CatBoost\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-07T14:08:43.068882Z","iopub.execute_input":"2025-01-07T14:08:43.069425Z","iopub.status.idle":"2025-01-07T14:08:43.087912Z","shell.execute_reply.started":"2025-01-07T14:08:43.069378Z","shell.execute_reply":"2025-01-07T14:08:43.086876Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"xgb_model.fit(X,y)\nfeature_imp = pd.Series(xgb_model.feature_importances_,index=X.columns).sort_values(ascending=False)\nsns.barplot(x=feature_imp, y=feature_imp.index)\nplt.xlabel('Feature Importance Score')\nplt.title(\"xgb_model\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-07T14:08:43.089467Z","iopub.execute_input":"2025-01-07T14:08:43.089838Z","iopub.status.idle":"2025-01-07T14:08:43.738214Z","shell.execute_reply.started":"2025-01-07T14:08:43.089807Z","shell.execute_reply":"2025-01-07T14:08:43.736918Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"light_model.fit(X,y)\nfeature_imp = pd.Series(light_model.feature_importances_,index=X.columns).sort_values(ascending=False)\nsns.barplot(x=feature_imp, y=feature_imp.index)\nplt.xlabel('Feature Importance Score')\nplt.title(\"light_model\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-07T14:08:43.739679Z","iopub.execute_input":"2025-01-07T14:08:43.740023Z","iopub.status.idle":"2025-01-07T14:08:44.580705Z","shell.execute_reply.started":"2025-01-07T14:08:43.739990Z","shell.execute_reply":"2025-01-07T14:08:44.579509Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"catboost_model.fit(X,y)\nfeature_imp = pd.Series(catboost_model.feature_importances_,index=X.columns).sort_values(ascending=False)\nsns.barplot(x=feature_imp, y=feature_imp.index)\nplt.xlabel('Feature Importance Score')\nplt.title(\"catboost_model\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-07T14:08:44.581972Z","iopub.execute_input":"2025-01-07T14:08:44.582269Z","iopub.status.idle":"2025-01-07T14:08:45.354835Z","shell.execute_reply.started":"2025-01-07T14:08:44.582239Z","shell.execute_reply":"2025-01-07T14:08:45.353724Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"voting_regressor.fit(X, y)\ny_pre = voting_regressor.predict(X) \n\nqwk_score = quadratic_kappa(y, y_pre)\nprint(\"QWK no Optimal Score:\", qwk_score)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-07T14:09:09.295393Z","iopub.execute_input":"2025-01-07T14:09:09.295843Z","iopub.status.idle":"2025-01-07T14:09:09.906783Z","shell.execute_reply.started":"2025-01-07T14:09:09.295807Z","shell.execute_reply":"2025-01-07T14:09:09.905621Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"scores = cross_val_score(voting_regressor, X, y, cv=skf, scoring=kappa_scorer)\nprint(\"QWK no Optimal Scores:\", scores)\nprint(\"Mean QWK no Optimal Score:\", np.mean(scores))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-07T14:08:45.956195Z","iopub.execute_input":"2025-01-07T14:08:45.956617Z","iopub.status.idle":"2025-01-07T14:08:51.568161Z","shell.execute_reply.started":"2025-01-07T14:08:45.956583Z","shell.execute_reply":"2025-01-07T14:08:51.567025Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def new_convert(oof_non_rounded, thresholds):\n    return np.where(oof_non_rounded < thresholds[0], 0.0,\n                    np.where(oof_non_rounded < thresholds[1], 1.0,\n                             np.where(oof_non_rounded < thresholds[2], 2.0, 3.0)))\ndef evaluate_predictions(thresholds, y_true, oof_non_rounded):\n    y_true_cat = convert(y_true)\n    rounded_p = new_convert(oof_non_rounded, thresholds)\n    return -cohen_kappa_score(y_true_cat, rounded_p)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-07T14:08:51.569415Z","iopub.execute_input":"2025-01-07T14:08:51.569856Z","iopub.status.idle":"2025-01-07T14:08:51.577566Z","shell.execute_reply.started":"2025-01-07T14:08:51.569810Z","shell.execute_reply":"2025-01-07T14:08:51.576465Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_train = X\ny_train = y\nvoting_regressor.fit(X_train, y_train)\n\ny_pre = voting_regressor.predict(X_train) \n\n# Tối ưu hóa các ngưỡng trên toàn bộ tập dữ liệu\nKappaOptimizer = minimize(evaluate_predictions, \n                          x0=[30, 50, 80], \n                          args=(y_train, y_pre), \n                          method='Nelder-Mead')\n\nassert KappaOptimizer.success, \"Optimization did not converge.\"\n\n# Lưu trữ các ngưỡng tối ưu\noptimal_thresholds = KappaOptimizer.x\n\n# In các ngưỡng tối ưu\nprint(\"Optimal thresholds:\", optimal_thresholds)\n\ny_pre = new_convert(y_pre, optimal_thresholds) #Quy đổi theo ngưỡng tối ưu\ny_train = convert(y_train) #Quy đổi theo ngưỡng gốc\nqwk_score = cohen_kappa_score(y_train, y_pre, weights='quadratic')\nprint(\"QWK Optimal Score:\", qwk_score)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-07T14:08:51.581078Z","iopub.execute_input":"2025-01-07T14:08:51.581394Z","iopub.status.idle":"2025-01-07T14:08:52.522769Z","shell.execute_reply.started":"2025-01-07T14:08:51.581365Z","shell.execute_reply":"2025-01-07T14:08:52.521736Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <p style=\"padding:15px; background-color:chocolate; font-family:arial; font-weight:bold; color:white; font-size:100%; letter-spacing: 2px; text-align:left; border-radius: 10px 10px\">12/ Submission</p>","metadata":{}},{"cell_type":"code","source":"voting_regressor.fit(X,y)\npreds = voting_regressor.predict(test)\npreds = convert2(preds)\npreds = pd.Series(preds)\npreds.index = test.index\npreds.to_csv('submission.csv')\npreds","metadata":{"execution":{"iopub.status.busy":"2025-01-07T14:08:52.524200Z","iopub.execute_input":"2025-01-07T14:08:52.524579Z","iopub.status.idle":"2025-01-07T14:08:53.094251Z","shell.execute_reply.started":"2025-01-07T14:08:52.524532Z","shell.execute_reply":"2025-01-07T14:08:53.093172Z"},"trusted":true},"outputs":[],"execution_count":null}]}