{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30761,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## CMI 2024 - Kaggle Addiction?\n\nThe [CMI 2024 Problematic Internet Use](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use) competition(*) includes both tabular data (CSV file) and time series data (parquet files from a Wrist-worn accelerometer.) The target to learn (and predict) is the \"sii\" value (0, 1, 2 or 3) which is directly related to the PCIAT Total value (0--100); the [Questioning the Task](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535525) disscussion topic explores many aspects of the target.\n\nThe main challenges are i) dealing with the many missing CSV data values, and ii) finding a way to use the times series data to create additional features that help in the machine learning. Here are some bullet points of what I tried in this notebook:\n<ul>\n<li>Used XGB regression with y = PCIAT Total; its values were 'squished' to have gaps at the sii thresholds. </li>\n<li>Replaced missing Weights, Heights, and BMIs with functions of age.</li>\n<li>Made adjustments to some CSV values, like clipping between 0 and 99-th percentile.</li>\n<li>Flagged ids that have parquet files, 996 of them out of 2736 with sii values.</li>\n<li>Created and added \"Wrist-\" features based on the time series, for examples: percent active, bedtime, and a measure of enmo-times-light during weekend days. Required at least a week of good data so only 777 ids had these features.</li>\n<li>Looked at the correlation of the features with Total and with each other, keeping higher-correlation, non-redundant features. (Note that assuming the truth of: \"high correlation --implies--> maybe a useful feature\" does not mean assuming that its inverse is true: \"low correlation --implies--> not useful feature\".)</li> \n<li>Down selected to 14 CSV and 3 Wrist features that seemed most useful/important, and filled remaining NaN values with their mean or median.</li> \n<li>Did stratified (by parquet-sii) GSCV on 5 shuffled XGB models over a limited number of hyperparameters, most importantly the learning rate and number of estimators; weighted sample i by (50+Total[i])/150.</li>\n<li>Applied quantile adjustments to the model y values to match the known quantiles at the sii thresholds.</li>\n<li>Considered using two models: one to fit all 2736 samples with CSV features, and another to fit just the 777 that had valid Wrist features using CSV and Wrist features; did not get a two-model version to have an LB score greater than the one-model version.</li>\n<li>Decided which features to keep by paying a lot of attention to the eli5 Permutation Importance.</li>\n<li>Decided to use one model on all 2736 ids, using 9 CSV features and 1 Wrist feature, wkend_enmXlight.</li>\n<li>Looked at the correlations of the 20 individual PCIAT questions, Total, and the features. Questions 4, 7, and 12 standout as being much less correlated with the other questions.</li>\n</ul>\n\n<HR width=250>\n(*) \"Problematic Internet Use\" is an ironic theme for this competition given how 'addicated' I can get with Kaggle competitions, constantly tinkering with the code. My behavior is like a \"kaggle gambler\" forever in a loop: thinking up a new scheme, placing bets (coding), going broke (low CV/LB), deciding I'm done with it, then reading a discussion post or having an \"ah-ha\" thought that sends me back to the beginning of the loop.","metadata":{}},{"cell_type":"code","source":"# The usual things\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom time import time\n\n# heat map\nimport seaborn as sns\n\n# linear regression\nfrom sklearn import linear_model\n\n# XGBoost\nimport xgboost as xgb\n\n# Train-Test, etc.\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import KFold\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.metrics import make_scorer\nfrom sklearn.utils import shuffle as skl_shuffle\n\n# Fitting routine\nfrom scipy.optimize import curve_fit\n\n# To test importance of X features - Thanks to @gkitchen\nimport eli5\nfrom eli5.sklearn import PermutationImportance\n\nimport warnings\nwarnings.simplefilter('ignore')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-12-20T21:02:09.226760Z","iopub.execute_input":"2024-12-20T21:02:09.227647Z","iopub.status.idle":"2024-12-20T21:02:14.678881Z","shell.execute_reply.started":"2024-12-20T21:02:09.227574Z","shell.execute_reply":"2024-12-20T21:02:14.677595Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Some options\nUSE_PARQUET = True         # True = parquet-based Wrist- features are made available & used\nTWO_MODELS = False         # True = Use a 2nd model on just the any(996)/valid(777) parquet ids\n\nGSCV_EXPLORE = True        # Explore one 2D grid of parameters, e.g., learning rate and n_estim.s\nSHOW_STUFF = True          # True = show more plots/output than just key ones\n\nDELTA_SEEDS = 0            # An integer added to all seeds in the code to check seed-related variability.\n\n# Other options - left as-is, usually\nMODIFY_TOTAL = True        # Adjust the PCIAT Total values to have 'regress-ification' gaps at the sii boundaries. \nBLUR_AGE = False           # True = add a +/-0.4 blur to the ages (0.002 LB improvement with it off)\nFILL_NANS = True           # True = fill remaining NaNs with train means\nQUANTILE_SCALE = True      # True = apply quantile-intervals mapping to predicted Total\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Useful routines, roughly in order of use\n\n# Use medians at ages to fit Weight and Height equations  Got medians from:\n# for iage in range(5,21):  # copy these 3 lines to a cell\n#    print(train.loc[(train['Basic_Demos-Age'] > iage-1) & (train['Basic_Demos-Age'] < iage+1),\n#                          'Physical-Weight'].median())\n# Weight: use medians from 5 to 20, fit to linear eq\ndef weight_of_age(age):\n    return 8.38537*age + -1.21397\n# Height: use medians from 5 to 20, fit to quadratic eq\ndef height_of_age(age):\n    return -0.0809874*age**2 + 3.64542*age + 28.0763\n\ndef sii_from_total(totalvals):\n    sii_out = 3.0 + np.zeros(len(totalvals))\n    sii_out[totalvals < 79.5] = 2.0\n    sii_out[totalvals < 49.5] = 1.0\n    sii_out[totalvals < 30.5] = 0.0\n    return sii_out\n\ndef metric_from_siis(sii_pred, sii_real, verbose=1):\n    # Crosstab between sii_hat (validation) and real sii of validation \n    sii_crtab = pd.crosstab(sii_pred,sii_real)\n    margcol = sii_crtab.sum(axis=0)\n    margrow = sii_crtab.sum(axis=1)\n    sumsum = margcol.sum()\n    sum_EO = 0.0\n    sum_WO = 0.0\n    for irow in sii_crtab.index:\n        for icol in sii_crtab.columns:\n            erowcol = margcol[icol]*margrow[irow]/sumsum\n            ##print(irow,icol,sii_crtab.loc[irow,icol],erowcol)\n            sum_WO += (irow-icol)**2 * sii_crtab.loc[irow,icol]\n            sum_EO += (irow-icol)**2 * erowcol\n    if verbose > 0:\n        print(\"Sum of W O =\", sum_WO, \"(want low value)\" +\n              \"   Sum of E O = {:.1f}\".format(sum_EO))\n    return 1.0 - sum_WO/sum_EO\n\n# Make a GSCV scorer based on the metric\ndef qwk_score(tot_pred, tot_real):\n    sii_pred = sii_from_total(tot_pred)\n    sii_real = sii_from_total(tot_real)\n    metric_out = metric_from_siis(sii_pred, sii_real, verbose=0)\n    return metric_out\n\n# Create the scorer:\nqwk_scorer = make_scorer(qwk_score, greater_is_better=True)\n\ndef calc_test_sterr(df_in, n_folds):\n    # Used when looking at GSCV output dataframe.\n    # Ignore large systematic variations between the splits' scores\n    # to calculate a better estimate of the test sterr.\n    gscv_stats = df_in.describe()\n    std_split = 0.0\n    for isplit in range(n_folds):\n        std_split += (gscv_stats.loc['std','split'+str(isplit)+'_test_score'])**2\n    std_split = np.sqrt(std_split/n_folds)\n    # The sterr of the test means is then:\n    sterr_splits = std_split/np.sqrt(n_folds)\n    # Add as a sterr_test_score column\n    return sterr_splits\n\ndef adjust_intervals(y_pred_in, boundaries):\n    # Scale input intervals between: min, boundaries, max\n    # linearly to intervals between: 0, 30, 49, 79, 96\n    # Include a +/- gap in output values to show boundaries.\n    in_inters = [min(y_pred_in)] + boundaries + [max(y_pred_in)]\n    out_inters = [0.0, 30.0, 49.0, 79.0, 96.0]\n    y_pred_out = y_pred_in.copy()\n    for iint in range(4):\n        inlow = in_inters[iint]; inhigh = in_inters[iint + 1]\n        outlow = out_inters[iint]+3; outhigh = out_inters[iint + 1]-2\n        selint = (y_pred_in >= inlow) & (y_pred_in <= inhigh)\n        if len(selint) > 0:\n            y_pred_out[selint] = ((y_pred_in[selint] - inlow) / (inhigh - inlow)\n                             ) * (outhigh - outlow) + outlow\n    return y_pred_out\n\n# Draw a rectangle\ndef draw_rect(left, right, lower, upper, color):\n    plt.plot([left,right],[lower,lower],c=color)\n    plt.plot([right,right],[lower,upper],c=color)\n    plt.plot([right,left],[upper,upper],c=color)\n    plt.plot([left,left],[upper,lower],c=color)\n    return\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Read in and adjust the CSVs\nGenerally train and test are modified at the same time.","metadata":{}},{"cell_type":"code","source":"# Directory prefix for the data\nabove_dir = \"../input/child-mind-institute-problematic-internet-use/\"\n\n# Train\ntrain = pd.read_csv(above_dir+\"train.csv\")\n# Test\ntest = pd.read_csv(above_dir+\"test.csv\")\n\n# _ _ _ Train specific modifications:\n# Downselect train to just the rows with an sii (target) value\ntrain = train[train[\"sii\"] > -0.5]\n\n# Rename/shorten the train PCIAT-PCIAT_Total column\ntrain = train.rename(columns={\"PCIAT-PCIAT_Total\":\"Total\"})\nif MODIFY_TOTAL:\n    # Adjust the PCIAT Total values to provide gaps for doing \"regres-ification\"\n    # Total=0 Adjustment: Total=0 becomes -4\n    tot_zeros = train[\"Total\"] == 0\n    train.loc[tot_zeros, \"Total\"] = 0 # will shift to -5\n    # All sii=0 range: Add a gap around 30 -- Shift the sii==0 ones down additional amount\n    tot_sii0 = train[\"Total\"] < 31\n    train.loc[tot_sii0, \"Total\"] -= 5\n    # sii=1 range: Squish 31 to 49 to 36 to 46\n    tot_sii1 = (train[\"Total\"] > 30) & (train[\"Total\"] < 50)\n    train.loc[tot_sii1, \"Total\"] = (10/18)*(\n            train.loc[tot_sii1, \"Total\"] - 49) + 46\n    # sii=2 range: Squish the 50 to 79 values to 53 to 76\n    tot_sii2 = (train[\"Total\"] > 49) & (train[\"Total\"] < 80)\n    train.loc[tot_sii2, \"Total\"] = (23/29)*(\n            train.loc[tot_sii2, \"Total\"] - 79) + 76\n    # sii=3 range: Add 3 to these\n    tot_sii3 = train[\"Total\"] > 79\n    train.loc[tot_sii3, \"Total\"] += 3\n\n\n# _ _ _ Adjustments to Weight, Height, BMI, and Age\n# Replace 0 values in Physical Height, Weight, and BMI with NaNs\nfor this_col in [\"Physical-BMI\", \"Physical-Height\", \"Physical-Weight\"]:\n    train.loc[train[this_col] < 5.0, this_col] = np.nan\n    test.loc[test[this_col] < 5.0, this_col] = np.nan\n\n# Add +/-0.4 blur to Age?\nif BLUR_AGE:\n    np.random.seed(17 + DELTA_SEEDS)\n    for this_col in [\"Basic_Demos-Age\"]:\n        train[this_col] += 0.8*np.random.rand(len(train)) - 0.4\n        test[this_col] += 0.8*np.random.rand(len(test)) - 0.4\n\nif SHOW_STUFF:\n    plot_ages = np.linspace(4.0,20.0,100)\n    plt.figure(figsize=(5,3))\n    plt.scatter(train['Basic_Demos-Age'], train['Physical-Weight'], s=10, alpha=0.2)\n    weight_equ = weight_of_age(plot_ages)\n    plt.plot(plot_ages,weight_equ,\".r\",markersize=2)\n    plt.title('Weight vs Age with linear median approx')\n    plt.savefig('Weight_vs_Age.png')\n    plt.show()\n    #\n    plt.figure(figsize=(5,3))\n    plt.scatter(train['Basic_Demos-Age'], train['Physical-Height'], s=10, alpha=0.2)\n    height_equ = height_of_age(plot_ages)\n    plt.plot(plot_ages,height_equ,\".r\",markersize=2)\n    plt.title('Height vs Age with quadratic median approx')\n    plt.savefig('Height_vs_Age.png')\n    plt.show()\n    \n# Use model to fill ones without Weight\nnoWsel = train['Physical-Weight'].isnull()\ntrain.loc[noWsel,'Physical-Weight'] = weight_of_age(\n                        train.loc[noWsel,'Basic_Demos-Age'])\n# And for test too\nnoWsel = test['Physical-Weight'].isnull()\ntest.loc[noWsel,'Physical-Weight'] = weight_of_age(\n                        test.loc[noWsel,'Basic_Demos-Age'])\n\n# Use model to fill ones without Height\nnoWsel = train['Physical-Height'].isnull()\ntrain.loc[noWsel,'Physical-Height'] = height_of_age(\n                        train.loc[noWsel,'Basic_Demos-Age'])\n# And for test too\nnoWsel = test['Physical-Height'].isnull()\ntest.loc[noWsel,'Physical-Height'] = height_of_age(\n                        test.loc[noWsel,'Basic_Demos-Age'])\n\n# Fill missing BMIs from the equation\nnoWsel = train['Physical-BMI'].isnull()\ntrain.loc[noWsel,'Physical-BMI'] = 703*(train.loc[noWsel,'Physical-Weight'] /\n                            train.loc[noWsel,'Physical-Height']**2)\n# And for test too\nnoWsel = test['Physical-BMI'].isnull()\ntest.loc[noWsel,'Physical-BMI'] = 703*(test.loc[noWsel,'Physical-Weight'] /\n                            test.loc[noWsel,'Physical-Height']**2)\n\n# Replace W[H]eight with (W[H]eight - w[h]eight_equ(Age)) / w[h]eight_equ(Age):\nif True:\n    equ_values = weight_of_age(train['Basic_Demos-Age'])\n    train['Physical-Weight'] = (train['Physical-Weight'] - equ_values)/equ_values\n    equ_values = weight_of_age(test['Basic_Demos-Age'])\n    test['Physical-Weight'] = (test['Physical-Weight'] - equ_values)/equ_values\n    #\n    equ_values = height_of_age(train['Basic_Demos-Age'])\n    train['Physical-Height'] = (train['Physical-Height'] - equ_values)/equ_values\n    equ_values = height_of_age(test['Basic_Demos-Age'])\n    test['Physical-Height'] = (test['Physical-Height'] - equ_values)/equ_values\n    \n# Continue in next cell...","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# _ _ _ \n# Make Feature modifications and additions to Train and Test together\n# Test has same columns, not including 22 PCIATs (54-75) and sii (81)\n\n# Shorten name\ntrain = train.rename(columns={\"PreInt_EduHx-computerinternet_hoursday\":\n                              \"PreInt_EduHx_hoursday\"})\ntest = test.rename(columns={\"PreInt_EduHx-computerinternet_hoursday\":\n                              \"PreInt_EduHx_hoursday\"})\n# Fill the few NaNs with 0\ntrain[\"PreInt_EduHx_hoursday\"] = train[\"PreInt_EduHx_hoursday\"].fillna(0.0)\ntest[\"PreInt_EduHx_hoursday\"] = test[\"PreInt_EduHx_hoursday\"].fillna(0.0)\n\n# Most BIAs and some FGCs have outlier issues... Clip these between 0 and 99%-tile\nclip_cols = [\"BIA-BIA_FFMI\",\"BIA-BIA_FMI\",\"BIA-BIA_LST\",\n                \"BIA-BIA_DEE\",\"BIA-BIA_ICW\",\"BIA-BIA_SMM\",\n                 \"FGC-FGC_CU\",\"FGC-FGC_GSND\",\"FGC-FGC_GSD\"]\nfor this_col in clip_cols:\n    clip99 = train[this_col].quantile(0.99)\n    train[this_col] = train[this_col].clip(0.0,clip99)\n    test[this_col] = test[this_col].clip(0.0,clip99)\n    \n# Make PAQ-PAQ_Total by merging PAQ_A-PAQ_A_Total and PAQ_C-PAQ_C_Total; leave NaNs\ntrain[\"PAQ-PAQ_Total\"] = train[\"PAQ_C-PAQ_C_Total\"]\npaq_valid = train[\"PAQ_A-PAQ_A_Total\"] > -100\ntrain.loc[paq_valid,\"PAQ-PAQ_Total\"] = train.loc[paq_valid,\"PAQ_A-PAQ_A_Total\"]\n# Same thing for the PAQ Season\ntrain[\"PAQ-PAQ_Season\"] = train[\"PAQ_C-Season\"]\npaq_valid = train[\"PAQ_A-Season\"].notnull()\ntrain.loc[paq_valid,\"PAQ-PAQ_Season\"] = train.loc[paq_valid,\"PAQ_A-Season\"]\n# Do the same things to Test\ntest[\"PAQ-PAQ_Total\"] = test[\"PAQ_C-PAQ_C_Total\"]\npaq_valid = test[\"PAQ_A-PAQ_A_Total\"] > -100\ntest.loc[paq_valid,\"PAQ-PAQ_Total\"] = test.loc[paq_valid,\"PAQ_A-PAQ_A_Total\"]\ntest[\"PAQ-PAQ_Season\"] = test[\"PAQ_C-Season\"]\npaq_valid = test[\"PAQ_A-Season\"].notnull()\ntest.loc[paq_valid,\"PAQ-PAQ_Season\"] = test.loc[paq_valid,\"PAQ_A-Season\"]\n\n# Categorical predictors - all are seasons\n# From: https://www.kaggle.com/code/gkitchen/introduction-to-problematic-internet-use\ncat_cols = list(train.columns[[1,4,6,14,18,33,50,52,76,79,83]]) # incl the new PAQ Season\n# Assign numbers in increasing amount of light/temperature, NaNs are set to the middle\nfor season in cat_cols:\n    train[season] = train[season].fillna(2.5)\n    train[season] = train[season].replace({'Spring':3, 'Summer':4, 'Fall':2, 'Winter':1})\n    test[season] = test[season].fillna(2.5)\n    test[season] = test[season].replace({'Spring':3, 'Summer':4, 'Fall':2, 'Winter':1})\n\n# Done with csv features, show the df\ntrain","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Look at the Targets: sii and Modified Total","metadata":{}},{"cell_type":"code","source":"# Total and sii are the targets\nmin_train_total = train['Total'].min()\n\nif SHOW_STUFF:\n    train.plot.scatter(\"Total\",\"sii\",figsize=(6,3),alpha=0.1)\n    plt.xlim(min_train_total-6, 100)\n    plt.title(\"Sii class vs PCIAT Total (modified to have gaps)\")\n    plt.show()\n\n    print(\"Boundaries are  0:<=30; 1:<=49; 2:<=79;  3:>=80\")\n    print(\"Number of samples in each sii bin:\",\n              train[\"sii\"].value_counts().values)\n# Number of samples in each sii bin: 1594  730  378   34\n# Bin boundaries as quantiles: 0.5826, 0.8494, 0.9876","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Look at the distribution of Total -- All that have values.\npciat_valcnts = train[\"Total\"].value_counts()\nprint(\"\\n     The mode is at Total =\", pciat_valcnts.index[0],\n                      \"with\", pciat_valcnts.values[0], \"counts.\")\n\nplt.figure(figsize=(6,3))\nplt.scatter(pciat_valcnts.index,pciat_valcnts.values)\nfor bound in [30.5, 49.5, 79.5]:\n    plt.plot([bound,bound],[0,100],c=\"gray\")\nplt.title(\"Distribution of the modified PCIAT Total\")\nplt.ylim(0.0,80.0); plt.ylabel(\"Number of Values, \"+\n                    str(pciat_valcnts.values[0])+\" are in mode.\")\nplt.xlim(min_train_total-2,100.0); plt.xlabel(\"PCIAT Total\")\nplt.savefig('Modified_Total_dist.png')\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Flag ids having Parquet Data Files","metadata":{}},{"cell_type":"code","source":"def read_parquet_file(this_id, prq_source):\n    # Read the parquet file\n    above_dir = \"../input/child-mind-institute-problematic-internet-use/\"\n    prqdf = pd.read_parquet(above_dir + 'series_'+prq_source+\n                        '.parquet/id='+this_id+'/part-0.parquet')\n    return prqdf","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if True:\n    # Even if not processing/using the parquet features,\n    # Find out and flag which ids have parquet data.\n    \n    # Train has 996 ids with parquet files, out of 2736 with sii.\n    prq_source = \"train\"  # train or test\n    # Add this column: 0.0 if not parquet data, 1.0 if parquet data\n    train['Wrist-file_days'] = 0.0\n    print(\"\\nGoing through \"+prq_source+\" parquet files ...\")\n    # Go through the rows\n    t_start = time()\n    n_done = 0\n    for this_row in train.loc[0: ].index:\n        this_id = train.loc[this_row,\"id\"]\n        try:\n            prqdf = read_parquet_file(this_id, prq_source)\n            got_it = True\n        except:\n            got_it = False\n        if got_it:\n            # Note that it is there and count number found\n            train.loc[this_row,\"Wrist-file_days\"] = 1.0\n            n_done += 1   # Monitor progress\n            if (n_done % 50) == 0:\n                print(\" ...\",n_done,\"have parquet so far ...\")\n    print(40*\" \"+\"... took {:.2f} seconds\".format(time() - t_start))\n    print(\"{} of the \".format(n_done)+prq_source+\" ids have parquet data.\\n\")\n\n    # Same for Test\n    prq_source = \"test\"  # train or test\n    # Add this column: 0.0 if not parquet data, 1.0 if parquet data\n    test['Wrist-file_days'] = 0.0\n    print(\"\\nGoing through \"+prq_source+\" parquet files ...\")\n    # Go through the rows\n    t_start = time()\n    n_done = 0\n    for this_row in test.loc[0: ].index:\n        this_id = test.loc[this_row,\"id\"]\n        try:\n            prqdf = read_parquet_file(this_id, prq_source)\n            got_it = True\n        except:\n            got_it = False\n        if got_it:\n            # Note that it is there and count number found\n            test.loc[this_row,\"Wrist-file_days\"] = 1.0\n            n_done += 1   # Monitor progress\n            if (n_done % 50) == 0:\n                print(\" ...\",n_done,\"have parquet so far ...\")\n    print(40*\" \"+\"... took {:.2f} seconds\".format(time() - t_start))\n    print(\"{} of the \".format(n_done)+prq_source+\" ids have parquet data.\\n\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Routines to process parquet data","metadata":{}},{"cell_type":"code","source":"# Actigraphy - Objective measure of ecological physical activity with a biotracker.\n#     Column                   Dtype  \n# ---  ------                  -----  \n#  0   step                    uint32 \n#  1   X                       float32\n#  2   Y                       float32\n#  3   Z                       float32\n#  4   enmo                    float64\n#  5   anglez                  float32\n#  6   non-wear_flag           float32\n#  7   light                   float32\n#  8   battery_voltage         float32\n#  9   time_of_day             int64  \n#  10  weekday                 int8   \n#  11  quarter                 int8   \n#  12  relative_date_PCIAT     float32\n\n# Columns/series added by parquet_feats_prune()\n# ---  ------                  ----- \n#  13  hour_of_day             float64\n#  14  time_in_days            float64\n#  15  enmoMed                 float64\n#  16  enmoMedLog              float64\n#  17  anglezMed               float64\n#  18  lightLog                float32\n#  19  lightLogMed             float64\n#  20  enmoStd                 float64\n#  21  anglezStd               float64\n#  22  lightLogStd             float64\n#  23  enmoMLSmooth\n#  24  lightLMSmooth\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Add additional series/columns to a parquet series dataframe.\n# Also down-sample it to wind_width x 5s bins, e.g., 10x12 --> 10 min samples.\n# Add further smoothed versions of enmo and light.\ndef parquet_feats_prune(this_series, wind_width, smoo_width):\n    \n    # Median of enmo in the window (wind_width = 12 * #minutes)\n    # Clip enmo at the 99.5% value\n    enmo995 = this_series[\"enmo\"].quantile(0.995)\n    this_series['enmo'] = this_series.enmo.clip(0.0,enmo995)\n    this_series['enmoMed'] = this_series.enmo.rolling(\n                            wind_width+1, center=True, min_periods=1).median()##.bfill()\n    this_series['enmoMedLog'] = 2.0 + np.log10(0.01 + this_series['enmoMed'])\n    \n    # Median of anglez in the window\n    this_series['anglezMed'] = this_series.anglez.rolling(\n                            wind_width+1, center=True, min_periods=1).median()##.bfill()\n    # Median of log light in the window\n    this_series['lightLog'] = 0.0 + np.log10(1.0 + this_series['light'])\n    this_series['lightLogMed'] = this_series.lightLog.rolling(\n                            wind_width+1, center=True, min_periods=1).median()##.bfill()\n    \n    # Calculate the std of enmo in the window\n    this_series['enmoStd'] = this_series.enmo.rolling(\n                            wind_width+1, center=True, min_periods=3\n                            ).std()##.bfill()\n    # Calculate the std of anglez in the window\n    this_series['anglezStd'] = this_series.anglez.rolling(\n                            wind_width+1, center=True, min_periods=3\n                            ).std()##.bfill()\n    # Calculate the std of light in the window\n    this_series['lightLogStd'] = this_series.lightLog.rolling(\n                            wind_width+1, center=True, min_periods=3\n                            ).std()##.bfill()\n    \n    # Get the series down-sampled to once every window width\n    this_series = this_series.iloc[\n                range(int(wind_width/2), len(this_series), wind_width)].copy()\n    \n    # Make triangle-smoothed enmo (could be in parquet_feats_prune())\n    this_series['enmoMLSmooth'] = this_series.enmoMedLog.rolling(\n                smoo_width+1, center=True, min_periods=3).mean()\n    this_series['enmoMLSmooth'] = this_series.enmoMLSmooth.rolling(\n                smoo_width+1, center=True, min_periods=3).mean()\n    # Make triangle-smoothed light\n    this_series['lightLMSmooth'] = this_series.lightLogMed.rolling(\n                smoo_width+1, center=True, min_periods=3).mean()\n    this_series['lightLMSmooth'] = this_series.lightLMSmooth.rolling(\n                smoo_width+1, center=True, min_periods=3).mean()\n\n    return this_series","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def func_sleep_wake(x, sleephr, awakeval, wakehr):\n    '''\n    # Fit Function used to fit bedtime and waketime.\n    # Returns a function of x (time in hours) with three regions:\n    #    -------\\________/------------\n    #  awakeval  sleepval    awakeval\n    # sleepval and a ramp transistion time are fixed in the code.\n    # Most useful if the x values go from awake time to awake time,\n    # e.g., 3 pm (x=15 hours) to the next 3 pm (x=15+24 hours).\n    '''\n    sleepval = 0.05\n    ramp = 4.0*(1/6)  # hours\n    out = (sleepval + (awakeval - sleepval) *\n                            (1.0 - np.clip((x - (sleephr - 0*ramp))/ramp, 0.0,1.0) *\n                                 np.clip(((wakehr + ramp/2) - x)/ramp, 0.0,1.0)))\n    return out","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def parquet_feats_into_csvdf(prqdf, csvdf=[], this_row=[]):\n    # When csvdf is provided, the prqdf is processed and\n    # Wrist- features are added to this_row of csvdf.\n    # Without a csvdf given, the parquet is processed and\n    # Wrist- feature values are printed and the processed parqdf is returned;\n    # this is useful for seeing details of the parquet data and feature creation for an id.\n    got_csvdf = len(csvdf) > 0\n    \n    # Some constants\n    window_width = 10*12    # 10 minutes\n    smooth_width = 6          # 1 hr\n    min_battery_voltage = 3549   # only use data above this battery level\n    # Define levels for idle and active\n    enmoStd_idle = 0.0055   # seems to roughly mimic the idle mode\n    # Conversion: enmoMedLog = 2.0 + log10(0.01 + enmoMed)\n    enmoML_active = 0.50    # fyi, 0.50 is equiv to enmoMed of 0.0216\n    \n    prqdf['hour_of_day'] = prqdf['time_of_day'] / 1e9 / 3600   # 0 to 23.9999 \n    # Add time in days\n    prqdf['time_in_days'] = prqdf['relative_date_PCIAT'] + (\n                            prqdf['hour_of_day']/24.0)\n    file_days = (prqdf['time_in_days'].max()-\n                                  prqdf['time_in_days'].min())\n    # ---84--- Wrist-file_days   -  save this value for any id that has a parquet file\n    if got_csvdf: csvdf.loc[this_row, \"Wrist-file_days\"] = file_days\n    else: print(\"\\nWrist-file_days:\", file_days)\n        \n    # ---85--- Wrist-idle_mode\n    # Are there non-wear_flag = 1 values? (if so then not in idle_sleep_mode)\n    idle_mode = 1  # in idle mode\n    if sum(prqdf[\"non-wear_flag\"] > 0.5) > 0: idle_mode = 0\n    if not got_csvdf: print(\"Wrist-idle_mode:\", idle_mode)\n        \n    # Keep the battery_voltage > min_battery_voltage\n    prqdf = prqdf[prqdf[\"battery_voltage\"] > min_battery_voltage]\n    # Keep at most 31 days of data (allows 4 weeks)\n    prqdf = prqdf[prqdf[\"time_in_days\"] < (31.0 + prqdf[\"time_in_days\"].min())]\n    # Remove any large trailing gaps by removing the last 1/2 day equiv of 5s rows\n    prqdf = prqdf.iloc[0:(len(prqdf)-8640)]\n    \n    # Return if the remaining file is short\n    # ---86--- Wrist-battery_days\n    if len(prqdf) > 0:\n        battery_days = prqdf['time_in_days'].max() - prqdf['time_in_days'].min()\n        if not got_csvdf: print(\"Wrist-battery_days:\", battery_days)\n    else:  battery_days = 0.0\n    if battery_days < 8.0:\n        ####print(this_id,\"Rejected:     days<8\")\n        if got_csvdf:\n            return csvdf   # Not using, battery days will be kept at default of 0.0\n        else:\n            print(\"!!! Battery days < 8 -- no processing or features\")\n            return prqdf\n\n    # Mimic the idle_sleep_mode (for idle_mode=0 ones): keep sections with some activity\n    enmoStd_short = prqdf.enmo.rolling(\n                            12+1, center=True, min_periods=3).std()\n    enmoStd_1pct = enmoStd_short.quantile(0.01)\n    ##if not got_csvdf: print(\"enmoStd_short 1% quantile: {:.6f}\".format(enmoStd_1pct))\n    # Apply approximation to idle mode if either of these are true:\n    apply_idle = (idle_mode == 0) or (enmoStd_1pct < 0.000010)\n    if apply_idle:   # keep ones above the idle level, 1 min window\n        prqdf = prqdf[enmoStd_short > enmoStd_idle]\n        # Also, Remove any remaining non-wear flagged ones\n        prqdf = prqdf[prqdf[\"non-wear_flag\"] < 0.5]\n    post_idle_days = (prqdf['time_in_days'].max() - \n                                    prqdf['time_in_days'].min())\n    # Check for jumps > 5s in time_in_days  5s = 57.870370 micro days\n    delta_times = prqdf[\"time_in_days\"].diff()   # in days\n    missing_time = sum(delta_times[1:] - 57.870370e-6)  # in days\n    # Include the time lost on the ends (it's not in the delta-made missing):\n    all_missing_time = missing_time + battery_days - post_idle_days\n    if not got_csvdf: print(\"Battery days has total no-signal time =\",\n                                \"{:.4f} days\".format(all_missing_time))\n    # ---87--- Wrist-pct_signal\n    # Calculate the \"percent signal\" for the region available after idle mode applied.\n    percent_signal = 1.0 - missing_time/post_idle_days\n    if not got_csvdf: print('Wrist-pct_signal = {:.4f} in the {:.4f} days'.format(\n                            percent_signal, post_idle_days))\n    if (percent_signal < 0.1) or (post_idle_days < 7.0):\n        ####print(this_id,\"Rejected:                 >90% missing\")\n        if got_csvdf:\n            return csvdf   # too much missing time, this one's no good\n        else:\n            print(\"!!! too much missing time --- no processing or features \")\n            return prqdf\n    \n    # _ _ _\n    # Add/modify time series and down-sample them to one sample per wind_width\n    prqdf = parquet_feats_prune(prqdf, wind_width=window_width, smoo_width=smooth_width)\n    # _ _ _\n  \n    # For the feature averages below, use full days from midnight to midnight\n    day_start = int(prqdf[\"time_in_days\"].min() + 0.999)  # round up\n    day_stop = int(prqdf[\"time_in_days\"].max())   # round down\n    if day_stop - day_start < 7:     # This only remove 3 that had made it this far\n        ####print(this_id,\"Rejected:                                  <7 full days\")\n        if got_csvdf:\n            return csvdf   # battery days stays at 0, others at defaults\n        else:\n            print(\"!!! Less than 7 full days of useful data\")\n            return prqdf\n    # Down select to a whole number of days\n    prqdf = prqdf[ ((prqdf[\"time_in_days\"] > day_start) &\n                              (prqdf[\"time_in_days\"] < day_stop)) ]\n    # _ _ _\n    # Create features from the parquet series\n    \n    # ---88--- Wrist-ave_PCIAT_date, clipped to -2 to +6 months\n    ave_PCIAT_date = np.clip(prqdf['relative_date_PCIAT'].mean(),-60.0,180.0)\n    \n    # ---89, 90--- Wrist-bedtime, Wrist-sleep_hours\n    # Fit a ----\\___/---- function to determine average bedtime and waketime.\n    # (Could select just weekday sleep/wakes, e.g., Mon 3pm to Fri 3pm.)\n    # Sleep/wake times can vary, but person is almost always awake at 3pm:\n    # Have time go from 15.0 to 24+15 hours - just change the hours, not rearrange\n    wrap_hour = prqdf['hour_of_day'].copy()\n    sel_wrap = (prqdf['hour_of_day'] < 15.0)\n    wrap_hour[sel_wrap] += 24.0\n    wrap_enmo = prqdf['enmoMedLog'] # values and locations don't change\n    #\n    lower_bounds = [ 16.0,       0.2,   5.0+24.0]\n    upper_bounds = [  5.0+24.0,  3.0,  12.0+24.0]\n    try:\n        # curve_fit from scipy.optimize\n        bedtime, wake_level, waketime = curve_fit( func_sleep_wake, \n                    wrap_hour, wrap_enmo, sigma=0.05, nan_policy='raise',\n                    x_scale=[5,0.5,5],\n                            bounds=(lower_bounds, upper_bounds)  )[0]\n        sleep_hours = waketime - bedtime\n        good_sleep_fit = True\n    except:\n        # Catch a problem with the fitting - doesn't happen with train\n        print(20*\"* \",\" --> Oops! Sleep-Wake fit failed!  row=\",this_row)\n        bedtime = 21.0; wake_level=0.6; waketime=24.0+6.0; sleep_hours = 9.0\n        good_sleep_fit = False\n    \n    # Select sleeping and waking regions of the data\n    sel_sleeping = (wrap_hour > bedtime) & (wrap_hour < waketime)\n    sel_waking = (wrap_hour < bedtime) | (wrap_hour > waketime)\n    \n    # ---91--- Wrist-pct_active when awake\n    active_hours = 5.0*window_width*sum(\n                    prqdf.loc[sel_waking,\"enmoMedLog\"] > enmoML_active)/3600\n    percent_active = active_hours/((24.0 - sleep_hours)*(day_stop - day_start))\n    \n    # ---92--- Wrist-enmoMLSmean\n    # ---93--- Wrist-enmoMLSstd\n    # ---94--- Wrist-lightLMSmean\n    # ---95--- Wrist-lightLMSstd\n    # Means and stds of smoothed while awake:\n    enmoMLSmean = prqdf.loc[sel_waking,'enmoMLSmooth'].mean()\n    enmoMLSstd = prqdf.loc[sel_waking,'enmoMLSmooth'].std()\n    lightLMSmean = prqdf.loc[sel_waking,'lightLMSmooth'].mean()\n    lightLMSstd = prqdf.loc[sel_waking,'lightLMSmooth'].std()\n    \n    # ---96--- Wrist-enmoXlight\n    enmoXlight = (prqdf.loc[sel_waking,'enmoMLSmooth'] *\n                  prqdf.loc[sel_waking,'lightLMSmooth']).mean()\n    \n    # ---XX--- Tried making wkendEnmo - wkdayEnmo difference - not useful, corr=-0.019\n    # ---97--- Wrist-wkend_enmoXlight\n    weekend_day = (prqdf['weekday'] > 5.5) & sel_waking\n    wkend_enmoXlight = (prqdf.loc[weekend_day,'enmoMLSmooth'] *\n                          prqdf.loc[weekend_day,'lightLMSmooth']).mean()\n    # A few are missing weekends (?), use enmoXlight in those cases\n    if np.isnan(wkend_enmoXlight): wkend_enmoXlight = 1.0 * enmoXlight\n    \n    # _ _ _\n    # Save (or show) the created Wrist- features:  \n    if got_csvdf:   # fill the values at one time\n        csvdf.loc[this_row, [\"Wrist-idle_mode\",\"Wrist-battery_days\",\n                             \"Wrist-pct_signal\",\"Wrist-ave_PCIAT_date\",\n                             \"Wrist-bedtime\",\"Wrist-sleep_hours\",\n                             \"Wrist-pct_active\",\n                             \"Wrist-enmoMLSmean\",\"Wrist-enmoMLSstd\",\n                             \"Wrist-lightLMSmean\",\"Wrist-lightLMSstd\",\n                             \"Wrist-enmoXlight\", \"Wrist-wkend_enmoXlight\"\n                            ]] = [\n                                idle_mode, battery_days,\n                                percent_signal, ave_PCIAT_date,\n                                bedtime, sleep_hours,\n                                percent_active,\n                                enmoMLSmean, enmoMLSstd,\n                                lightLMSmean, lightLMSstd,\n                                enmoXlight, wkend_enmoXlight]\n    else:\n        print(\"Wrist-ave_PCIAT_date:\", ave_PCIAT_date)\n        print(\"Wrist-bedtime:\", bedtime, \"  (waketime = \", waketime - 24.0, \"+ 24)\")\n        print(\"Wrist-sleep_hours:\",sleep_hours)\n        print(\"Wrist-pct_active:\", percent_active)\n        print(\"Wrist-enmoMLSmean:\", enmoMLSmean)\n        print(\"Wrist-enmoMLSstd:\", enmoMLSstd)\n        print(\"Wrist-lightLMSmean:\", lightLMSmean)\n        print(\"Wrist-lightLMSstd:\", lightLMSstd)\n        print(\"Wrist-enmoXlight:\", enmoXlight) \n        print(\"Wrist-wkend_enmoXlight:\", wkend_enmoXlight)\n\n        # Plot the sleep-wake values\n        plt.figure(figsize=(8,3))\n        plt.plot(wrap_hour, wrap_enmo,'.',markersize=2)\n        plt.xlabel(\"Hour of Day  (wrapped at 3 pm)\"); plt.ylabel(\"enmoMedLog\")\n        if good_sleep_fit:\n            sleep_wake = func_sleep_wake(wrap_hour, bedtime, wake_level, waketime)\n            plt.plot(wrap_hour,sleep_wake,'.r',markersize=2)\n            plt.title(\"Enmo (10 min bins) with Awake/Asleep fit function\")\n        plt.savefig('Sleep_Awake_times.png')\n        plt.show()\n        # Plot the smoothed enmo series with mean+/-std lines\n        plt.figure(figsize=(8,3))\n        plt.plot(wrap_hour, prqdf['enmoMLSmooth'],'.',markersize=2)\n        plt.xlabel(\"Hour of Day  (wrapped at 3 pm)\"); plt.ylabel(\"enmoMLSmooth\")\n        plt.title(\"Smoothed (~ 1hr) Enmo with Awake Mean +/- Std lines\")\n        for use_xs in [ [15.0,bedtime], [waketime, 24+15] ]:\n            plt.plot(use_xs, 2*[enmoMLSmean],c='black',linewidth=2)\n            plt.plot(use_xs, 2*[enmoMLSmean + enmoMLSstd],c='orange',linewidth=2)\n            plt.plot(use_xs, 2*[enmoMLSmean - enmoMLSstd],c='orange',linewidth=2)\n        plt.show()\n    \n    # All done\n    if got_csvdf:\n        return csvdf\n    else:\n        return prqdf\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Fill Train and Test with the parquet features","metadata":{"execution":{"iopub.status.busy":"2024-11-07T01:54:34.020932Z","iopub.execute_input":"2024-11-07T01:54:34.022315Z","iopub.status.idle":"2024-11-07T01:54:34.039292Z","shell.execute_reply.started":"2024-11-07T01:54:34.022261Z","shell.execute_reply":"2024-11-07T01:54:34.037769Z"}}},{"cell_type":"code","source":"if USE_PARQUET:\n    # Go through the parquet files...   Train has 996 with parquet files (of 2736)\n    prq_source = \"train\"  # train or test\n    # Add some columns - set some defaults and set others explicitly to nan for filling later.\n    train['Wrist-file_days'] = 0.0;  train['Wrist-idle_mode'] = np.nan\n    train['Wrist-battery_days'] = 0.0;  train['Wrist-pct_signal'] = np.nan\n    train['Wrist-ave_PCIAT_date'] = np.nan\n    train['Wrist-bedtime'] = np.nan; train['Wrist-sleep_hours'] = np.nan\n    train[\"Wrist-pct_active\"] = np.nan\n    train['Wrist-enmoMLSmean'] = np.nan;  train['Wrist-enmoMLSstd'] = np.nan\n    train['Wrist-lightLMSmean'] = np.nan;  train['Wrist-lightLMSstd'] = np.nan\n    train['Wrist-enmoXlight'] = np.nan; train['Wrist-wkend_enmoXlight'] = np.nan\n\n    print(\"\\nGoing through \"+prq_source+\" parquet files ...\")\n    # Go through the rows\n    t_start = time()\n    n_done = 0\n    for this_row in train.loc[0: ].index:\n        this_id = train.loc[this_row,\"id\"]\n        try:\n            prqdf = read_parquet_file(this_id, prq_source)\n            got_it = True\n        except:\n            got_it = False\n        if got_it:\n            # Process it putting feature values into train\n            train = parquet_feats_into_csvdf(prqdf, train, this_row) \n            n_done += 1   # Monitor progress\n            if (n_done % 50) == 0:\n                print(\" ...\",n_done,\"processed ...\")\n\n    print(40*\" \"+\"... took {:.2f} seconds\".format(time() - t_start))\n\n    num_with_parquet = sum(train[\"Wrist-file_days\"] > 0.0)\n    num_idle_on = sum(train[\"Wrist-idle_mode\"] == 1)\n    print(\"Number of ids with parquet files:\", num_with_parquet)\n    print(\"Number of ids with idle_sleep_mode=ON:\", num_idle_on)\n    print(\"Number of ids with all Wrist- vals:\", sum(train[\"Wrist-pct_signal\"] > 0.0))\n\n    # Save this train df to a csv file:\n    train.to_csv(\"train_processed.csv\")\n\n# fyi, Checked if the sii categories are independent of the idle_sleep_mode -- YES:\n# The Chi-square statistic, 5.40, for 3 dof is not significant at 10% significance.","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if USE_PARQUET:\n    # Test: just copy/paste the above cell changing train to test...\n    prq_source = \"test\"  # train or test\n    # Add some columns - set some defaults and set others explicitly to nan for filling later.\n    test['Wrist-file_days'] = 0.0;  test['Wrist-idle_mode'] = np.nan\n    test['Wrist-battery_days'] = 0.0;  test['Wrist-pct_signal'] = np.nan\n    test['Wrist-ave_PCIAT_date'] = np.nan\n    test['Wrist-bedtime'] = np.nan; test['Wrist-sleep_hours'] = np.nan\n    test[\"Wrist-pct_active\"] = np.nan\n    test['Wrist-enmoMLSmean'] = np.nan;  test['Wrist-enmoMLSstd'] = np.nan\n    test['Wrist-lightLMSmean'] = np.nan;  test['Wrist-lightLMSstd'] = np.nan\n    test['Wrist-enmoXlight'] = np.nan; test['Wrist-wkend_enmoXlight'] = np.nan\n    \n    print(\"\\nGoing through \"+prq_source+\" parquet files ...\")\n    # Go through the rows\n    t_start = time()\n    n_done = 0\n    for this_row in test.loc[0: ].index:\n        this_id = test.loc[this_row,\"id\"]\n        try:\n            prqdf = read_parquet_file(this_id, prq_source)\n            got_it = True\n        except:\n            got_it = False\n        if got_it:\n            # Process it putting feature values into train\n            test = parquet_feats_into_csvdf(prqdf, test, this_row) \n            n_done += 1   # Monitor progress\n            if (n_done % 50) == 0:\n                print(\" ...\",n_done,\"processed ...\")\n\n    print(40*\" \"+\"... took {:.2f} seconds\".format(time() - t_start))\n\n    print(\"Number of ids with parquet files:\", sum(test[\"Wrist-file_days\"] > 0.0))\n    print(\"Number of ids with idle_sleep_mode=ON:\", sum(test[\"Wrist-idle_mode\"] == 1))\n    print(\"Number of ids with all Wrist- vals:\", sum(test[\"Wrist-pct_signal\"] > 0.0))\n\n    # Save this test df to a csv file:\n    test.to_csv(\"test_processed.csv\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if USE_PARQUET:\n    # List the wrist columns and how many non-null values\n    wrist_cols_indices = [84,85,86,87,88,89,90,91,92,93,94,95,96,97]\n    train[train.columns[wrist_cols_indices]].info()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if USE_PARQUET and SHOW_STUFF:\n    # Scatter plots of Total vs the Wrist features\n    prq_select = (train[\"Wrist-battery_days\"] > 0)\n    for this_col in train.columns[wrist_cols_indices]:\n        this_title = \"Total vs \" + this_col\n        if this_col == 'Wrist-idle_mode':\n            this_title = (\"Total vs Wrist Idle Sleep Mode \" +\n                    \"({} have Idle Mode = 1)\".format(num_idle_on))\n        train[prq_select].plot.scatter(this_col,\"Total\", figsize=(6,3),\n                                c=\"#0000FF20\", title=this_title)\n        plt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Look at time series in a selected parquet file","metadata":{}},{"cell_type":"code","source":"# Find ids with particular properties to then look at one in detail\nif USE_PARQUET:\n    prq_select = ((train[\"Wrist-pct_active\"] > 0.3) &    # selects valid ones\n              (train[\"Total\"] > 60)  &  # Range of Total\n              (train[\"Wrist-bedtime\"] > 22) &       #\n              (train[\"Basic_Demos-Age\"] >11.5))\n    print(train.loc[prq_select, [\"id\",\"Wrist-idle_mode\",\"Wrist-pct_active\",\n                       \"Wrist-sleep_hours\",\"Wrist-bedtime\",\"Total\",\"sii\"]])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Pick a train id to look at...\n\n# High sii\n#           id     Wrist-idle_mode  Wrist-pct_active  Wrist-sleep_hours  Wrist-bedtime    Total  sii \n# 3460  df556fd2        0.0            0.319515          7.009982           22.330500      88.0  3.0\nthis_id = 'df556fd2'\n\n# Antonina's id, https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-actigraphy-data-eda\n##this_id = '0417c91e'\n\n\nif USE_PARQUET:\n    # _ _ _ Read and process the selected parquet file _ _ _\n    # Read the parquet file\n    prq_source = 'train'\n    prqdf = read_parquet_file(this_id, prq_source)\n    # Process it, verbosely, and\n    # Returns the parquet data with added-features and down sampled  \n    prqdf = parquet_feats_into_csvdf(prqdf)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Look at the various time series in the parqdf\nif USE_PARQUET and SHOW_STUFF:\n    show_cols = [\"weekday\", \"battery_voltage\",\n                 \"enmo\", \"enmoMedLog\", \"enmoStd\", \"enmoMLSmooth\",\n                 \"light\", \"lightLogMed\", \"lightLogStd\", 'lightLMSmooth']\n                 # ignoring these, not informative? \"anglez\", \"anglezMed\", \"anglezStd\", ]\n    # Color-code by something\n    prqdf['plot_color'] = \"#0000FF60\" # blue w/alpha\n    prqdf.loc[(prqdf['weekday'] > 5.5),'plot_color'] = \"#FF000060\"  # red w/alpha\n    print(\"\\n          Color-coded by M,T,W,Th,F (blue) and Sat,Sun (red)\\n\")\n    for this_col in show_cols:\n        if True:  # True = show plots vs Hour of Day\n            prqdf.plot.scatter(\"hour_of_day\",this_col,s=2,c=\"plot_color\",\n                            figsize=(9,3)); plt.xlabel(\"Hour of Day\")\n            use_xs = [0.0,24.0]\n        else:\n            prqdf.plot.scatter(\"time_in_days\",this_col,s=2,c=\"#0000FF60\",\n                       figsize=(9,3)); plt.xlabel(\"Time (days)\")\n            use_xs = [prqdf[\"time_in_days\"].min(), prqdf[\"time_in_days\"].max()]\n        # Show levels used in making features:\n        if this_col == \"enmoMedLog\":\n            plt.plot(use_xs,2*[0.35],c='lime')    # enmoML_active = 0.35\n        if this_col == \"enmoStd\":\n            plt.plot(use_xs,2*[0.0055],c='lime')  # enmoStd_idle = 0.0055\n        plt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Features and their Correlation with PCIAT Total","metadata":{}},{"cell_type":"code","source":"# Show the columns and their number of non-null values:\n##print(train.info())","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# All columns in the train df:  non-null numbers are before a final (or not) NaN replacement\n\n# Demographics - Information about age and sex of participants.\n#   0  id                                      2736 non-null   object \n#   1  Basic_Demos-Enroll_Season               2736 non-null   object \n#   2  Basic_Demos-Age                         2736 non-null   float64\n#   3  Basic_Demos-Sex                         2736 non-null   int64  \n\n# CGAS - Children's Global Assessment Scale - Numeric scale used by mental health clinicians\n#        to rate the general functioning of youths under the age of 18.\n#   4  CGAS-Season                             2736 non-null   object \n#   5  CGAS-CGAS_Score                         2342 non-null   float64\n\n# Physical Measures - blood pressure, heart rate, height, weight and waist, and hip meas.s.\n#   6  Physical-Season                         2736 non-null   object \n#   7  Physical-BMI                            2736 non-null   float64\n#   8  Physical-Height                         2736 non-null   float64\n#   9  Physical-Weight                         2736 non-null   float64\n#  10  Physical-Waist_Circumference             483 non-null    float64\n#  11  Physical-Diastolic_BP                   2478 non-null   float64\n#  12  Physical-HeartRate                      2486 non-null   float64\n#  13  Physical-Systolic_BP                    2478 non-null   float64\n\n# FitnessGram Vitals and Treadmill - Measurements of cardiovascular fitness assessed using the NHANES treadmill protocol.\n#  14  Fitness_Endurance-Season                2736 non-null   object \n#  15  Fitness_Endurance-Max_Stage              731 non-null    float64\n#  16  Fitness_Endurance-Time_Mins              728 non-null    float64\n#  17  Fitness_Endurance-Time_Sec               728 non-null    float64\n\n# FGC - FitnessGram Child - Health related physical fitness assessment\n#   measuring five different parameters including\n#   aerobic capacity, muscular strength, muscular endurance, flexibility, body composition.\n#  18  FGC-Season                              2736 non-null   object \n#  19  FGC-FGC_CU                              1919 non-null   float64\n#  20  FGC-FGC_CU_Zone                         1884 non-null   float64\n#  21  FGC-FGC_GSND                             872 non-null    float64\n#  22  FGC-FGC_GSND_Zone                        864 non-null    float64\n#  23  FGC-FGC_GSD                              871 non-null    float64\n#  24  FGC-FGC_GSD_Zone                         864 non-null    float64\n#  25  FGC-FGC_PU                              1909 non-null   float64\n#  26  FGC-FGC_PU_Zone                         1875 non-null   float64\n#  27  FGC-FGC_SRL                             1911 non-null   float64\n#  28  FGC-FGC_SRL_Zone                        1877 non-null   float64\n#  29  FGC-FGC_SRR                             1913 non-null   float64\n#  30  FGC-FGC_SRR_Zone                        1879 non-null   float64\n#  31  FGC-FGC_TL                              1919 non-null   float64 \"Trunk Lift Total\"\n#  32  FGC-FGC_TL_Zone                         1885 non-null   float64\n\n# BIA - Bio-electric Impedance Analysis - Measure of key body composition elements:\n#       including BMI, fat, muscle, and water content.\n#  33  BIA-Season                              2736 non-null   object \n#  34  BIA-BIA_Activity_Level_num              1813 non-null   float64\n#  35  BIA-BIA_BMC                             1813 non-null   float64\n#  36  BIA-BIA_BMI                             1813 non-null   float64\n#  37  BIA-BIA_BMR                             1813 non-null   float64\n#  38  BIA-BIA_DEE                             1813 non-null   float64\n#  39  BIA-BIA_ECW                             1813 non-null   float64\n#  40  BIA-BIA_FFM                             1813 non-null   float64\n#  41  BIA-BIA_FFMI                            1813 non-null   float64\n#  42  BIA-BIA_FMI                             1813 non-null   float64\n#  43  BIA-BIA_Fat                             1813 non-null   float64\n#  44  BIA-BIA_Frame_num                       1813 non-null   float64\n#  45  BIA-BIA_ICW                             1813 non-null   float64\n#  46  BIA-BIA_LDM                             1813 non-null   float64\n#  47  BIA-BIA_LST                             1813 non-null   float64\n#  48  BIA-BIA_SMM                             1813 non-null   float64\n#  49  BIA-BIA_TBW                             1813 non-null   float64\n\n# PAQ - Physical Activity Questionnaire - Information about \n#       children's participation in vigorous activities over the last 7 days.\n#  50  PAQ_A-Season                            2736 non-null    object \n#  51  PAQ_A-PAQ_A_Total      Adolecents        363 non-null    float64\n#  52  PAQ_C-Season                            2736 non-null   object \n#  53  PAQ_C-PAQ_C_Total      Children         1440 non-null   float64\n\n# PCIAT - Parent-Child Internet Addiction Test - 20-item scale that\n#         measures characteristics and behaviors associated with compulsive use\n#         of the Internet including compulsivity, escapism, and dependency.\n#  54  PCIAT-Season                            2736 non-null   object \n#  55  PCIAT-PCIAT_01                          2733 non-null   float64\n#  56  PCIAT-PCIAT_02                          2734 non-null   float64\n#  57  PCIAT-PCIAT_03                          2731 non-null   float64\n#  58  PCIAT-PCIAT_04                          2731 non-null   float64\n#  59  PCIAT-PCIAT_05                          2729 non-null   float64\n#  60  PCIAT-PCIAT_06                          2732 non-null   float64\n#  61  PCIAT-PCIAT_07                          2729 non-null   float64\n#  62  PCIAT-PCIAT_08                          2730 non-null   float64\n#  63  PCIAT-PCIAT_09                          2730 non-null   float64\n#  64  PCIAT-PCIAT_10                          2733 non-null   float64\n#  65  PCIAT-PCIAT_11                          2734 non-null   float64\n#  66  PCIAT-PCIAT_12                          2731 non-null   float64\n#  67  PCIAT-PCIAT_13                          2729 non-null   float64\n#  68  PCIAT-PCIAT_14                          2732 non-null   float64\n#  69  PCIAT-PCIAT_15                          2730 non-null   float64\n#  70  PCIAT-PCIAT_16                          2728 non-null   float64\n#  71  PCIAT-PCIAT_17                          2725 non-null   float64\n#  72  PCIAT-PCIAT_18                          2728 non-null   float64\n#  73  PCIAT-PCIAT_19                          2730 non-null   float64\n#  74  PCIAT-PCIAT_20                          2733 non-null   float64\n#  75  PCIAT-PCIAT_Total                       2736 non-null   float64\n\n# SDS - Sleep Disturbance Scale - Scale to categorize sleep disorders in children.\n#  76  SDS-Season                              2736 non-null   object \n#  77  SDS-SDS_Total_Raw                       2527 non-null   float64\n#  78  SDS-SDS_Total_T                         2525 non-null   float64\n    \n# Internet Use - Number of hours of using computer/internet per day.\n#  79  PreInt_EduHx-Season                     2736 non-null   object \n#  80  PreInt_EduHx_hoursday                   2654 non-null   float64\n\n# The target: 0, 1, 2, 3 <- directly realted to PCIAT-PCIAT_Total.\n#  81  sii                                     2736 non-null   float64\n\n# Added these related columns:\n#  82  PAQ-PAQ_Total                 1802 non-null   float64\n#  83  PAQ-PAQ_Season                2736 non-null   float64\n\n# Column added to indicate the id has a parquet file\n#  84  Wrist-file_days               2736  day equiv. of #rows of data; 0 if no parquet data\n# Columns added to train & test by processing the parquet data:\n#  85  Wrist-idle_mode                777  1 = in sleep idle mode, 0 = not idle mode\n#  86  Wrist-battery_days            2736  day equiv. of #rows w/good battery, etc;    \"\n#  87  Wrist-pct_signal               777  fraction of rows with enmo values\n#  88  Wrist-ave_PCIAT_date           777  average of the PCIAT days values\n#  89  Wrist-bedtime                  777  average bedtime given between 15 (3pm) and 39 (next 3pm)\n#  90  Wrist-sleep_hours              777  average time asleep\n#  91  Wrist-pct_active               777  \n#  92  Wrist-enmoMLSmean              777   \n#  93  Wrist-enmoMLSstd               777  \n#  94  Wrist-lightLMSmean             777   \n#  95  Wrist-lightLMSstd              777  \n#  96  Wrist-enmoXlight               777\n#  97  Wrist-wkend_enmoXlight         772\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Numerical predictors\n# Not including:\n#  circumference(10), endurance(14-17),\n#  BIA-BIA_BMI(36) <-- Physical is similar and has more values\n#  the separate PAQ A,C values(50-53) --> Use combined ones (82,83)\nnumer_cols_indices = (list(range(1,9+1)) + [11,12,13] + list(range(18,35+1)) +\n                         list(range(37,49+1)) + [76,77,78,79,80,  82,83])\n\nif USE_PARQUET:\n    # Include the Wrist- potential features\n    numer_cols_indices += wrist_cols_indices\nelse:\n    # Include the Wrist-file_days by itself\n    numer_cols_indices += [84]\n\nnumer_cols = list(train.columns[ numer_cols_indices ])\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if False:\n    # Useful to see the Correlations of Numerics with another numeric\n    corr_with = \"Wrist-file_days\"\n    corr_out = train[numer_cols].corr()[corr_with]\n    # Sorted by abs(), high to low\n    corr_sortabs = corr_out.abs().sort_values(ascending=False)\n    print(corr_sortabs[0:35])\n    print(corr_sortabs[35:])\n    plt.figure(figsize=(6,3))\n    plt.hist(np.abs(corr_sortabs.values[1:]), bins=200)\n    plt.title(\"abs( Correlation of Numerics with \"+corr_with+\" )\")\n    plt.show()\n\n# Wrist-file_days: Highest correlation of Wearing Monitor among the CSV features used \n#   CGAS-CGAS_Score               0.084543  (NaNs included)\n#   FGC-FGC_CU                    0.054564  ( \" )\n#   PAQ-PAQ_Season                0.049813  ( \" )\n\n# Wrist-idle_mode: Highest correlation of being in Idle Mode among the CSV features \n#   CGAS-CGAS_Score               0.192425  (NaNs included)\n#   FGC-FGC_CU                    0.147676\n#   FGC-FGC_TL_Zone               0.114018","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Correlation of the numeric variables with Total\ncorr_with = \"Total\" # or \"PCIAT-PCIAT_04\" for with one question\ncorr_out = train[[corr_with]+numer_cols].corr()[corr_with]\n# Sorted by abs(), high to low\ncorr_sortabs = corr_out.abs().sort_values(ascending=False)\n# Get numeric cols and corr_cols in sorted-abs order:\nnumer_cols = list(corr_sortabs.keys()[1:])\ncorr_out = train[[corr_with]+numer_cols].corr()[corr_with]\n\nif SHOW_STUFF:\n    print(\"\\nCorrelation of Features with Total:   (in decreasing abs(corr) order)\\n\")\n    print(corr_out[0:35])\n    print(corr_out[35:])\n    print(\"\")\n    plt.figure(figsize=(6,3))\n    plt.hist(np.abs(corr_out.values[1:]), bins=200)\n    plt.title(\"abs( Correlation with Total )\")\n    plt.show()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Find the features with abs(correlation) above a threshold\n# and show a Scatter plot of Total vs each of those features\n# For large n, an s-sigma \"correlation exists\" threshhold is\n#    r > s/sqrt(n)\ns_sigma = 2.6  # ~ 99%\ncorr_cols = []\n# Add a color column - based on parquet data or not\ntrain[\"color\"] = \"blue\"\ntrain.loc[train['Wrist-file_days'] > 0, \"color\"] = \"red\"\nprint(\"\\n      Color-coded by ids with (red) and without (blue) parquet files.\\n\")\nfor this_var in numer_cols:\n    if (np.abs(corr_out[this_var]) > s_sigma/np.sqrt(np.sum(train[this_var].notnull()))):\n        corr_cols.append(this_var)\n        if SHOW_STUFF:\n            plt.figure(figsize=(5,3))\n            plt.scatter(train[this_var],train[\"Total\"],alpha=0.1,c=train[\"color\"])\n            plt.title(\"Total vs   \"+this_var+\"   corr={:.4f}\".format(corr_out[this_var]))\n            if this_var =='Wrist-wkend_enmoXlight':\n                plt.savefig('Total_vs_wkend_enmoXlight.png')\n            plt.show()\n\nprint(\"\\n      Color-coded by ids with (red) and without (blue) parquet files.\\n\")\nprint(\"\\nFound {} features with > {:.1f}-sigma correlations.\\n\".format(\n                    len(corr_cols), s_sigma))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Choose the Features for the Model(s)","metadata":{}},{"cell_type":"code","source":"# Choose the predictors to use\n\n# Try to be systematic: start with all the significant correlation columns:\npreds = corr_cols.copy()\n# and remove ones that are very correlated with others, etc.:\n# These are all very highly correlated except a bit less between DEE and FFMI\n# So keep just DEE and FFMI\n#   BIA-BIA_LST, BIA-BIA_ICW, BIA-BIA_SMM, BIA-BIA_DEE, BIA-BIA_FFMI\nremove_these = [\"BIA-BIA_LST\", \"BIA-BIA_ICW\", \"BIA-BIA_SMM\", \"BIA-BIA_DEE\", \"BIA-BIA_FFMI\"]\nfor this_col in remove_these:\n    preds.remove(this_col)\n# Keep just the T version:\npreds.remove(\"SDS-SDS_Total_Raw\")\n# FMI is very corrleated with BMI, but not identical, keep for now\n##preds.remove(\"BIA-BIA_FMI\")\n# GSD and GSND are basically identical, remove one:\npreds.remove(\"FGC-FGC_GSD\")\n# FGC PU Zone: very correlated with PU\npreds.remove(\"FGC-FGC_PU_Zone\")\n# FGC Zones: keep SRL, remove SRR\npreds.remove(\"FGC-FGC_SRR_Zone\")\n# Weight is correlated with BMI, but not identical, keep for now\n# FGC SRR and SRL very similar, keep the SRR (opposite from Zone)\npreds.remove(\"FGC-FGC_SRL\")\n#\n# Now, do model fits and look at Permutation Importance and remove one by one\n# Systolic and Diastolic have very low Permu Import\npreds.remove(\"Physical-Systolic_BP\")\npreds.remove(\"Physical-Diastolic_BP\")\n# BIA Activity Level has low Permu Import\npreds.remove(\"BIA-BIA_Activity_Level_num\")\n# BIA FMI has low Permu Import\npreds.remove(\"BIA-BIA_FMI\")\n# Very correlated with BMI and low Permu Import\npreds.remove(\"Physical-Weight\")\n# Not convincing in Permu Import\npreds.remove(\"BIA-BIA_Frame_num\")\n# No more 0 or negative Permu Imports\n\n# Remove some Wrist predictors (5) from all the corr selected ones (8)    \nif USE_PARQUET:\n    # Decide which of the 7 to not use\n    # enmoMLSmean and pct_active are very similar but not identical\n    # pct_active overlaps with pct_signal too, remove active\n    preds.remove('Wrist-enmoMLSmean')\n    ##preds.remove('Wrist-pct_active') #            <-- keep\n    # These two are very similar... Keep the weekend one\n    preds.remove('Wrist-enmoXlight')\n    ##preds.remove('Wrist-wkend_enmoXlight') #      <-- keep\n    # These two are very anti-correlated, keep bedtime\n    ##preds.remove('Wrist-bedtime') #               <-- keep\n    preds.remove('Wrist-sleep_hours')\n    # Overlap with pct_acitive, but active is like enmoMLSmean\n    preds.remove('Wrist-pct_signal')\n    # Just passed significance but low Perm Import\n    preds.remove('Wrist-idle_mode')\n    # These didn't pass the corr significance level\n    ##preds.append('Wrist-file_days')\n    ##preds.append('Wrist-battery_days')\n    ##preds.append('Wrist-ave_PCIAT_date') \n    ##preds.append('Wrist-enmoMLSstd')\n    ##preds.append('Wrist-lightLMSmean')\n    ##preds.append('Wrist-lightLMSstd')\n    \n\nif SHOW_STUFF or True:\n    # Look at correlations among all the selected predictors (not diluted w/filled NaNs)\n    # calc correlation matrices\n    corr_preds = train[['Total'] + preds].corr()\n\n    # Correlation Heatmap, code from: \n    #   https://www.kaggle.com/code/docxian/cmi-problematic-internet-use-visual-starter\n    plt.figure(figsize=(7.5,6))\n    sns.heatmap(corr_preds, annot=False, cmap='RdYlGn',\n            fmt='.2f', linecolor='black', linewidths=0.5,\n            vmin=-0.8, vmax=+0.8)\n    plt.title('Correlations among Selected Predictors and Total')\n    plt.savefig('Features_correlations.png')\n    plt.show()\n\n# See the numerical corr values:\n##corr_preds","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Fill any remaining NaN values with the mean? median? of each numeric column\n# _ Mean will produce a new category value for discrete variables.\n# For example, Wrist-idle mode has 0, 1, and NaNs; the NaNs will become a value like 0.33.\n# _ Median for a continuous/skewed variable might give a more \"typical\" value than the mean.\n# Wanna keep it simple too.\nif FILL_NANS:\n    # Fill the discrete columns using their means:\n    fill_w_means = ['BIA-BIA_Frame_num', 'FGC-FGC_SRL_Zone', 'FGC-FGC_SRR_Zone',\n                     'BIA-BIA_Activity_Level_num', 'PAQ-PAQ_Season', 'FGC-FGC_PU_Zone']\n    if USE_PARQUET:\n        fill_w_means.append('Wrist-idle_mode')\n    for this_col in fill_w_means:\n        this_mean = train[this_col].mean()\n        train[this_col] = train[this_col].fillna(this_mean)\n        # same fill value for Test\n        test[this_col] = test[this_col].fillna(this_mean)\n    # Fill all remaining (continuous) with their median\n    trainfills = train[numer_cols].median()\n    train.fillna(trainfills, inplace=True)\n    test.fillna(trainfills, inplace=True)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Scatter Matrix Comparison of most correlated CSV and Parquet features\n# Color-coding is by sii categories 0, 1, and 2.\nif USE_PARQUET and SHOW_STUFF and FILL_NANS:\n    print(\"\\nColor-coding is by sii categories 0-yellow, 1-bluish, and 2-red.\\n\")\n    # Add color to train\n    train[\"color\"] = \"yellow\"\n    train.loc[train[\"Total\"] > 30.5, \"color\"] = \"cyan\"\n    train.loc[train[\"Total\"] > 49.5, \"color\"] = \"red\" #\"magenta\"\n    # T/F (1/0) for which have valid Wrist values\n    row_has_wrist = np.array(train[\"Wrist-battery_days\"] > 0)\n    pd.plotting.scatter_matrix(train.loc[row_has_wrist,['Total', \n                    'Basic_Demos-Age', 'PreInt_EduHx_hoursday', 'SDS-SDS_Total_T',\n                    'Wrist-pct_active','Wrist-wkend_enmoXlight','Wrist-bedtime',\n                                        ]],\n                    figsize=(10,10), s=50, alpha=0.2, c=train.loc[row_has_wrist,\"color\"])\n    plt.savefig('Features_scatter_matrix.png')\n    plt.show()\n    print(\"\\nColor-coding is by sii categories 0-yellow, 1-bluish, and 2-red.\\n\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Create predictor sets for i) All-ids models and ii) w/parquet-ids models ","metadata":{}},{"cell_type":"code","source":"# Creating preds_csv for the All-ids model\n# Start with all available\npreds_csv = preds.copy()  # called \"csv\" but really means first model\n\nif True:  # use only 9 of the CSVs, remove:\n    remove_these = ['FGC-FGC_PU', 'FGC-FGC_GSND', 'FGC-FGC_TL',\n                    'FGC-FGC_SRR', 'CGAS-CGAS_Score']\n    for this_col in remove_these:\n        preds_csv.remove(this_col)\n\n# Make adjustments for the different possibilities:\nif not USE_PARQUET and not TWO_MODELS:\n    # keep selected (CSV) features\n    pass\nif not USE_PARQUET and TWO_MODELS:\n    # keep selected (CSV) features\n    pass\nif USE_PARQUET and not TWO_MODELS:\n    # One model and parquet Wrist features are included, only wkend_enmoXlight is useful?\n    preds_csv.remove('Wrist-pct_active')\n    preds_csv.remove('Wrist-bedtime')\nif USE_PARQUET and TWO_MODELS:\n    # 2 models, so don't include Wrist features here, they will be in the 2nd model\n    remove_these = ['Wrist-pct_active','Wrist-bedtime',\n                    'Wrist-wkend_enmoXlight']\n    for this_col in remove_these:\n        preds_csv.remove(this_col)\n        \nprint(\"\\n{} CSV Predictors for models using All ids:\\n\".format(len(preds_csv)),preds_csv)\n\n\n# Predictors for the 2nd, w/parquet-ids model\n# Create the preds_wprq even if TWO_MODELS is False, since it's only used if True.\nif USE_PARQUET:\n    # Start with all, includes the 3 Wrists\n    preds_wprq = preds.copy()\n    # v150...\n    # Remove some CSV features:\n    remove_these = ['FGC-FGC_PU', 'FGC-FGC_GSND', 'FGC-FGC_TL',\n                    'FGC-FGC_SRR', 'CGAS-CGAS_Score','FGC-FGC_SRL_Zone','Physical-Height']\n    for this_col in remove_these:\n        preds_wprq.remove(this_col)\n    # Remove one Wrist feature, leaving 2: 'Wrist-pct_active','Wrist-wkend_enmoXlight'\n    preds_wprq.remove('Wrist-bedtime')\n    print(\"\\n{} Predictors for the 2nd model (if used), fits only ids w/parquet features:\\n\".format(\n                                                        len(preds_wprq)),preds_wprq)\nelse:\n    # If there are no parquet Wrist features, start with all csv ones,\n    # remove some that have low Permu Import when fitting the 996 ids\n    preds_wprq = preds.copy()\n    preds_wprq.remove('FGC-FGC_PU')\n    preds_wprq.remove('FGC-FGC_SRL_Zone')\n    preds_wprq.remove('FGC-FGC_GSND')\n    preds_wprq.remove('FGC-FGC_SRR')\n    print(\"\\n{} Predictors for the 2nd model (if used), fits only ids w/parquet files:\\n\".format(\n                                                        len(preds_wprq)),preds_wprq)\n\n### - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -\n### Can skip GSCV in the next ~ 5 cells\n### - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Do GSCV on the XGBoost Regression models\nBasics of XGBoost at: \nhttps://www.kaggle.com/code/prashant111/xgboost-k-fold-cv-feature-importance","metadata":{}},{"cell_type":"code","source":"# Setup for GSCV\n\n# Get a combined X, y dataframe that includes all predictors, Total and sii.\n# Also, reset indices to 0 ... 2735\nXydf = train[[\"Total\",\"sii\"]+preds].copy().reset_index().drop(columns=['index'])\n\n# Which model to do GSCV on, select based on one or two models\ngscv_all = not TWO_MODELS\n\n# Do GSCV for the all-the-ids model\npreds_gscv = preds_csv.copy()\nif USE_PARQUET:\n    row_has_wrist = np.array(train[\"Wrist-battery_days\"] > 0)\nelse:\n    row_has_wrist = np.array(train[\"Wrist-file_days\"] > 0)\n# Select number of folds to be used:\nn_folds = 10\n\n# Or, for the w/parquet model\nif not gscv_all:\n    # Do GSCV for the 2nd model on any/valid parquet data\n    # Indices reset to 0 ... 776 (or ...995)\n    Xydf = train.loc[row_has_wrist,[\"Total\",\"sii\"]+preds_wprq].copy().reset_index().drop(columns=['index'])\n    preds_gscv = preds_wprq.copy()\n    # Set to all true\n    row_has_wrist = np.array(train.loc[row_has_wrist,\"Wrist-file_days\"] > 0)\n    # Select number of folds to be used:\n    n_folds = 10","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Get the X, y for GSCV\nX = Xydf[preds_gscv]\ny = Xydf[\"Total\"]\n# Make fit_weights to favor the higher Total values:\nfit_weights = (50.0+Xydf[\"Total\"])/150.0\n\n# Stratify the folds using sii\nsii_strata = Xydf[\"sii\"]\n# and stratify by has/not Wrist values, but keep sii=3 as one strata\nsii_strata = sii_strata + 2*row_has_wrist*(3 - sii_strata)\nprint(\"The modified sii stratification values by has/not parquet file or features:\\n\",\n              (pd.crosstab(sii_strata, row_has_wrist)).T)\n\n# Will do GSCVs with different shufflings with matching weights and siis.\n# So, don't need to shuffle the StratifiedKFolds below.\n# Have to reset the indices of the X,y, weights, and sii otherwise the shuffling is undone.\n# Yeesh what a mess...\nXA, yA, weightsA, siiA  = skl_shuffle(X, y, fit_weights, sii_strata,  random_state=17 + DELTA_SEEDS)\nXA = XA.reset_index().drop(columns=['index'])\nyA = np.array(yA); weightsA = np.array(weightsA); siiA = np.array(siiA)\n\nXB, yB, weightsB, siiB = skl_shuffle(X, y, fit_weights, sii_strata, random_state=34 + DELTA_SEEDS)\nXB = XB.reset_index().drop(columns=['index'])\nyB = np.array(yB); weightsB = np.array(weightsB); siiB = np.array(siiB)\n\nXC, yC, weightsC, siiC  = skl_shuffle(X, y, fit_weights, sii_strata,  random_state=51 + DELTA_SEEDS)\nXC = XC.reset_index().drop(columns=['index'])\nyC = np.array(yC); weightsC = np.array(weightsC); siiC = np.array(siiC)\n\nXD, yD, weightsD, siiD  = skl_shuffle(X, y, fit_weights, sii_strata,  random_state=68 + DELTA_SEEDS)\nXD = XD.reset_index().drop(columns=['index'])\nyD = np.array(yD); weightsD = np.array(weightsD); siiD = np.array(siiD)\n\nXE, yE, weightsE, siiE  = skl_shuffle(X, y, fit_weights, sii_strata,  random_state=85 + DELTA_SEEDS)\nXE = XE.reset_index().drop(columns=['index'])\nyE = np.array(yE); weightsE = np.array(weightsE); siiE = np.array(siiE)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Choose hyper-parameters for GSCV\n# v77+: start by using gkitchen's values:\n#           params = {'max_depth': 3, 'n_estimators': 59, 'learning_rate': 0.0733,\n#                        'subsample': 0.5968, 'colsample_bytree': 0.9124}\n# Other parameters that could be included:  lambda, min_child_weight\n\n# Explore parameter pairs...\nrate_n_ests_pair = True\n\n# Grid of 'learning_rate' and 'n_estimators'\nif rate_n_ests_pair and GSCV_EXPLORE:\n    if gscv_all:  # For all ids\n        xgb_grid = {  #\n            'colsample_bytree': [0.81,0.83],\n              'learning_rate':  list(np.linspace(0.045,0.081,13)),\n            'max_depth': [3],\n              'n_estimators': list(range(60, 140+1, 5)),\n            'subsample': [0.51,0.53]\n              }\n        param_combinations = 2 * 13 * 1 * 17 * 2; EXPLORE = True\n    else:   # For 996/777 w/parquet features\n        xgb_grid = {  #\n            'colsample_bytree':  [0.93,0.95],  \n              'learning_rate':  list(np.linspace(0.045, 0.081, 13)),\n            'max_depth': [2],\n              'n_estimators': list(range(55, 135+1, 5)),\n            'subsample':   [0.31,0.33]\n              }\n        param_combinations = 2 * 13 * 1 * 17 * 2; EXPLORE = True\n    # Names of the two params for the GSCV grid\n    xparam = 'param_learning_rate'\n    yparam = 'param_n_estimators'\n    ##gscv_scorer = qwk_scorer; scorer_str = \"QWK\"\n    gscv_scorer = None; scorer_str = \"R$^2$\"\n    \n# Grid of 'colsample_bytree' and 'subsample'\nif not rate_n_ests_pair and GSCV_EXPLORE:\n    if gscv_all:  # For All ids\n        xgb_grid = {\n              'colsample_bytree': list(np.linspace(0.74,1.00, 14)),\n            'learning_rate': [0.050,0.052],\n            'max_depth': [3],\n            'n_estimators':[114,116],\n              'subsample': list(np.linspace(0.40,0.75, 15)),\n              }\n    else:   # For 996/777 w/parquet features\n        xgb_grid = {\n              'colsample_bytree': list(np.linspace(0.74,1.00, 14)),\n            'learning_rate':  [0.047,0.049],\n            'max_depth': [2],\n            'n_estimators':   [129,131],\n              'subsample': list(np.linspace(0.30,0.65, 15)),\n              }\n    param_combinations = 14 * 2 * 1 * 2 * 15; EXPLORE = True\n    # Names of the two params for the GSCV grid\n    xparam = 'param_subsample'\n    yparam = 'param_colsample_bytree'\n    ##gscv_scorer = qwk_scorer; scorer_str = \"QWK\"\n    gscv_scorer = None; scorer_str = \"R$^2$\"\n    \n\n\n# If not exploring, use the specific/best values, average over some similar sets\nif not GSCV_EXPLORE:\n    if gscv_all:   # Model 1, All-ids\n        if USE_PARQUET and not TWO_MODELS:\n            # All-ids model only, with Wrist feature(s):\n            xgb_grid = {\n                'colsample_bytree': [0.77,0.78],\n                'learning_rate': [0.050,0.052],\n                'max_depth': [3],\n                'n_estimators': [114,116],\n                'subsample': [0.60,0.61],\n                    }\n        else:\n            # All-ids model has no Wrist features (aren't any or model 2 has them.)\n            xgb_grid = {\n                'colsample_bytree': [0.81,0.83],\n                'learning_rate': [0.059,0.061],\n                'max_depth': [3],\n                'n_estimators': [89,91],\n                'subsample': [0.51,0.53],\n                  }\n\n    else:   # Model 2 for the 777 w/parquet features (or 996)\n        xgb_grid = {\n            'colsample_bytree': [0.78,0.79],\n            'learning_rate': [0.047,0.049],\n            'max_depth': [2], \n            'n_estimators': [129,131],\n            'subsample': [0.53,0.54],\n              }\n    param_combinations = 2 * 2 * 1 * 2 * 2; EXPLORE = False\n    xparam = 'param_subsample'\n    yparam = 'param_learning_rate'\n    gscv_scorer = qwk_scorer; scorer_str = \"QWK\"\n\n\n# Set up the GSCVs\ngscvA = GridSearchCV(estimator=xgb.XGBRegressor(seed=1 + DELTA_SEEDS), param_grid=xgb_grid,\n            scoring=gscv_scorer,\n            n_jobs=-1,  # all CPUs\n            refit=True,\n            cv=StratifiedKFold(n_splits=n_folds, shuffle=False).split(XA, siiA),\n            verbose=1,   # 2 - one line for each fold fit\n            return_train_score=True)\ngscvB =  GridSearchCV(estimator=xgb.XGBRegressor(seed=3 + DELTA_SEEDS), param_grid=xgb_grid,\n            scoring=gscv_scorer,\n            n_jobs=-1,  # all CPUs\n            refit=True,\n            cv=StratifiedKFold(n_splits=n_folds, shuffle=False).split(XB, siiB),\n            verbose=1,   # 2 - one line for each fold fit\n            return_train_score=True)\ngscvC =  GridSearchCV(estimator=xgb.XGBRegressor(seed=9 + DELTA_SEEDS), param_grid=xgb_grid,\n            scoring=gscv_scorer,\n            n_jobs=-1,  # all CPUs\n            refit=True,\n            cv=StratifiedKFold(n_splits=n_folds, shuffle=False).split(XC, siiC),\n            verbose=1,   # 2 - one line for each fold fit\n            return_train_score=True)\ngscvD =  GridSearchCV(estimator=xgb.XGBRegressor(seed=27 + DELTA_SEEDS), param_grid=xgb_grid,\n            scoring=gscv_scorer,\n            n_jobs=-1,  # all CPUs\n            refit=True,\n            cv=StratifiedKFold(n_splits=n_folds, shuffle=False).split(XD, siiD),\n            verbose=1,   # 2 - one line for each fold fit\n            return_train_score=True)\ngscvE =  GridSearchCV(estimator=xgb.XGBRegressor(seed=81 + DELTA_SEEDS), param_grid=xgb_grid,\n            scoring=gscv_scorer,\n            n_jobs=-1,  # all CPUs\n            refit=True,\n            cv=StratifiedKFold(n_splits=n_folds, shuffle=False).split(XE, siiE),\n            verbose=1,   # 2 - one line for each fold fit\n            return_train_score=True)\n\n\n# _ _ _\nif True:\n    # Do the fits - including sample weights\n    print(\"\\nFitting {} data values;  \".format(len(y)) +\n          \"expected to take ~ {:.1f} minutes x 5.\\n\".format(\n              0.0358*n_folds/60.0 * param_combinations)) # rough calibration (depth=3)\n    t_start = time()\n    # First one\n    _dummyA = gscvA.fit(XA, yA, sample_weight=weightsA)\n    df_gscv = pd.DataFrame.from_dict(gscvA.cv_results_)\n    df_gscv['sterr_test_score'] = calc_test_sterr(df_gscv, n_folds)\n    xgbregrA = gscvA.best_estimator_\n    print(\"Params for the best \"+scorer_str+\"-Score {:.4f} :\\n\".format(\n            gscvA.best_score_), gscvA.best_params_)\n    print(\"GridSearchCV took {:.2f} cumul. minutes\\n\".format((time() - t_start)/60.0))\n    #\n    _dummyB = gscvB.fit(XB, yB, sample_weight=weightsB)\n    df_addon = pd.DataFrame.from_dict(gscvB.cv_results_)\n    df_addon['sterr_test_score'] = calc_test_sterr(df_addon, n_folds)\n    df_gscv = pd.concat([df_gscv, df_addon], ignore_index=True )\n    xgbregrB = gscvB.best_estimator_\n    print(\"Params for the best \"+scorer_str+\"-Score {:.4f} :\\n\".format(\n            gscvB.best_score_), gscvB.best_params_)\n    print(\"GridSearchCV took {:.2f} cumul. minutes\\n\".format((time() - t_start)/60.0))\n    #\n    _dummyC = gscvC.fit(XC, yC, sample_weight=weightsC)\n    df_addon = pd.DataFrame.from_dict(gscvC.cv_results_)\n    df_addon['sterr_test_score'] = calc_test_sterr(df_addon, n_folds)\n    df_gscv = pd.concat([df_gscv, df_addon], ignore_index=True )\n    xgbregrC = gscvC.best_estimator_\n    print(\"Params for the best \"+scorer_str+\"-Score {:.4f} :\\n\".format(\n            gscvC.best_score_), gscvC.best_params_)\n    print(\"GridSearchCV took {:.2f} cumul. minutes\\n\".format((time() - t_start)/60.0))\n    #\n    _dummyD = gscvD.fit(XD, yD, sample_weight=weightsD)\n    df_addon = pd.DataFrame.from_dict(gscvD.cv_results_)\n    df_addon['sterr_test_score'] = calc_test_sterr(df_addon, n_folds)\n    df_gscv = pd.concat([df_gscv, df_addon], ignore_index=True )\n    xgbregrD = gscvD.best_estimator_\n    print(\"Params for the best \"+scorer_str+\"-Score {:.4f} :\\n\".format(\n            gscvD.best_score_), gscvD.best_params_)\n    print(\"GridSearchCV took {:.2f} cumul. minutes\\n\".format((time() - t_start)/60.0))\n    #\n    _dummyE = gscvE.fit(XE, yE, sample_weight=weightsE)\n    df_addon = pd.DataFrame.from_dict(gscvE.cv_results_)\n    df_addon['sterr_test_score'] = calc_test_sterr(df_addon, n_folds)\n    df_gscv = pd.concat([df_gscv, df_addon], ignore_index=True )\n    xgbregrE = gscvE.best_estimator_\n    print(\"Params for the best \"+scorer_str+\"-Score {:.4f} :\\n\".format(\n            gscvE.best_score_), gscvE.best_params_)\n    print(\"GridSearchCV took {:.2f} cumul. minutes\\n\".format((time() - t_start)/60.0))\n    #\n    # _ _ _","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if True:\n    # Show the GSCV results; start by sorting them.\n    df_gscv = df_gscv.sort_values(by='mean_test_score',ascending=False)\n    # Put the original order in an \"index\" column:\n    if 'index' not in df_gscv.columns:  # Don't do again if re-running this cell.\n        df_gscv = df_gscv.reset_index()\n    # Save this to a csv file:\n    df_gscv.to_csv(\"gscv_results.csv\",index=False)\n\n    # Can look at the best/worse results, showing just the relevant columns:\n    pcols = (['index','mean_test_score','mean_train_score'] +\n             [ ix for ix in df_gscv.columns.values[5:5+len(xgb_grid)] ])\n    print(\"The top scores ({} in all):\".format(len(df_gscv)))\n    print(df_gscv[pcols].head(9))\n    ##print(df_gscv[pcols].tail(9))\n\n    # The overall mean and std of test scores\n    gscv_stats = df_gscv.describe()\n    if True: print(\"\\nThe median \" + scorer_str+\"-Score \" +\n                \"over the grid is {:.4f} +/- {:.4f}\\n\".format(\n                    gscv_stats.loc[\"50%\",\"mean_test_score\"],\n                        gscv_stats.loc[\"std\",\"mean_test_score\"]))\n    \n    # Show test score with error bars vs the GS original order (index)\n    # Include a line showing the median score level:\n    reference_score = df_gscv[\"mean_test_score\"].median()\n    df_gscv.plot.scatter('index','mean_test_score',\n                        yerr='sterr_test_score',\n                        figsize=(15,4),\n                        title='Test '+scorer_str+'-Score vs Grid Search index')\n    plt.plot([0,len(df_gscv)],2*[reference_score],alpha=0.5,c=\"yellow\")\n    plt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# If a 2-parameter grid exploration was done then show grouped results.\n# Group-average the results by the 2 grid parameters\nif GSCV_EXPLORE:\n    df_grouped = df_gscv[['mean_test_score', 'sterr_test_score', 'mean_train_score',\n                                  xparam, yparam]].groupby(\n                    [xparam,yparam],as_index=False).mean().sort_values(\n                                        by='mean_test_score',ascending=False)\n    # Adjust the sterrs for the averaging:\n    df_grouped['sterr_test_score'] /= np.sqrt(len(df_gscv)/len(df_grouped))\n\n    print(\"Results group-averaged by parameter combinations: ({} in all):\".format(len(df_grouped)))\n    print(df_grouped.head(6))\n\n    # Plot the group-average Test Score vs Train Score\n    df_grouped.plot.scatter('mean_train_score', 'mean_test_score',\n                        yerr='sterr_test_score',\n                        figsize=(6,3),\n                        title='Test '+scorer_str+'-Score vs Train '+scorer_str+'-Score')\n    # Quantilkes used for color coding in next plot too\n    quantLow = df_grouped[\"mean_test_score\"].quantile(0.50)\n    quantMid = df_grouped[\"mean_test_score\"].quantile(0.75)\n    quant90 = df_grouped[\"mean_test_score\"].quantile(0.90)\n    train_range = [df_grouped['mean_train_score'].min(),\n                   df_grouped['mean_train_score'].max()]\n    plt.plot(train_range,2*[quantLow],alpha=0.5,c=\"yellow\")\n    plt.plot(train_range,2*[quantMid],alpha=0.5,c=\"orange\")\n    plt.plot(train_range,2*[quant90],alpha=0.5,c=\"red\")\n    plt.savefig('Test_vs_Train_scores.png')\n    plt.show()\n\n    # Plot showing score-color for each param pair combination\n    df_grouped.loc[df_grouped[\"mean_test_score\"] <= quantLow, \"clr\"] = \"white\"\n    df_grouped.loc[df_grouped[\"mean_test_score\"] > quantLow, \"clr\"] = \"yellow\"\n    df_grouped.loc[df_grouped[\"mean_test_score\"] > quantMid, \"clr\"] = \"orange\"\n    df_grouped.loc[df_grouped[\"mean_test_score\"] > quant90, \"clr\"] = \"red\"\n    if True: \n        log_axes = False\n        df_grouped.plot.scatter(xparam , yparam, c=\"clr\", s=100,\n                    logx=log_axes, logy=log_axes, alpha=0.35, figsize=(5,5))\n        plt.savefig('param_param_region.png')\n        plt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -\n### End of GSCV\n### - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -","metadata":{}},{"cell_type":"markdown","source":"## Fit all ids with the \"All\" Model(s)\nThe parameters for the 5 XGB models are manually selected based on the GSCV output and then put into the code below .","metadata":{}},{"cell_type":"code","source":"# Get a combined X, y dataframe that includes all predictors, Total and sii.\n# Also, reset indices to 0 ... 2735\nXydf = train[[\"Total\",\"sii\"]+preds].copy().reset_index().drop(columns=['index'])\n\n# Make fit_weights to favor the higher Total values:\nfit_weights = (50.0+Xydf[\"Total\"])/150.0\n\n# Create row_has_wrist boolean here and used variously below\nif USE_PARQUET:   # 777 ids that have valid Wrist features\n    row_has_wrist = np.array(train[\"Wrist-battery_days\"] > 0)\nelse:             # 996 ids that have parquet files\n    row_has_wrist = np.array(train[\"Wrist-file_days\"] > 0)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Fit all the data using the GSCV best regressors\nX = Xydf[preds_csv]\ny = Xydf[\"Total\"]\nprint(\"\\nFitting {} data values.\".format(len(y)))\n\n# Keep the parameter lines shorter\nxp1='colsample_bytree'; xp2='learning_rate'; xp3='max_depth'\nxp4='n_estimators'; xp5='subsample'\n\n# Set default parameters for the All-ids models\nif True:\n    paramsA = {xp1: 0.82, xp2: 0.054, xp3: 3, xp4: 105, xp5: 0.52}\n    paramsB = {xp1: 0.82, xp2: 0.057, xp3: 3, xp4:  95, xp5: 0.52}\n    paramsC = {xp1: 0.82, xp2: 0.060, xp3: 3, xp4:  90, xp5: 0.52}\n    paramsD = {xp1: 0.82, xp2: 0.063, xp3: 3, xp4:  90, xp5: 0.52}\n    paramsE = {xp1: 0.82, xp2: 0.066, xp3: 3, xp4:  90, xp5: 0.52}\n# Different parameters for just the All-ids model and with Wrist feature(s)\nif not TWO_MODELS and USE_PARQUET:\n    # One Model for all-ids CSV with Wrist feature(s). v152, from v150+ GSCV\n    paramsA = {xp1: 0.77, xp2: 0.045, xp3: 3, xp4: 120, xp5: 0.60}\n    paramsB = {xp1: 0.77, xp2: 0.048, xp3: 3, xp4: 115, xp5: 0.60}\n    paramsC = {xp1: 0.77, xp2: 0.051, xp3: 3, xp4: 115, xp5: 0.60}\n    paramsD = {xp1: 0.77, xp2: 0.054, xp3: 3, xp4: 115, xp5: 0.60}\n    paramsE = {xp1: 0.77, xp2: 0.057, xp3: 3, xp4: 115, xp5: 0.60}\n\n\nxgbregrA = xgb.XGBRegressor(seed=1 + DELTA_SEEDS, **paramsA)\nxgbregrB = xgb.XGBRegressor(seed=3 + DELTA_SEEDS, **paramsB)\nxgbregrC = xgb.XGBRegressor(seed=9 + DELTA_SEEDS, **paramsC)\nxgbregrD = xgb.XGBRegressor(seed=27 + DELTA_SEEDS, **paramsD)\nxgbregrE = xgb.XGBRegressor(seed=81 + DELTA_SEEDS, **paramsE)\n\n# Do the fit - include sample weights\nxgbregrA.fit(X, y, sample_weight=fit_weights)\nprint(\"\\n  Training R$^2$-Score: {:.4f}\".format(xgbregrA.score(X,y)))\nxgbregrB.fit(X, y, sample_weight=fit_weights)\nprint(\"  Training R$^2$-Score: {:.4f}\".format(xgbregrB.score(X,y)))\nxgbregrC.fit(X, y, sample_weight=fit_weights)\nprint(\"  Training R$^2$-Score: {:.4f}\".format(xgbregrC.score(X,y)))\nxgbregrD.fit(X, y, sample_weight=fit_weights)\nprint(\"  Training R$^2$-Score: {:.4f}\".format(xgbregrD.score(X,y)))\nxgbregrE.fit(X, y, sample_weight=fit_weights)\nprint(\"  Training R$^2$-Score: {:.4f}\".format(xgbregrE.score(X,y)))\n# Average them\ny_hat_0 =( xgbregrA.predict(X) + xgbregrB.predict(X) +\n        xgbregrC.predict(X) + xgbregrD.predict(X) + xgbregrE.predict(X))/5.0\n\n# Unadjusted QWK\ny_hat = 1.0 * y_hat_0\n# The predicted siis\nsii_hat = sii_from_total(y_hat)\ncomp_metric = metric_from_siis(sii_hat, sii_from_total(y), verbose=0)\nprint(\"\\n       Original fit values: QWK metric = {:.4f}\".format(\n                comp_metric))\n\n# Show the metric for just the ones with any or valid parquet data\nparquet_metric = metric_from_siis(sii_hat[row_has_wrist],\n                                    sii_from_total(y[row_has_wrist]), verbose=0)\nprint(\"         for {} w/parquet: QWK metric = {:.4f}\".format(\n                            sum(row_has_wrist), parquet_metric))\nprint()\n# \"Optimize\" by transforming intervals based on the train quantiles\n# Can manually adjust them by looping over adjustments\ndeltaq=0.0\n##for deltaq in np.linspace(-0.05, 0.05, 20+1):\nif True:\n    q1adj =  0.0 #deltaq   # 0s seem ~ best\n    q2adj =  0.0 #deltaq\n    q3adj =  0.0 #deltaq\n    train_boundaries = ([np.quantile(y_hat_0, 0.5826 + q1adj),\n                 np.quantile(y_hat_0, 0.8494 + q2adj),\n                 np.quantile(y_hat_0, 0.9876 + q3adj)])\n    y_hat = adjust_intervals(y_hat_0, train_boundaries)\n    # The predicted siis\n    sii_hat = sii_from_total(y_hat)\n    # Final QWK\n    comp_metric = metric_from_siis(sii_hat, sii_from_total(y), verbose=0)\n    print(\"{:.3f}  Quantile optimization: QWK metric = {:.4f}\".format(\n                deltaq, comp_metric), \"     Boundaries: {:.2f}, {:.2f}, {:.2f}\".format(\n                    train_boundaries[0], train_boundaries[1], train_boundaries[2]))\n    parquet_metric = metric_from_siis(sii_hat[row_has_wrist],\n                                    sii_from_total(y[row_has_wrist]), verbose=0)\n    print(\"           the {} optimized: QWK metric = {:.4f}\".format(\n                            sum(row_has_wrist), parquet_metric))\nprint(\"\")\nif not QUANTILE_SCALE:\n    # Use the un-adjusted predictions\n    y_hat = 1.0 * y_hat_0","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Fit 2$^{nd}$ Model on just the w/parquet ids (996 or 777)","metadata":{}},{"cell_type":"code","source":"if TWO_MODELS:   # A separate model for the 996/777 ids with any/valid parquet data\n    # Get selected X, y, and weights values \n    Xprq = Xydf.loc[row_has_wrist, preds_wprq]\n    yprq = Xydf.loc[row_has_wrist,\"Total\"]\n    prq_weights = fit_weights[row_has_wrist]\n    print(\"\\nFitting {} data values.\".format(len(yprq)))\n    \n    if USE_PARQUET:   # for v150 w/ 7 CSVs + 2 Wr, params in-the-red\n        # Fitting 2nd model on 777 ids\n        paramsP = {xp1: 0.94, xp2: 0.045, xp3: 2, xp4:130, xp5: 0.32}\n        paramsQ = {xp1: 0.94, xp2: 0.048, xp3: 2, xp4:130, xp5: 0.32}\n        paramsR = {xp1: 0.94, xp2: 0.051, xp3: 2, xp4:130, xp5: 0.32}\n    else:\n        # Fitting the 996 w/parquet-file ids w/csv feat.s (no parquet Wrist features)\n        # v136, Different but similar pairs for rate,n_est.s: 0.072,65; 0.060,75; 0.077,55 \n        paramsP = {xp1: 0.78, xp2: 0.072, xp3: 2, xp4: 65, xp5: 0.53}\n        paramsQ = {xp1: 0.78, xp2: 0.060, xp3: 2, xp4: 75, xp5: 0.53}\n        paramsR = {xp1: 0.78, xp2: 0.077, xp3: 2, xp4: 55, xp5: 0.53}\n\n    xgbregrP = xgb.XGBRegressor(seed=11 + DELTA_SEEDS, **paramsP)\n    XP, yP, weightsP = skl_shuffle(Xprq, yprq, prq_weights, random_state=17)\n    xgbregrP.fit(XP, yP, sample_weight=weightsP)\n    print(\"\\n  P - Just Parquet Training R$^2$-Score: {:.4f}\".format(xgbregrP.score(Xprq,yprq)))\n    #\n    xgbregrQ = xgb.XGBRegressor(seed=23 + DELTA_SEEDS, **paramsQ)\n    XQ, yQ, weightsQ = skl_shuffle(Xprq, yprq, prq_weights, random_state=34)\n    xgbregrQ.fit(XQ, yQ, sample_weight=weightsQ)\n    print(  \"  Q - Just Parquet Training R$^2$-Score: {:.4f}\".format(xgbregrQ.score(XQ,yQ)))\n    #\n    xgbregrR = xgb.XGBRegressor(seed=37 + DELTA_SEEDS, **paramsR)\n    XR, yR, weightsR = skl_shuffle(Xprq, yprq, prq_weights, random_state=51)\n    xgbregrR.fit(XR, yR, sample_weight=weightsR)\n    print(  \"  R - Just Parquet Training R$^2$-Score: {:.4f}\".format(xgbregrR.score(XR,yR)))\n    #\n    y_hat_0prq = (xgbregrP.predict(Xprq) + xgbregrQ.predict(Xprq) +\n                                              xgbregrR.predict(Xprq))/3.0\n    \n    # Unadjusted QWK\n    y_hat_prq = 1.0 * y_hat_0prq\n    # The predicted siis\n    sii_hat_prq = sii_from_total(y_hat_prq)\n    comp_metric = metric_from_siis(sii_hat_prq, sii_from_total(yprq), verbose=0)\n    print(\"\\n  Just Parquet Original fit values: QWK metric = {:.4f}\".format(\n                comp_metric))\n    if True:\n        # Put in the quantile fractions for the 777 ids. From doing:\n        #   Xyprq.loc[row_has_wrist,\"sii\"].value_counts().values\n        # Number of samples in each sii bin: [468 204 100   5] = 777\n        # Bin boundaries as quantiles: 0.6023, 0.8649, 0.9935\n        if USE_PARQUET:\n            prq_boundaries = [np.quantile(y_hat_0prq, 0.6023),\n                             np.quantile(y_hat_0prq, 0.8649),\n                                 np.quantile(y_hat_0prq, 0.9935)]\n        else: prq_boundaries = train_boundaries\n        y_hat_prq = adjust_intervals(y_hat_0prq, prq_boundaries)\n        # The predicted siis\n        sii_hat_prq = sii_from_total(y_hat_prq)\n        # Quantile-optimized QWK\n        comp_metric = metric_from_siis(sii_hat_prq, sii_from_total(yprq), verbose=0)\n        print(\"  Just Parquet Quantile optimization: QWK metric = {:.4f}\".format(\n                comp_metric), \"     Boundaries: {:.2f}, {:.2f}, {:.2f}\".format(\n                    prq_boundaries[0], prq_boundaries[1], prq_boundaries[2]))\n    if not QUANTILE_SCALE:\n        y_hat_prq = 1.0 * y_hat_0prq\n\n    # Put these y_hat_prq values into the all-ids y_hat\n    y_hat[row_has_wrist] = y_hat_prq","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Look at the final predicted values","metadata":{}},{"cell_type":"code","source":"sii_hat = sii_from_total(y_hat)    \n# QWK with parquet ids' values updated\ncomp_metric = metric_from_siis(sii_hat, sii_from_total(y), verbose=0)\nprint(\"\\n  Final QWK metric = {:.4f}\".format(comp_metric))\n\nparquet_metric = metric_from_siis(sii_hat[row_has_wrist],\n                                    sii_from_total(y[row_has_wrist]), verbose=0)\nprint(\"    {} QWK metric = {:.4f}\".format(\n                            sum(row_has_wrist), parquet_metric))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Show how the Predicted Total values, y_hat, compare with the Actual Total values, y.\n# Output a crosstab of the sii_hat vs sii values.\n\n# Add color-code by gender\nX[\"clr\"] = \"blue\"\nX.loc[X[\"Basic_Demos-Sex\"] == 1, \"clr\"] = \"red\"\nplt.figure(figsize=(6,6))\nplt.scatter(y,y_hat,alpha=0.1,c=X[\"clr\"])\nlast_bound = min_train_total - 3\nfor bound in [30, 49, 79, 100]:\n    # Show the target boxes:\n    draw_rect(last_bound+1,bound,last_bound+1,bound, \"lime\")\n    # And the four that have (i-j)^2 = 4\n    if bound == 30:\n        draw_rect(last_bound+1,bound,49+1,79, \"#C0303070\")\n    if bound == 49:\n        draw_rect(last_bound+1,bound,79+1,100, \"#C0303070\")\n    if bound == 79:\n        draw_rect(last_bound+1,bound,min_train_total-1, 30, \"#C0303070\")\n    if bound == 100:\n        draw_rect(last_bound+1,bound,30+1,49, \"#C0303070\")\n    last_bound = bound\n    \nplt.title(\"Predicted vs Actual Totals, with target rectangles\")\nplt.xlabel(\"Actual Total\"); plt.ylabel(\"Predicted Total  (w/interval scaling)\")\nplt.xlim(min_train_total - 5, 102); plt.ylim(min_train_total - 5, 102)\nplt.savefig('Predicted_vs_Actual_Total.png')\nplt.show()\n# kind of kludgy... remove clr so this cell can be re-run\nX = X.drop(columns=[\"clr\"])\n\nprint(\"\\nCounts in the different sii regions oriented like the plot:\")\nprint(pd.crosstab(-1*sii_hat,sii_from_total(y)))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Feature importance\nfor this_model in [xgbregrA,xgbregrB,xgbregrC,xgbregrD,xgbregrE]:\n    xgb.plot_importance(this_model)\n    plt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Evaluate Features with permutations  https://eli5.readthedocs.io/en/latest/overview.html\n# PermutationImportance(estimator, scoring=None,\n#                        n_iter=5, random_state=None, cv='prefit', refit=True)\n# cv choices: (int, cross-validation generator, iterable or “prefit”)\n# Use CV to get \"generalization\" importance\nperm = PermutationImportance(xgbregrA, scoring=None, n_iter=9, cv=6,  # cv='prefit' or number\n                             random_state=None, refit=True).fit(X,y)\neli5.show_weights(perm, top=None, feature_names = X.columns.tolist())","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"perm = PermutationImportance(xgbregrB, scoring=None, n_iter=9, cv=6,  # cv='prefit' or number\n                             random_state=None, refit=True).fit(X,y)\neli5.show_weights(perm, top=None, feature_names = X.columns.tolist())","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"perm = PermutationImportance(xgbregrC, scoring=None, n_iter=9, cv=6,  # cv='prefit' or number\n                             random_state=None, refit=True).fit(X,y)\neli5.show_weights(perm, top=None, feature_names = X.columns.tolist())","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"perm = PermutationImportance(xgbregrD, scoring=None, n_iter=9, cv=6,  # cv='prefit' or number\n                             random_state=None, refit=True).fit(X,y)\neli5.show_weights(perm, top=None, feature_names = X.columns.tolist())","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"perm = PermutationImportance(xgbregrE, scoring=None, n_iter=9, cv=6,  # cv='prefit' or number\n                             random_state=None, refit=True).fit(X,y)\neli5.show_weights(perm, top=None, feature_names = X.columns.tolist())","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if TWO_MODELS:\n    print(\"\\nThe 2nd model(s) fit to only the w/parquet ids:\")\n    for this_model in [xgbregrP,xgbregrQ,xgbregrR]:\n        xgb.plot_importance(this_model)\n        plt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# *** If not using a second model COMMENT OUT the LAST LINE ***\nif TWO_MODELS:\n    # The model using just ids w/parquet values\n    # Have to reset the indices to go from 0 to 776/996 (for the CV portion I think.)\n    perm = PermutationImportance(xgbregrP, scoring=None, n_iter=9, cv=6,  # cv='prefit' or number\n            random_state=None, refit=True).fit(\n                        Xprq.reset_index().drop(columns=['index']), np.array(yprq))\n# Have to run this not indented to see the output (?!?)\n##eli5.show_weights(perm, top=None, feature_names = Xprq.columns.tolist())","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# *** If not using a second model COMMENT OUT the LAST LINE ***\nif TWO_MODELS:\n    perm = PermutationImportance(xgbregrQ, scoring=None, n_iter=9, cv=6,  # cv='prefit' or number\n            random_state=None, refit=True).fit(\n                        Xprq.reset_index().drop(columns=['index']), np.array(yprq))\n##eli5.show_weights(perm, top=None, feature_names = Xprq.columns.tolist())","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# *** If not using a second model COMMENT OUT the LAST LINE ***\nif TWO_MODELS:\n    perm = PermutationImportance(xgbregrR, scoring=None, n_iter=9, cv=6,  # cv='prefit' or number\n            random_state=None, refit=True).fit(\n                        Xprq.reset_index().drop(columns=['index']), np.array(yprq))\n##eli5.show_weights(perm, top=None, feature_names = Xprq.columns.tolist())","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Make predictions for Test","metadata":{}},{"cell_type":"code","source":"# First, make the Test w/csv-only predictions\ntestX = test[preds_csv].copy()\n# Predict with CSV and the 5 XGBs\ntest_hat_0 = ( xgbregrA.predict(testX) + xgbregrB.predict(testX) +\n            xgbregrC.predict(testX) + xgbregrD.predict(testX) + xgbregrE.predict(testX))/5.0\n# Add quantile-intervals adjustment\nif QUANTILE_SCALE:\n    test_hat = adjust_intervals(test_hat_0, train_boundaries)\nelse:\n    test_hat = 1.0 * test_hat_0\n\nif TWO_MODELS:\n    # Do the 2nd model predictions\n    if USE_PARQUET:\n        test_has_wrist = test[\"Wrist-battery_days\"] > 0\n    else:\n        test_has_wrist = test[\"Wrist-file_days\"] > 0\n    testXprq = test.loc[test_has_wrist, preds_wprq]\n    test_hat_0prq = (xgbregrP.predict(testXprq) + xgbregrQ.predict(testXprq) +\n                             xgbregrR.predict(testXprq))/3.0\n    if QUANTILE_SCALE:\n        test_hat_prq = adjust_intervals(test_hat_0prq, prq_boundaries)\n    else:\n        test_hat_prq = 1.0 * test_hat_0prq\n    # Merge prq predictions with the csv ones, directly or averaging.\n    test_hat[test_has_wrist] = test_hat_prq\n\n# Get the sii values from test_hat    \nsii_hat = sii_from_total(test_hat)\n     \nplt.figure(figsize=(5,3))\nplt.scatter(range(len(test_hat)), test_hat, alpha=0.1)\nfor bound in [30, 49,79]:\n    plt.plot([0,len(test_hat)],[bound,bound],c=\"gray\")\nplt.title(\"Predicted Total values for Test\")\nplt.ylabel(\"Predicted Total\"); plt.ylim(0,100)\nplt.show()\n\n# Add the id to testX\ntestX[\"id\"] = test[\"id\"]\n# Add the sii values\ntestX[\"sii\"] = sii_from_total(test_hat)\n\ntestX[\"sii\"].value_counts()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Output the submission\ntestX[[\"id\",\"sii\"]].to_csv(\"submission.csv\",index=False)\n\n##!more submission.csv","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Correlations with the Question Answers","metadata":{}},{"cell_type":"code","source":"# Look at the correlation of individual question answers with the predictors\nquest_cols = list(train.columns[list(range(55,74+1))])\n# Correlation with all significant predictors\ncorr_quest = train[quest_cols + preds + ['Total']].corr()  # model predictors\n\n# Correlation Heatmap, code from: \n#   https://www.kaggle.com/code/docxian/cmi-problematic-internet-use-visual-starter\nplt.figure(figsize=(12,10))\nsns.heatmap(corr_quest, annot=False, cmap='RdYlGn',\n            fmt='.2f', linecolor='black', linewidths=0.5,\n            vmin=-0.8, vmax=+0.8)\nplt.title('Question Correlations with the Predictors and Total')\nplt.savefig('Questions_Correlations.png')\nplt.show()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# See the numerical corr values:\nif False:\n    this_col = \"FGC-FGC_SRL_Zone\"\n    print(\"\\nCorrelation of \"+this_col+\" with the Questions and other Features:\")\n    corr_quest[this_col]","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Look at the answer distribution for the questions\nfor this_col in quest_cols:\n    print(this_col, \":\", train[this_col].value_counts().values)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Look at the 20 PCIAT questions -- do they have signatures in the Wrist data?\n# 01 How often ... disobey time limits you set for online use?\n# 02 How often ... neglect household chores to spend more time online?\n# 03 How often ... prefer to spend time online rather than with the rest of your family?\n# 04 How often ... form new relationships with fellow online users?\n# 05 How often do you complain about the amount of time your child spends online?\n# 06 How often ... grades suffer because of the amount of time he or she spends online?\n# 07 How often ... check his or her e-mail before doing something else?\n# 08 How often ... seem withdrawn from others since discovering the Internet?\n# 09 How often ... become defensive or secretive when asked what he or she does online?\n# 10 How often have you caught your child sneaking online against your wishes?\n# 11 How often ... spend time alone in his or her room playing on the computer?\n# 12 How often ... receive strange phone calls from new \"online\" friends?\n# 13 How often ... snap, yell, or act annoyed if bothered while online?\n# 14 How often ... seem more tired and fatigued than he or she did before the Internet?\n# 15 How often ... seem preoccupied with being back online when off-line?\n# 16 How often ... throw tantrums when you interfere with how long he or she spends online?\n# 17 How often ... choose to spend time online rather than doing once enjoyed hobbies, etc.?\n# 18 How often ... become angry or belligerent when you place time limits on online time?\n# 19 How often ... choose to spend more time online than going out with friends?\n# 20 How often ... feel depressed, moody, or nervous when off-line and goes away once online?\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#   Summary of versions\n# v1  0.247 LinRegr\n# v2  0.273 LinRegr w/dummy: Female\n# v3  0.305 LinRegr w/Frame dummies added: S, L, S-F, L-F\n# v4  0.296 LinRegr w/Season dummies added Seas-[Spring|Summer|Winter]\n#           Separated male and female LinRegr fits, little change.\n# v5  0.281 Use HGBRegressor - looks too good, it's got to be overfitting? Yup.\n# v6  0.294 HGBRegressor with more regularization, don't separate M/F.\n# v7        Random cleaning up, predictor adjustments; Calculated metric.\n#     0.299 LinRegr like v3 with y_hat offset of 4.75 (improved \"CV\" to ~0.36)\n# v8        Split into train and vaidation (65:35) and adjusted HGBR params.\n#     0.326 Could not get much above 0.32 for \"CV\" (offset=2.0) - good.\n# v9        Use XGBoost based on:\n#     -.---    https://www.kaggle.com/code/prashant111/xgboost-k-fold-cv-feature-importance\n# v10       XGB: Remove dummies for Frame and Season (not important based on XGB)\n#     0.334 Some other minor feature changes; magic offset of 3.0.\n# v11 0.330 XGB with parameters tuned via GSCV and some cleaning up.\n# v12 -.--- More cleaning up, etc.\n# v13 0.313 Did GSCV with scorer based on the QWK metric. max_depth set to 4.\n# v14 0.319 Try max_depth=5 and include offset=6.\n# v15 0.319 Probably this is about the best given just these predictors...\n# v16 0.334 Added sample_weight=(50.0+y)/150.0 weighting larger Total more. \n# v18       Found I'd left out two useful predictors:\n#     0.412 SDS-SDS_Total_T and PreInt_EduHx-computerinternet_hoursday\n# v20 0.416 Include SDS-Season dummy vars (1-hots), not important but keep around.\n# v21 0.401 Created PAQ-PAQ_Total by merging 'A and 'C values\n#     0.412 Re-submitted by accident\n# v23 0.408 Did GSCV over depth = 3,4,5. Trades off depth and other params.\n# v25 0.426 First look at the Wrist data files: some strange time_of_day & weekday values.\n# v27       Add some simple \"Wrist-\" features from the parquet files...\n#           \"Wrist-\"s are NaN for rows without parquet data, filled before fitting.\n#     0.436 (Notebook params: 800, 0.22, 5, 57. It's random: submission values could differ.) \n# v29       Useful: https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-actigraphy-data-eda\n#           From Antonina's notebook decide to select battery_voltage > 3800 for better data.\n#     0.432 Made other little changes to Wrist- features.\n# v30 0.419 Other little adjustments...(Oops?) Started reading Discussion posts...\n# v31 0.426 Realized that about 350 parquet files have had the non-wear_flag=1 rows removed...\n#           Re-do some of my Wrist- features (in particular pct_active) to work either way.\n# v33 0.422 Improved consistency of Wrist- values by removing non-worn (and equiv) from all.\n# v34 0.415 Same as v33 but scorer for Grid Search is R^2 instead of QWK.\n# v35 0.435 Weighted ones with Wrist- data higher than others; ignore short parquets.\n# v38       Changed how the idle_sleep_mode was applied to files without non-wear flags;\n#           also applied to ones with flags=0 that did not have idle times removed.\n#     0.439 Some cleaning up, but it is still too long-winded :(\n# v39       Same as v38 but ignore the Total==0 samples. Goes from 2736 with 996 parquet\n#     0.401 to 2423 with 879 parquet (~ proportional.)  Adjusted features to use based on corr. \n# v40 0.422 Same as v38,39 but now Total=0 are randomly assigned 1 to 15. Adjust features to use.\n# v41 0.437 Same as v38,39,40 now Total=0 are randomly assigned -1 to -15. Adjusted features to use.\n# v43 0.453 Bingo! Same as v38-41, now added a gap of 10 at Total=30 and 0s go from -5 to -13.\n# v44 0.457 Going too far?: added other gaps in Total - it's \"Regres-ification\"\n#           Will keep these gaps, squish the sii=3 one to 90-100, and leave it for now.\n# v46 0.451 Look through: https://www.kaggle.com/code/abdmental01/cmi-best-single-model\n#           for any ideas... They used stats of the parquet files as features...\n#           Set XGB parameters similar to LGBM ones in the notebook. Hmmm...\n# v47 0.453 Little cleanups; auto-offset; Use previous (v44) XGB grid w/min_child_weight=10\n# v48 0.432 Add 3 to the sii=3 values; change 0s in Physicals to NaN; NaN fill w/means by sex\n# v49 -.--- I think the NaNs=mean by sex messed the score up? ... \n# v50 0.413 I think I broke the code :(  Has NaNs = 0. \n# v51 0.435 Compare with v44 and remove some changes (nothing dramatic).\n#           Set seeds so not random; fix XGB params at v44's\n# v52       fyi, article mentioned in discussion: Modeling Approach: Semi-Supervised Learning\n#              https://www.ibm.com/topics/semi-supervised-learning\n#           Option to use/not-use parquet; remove feature and fitting randomness;\n#           Do a very wide GSCV search to find best XGB params.\n# v53 0.449 Fit just 20 of the csv features, find good GSCV param range based on QWK.\n# v54 0.449 The exact same as v53 but USE_PARQUET=True to include the Wrist- variables.\n#           Strange it's the exact same score... Different QWK CVs: v53:0.4478, v54:0.4563.\n# v56       Look into using gamma with a smaller lambda, or just gamma... Not promising.\n# v57 0.448 Back to lambda and learning rate around 450 and 0.145, QWK CV/Train: 0.4495/0.5722\n# v58 0.446 Same as v57, with Wrist features. QWK CV/Train: 0.4634/0.5484\n# v59 0.450 Same as v57 w/train-quantile \"optimization\"; QWK CV/Train: 0.4495/0.5890\n# v60 0.433 Same as v58 (w/Wrist), GS grid increased a little. QWK CV/Train: 0.4634/0.5678\n# v61 0.442 OK, need to settle a clear way to select params and assess performance\n# v62       --> Use GSCV with QWK for param selection, use mini-GSCV to assess CV Test QWK.\n#     0.442 Made/Adjusted csv features including PAQ-PAQ_Season, filling missing Weights...\n# v63 0.457 Same as v62 with Wrist- features for fun.\n# v64       Yeesh... Filled the weights before the ssi down-select to use all the data.\n#     0.432 Search parameters with Wrist- features, use lambda=640, learn_rate=0.330    \n# v65 0.439 Used 717, 0.232 Yeesh II.\n# v66 0.437 Used 300, 0.155 as in v63. Added correlations with the questions' answers.\n# v67       Yeesh... Based on Stuart Houghton's post in discussion:\n#              https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535023\n#           I'm going back to avoiding the unlabeled (no sii/Total) data.\n#           Fill Weight nans from equation. More 'cleaning up':\n#           make idle_mode feature instead of num_non_wear. PCIAT ave date too.\n#           Make correlations heatmap for PCIAT individual questions with csv (and Wrist).\n# v69 0.425 No-Wrist: Do a wide CV grid for lambda and learning_rate, used 550, 0.150\n# v70 0.434 No-Wrist: Try the low lambda-rate good fit found at: 50, 0.048\n# v71 0.433 No-Wrist: GSCV scan 300-1200, 0.10-0.35. Bests: 1028, 0.265; 476, 0.175\n# v72 0.422 +Wrist and adjusted Weight,Height. GSCV scan 300-1200, 0.10-0.35: used ~ 450, 0.190\n# v73 0.420 More little changes, used ~ 600, 0.350 \n# v74 0.428 Weight, Height are relatve, Height filled from equ. Best~ 365, 0.300\n# v75 0.433 v74 accidentally included BIA Fat and TBW (very anti-corr w/eachother)\n# v76       Look at other notebooks: what feaures, models, params?\n#           https://www.kaggle.com/code/gkitchen/introduction-to-problematic-internet-use\n#           Above used XGB Regressor with 16 feat.s (corr>0.1, <1/2 missing) and params: \n#           params = {'max_depth': 3, 'n_estimators': 59, 'learning_rate': 0.0733,\n#                        'subsample': 0.5968, 'colsample_bytree': 0.9124}\n#           Also uses ELI5:\n#                 perm = PermutationImportance(model, random_state=1).fit(X,y)\n#                 eli5.show_weights(perm, feature_names = X.columns.tolist())\n#     0.435 Use gkitchen's Seasons encoding (0,1,2,3,4) for PAQ-PAQ_Season\n# v77 0.455 Use gkitchen's XGB parameters... good! But had very low CV score 0.4177?\n# v78       Use PermutationImportance with CV to remove 5 csv features.\n#     0.442 GSCV by colsample_bytree and subsample, 15 csv + 4 Wrist-s included.\n# v80       Set bytree=0.93 and subsamp=0.44 from scan. Do a scan for learning and n_estim.s\n# v81 0.432 Set the 5 parameters to: 0.93, 0.075, 3, 94, 0.44 based on GSCVs in v78, v80.\n# v82 0.428 Same as v81 but using train_boundaries for test optimization (vs forcing test quantiles.)\n# v83 0.426 That was disappointing... Do v82 with the gkitchen params...\n# v84       Yeesh, I messed up adjusting the predictors and NaNs?\n#     0.452 Use gkitchen's values on just 17 csv features. Features are OK, Wrist- the problem?\n# v85 0.450 OK, same but use the 0.93, 0.075, 3, 94, 0.44 parameters.\n# v86 0.431 As v85, but no NaN fill -- it's better to fill :)\n# v87 0.448 As v84 gkitchen's, but no NaN fill - yeesh, nothing is consistent...\n# v88       Want to reduce randomness of model, GSCV uses one split for all folds/params\n#           Try fitting a number of GSCV/XGBs each with different random seed, average them.\n#     0.430 Same as v84, has NaNs filled. Averaging a stand of 5 trees instead of a forest :)\n# v89 0.428 Same, only use Model B for predictions\n# v90 0.423 Same, only use Model D for predictions\n# v92       GSCV by colsample_bytree and subsample, 17 csv, w/NaNs. Use 0.79 and 0.41\n# v93       GSCV by learning_rate and n_estimators, 17 csv, w/NaNs. Use 0.0675 and 109\n# v94 0.439 Using params: 0.79, 0.0675, 3, 109, 0.41 . (Weights error was in v88--v93)\n# v95 0.401 As v94 but no interval adjustment to predictions. Adjustments are good :)\n# v96 0.437 Using params: 0.79, 0.05, 3, 59, 0.41  underfitting a little?\n# v97 0.458 Using params: 0.85, 0.07, 3, 80, 0.50. That's nice.\n# v98       OK, decide to 'lockdown' the v97 csv features and XGB params.\n#           Turn attention to the parquet files and Wrist- features...  v98 as check point.\n# v99       Continued adjusting parquet processing, 771 'useful', > 7 days, samples.\n#     0.448 Submitted for sanity check. Include 2 of the Wrist- features: pct_active and ave_enmo.\n# 100 0.443 Submitted as check point. Still working on parquet features...\n# 101       Adjust enmoStd_idle a little to 0.0055 to get same values for idle=0 and 1 of:\n#           Wrist-11pm5am_active_hrs for idle=0: 0.551, idle=1: 0.552\n#     0.439 Choose mean fill for discrete features, and median for continuous.\n# 102 0.430 Added Wrist- bedtime and sleep_hours, though not very correlated with Total.\n#           Score keeps going down with the/more Wrist- predictors added - expected that.\n# 103 0.443 Improved the bedtime/sleep_hours fitting, corr.s w/Total of 0.0900/-0.0827.\n# 104 0.445 For fun (and improvement?) Do the 5 XGBs as: 3 with Wrist feat.s and 2 without.\n# 105       Using StratifiedKFold on sii but fitting y; 5 models w/Xy shuffled; XGB seed used.\n#     0.442 Similar to v97: 16 features, no Wrists, params: 0.85, 0.07, 3, 80, 0.50.\n# 106       Slight changes to ave_enmo and \"11pm5am[sleep]_active_hours\" Wrist features.\n#     0.434 <- Overfitting: For fun: all 16 csvs and all Wrist-\n# 107       Use Permutations in v106 to keep only 4 Wrist feats and remove 4 CSV ones.\n#     0.442 <- Disappointing. Showing raw QWK score of the 777 w/parquet: higher than the all score.\n# 108 0.442 Same as 107 but NaN-remaining not filled.\n# 109       Rename, organize, adjust the Wrist features...\n#     0.437 <-Oops, error? Use 12 CSV w/two new Wrists: 'quant40_enmo and 'quant85_enmo. NaN filled.\n# 110 0.432 Same as 109 but NaN-remaining not filled.\n# 111       12 CSV + '40enmo,'85enmo and quantHi_light.\n# 112 0.444 Don't weigh the parquet samples more than others (of course, yeesh.)\n#           12 CSVs + 4 Wrist. Stratify by combined sii and has/not parquet data.\n# 113       Add new Wrist features: means, stds, and product of smoothed enmo and light.\n#     0.448 12 CSVs + 3 Wrist: sleep_hours, enmoMLSmean, enmoXlight \n# 114       Better see importance of Wrist features, fit only the 777 w/Wrist rows\n#           and remove ones by permutation importance, leaving:\n#     0.440 9 CSVs + 7 Wrist, then use these to fit on all data.\n# 115       Tried making weekend - weekday enmo mean difference - not useful, corr=-0.019\n#           Make weekend days only enmoXlight feature.\n#     0.434 <- Overfit 12 CSVs + 5 Wrist: sleep_hours, enmoMLSmean, [wkend_]enmoXlight, pct_active\n# 116       Re-assess the CSVs: select corr > 3/sqrt(n_this_one), then fill NaN.\n#     0.453 13 CSVs selected/left after going though systematically w.Permu Import.\n# 117       Add and select Wrist features... Keep just 'pct_active, 'bedtime, 'wkend_enmoXlight\n#     0.447 13 CSVs + 3 Wrist: pct_active, wkend_enmoXlight, bedtime\n# 118 0.440 Same as 117 but the 777 are fit separately by one XGB model.\n# 119       Ooops - too kludgy: gotta fit all data not using parquet features, as in v116,\n#           then do a separate fit of the 777 w/parquet ids and paste those predictions into y_hat.\n#     0.448 Cleaned up, should be similar to v116 even though parquet are processed - just checking.\n# 120       Only difference v119 and v116 was including parquet w/sii in the stratified folding.\n#     0.438 Add a model for just the with-parquet ids and insert its predictions in y_hat.\n# 121       Well, that was not good: overfitting the 777? or messed something up splitting models?\n#           Use split models but use the preds_csv for both, looks like logic is OK.\n#           The quantile-interval adjustments for 777 are slightly different, but\n#           the 777 model is over-fitting with a better training score.\n#     0.456 <-good. Keep split model, w/csv preds for both, reduced tree depth on 777 from 3 to 2.           \n# 122 0.449 Same as 121 but use the preds_wprq predictors for the 777.\n# 123 0.448 Same as 122 but leave Height out of 777-fit features\n# 125 0.444 Same as 123 but weighted average of the two (0.82prq+0.18csv), dropped 'SLR_Zone from prq\n# 126       Add TWO_MODELS option. Add code to compare the csv and prq models on the 777.\n#           (fyi, using only the csv model is ~ same as v119.)\n# 127 0.445 like v119 but added the wkend_enmoXlight feature to the csvs model, TWO_MODELS=False.\n# 128 0.443 13 CSVs, no Wrist, one model w/params: 0.95, 0.070, 3, 95, 0.42 .     \n# 129       Trying to improve over v121 [which uses 13 CSVs, no Wrist, and TWO_MODELS=True]\n#           by including Wrist features... Did GSCV to update the all and 777 models params.\n#     0.436 Overfitting the models a problem? The 777 used 8 CSVs + 3 Wrist features.\n# 130       Should be like v121 (2 models, csv feats only)\n#     0.442 Params all:0.95,0.07,3,95,0.42 and prq-only:0.95,0.075,2,95,0.38\n# 131       Similar to v130: 2 models(all,996), csv feats only,\n#           but 3 prq models, drop PU, GSND from prq preds, and\n#     0.440 Params from v97/v121: 0.85, 0.07, 3, 80, 0.50 for both models.\n# 133       Clean some things up, added PAQ-PAQ_Season: 14 CSVs and 3 Wrist.\n#     0.434 Use most convervative params so far for the models.\n# 134 0.438 v133 without age blur\n# 135 0.449 One Model, No-parq: 14 CSVs, GSCV w/R^2, choose params on underfit side.\n# 136       Two Models, No-parq: All w/14CSVs, and 996 w/10 CSVs.\n#     0.444 Not too bad since wouldn't expect improvement using a smaller model for some ids.\n# 137       One Model, W-parq: 14 CSVs + 1 Wrist feature\n#     0.441 params same as v135: ~ {xp1: 0.8, xp2: 0.058, xp3: 3, xp4: 90, xp5: 0.55}\n# 138 0.440 One Model, W-parq: same as v137 but using assorted underfit param combinations.\n#           --> Adding the one Wrist feature, valid in 777 of the 2736 ids, doesn't help.\n# 139 0.443 One Model, No-parq: 14 CSVs, v135 but with GSCV w/R^2 assorted underfit params - Hmmm, similar.  \n# 141       Two Models, W-parq: all w/14CSVs, and 777 w/3Wr+9CSVs. \n#     0.433 Yeesh...\n# 142 0.433 Two Models, W-parq: All w/14CSVs, and 777 w/ 3Wr+10CSVs. <-- oops, wanted just 10CSVs.\n#           The 141,142 GSCV regions are different even though mostly the same features,\n#           3Wr+8CSV in common and extras: 141(+GSND), 142(+Height,+CGAS_Score)\n# 143 0.431 Two Models, W-parq: All w/14CSVs, and 777 w/ 7 CSVs. Expected similar to v136(-3CSVs)\n# 144 0.432     Same as 143 w/different rates,n_est.s from v143 GSCV.\n# 145 0.424 One Model, No-parq: 7 CSVs, GSCV w/R^2, choose params on underfit side.\n# 146 0.436 One Model, No-parq: 7 CSVs, (GSCV same as 145) using params near best GSCV score\n# 147 0.440 One Model, No-parq: 9 CSVs, GSCV w/R^2, using params near best GSCV score\n# 148 0.435 One Model, W-parq: 9 CSVs + 1 Wr, GSCV done, but params from 147.\n# 149 0.440 One Model, W-parq: 9 CSVs + 1 Wr, GSCV v148 used to set params.\n# 150 0.436 Two Models, W-parq:  All w/9 CSVs(~v152+), and 777 w/ 7 CSVs + 2 Wr.\n# 152 0.447 One Model, W-parq: 9 CSVs + 1 Wr, GSCV v150+ used to set params.\n# 153 0.436 One Model, No-parq: 9 CSVs, GSCV w/R^2, using in-red params. Like v147\n# 154 0.432 Two Models, W-parq:  All w/9 CSVs(~v152+), and 777 w/ 7 CSVs + 2 Wr.\n# 155       Hmmm, can't seem to get a 2 Model W-parq into the LB 0.440s. (v136 2 Model No-parq was 0.444)\n#     0.449 Turned off the Age Blur  <-- v152 + checking some things\n# 156 0.447 Age blur back on, v152 (check repeatable, getting parnoid.)\n# 157 0.443 Turned off additional NaN filling  <-- v152 + checking some things\n# 158 0.446 Same as v155 but the Total is not modified\n# 159       Well, v155 is a good stopping point & submission to end this addiction :)\n#           Clean things up a bit, output some plots.\n# 160       Postmortem :)   Clean up a little; Added LB plots below.","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Plot Public and Private LB scores","metadata":{}},{"cell_type":"code","source":"# Copied and Pasted from my Submissions web page:\nmy_scores = '''\nCMI 2024 - Kaggle Addiction ? - Version 159\nSucceeded · 11d ago\n0.429\n\n0.449\n\nCMI 2024 - Kaggle Addiction ? - Version 158\nSucceeded · 12d ago\n0.432\n\n0.446\n\nCMI 2024 - Kaggle Addiction ? - Version 157\nSucceeded · 12d ago\n0.436\n\n0.443\n\nCMI 2024 - Kaggle Addiction ? - Version 155\nSucceeded · 12d ago\n0.429\n\n0.449\n\nCMI 2024 - Kaggle Addiction ? - Version 156\nSucceeded · 12d ago\n0.426\n\n0.447\n\nCMI 2024 - Kaggle Addiction ? - Version 154\nSucceeded · 13d ago\n0.421\n\n0.432\n\nCMI 2024 - Kaggle Addiction ? - Version 153\nSucceeded · 13d ago\n0.420\n\n0.436\n\nCMI 2024 - Kaggle Addiction ? - Version 152\nSucceeded · 13d ago\n0.426\n\n0.447\n\nCMI 2024 - Kaggle Addiction ? - Version 150\nSucceeded · 15d ago\n0.417\n\n0.436\n\nCMI 2024 - Kaggle Addiction ? - Version 149\nSucceeded · 15d ago\n0.426\n\n0.440\n\nCMI 2024 - Kaggle Addiction ? - Version 148\nSucceeded · 15d ago\n0.425\n\n0.435\n\nCMI 2024 - Kaggle Addiction ? - Version 147\nSucceeded · 15d ago\n0.423\n\n0.440\n\nCMI 2024 - Kaggle Addiction ? - Version 146\nSucceeded · 16d ago\n0.425\n\n0.436\n\nCMI 2024 - Kaggle Addiction ? - Version 145\nSucceeded · 16d ago\n0.419\n\n0.424\n\nCMI 2024 - Kaggle Addiction ? - Version 144\nSucceeded · 16d ago\n0.434\n\n0.432\n\nCMI 2024 - Kaggle Addiction ? - Version 143\nSucceeded · 16d ago\n0.433\n\n0.431\n\nCMI 2024 - Kaggle Addiction ? - Version 142\nSucceeded · 16d ago\n0.439\n\n0.433\n\nCMI 2024 - Kaggle Addiction ? - Version 141\nSucceeded · 17d ago\n0.440\n\n0.433\n\nCMI 2024 - Kaggle Addiction ? - Version 139\nSucceeded · 17d ago\n0.443\n\n0.443\n\nCMI 2024 - Kaggle Addiction ? - Version 138\nSucceeded · 17d ago\n0.430\n\n0.440\n\nCMI 2024 - Kaggle Addiction ? - Version 137\nSucceeded · 18d ago\n0.434\n\n0.441\n\nCMI 2024 - Kaggle Addiction ? - Version 136\nSucceeded · 18d ago\n0.432\n\n0.444\n\nCMI 2024 - Kaggle Addiction ? - Version 135\nSucceeded · 19d ago\n0.444\n\n0.449\n\nCMI 2024 - Kaggle Addiction ? - Version 134\nSucceeded · 20d ago\n0.426\n\n0.438\n\nCMI 2024 - Kaggle Addiction ? - Version 133\nSucceeded · 20d ago\n0.426\n\n0.434\n\nCMI 2024 - Kaggle Addiction ? - Version 131\nSucceeded · 1mo ago\n0.429\n\n0.440\n\nCMI 2024 - Kaggle Addiction ? - Version 130\nSucceeded · 1mo ago\n0.419\n\n0.442\n\nCMI 2024 - Kaggle Addiction ? - Version 129\nSucceeded · 1mo ago\n0.428\n\n0.436\n\nCMI 2024 - Kaggle Addiction ? - Version 128\nSucceeded · 1mo ago\n0.435\n\n0.443\n\nCMI 2024 - Kaggle Addiction ? - Version 127\nSucceeded · 1mo ago\n0.445\n\n0.445\n\nCMI 2024 - Kaggle Addiction ? - Version 126\nSucceeded · 1mo ago\n0.441\n\n0.438\n\nCMI 2024 - Kaggle Addiction ? - Version 125\nSucceeded · 1mo ago\n0.431\n\n0.444\n\nCMI 2024 - Kaggle Addiction ? - Version 123\nSucceeded · 1mo ago\n0.431\n\n0.448\n\nCMI 2024 - Kaggle Addiction ? - Version 122\nSucceeded · 1mo ago\n0.432\n\n0.449\n\nCMI 2024 - Kaggle Addiction ? - Version 121\nSucceeded · 1mo ago\n0.429\n\n0.456\n\nCMI 2024 - Kaggle Addiction ? - Version 120\nSucceeded · 1mo ago\n0.431\n\n0.438\n\nCMI 2024 - Kaggle Addiction ? - Version 119\nSucceeded · 1mo ago\n0.434\n\n0.448\n\nCMI 2024 - Kaggle Addiction ? - Version 118\nSucceeded · 1mo ago\n0.437\n\n0.440\n\nCMI 2024 - Kaggle Addiction ? - Version 117\nSucceeded · 1mo ago\n0.448\n\n0.447\n\nCMI 2024 - Kaggle Addiction ? - Version 116\nSucceeded · 1mo ago\n0.432\n\n0.453\n\nCMI 2024 - Kaggle Addiction ? - Version 115\nSucceeded · 1mo ago\n0.444\n\n0.434\n\nCMI 2024 - Kaggle Addiction ? - Version 114\nSucceeded · 1mo ago\n0.442\n\n0.440\n\nCMI 2024 - Kaggle Addiction ? - Version 113\nSucceeded · 1mo ago\n0.445\n\n0.448\n\nCMI 2024 - Kaggle Addiction ? - Version 112\nSucceeded · 1mo ago\n0.444\n\n0.444\n\nCMI 2024 - Kaggle Addiction ? - Version 110\nSucceeded · 1mo ago\n0.435\n\n0.432\n\nCMI 2024 - Kaggle Addiction ? - Version 109\nSucceeded · 1mo ago\n0.429\n\n0.437\n\nCMI 2024 - Kaggle Addiction ? - Version 108\nSucceeded · 1mo ago\n0.448\n\n0.442\n\nCMI 2024 - Kaggle Addiction ? - Version 107\nSucceeded · 1mo ago\n0.432\n\n0.442\n\nCMI 2024 - Kaggle Addiction ? - Version 106\nSucceeded · 2mo ago\n0.424\n\n0.434\n\nCMI 2024 - Kaggle Addiction ? - Version 105\nSucceeded · 2mo ago\n0.444\n\n0.442\n\nCMI 2024 - Kaggle Addiction ? - Version 104\nSucceeded · 2mo ago\n0.418\n\n0.445\n\nCMI 2024 - Kaggle Addiction ? - Version 103\nSucceeded · 2mo ago\n0.426\n\n0.443\n\nCMI 2024 - Kaggle Addiction ? - Version 102\nSucceeded · 2mo ago\n0.413\n\n0.430\n\nCMI 2024 - Kaggle Addiction ? - Version 101\nSucceeded · 2mo ago\n0.421\n\n0.439\n\nCMI 2024 - Kaggle Addiction ? - Version 100\nSucceeded · 2mo ago\n0.424\n\n0.443\n\nCMI 2024 - Kaggle Addiction ? - Version 99\nSucceeded · 2mo ago\n0.417\n\n0.448\n\nCMI 2024 - Kaggle Addiction ? - Version 97\nSucceeded · 2mo ago\n0.431\n\n0.458\n\nCMI 2024 - Kaggle Addiction ? - Version 96\nSucceeded · 2mo ago\n0.412\n\n0.437\n\nCMI 2024 - Kaggle Addiction ? - Version 95\nSucceeded · 2mo ago\n0.409\n\n0.401\n\nCMI 2024 - Kaggle Addiction ? - Version 94\nSucceeded · 2mo ago\n0.436\n\n0.439\n\nCMI 2024 - Kaggle Addiction ? - Version 89\nSucceeded · 2mo ago\n0.421\n\n0.428\n\nCMI 2024 - Kaggle Addiction ? - Version 90\nSucceeded · 2mo ago\n0.422\n\n0.423\n\nCMI 2024 - Kaggle Addiction ? - Version 88\nSucceeded · 2mo ago\n0.424\n\n0.430\n\nCMI 2024 - Kaggle Addiction ? - Version 87\nSucceeded · 2mo ago\n0.436\n\n0.448\n\nCMI 2024 - Kaggle Addiction ? - Version 86\nSucceeded · 2mo ago\n0.425\n\n0.431\n\nCMI 2024 - Kaggle Addiction ? - Version 85\nSucceeded · 2mo ago\n0.427\n\n0.450\n\nCMI 2024 - Kaggle Addiction ? - Version 84\nSucceeded · 2mo ago\n0.421\n\n0.452\n\nCMI 2024 - Kaggle Addiction ? - Version 83\nSucceeded · 2mo ago\n0.408\n\n0.426\n\nCMI 2024 - Kaggle Addiction ? - Version 82\nSucceeded · 2mo ago\n0.410\n\n0.428\n\nCMI 2024 - Kaggle Addiction ? - Version 81\nSucceeded · 2mo ago\n0.447\n\n0.432\n\nCMI 2024 - Kaggle Addiction ? - Version 78\nSucceeded · 2mo ago\n0.454\n\n0.442\n\nCMI 2024 - Kaggle Addiction ? - Version 77\nSucceeded · 2mo ago\n0.450\n\n0.454\n\nCMI 2024 - Kaggle Addiction ? - Version 76\nSucceeded · 2mo ago\n0.429\n\n0.435\n\nCMI 2024 - Kaggle Addiction ? - Version 75\nSucceeded · 2mo ago\n0.426\n\n0.433\n\nCMI 2024 - Kaggle Addiction ? - Version 74\nSucceeded · 2mo ago\n0.424\n\n0.428\n\nCMI 2024 - Kaggle Addiction ? - Version 73\nSucceeded · 2mo ago\n0.434\n\n0.420\n\nCMI 2024 - Kaggle Addiction ? - Version 72\nSucceeded · 2mo ago\n0.428\n\n0.422\n\nCMI 2024 - Kaggle Addiction ? - Version 71\nSucceeded · 2mo ago\n0.448\n\n0.433\n\nCMI 2024 - Kaggle Addiction ? - Version 70\nSucceeded · 2mo ago\n0.441\n\n0.434\n\nCMI 2024 - Kaggle Addiction ? - Version 69\nSucceeded · 2mo ago\n0.448\n\n0.425\n\nCMI 2024 - Kaggle Addiction ? - Version 66\nSucceeded · 2mo ago\n0.435\n\n0.437\n\nCMI 2024 - Kaggle Addiction ? - Version 65\nSucceeded · 2mo ago\n0.440\n\n0.439\n\nCMI 2024 - Kaggle Addiction ? - Version 64\nSucceeded · 2mo ago\n0.434\n\n0.432\n\nCMI 2024 - Kaggle Addiction ? - Version 63\nSucceeded · 2mo ago\n0.431\n\n0.457\n\nCMI 2024 - Kaggle Addiction ? - Version 62\nSucceeded · 2mo ago\n0.429\n\n0.442\n\nCMI 2024 - Kaggle Addiction ? - Version 61\nSucceeded · 2mo ago\n0.446\n\n0.442\n\nCMI 2024 - Kaggle Addiction ? - Version 60\nSucceeded · 2mo ago\n0.434\n\n0.433\n\nCMI 2024 - Kaggle Addiction ? - Version 59\nSucceeded · 2mo ago\n0.435\n\n0.450\n\nCMI 2024 - Kaggle Addiction ? - Version 58\nSucceeded · 2mo ago\n0.392\n\n0.446\n\nCMI 2024 - Kaggle Addiction ? - Version 57\nSucceeded · 2mo ago\n0.386\n\n0.448\n\nCMI 2024 - Kaggle Addiction ? - Version 54\nSucceeded · 2mo ago\n0.386\n\n0.449\n\nCMI 2024 - Kaggle Addiction ? - Version 53\nSucceeded · 2mo ago\n0.394\n\n0.449\n\nCMI 2024 - Kaggle Addiction ? - Version 51\nSucceeded · 2mo ago\n0.388\n\n0.435\n\nCMI 2024 - Kaggle Addiction ? - Version 50\nSucceeded · 2mo ago\n0.397\n\n0.413\n\nCMI 2024 - Kaggle Addiction ? - Version 48\nSucceeded · 2mo ago\n0.391\n\n0.432\n\nCMI 2024 - Kaggle Addiction ? - Version 47\nSucceeded · 2mo ago\n0.393\n\n0.453\n\nCMI 2024 - Kaggle Addiction ? - Version 46\nSucceeded · 2mo ago\n0.409\n\n0.451\n\nCMI 2024 - Kaggle Addiction ? - Version 44\nSucceeded · 2mo ago\n0.402\n\n0.457\n\nCMI 2024 - Kaggle Addiction ? - Version 43\nSucceeded · 2mo ago\n0.394\n\n0.453\n\nCMI 2024 - Kaggle Addiction ? - Version 41\nSucceeded · 2mo ago\n0.382\n\n0.437\n\nCMI 2024 - Kaggle Addiction ? - Version 40\nSucceeded · 2mo ago\n0.369\n\n0.422\n\nCMI 2024 - Kaggle Addiction ? - Version 39\nSucceeded · 2mo ago\n0.356\n\n0.401\n\nCMI 2024 - Kaggle Addiction ? - Version 38\nSucceeded · 3mo ago\n0.371\n\n0.439\n\nCMI 2024 - Kaggle Addiction ? - Version 35\nSucceeded · 3mo ago\n0.375\n\n0.435\n\nCMI 2024 - Kaggle Addiction ? - Version 34\nSucceeded · 3mo ago\n0.372\n\n0.415\n\nCMI 2024 - Kaggle Addiction ? - Version 33\nSucceeded · 3mo ago\n0.301\n\n0.422\n\nCMI 2024 - Kaggle Addiction ? - Version 31\nSucceeded · 3mo ago\n0.331\n\n0.426\n\nCMI 2024 - Kaggle Addiction ? - Version 30\nSucceeded · 3mo ago\n0.356\n\n0.419\n\nCMI 2024 - Kaggle Addiction ? - Version 29\nSucceeded · 3mo ago\n0.350\n\n0.432\n\nCMI 2024 - Kaggle Addiction ? - Version 27\nSucceeded · 3mo ago\n0.363\n\n0.436\n\nCMI 2024 - Kaggle Addiction ? - Version 25\nSucceeded · 3mo ago\n0.330\n\n0.426\n\nCMI 2024 - Kaggle Addiction ? - Version 23\nSucceeded · 3mo ago\n0.339\n\n0.408\n\nCMI 2024 - Kaggle Addiction ? - Version 21\nSucceeded · 3mo ago\n0.334\n\n0.412\n\nCMI 2024 - Kaggle Addiction ? - Version 21\nSucceeded · 3mo ago\n0.348\n\n0.401\n\nCMI 2024 - Kaggle Addiction ? - Version 20\nSucceeded · 3mo ago\n0.344\n\n0.416\n\nCMI 2024 - Kaggle Addiction ? - Version 18\nSucceeded · 3mo ago\n0.363\n\n0.412\n\nCMI 2024 - Kaggle Addiction ? - Version 16\nSucceeded · 3mo ago\n0.335\n\n0.326\n\nCMI 2024 - Kaggle Addiction ? - Version 15\nSucceeded · 3mo ago\n0.304\n\n0.319\n\nCMI 2024 - Kaggle Addiction ? - Version 14\nSucceeded · 3mo ago\n0.323\n\n0.319\n\nCMI 2024 - Kaggle Addiction ? - Version 13\nSucceeded · 3mo ago\n0.348\n\n0.313\n\nCMI 2024 - Kaggle Addiction ? - Version 11\nSucceeded · 3mo ago\n0.355\n\n0.330\n\nCMI 2024 - Kaggle Addiction ? - Version 10\nSucceeded · 3mo ago\n0.338\n\n0.334\n\nCMI 2024 - Kaggle Addiction ? - Version 8\nSucceeded · 3mo ago\n0.349\n\n0.326\n\nCMI 2024 - Kaggle Addiction ? - Version 7\nSucceeded · 3mo ago\n0.334\n\n0.299\n\nCMI 2024 - Kaggle Addiction ? - Version 6\nSucceeded · 3mo ago\n0.329\n\n0.294\n\nCMI 2024 - Kaggle Addiction ? - Version 5\nSucceeded · 3mo ago\n0.320\n\n0.281\n\nCMI 2024 - Kaggle Addiction ? - Version 4\nSucceeded · 3mo ago\n0.298\n\n0.296\n\nCMI 2024 - Kaggle Addiction ? - Version 3\nSucceeded · 3mo ago\n0.299\n\n0.305\n\nCMI 2024 - Kaggle Addiction ? - Version 2\nSucceeded · 3mo ago\n0.281\n\n0.273\n\nCMI 2024 - Kaggle Addiction ? - Version 1\nSucceeded · 3mo ago\n0.294\n\n0.247\n'''","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-20T21:02:27.537097Z","iopub.execute_input":"2024-12-20T21:02:27.537934Z","iopub.status.idle":"2024-12-20T21:02:27.553856Z","shell.execute_reply.started":"2024-12-20T21:02:27.537889Z","shell.execute_reply":"2024-12-20T21:02:27.552036Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Break it into pieces\nscore_pieces = my_scores.split()\nprint(score_pieces[0:45])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-20T21:02:27.556341Z","iopub.execute_input":"2024-12-20T21:02:27.556833Z","iopub.status.idle":"2024-12-20T21:02:27.579211Z","shell.execute_reply.started":"2024-12-20T21:02:27.556791Z","shell.execute_reply":"2024-12-20T21:02:27.577786Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# The pieces are periodic with period 15 -- depends on the notebook name\n# From index 0 the data are: [+8] is version number, [+13] is Private, [+14] is Public.\nvers = []; priv = []; publ = []\nfor izero in range(0, len(score_pieces), 15):\n    ##print(izero, score_pieces[izero+8])\n    vers.append(float(score_pieces[izero+8]))\n    priv.append(1000*float(score_pieces[izero+13]))\n    publ.append(1000*float(score_pieces[izero+14]))\nvers=np.array(vers); priv=np.array(priv); publ=np.array(publ); ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-20T21:02:27.580951Z","iopub.execute_input":"2024-12-20T21:02:27.581626Z","iopub.status.idle":"2024-12-20T21:02:27.598266Z","shell.execute_reply.started":"2024-12-20T21:02:27.581561Z","shell.execute_reply":"2024-12-20T21:02:27.596640Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Best Private score - 454 from version 78\nimax = np.argmax(priv)\nprint(\"Best Private score:\", priv[imax], \" from version\", vers[imax])\n\nplt.figure(figsize=(8,5))\nplt.scatter(vers,publ,c=\"orange\",s=7,label=\"Public LB\")\nplt.scatter(vers,priv,c=\"blue\",s=7,label=\"Private LB\")\n# Zoom in after version 60\n##plt.xlim(60,160); plt.ylim(400,470)\n# Show 430, 440, and 450\nfor score in [430, 440, 450]:\n    plt.plot([1,159],[score,score],c='gray',alpha=0.2)\nplt.legend(loc=\"lower right\")\nplt.xlabel(\"Version number\")\nplt.ylabel(\"LB Score x10$^3$\")\nplt.title(\"LB Scores vs Version number  (lines at 430, 440, 450)\")\nplt.savefig('Private_LB_vs_version.png')\nplt.show()\n\n# Public LB jumps at v18: had accidentally left out features\n#   SDS-SDS_Total_T and PreInt_EduHx-computerinternet_hoursday\n# Private LB jumps from 392 at version 58 (446 Pub) -->  435 at version 59 (450 Pub)\n#   Added train-quantile \"optimization\" of y_hat values.","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-20T21:04:36.283459Z","iopub.execute_input":"2024-12-20T21:04:36.284870Z","iopub.status.idle":"2024-12-20T21:04:36.763243Z","shell.execute_reply.started":"2024-12-20T21:04:36.284823Z","shell.execute_reply":"2024-12-20T21:04:36.761822Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(6,6))\nplt.scatter(publ,priv,s=20,c='blue',alpha=0.4)\n# Add the equal line\nplt.plot([400,465],[400,465],c='gray')\n# Uncomment next line to Zoom in on high scores region\n##plt.xlim(400,465); plt.ylim(400,465)\n#\nplt.xlabel(\"Public LB x10$^3$\")\nplt.ylabel(\"Private LB x10$^3$\")\nplt.title(\"Private LB vs Public LB  (equal on the gray line)\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-20T21:02:28.059986Z","iopub.execute_input":"2024-12-20T21:02:28.060450Z","iopub.status.idle":"2024-12-20T21:02:28.302313Z","shell.execute_reply.started":"2024-12-20T21:02:28.060410Z","shell.execute_reply":"2024-12-20T21:02:28.301007Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}