{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30775,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Preliminary Stuffs\n\nThis EDA will be a deep dive into the Child Internet addiction dataset. Both HBN Instruments and Actigraphy data will be examined closely. The interaction between both datasets will also be examined. \n\n# Loading the Data\n\n## Libraries being used\n\nlets load the necessary libraries and also take a look at the data dictionary to understand the dataset we will be playing with.","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport scipy\nimport pandas as pd \n\n\nimport os\n\nimport datetime\n\nfrom matplotlib import pyplot as plt\nimport seaborn as sns\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:20:57.957176Z","iopub.execute_input":"2024-11-18T02:20:57.957738Z","iopub.status.idle":"2024-11-18T02:21:01.185010Z","shell.execute_reply.started":"2024-11-18T02:20:57.957681Z","shell.execute_reply":"2024-11-18T02:21:01.183565Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## The Data Dictionary\n\nWe have been provided a data dictionary for the HBN Instruments dataset, so lets take a look at it!","metadata":{}},{"cell_type":"code","source":"data_dict = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/data_dictionary.csv')","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2024-11-18T02:21:01.187187Z","iopub.execute_input":"2024-11-18T02:21:01.187714Z","iopub.status.idle":"2024-11-18T02:21:01.209392Z","shell.execute_reply.started":"2024-11-18T02:21:01.187660Z","shell.execute_reply":"2024-11-18T02:21:01.207804Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#the style method conveniently shows all text within the cell \ndata_dict.style.set_properties(**{'width' : '100px'}).set_caption(\"HBN Instrument Data Dictionary\")","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2024-11-18T02:21:01.211079Z","iopub.execute_input":"2024-11-18T02:21:01.211575Z","iopub.status.idle":"2024-11-18T02:21:01.327842Z","shell.execute_reply.started":"2024-11-18T02:21:01.211522Z","shell.execute_reply":"2024-11-18T02:21:01.326715Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Reading in the Data\n\nLets start reading in the datasets. \n\n\n### HBN Instruments files\n\nlets load the non-parquet train and test data first. \n\nSince these processes do take a bit of time, I've added a progress bar for convenience and to maintain my sanity. \n\nwe will also apply the scale label from the data dictionary to the PCIAT total column to create a new \"Severity Rating\" column","metadata":{}},{"cell_type":"code","source":"# Print iterations progress. not necessary for analysis, but nice to see progress. \ndef printProgressBar (iteration, total, prefix = '', suffix = '', decimals = 1, length = 100, fill = '█', printEnd = \"\\r\"):\n    \n    percent = (\"{0:.\" + str(decimals) + \"f}\").format(100 * (iteration / float(total)))\n    filledLength = int(length * iteration // total)\n    bar = fill * filledLength + '-' * (length - filledLength)\n    print(f'\\r{prefix} |{bar}| {percent}% {suffix}', end = printEnd)\n    # Print New Line on Complete\n    if iteration == total: \n        print()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:21:01.330434Z","iopub.execute_input":"2024-11-18T02:21:01.330956Z","iopub.status.idle":"2024-11-18T02:21:01.338482Z","shell.execute_reply.started":"2024-11-18T02:21:01.330912Z","shell.execute_reply":"2024-11-18T02:21:01.337198Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_HBN = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\ntest_HBN = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:21:01.339996Z","iopub.execute_input":"2024-11-18T02:21:01.340375Z","iopub.status.idle":"2024-11-18T02:21:01.420507Z","shell.execute_reply.started":"2024-11-18T02:21:01.340337Z","shell.execute_reply":"2024-11-18T02:21:01.419270Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_HBN[\"PCIAT_Severity_Rating\"] = train_HBN[\"PCIAT-PCIAT_Total\"].apply(lambda x: \"None\" if x <= 30 else \n                                                                             \"Mild\" if x >= 31 and x <= 49 else \n                                                                             \"Moderate\" if x >= 50 and x <= 79 else\n                                                                             \"Severe\" if x >= 80 else \"error\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:21:01.422155Z","iopub.execute_input":"2024-11-18T02:21:01.422581Z","iopub.status.idle":"2024-11-18T02:21:01.432205Z","shell.execute_reply.started":"2024-11-18T02:21:01.422539Z","shell.execute_reply":"2024-11-18T02:21:01.430576Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Actigraphy files\n\nlets load the train parquet files for EDA.\n\nlets load each parquet into a dataframe. We will have a list of dataframes for each associated child ID. \n\nWe need to conserve space here, so we won't concatenate the list of dataframes. ","metadata":{}},{"cell_type":"code","source":"%%time\n\ndef pull_child_categorical(col = \"PCIAT_Severity_Rating\", df = train_HBN, child_id = \"00115b9f\"):\n    cat_val = df[df.id == child_id][col].reset_index(drop = True)[0]\n    \n    return(cat_val)\n \n\n#renaming the cat cols as old ones are obscure. building dict in case I want to rename in the train df too.\ncat_cols = [\"PCIAT_Severity_Rating\", \"Basic_Demos-Sex\", \"FGC-FGC_CU_Zone\", \"FGC-FGC_GSND_Zone\", \n\"FGC-FGC_GSD_Zone\",\"FGC-FGC_PU_Zone\", \n\"FGC-FGC_SRL_Zone\", \"FGC-FGC_SRR_Zone\", \"FGC-FGC_TL_Zone\",\n\"BIA-BIA_Activity_Level_num\", \"BIA-BIA_Frame_num\",\n\"PreInt_EduHx-computerinternet_hoursday\"]\n\ncat_cols_better_names = [\"internet_addiction_severity\", \"sex\", \"curl_up\", \"grip_non_dominant\", \"grip_dominant\",\n                       \"push_up\", \"left_sit_and_reach\", \n                       \"right_sit_and_reach\", \"trunk_lift\", \"activity_level\",\n                       \"body_frame\", \"Hours_of_daily_Internet\"]\n\nrename_dict = {old_name : new_name for old_name, new_name in zip(cat_cols, cat_cols_better_names)}\n\ndef Actigraphy_df_List(dir, cat_cols = cat_cols, cat_cols_better_names = cat_cols_better_names):\n\n    printProgressBar(0, len(os.listdir(dir)), \"Loading Data\", \"Complete\")\n\n    Train_Actigraphy = []\n    \n    #a list of the ids from the parquet files. build here for convenience\n    parquet_ids = []\n    \n    #build list of each parquet DF, add a few columns\n    for i, folder in enumerate(os.listdir(dir)):\n    \n        child_id = folder.replace(\"id=\", \"\")\n        \n        parquet_ids.append(child_id)\n        \n        mini_parquet = pd.read_parquet(dir + \"/\" + folder + \"/part-0.parquet\", engine = 'pyarrow')\n        \n        #change col types to category to save on space\n        \n        mini_parquet[\"id\"] = pd.Series(np.full(shape = mini_parquet.shape[0], fill_value = child_id), \n                                       dtype = \"category\")\n        \n        #squeeze in cat cols that may be useful for analysis, cat data type chosen to save RAM\n        \n        for cat in cat_cols:\n            \n            cat_val = pull_child_categorical(cat, train_HBN, child_id)\n            \n            mini_parquet[rename_dict[cat]] = pd.Series(np.full(shape=mini_parquet.shape[0],fill_value = cat_val), \n                                                       dtype = \"category\")\n    \n        \n        mini_parquet[\"time_of_day\"] = (mini_parquet.time_of_day/1000000000).astype('int32')\n        mini_parquet[\"Days\"] = (mini_parquet.relative_date_PCIAT - mini_parquet.relative_date_PCIAT.min()).astype('int8')\n        #the way in which total seconds is calculated assumed that the \"Days\" field represented continuous study PARTICIPATION\n        #this is not the case, so it is removed from future calculations (see section on Actigraphy data examination)\n        #mini_parquet[\"total_seconds\"] = mini_parquet.Days * (24 * 60 * 60) + mini_parquet.time_of_day\n        #mini_parquet[\"time_of_day_dt\"] = mini_parquet.apply(lambda x: datetime.datetime(2024, 1, 1, 0, 0, 0) + datetime.timedelta(x[\"Days\"], x[\"time_of_day\"]), axis = 1)\n        \n        mini_parquet = mini_parquet.drop([\"step\", \"X\", \"Y\", \"Z\", \"battery_voltage\", \"relative_date_PCIAT\"], axis = 1)\n        \n        \n        Train_Actigraphy.append(mini_parquet)\n        \n        printProgressBar(i + 1, len(os.listdir(dir)), \"Loading Data\", \"Complete\")\n\n    return(Train_Actigraphy, parquet_ids)\n\n\nTrain_Actigraphy, Train_Actigraphy_ids = Actigraphy_df_List('/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet')\n\nTest_Actigraphy, Test_Actigraphy_ids = Actigraphy_df_List('/kaggle/input/child-mind-institute-problematic-internet-use/series_test.parquet')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:21:01.434089Z","iopub.execute_input":"2024-11-18T02:21:01.434509Z","iopub.status.idle":"2024-11-18T02:25:51.438466Z","shell.execute_reply.started":"2024-11-18T02:21:01.434448Z","shell.execute_reply":"2024-11-18T02:25:51.437143Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Data examination and cleaning\n\nLets take a close look at the data. Get a sense of the overall dataset, Check for missing values, strange responses, and other oddities. \n\n## HBN Instruments examination and cleaning\n\n### HBN Instruments examination\n\n#### HBN Instruments - general record and field counting and comparisons\n\nlets get some quick information on the HBN dataset. lets compare the shapes of both the train and test sets. ","metadata":{}},{"cell_type":"code","source":"pd.DataFrame({\"Record Count\" : [train_HBN.shape[0], test_HBN.shape[0]], \"Field Count\" : [train_HBN.shape[1], test_HBN.shape[1]]}, index = [\"Train\", \"Test\"]).style.set_caption(\"Comparison of Train and Test dataset\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:25:51.440545Z","iopub.execute_input":"2024-11-18T02:25:51.441050Z","iopub.status.idle":"2024-11-18T02:25:51.452996Z","shell.execute_reply.started":"2024-11-18T02:25:51.440994Z","shell.execute_reply":"2024-11-18T02:25:51.451641Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There are less fields within the test set than anticipated. Lets examine this difference between test and training set. ","metadata":{}},{"cell_type":"code","source":"columns_missing_from_test = [f for f in train_HBN.columns if f not in test_HBN]\n\npd.DataFrame({\"Missing Fields in Test\" :  columns_missing_from_test}).set_index(\"Missing Fields in Test\").style.set_caption(\"Comparison of Train and Test dataset\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:25:51.454605Z","iopub.execute_input":"2024-11-18T02:25:51.455059Z","iopub.status.idle":"2024-11-18T02:25:51.469384Z","shell.execute_reply.started":"2024-11-18T02:25:51.455006Z","shell.execute_reply":"2024-11-18T02:25:51.467882Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Interesting that all of the PCIAT fields were removed from the test set. I can see to an extent why the majority of them would be removed, however the PCIAT-Season should have been kept. We will need to avoid using these fields in any supervised learning algorithms to follow. \n\n#### HBN Instruments - Actigraphy filtering\n\nWe have a list of ids from the actigraphy data; lets see where they line up with the HBN Instruments dataset.","metadata":{}},{"cell_type":"code","source":"for id_list, id_string in zip([Train_Actigraphy_ids, Test_Actigraphy_ids], [\"Train_Actigraphy_IDs\", \"Test_Actigraphy_IDs\"]):\n\n    for HBN_frame, HBN_string in zip([train_HBN, test_HBN], [\"Train_HBN\", \"Test_HBN\"]):\n\n        HBN_frame_id = HBN_frame.id.drop_duplicates().values\n\n        idlist_in_HBN = [id for id in id_list if id in HBN_frame_id]\n\n        idlist_prop = str(len(idlist_in_HBN)/len(id_list))\n\n        HBN_in_idlist = [id for id in HBN_frame_id if id in id_list]\n\n        HBN_prop = str(len(HBN_in_idlist)/len(HBN_frame_id))\n\n        print(\"The proportion of \" + id_string + \" found within the \" + HBN_string + \" dataframe is: \" + \n             idlist_prop + \"\\n\\nlikewise, the proportion of ids within \" + HBN_string + \" dataframe found within the \" + \n             id_string + \" list is \" + HBN_prop + \"\\n\\n\\n\")\n\nprint(\"it looks like both the Test Actigraphy IDs: \" + str([c for c in Test_Actigraphy_ids if c in Train_Actigraphy_ids]) +\n      \" are present in the training Actigraphy set.\\n\\n\")\n\nprint(\"it also looks like all \" + \n      str(len([c for c in test_HBN.id.drop_duplicates().values if c in train_HBN.id.drop_duplicates().values])) +\n      \" of the HBN test set ids are present within the HBN training set.\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:25:51.473634Z","iopub.execute_input":"2024-11-18T02:25:51.474033Z","iopub.status.idle":"2024-11-18T02:25:51.739777Z","shell.execute_reply.started":"2024-11-18T02:25:51.473994Z","shell.execute_reply":"2024-11-18T02:25:51.738568Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Looks like the actigraphy data ids are all accounted for within the HBN datasets. However it seems that these ids only account for a small proportion of the total HBN data itself. \n\nWell this is a bit awkward. it looks like there may be a mixture of test and training data between the HBN and actigraphy datasets. Will need to keep this in mind when it comes to any supervised learning exercises.\n","metadata":{}},{"cell_type":"markdown","source":"### HBN Instruments cleaning\n\n#### checking for missing values\n\nlets check for missing values for the full HBN table and also the Actigraphy-filtered table. I assume that the table with Actigraphy members would have fewer missing values.","metadata":{}},{"cell_type":"code","source":"train_HBN_Actigraphyids = train_HBN[train_HBN.id.apply(lambda x: x in Train_Actigraphy_ids)].reset_index(drop = True)\n\nplt.rcParams[\"figure.figsize\"] = (20, 10)\n\nsns.heatmap(train_HBN.isnull(), cbar = False)\n\nplt.title(\"Full train_HBN null check\")\n\nplt.show()\n\nplt.rcParams[\"figure.figsize\"] = (20, 10)\n\nsns.heatmap(train_HBN_Actigraphyids.isnull(), cbar = False)\nplt.title(\"Actigraphy_Filtered_HBN null check\")\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:25:51.741085Z","iopub.execute_input":"2024-11-18T02:25:51.741444Z","iopub.status.idle":"2024-11-18T02:25:55.106021Z","shell.execute_reply.started":"2024-11-18T02:25:51.741405Z","shell.execute_reply":"2024-11-18T02:25:55.104859Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Just as I assumed, there is a greater proportion of filled out data in the Actigraphy-filtered dataset. looks like a lot of kids didnt go through with the FitnessGram vitals and treadmill measures as the Fitness_Endurance-Max_Stage to Fitness_Endurance_Time_sec show a lot of missing data and appear to be associated with each other. This also is the case for Grip-Strength measurements. \n\nI should make a note for the missing data within the Physical Activity questionaire columns. These are separated into adolescent and child, therefore there is the possibility that if one column has a value the other must be null. This will need to be fixed for the dataset and the null values examined again...but first...I should confirm some assumptions.\n\n##### Physical Activity Questionaire Assumption Testing\n\nlets test 2 assumptions:\n1. If a Season field is null for either the adolescent or child category, then the associated Total field will also be null\n2. Assuming assumption 1 is true, Child and adolescent columns cannot both be non-null, but both sets can be null. ","metadata":{}},{"cell_type":"code","source":"adolescent_activity = [[\"PAQ_A-Season\"], [\"PAQ_A-PAQ_A_Total\"]]\nchild_activity = [[\"PAQ_C-Season\"], [\"PAQ_C-PAQ_C_Total\"]]\n\nidx_list = []\n\n#test if side a is null and side b is not null and vice versa. \n#catches if both sides are null or have data but does not determine\n#which case. \nfor a, c in zip(adolescent_activity, child_activity):\n    \n    i = 0\n    \n    activity_df = ((train_HBN[a].isnull()[a[i]] == train_HBN[c].notnull()[c[i]]) |\n                  (train_HBN[a].notnull()[a[i]] == train_HBN[c].isnull()[c[i]]))\n    \n    idx_list.append(train_HBN[activity_df].index)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:25:55.107552Z","iopub.execute_input":"2024-11-18T02:25:55.107906Z","iopub.status.idle":"2024-11-18T02:25:55.129069Z","shell.execute_reply.started":"2024-11-18T02:25:55.107869Z","shell.execute_reply":"2024-11-18T02:25:55.127962Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#assumption 1 test: if Season is null for adolescents, then Total will also be null. This would also be true\n#for children.\nprint(\"Assumption 1: \" + str(all(idx_list[0] == idx_list[1])) + \"\\nLooks like we are correct with the first assumption\\n\")\n\n\n#assumption 2 test: if a record has a value in the adolescent columns it will not\n#have a record in the child columns. take full table, take out records found from assumption 1 test,\n#take out records where both sets are null. \n\nassumption_2_test = train_HBN.shape[0] - train_HBN.iloc[idx_list[0]].shape[0] - train_HBN[train_HBN[[i[0] for i in adolescent_activity + child_activity]].isnull().sum(axis = 1) == 4].shape[0]\n\nprint(\"Assumption 2: \" + str(assumption_2_test == 0))\nprint(\"Looks like we are incorrect with assumption 2. There was \" + str(assumption_2_test) + \" record that has data for all tested columns\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:25:55.130854Z","iopub.execute_input":"2024-11-18T02:25:55.131292Z","iopub.status.idle":"2024-11-18T02:25:55.148083Z","shell.execute_reply.started":"2024-11-18T02:25:55.131239Z","shell.execute_reply":"2024-11-18T02:25:55.146533Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This is very interesting. Only one record within the dataset has data for both sets of child and adolescent columns. Lets take a look at the record's test values for those columns. The child may have aged into adolescence during the testing period, so lets also see if the child's age is 13. ","metadata":{}},{"cell_type":"code","source":"train_HBN[train_HBN[[i[0] for i in adolescent_activity + child_activity]].isnull().sum(axis = 1) == 0][[i[0] for i in adolescent_activity + child_activity] + [\"Basic_Demos-Age\"] + [\"id\"]]","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:25:55.149681Z","iopub.execute_input":"2024-11-18T02:25:55.150049Z","iopub.status.idle":"2024-11-18T02:25:55.173419Z","shell.execute_reply.started":"2024-11-18T02:25:55.150009Z","shell.execute_reply":"2024-11-18T02:25:55.172119Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It looks like the child is indeed 13, so they may have taken the child test at 12 in spring, then the adolescent test at 13 in the summer. I'll need to think on whether to include this record in the dataset. We will dig into the details of this child later. First lets continue examining the HBN dataset. \n\nWe can combine the child and Adolescent PAQ columns into a single set with a child/adolescent identifier to the side while we use a special identifier for our odd child. lets rerun the investigation of null values with this updated column type. ","metadata":{}},{"cell_type":"code","source":"train_HBN[\"PAQ_Season\"] = train_HBN.apply(lambda x: x[\"PAQ_A-Season\"] if pd.notnull(x[\"PAQ_A-Season\"]) else\n                                          x[\"PAQ_C-Season\"] if pd.notnull(x[\"PAQ_C-Season\"]) else\n                                          np.NaN, axis = 1)\n\ntrain_HBN[\"PAQ_Total\"] = train_HBN.apply(lambda x: x[\"PAQ_A-PAQ_A_Total\"] if pd.notnull(x[\"PAQ_A-PAQ_A_Total\"]) else\n                                         x[\"PAQ_C-PAQ_C_Total\"] if pd.notnull(x[\"PAQ_C-PAQ_C_Total\"]) else\n                                         np.NaN, axis = 1)\n\ntrain_HBN[\"PAQ_A_or_C\"] = train_HBN.apply(lambda x: \"d74e4d7c\" if pd.notnull(x[\"PAQ_A-Season\"]) and pd.notnull(x[\"PAQ_C-Season\"]) else\n                                          \"Adolescent\" if pd.notnull(x[\"PAQ_A-Season\"]) else \n                                          \"Child\" if pd.notnull(x[\"PAQ_C-Season\"]) else \n                                          np.NaN, axis = 1)\n\nPAQ_Index_Start = train_HBN.columns.get_loc('PAQ_A-Season')\n\n#instead of dropping the columns lets just push them to the end. \ntrain_HBN = train_HBN[[c for c in train_HBN.columns[0:PAQ_Index_Start]] + \n                      [\"PAQ_Season\", \"PAQ_Total\", \"PAQ_A_or_C\"] + \n                      [c for c in train_HBN.columns[PAQ_Index_Start:len(train_HBN.columns)] if c not in \n                       [\"PAQ_A-Season\", \"PAQ_C-Season\", \"PAQ_A-PAQ_A_Total\", \"PAQ_C-PAQ_C_Total\", \"PAQ_Season\", \"PAQ_Total\", \"PAQ_A_or_C\"]] +\n                      \n                      [\"PAQ_A-Season\", \"PAQ_C-Season\", \"PAQ_A-PAQ_A_Total\", \"PAQ_C-PAQ_C_Total\"]]\n\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:25:55.174802Z","iopub.execute_input":"2024-11-18T02:25:55.175151Z","iopub.status.idle":"2024-11-18T02:25:55.400875Z","shell.execute_reply.started":"2024-11-18T02:25:55.175112Z","shell.execute_reply":"2024-11-18T02:25:55.399750Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### Checking for missing values after updating the PAQ columns","metadata":{}},{"cell_type":"code","source":"plt.rcParams[\"figure.figsize\"] = (20, 10)\n\nsns.heatmap(train_HBN.isnull(), cbar = False)\n\nplt.title(\"Full train_HBN null check\")\n\nplt.show()\n\ntrain_HBN_Actigraphyids = train_HBN[train_HBN.id.apply(lambda x: x in Train_Actigraphy_ids)].reset_index(drop = True)\n\nplt.rcParams[\"figure.figsize\"] = (20, 10)\n\nsns.heatmap(train_HBN_Actigraphyids.isnull(), cbar = False)\nplt.title(\"Actigraphy_Filtered_HBN null check\")\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:25:55.402201Z","iopub.execute_input":"2024-11-18T02:25:55.402542Z","iopub.status.idle":"2024-11-18T02:25:58.763278Z","shell.execute_reply.started":"2024-11-18T02:25:55.402507Z","shell.execute_reply":"2024-11-18T02:25:58.762082Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The PAQ columns look much more populated now. Based on an eyeball test, I'd like to remove 3 sets of fields for having too many missing values: the waist circumference column, the Fitness_Endurance set of fields, and the FGC-GS fields. Lets first take a look at what proportion of records in the dataset have missing values in each of these fields:","metadata":{}},{"cell_type":"code","source":"Train_HBN_missing_most = [\"Physical-Waist_Circumference\", \"Fitness_Endurance-Max_Stage\", \"Fitness_Endurance-Time_Mins\", \"Fitness_Endurance-Time_Sec\",\n                         \"FGC-FGC_GSND\", \"FGC-FGC_GSD\", \"FGC-FGC_GSD_Zone\", \"Fitness_Endurance-Season\"]\n                         \nprint(\"Fields and their proportion of missing data\\n\")\n\nproportion_missing = pd.isnull(train_HBN).sum().sort_values(ascending = False) / train_HBN.shape[0]\n\nprint(proportion_missing[Train_HBN_missing_most])\n\n#generate list of columns to drop. keep the PAQ_A/C set of columns safe from being dropped. \ncolumns_to_drop = [c for c in proportion_missing[proportion_missing >= 0.70].index \n                   if c not in [i[0] for i in adolescent_activity + child_activity]] + [\"Fitness_Endurance-Season\"]\n\ntrain_HBN = train_HBN.drop(columns_to_drop, axis = 1)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:25:58.764879Z","iopub.execute_input":"2024-11-18T02:25:58.765385Z","iopub.status.idle":"2024-11-18T02:25:58.789014Z","shell.execute_reply.started":"2024-11-18T02:25:58.765335Z","shell.execute_reply":"2024-11-18T02:25:58.787511Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Based on the proportions found, I'll exclude any fields in the dataset that is missing 70% or more data, also the Fitness_Endurance-Season field as it is closely related to the other Fitness_Endurance fields. Note that we are keeping the PAQ_A/C columns at the end as they were used to define the general PAQ columns.\n\nLets take one last look at the missing records in the HBN Instruments dataset following all of these updates. ","metadata":{}},{"cell_type":"code","source":"plt.rcParams[\"figure.figsize\"] = (20, 10)\n\nsns.heatmap(train_HBN.isnull(), cbar = False)\n\nplt.title(\"Full train_HBN null check\")\n\nplt.show()\n\ntrain_HBN_Actigraphyids = train_HBN[train_HBN.id.apply(lambda x: x in Train_Actigraphy_ids)].reset_index(drop = True)\n\nplt.rcParams[\"figure.figsize\"] = (20, 10)\n\nsns.heatmap(train_HBN_Actigraphyids.isnull(), cbar = False)\nplt.title(\"Actigraphy_Filtered_HBN null check\")\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:25:58.790542Z","iopub.execute_input":"2024-11-18T02:25:58.791003Z","iopub.status.idle":"2024-11-18T02:26:02.202112Z","shell.execute_reply.started":"2024-11-18T02:25:58.790951Z","shell.execute_reply":"2024-11-18T02:26:02.200978Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"With the removal of these fields we will also need to redefine our categorical columns and dict","metadata":{}},{"cell_type":"code","source":"cat_cols = [\"PCIAT_Severity_Rating\", \"Basic_Demos-Sex\", \"FGC-FGC_CU_Zone\", \"FGC-FGC_PU_Zone\", \n\"FGC-FGC_SRL_Zone\", \"FGC-FGC_SRR_Zone\", \"FGC-FGC_TL_Zone\",\n\"BIA-BIA_Activity_Level_num\", \"BIA-BIA_Frame_num\",\n\"PreInt_EduHx-computerinternet_hoursday\"]\n\ncat_cols_better_names = [\"internet_addiction_severity\", \"sex\", \"curl_up\", \n                       \"push_up\", \"left_sit_and_reach\", \n                       \"right_sit_and_reach\", \"trunk_lift\", \"activity_level\",\n                       \"body_frame\", \"Hours_of_daily_Internet\"]\n\nrename_dict = {old_name : new_name for old_name, new_name in zip(cat_cols, cat_cols_better_names)}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-18T02:26:02.203558Z","iopub.execute_input":"2024-11-18T02:26:02.204323Z","iopub.status.idle":"2024-11-18T02:26:02.212030Z","shell.execute_reply.started":"2024-11-18T02:26:02.204249Z","shell.execute_reply":"2024-11-18T02:26:02.210642Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Actigraphy Data Examination and Cleaning\n\n#### Actigraphy Data Examination\n\nLets look at the data frame collection in the Train_Actigraphy list. Do they have the same number of fields? What is up with their record counts? ","metadata":{}},{"cell_type":"code","source":"Actigraphy_shapes = pd.DataFrame({\"record_count\" : [c.shape[0] for c in Train_Actigraphy], \"Field_Count\" : [c.shape[1] for c in Train_Actigraphy]})\n\nprint(\"Looking at all actigraphy data within the training group, all dataframes have the same number of fields: \" +\n      str(Actigraphy_shapes.Field_Count.drop_duplicates().shape[0] == 1))\n\nprint(\"\\nLets now look at the summarized record counts for these dataframes. A simple group and describe should give us a good idea of what things look like.\\n\\n\" +\n     \"Actigraphy Dataframes Record Count Summary\")\n\nprint(Actigraphy_shapes.describe()[\"record_count\"].apply(lambda x: \"{:,.2f}\".format(x)))\n\nprint(\"The dataset description states that the subjects were meant to wear the accelerometer for a continuous period of up to 30 days.\\n\" +\n      \"If each record was taken within 5 second increments of each other, then the total number of records required to represent 30 days is: \" +\n      \"{:,.0f}\".format((30 * 24 * 60 * 60 / 5)) + \" records\"\n      \n      )","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:26:02.213592Z","iopub.execute_input":"2024-11-18T02:26:02.214542Z","iopub.status.idle":"2024-11-18T02:26:02.244908Z","shell.execute_reply.started":"2024-11-18T02:26:02.214486Z","shell.execute_reply":"2024-11-18T02:26:02.243403Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"##### Actigraphy Data Record Count Deep Dive\n\nLets more closely examine how record count varries among the Actigraphy data. ","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize = (20, 10))\n\n#using the seaborne kde function would be more convenient, but we want to be able to identify\n#local mins and maximums in the density function. Therefore, we need to be able to access the\n#density function itself. \ndensity = scipy.stats.gaussian_kde(Actigraphy_shapes.record_count)\ndensity.covariance_factor = lambda: 0.25\ndensity._compute_covariance()\n\n#plt.figure(figsize=(14,8))\n\nxs = np.linspace(0, max(Actigraphy_shapes.record_count), 200)\n\n#collect local mins and maxes on the density line\nlocal_min_max = []\nfor i, val in enumerate(density(xs)):\n    \n    #case 1: at the beginning or at the end:\n    if i == 0 or i == len(density(xs)) - 1:\n        pass\n    #case 2: local max (left and right less than middle)\n    elif density(xs)[i - 1] < val and density(xs)[i + 1] < val:\n        local_min_max.append([\"max\", i])\n    #case 3: local min (left and right greater than middle)\n    elif density(xs)[i - 1] > val and density(xs)[i + 1] > val:\n        local_min_max.append([\"min\", i])\n#plot the data\nplt.fill_between( xs, density(xs), color=\"#69b3a2\", alpha=0.4)\n\n#annotations galore!\ndef record_count_to_days(record_count):\n    \"\"\"just an estimate of the number of days worth\n    of data that a dataframe contains\"\"\"\n\n    days = record_count / (24 * 60 * 60 / 5)\n\n    days = \"{:.2f}\".format(days)\n    return(days)\n\nviz_unit = max(density(xs))/100\n    \nfor max_min in local_min_max:\n\n    x = xs[max_min[1]]\n\n    y = density(xs)[max_min[1]]\n\n    msg = record_count_to_days(x) + \" days\"\n\n    \n    if max_min[0] == \"max\":\n        color = 'r'\n    else:\n        color = 'b'\n\n    plt.vlines(x, 0, y, color = color)\n\n    plt.annotate(msg, xy = (x, y + viz_unit * 2))\n\n\n\nplt.vlines(518400.0, 0, 5*10**(-6), color = 'r')\n\nbbox = dict(boxstyle = \"round\", fc = \"0.8\", color = 'w')\narrowprops = dict(\n    arrowstyle = \"->\",\n    connectionstyle = \"angle, angleA=90, angleB=0, rad=10\"\n)\n\nplt.annotate(\"30 days worth of records\", xy = (520400.0, 4*10**(-6)), xytext= (600000, 5.5*10**(-6)), \n             bbox = bbox, arrowprops = arrowprops) \nplt.title(\"Distribution of record counts for all Actigraphy dataframes (with day annotations)\")\nplt.ylabel(\"Relative Density\");","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:26:02.246891Z","iopub.execute_input":"2024-11-18T02:26:02.247373Z","iopub.status.idle":"2024-11-18T02:26:08.558438Z","shell.execute_reply.started":"2024-11-18T02:26:02.247321Z","shell.execute_reply":"2024-11-18T02:26:08.557358Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Grouping the Actigraphy data may help with understanding its nature. Accordingly, lets use the above visualization to split the dataset into three conceptual groups.  \n\nThe 3 groups of subjects with Actigraphy data are as follows: \n1. Those with less than 10 days worth of Actigraphy data\n2. Those with between 10 days and 30 days worth of Actigraphy data\n3. Those with greater than 30 days worth of Actigraphy data","metadata":{}},{"cell_type":"code","source":"record_days_10 = 10 * (24*60*60/5)\nrecord_days_30 = 30 * (24*60*60/5)\n\nrecord_less_10 = sum(Actigraphy_shapes.record_count <= record_days_10)\n\nrecord_between_10_30 = sum(Actigraphy_shapes.record_count <= record_days_30) - sum(Actigraphy_shapes.record_count < record_days_10)\n\nrecord_greater_30 = sum(Actigraphy_shapes.record_count > record_days_30)\n\nplt.bar([\"Group 1 (10 or less)\", \"Group 2 (between 10 and 30)\", \"Group 3 (30 or more)\"], [record_less_10, record_between_10_30, record_greater_30])\nplt.title(\"Artificial Groupings with record counts\")\n\nfor name, count in zip([\"Group 1 (10 or less)\", \"Group 2 (between 10 and 30)\", \"Group 3 (30 or more)\"], [record_less_10, record_between_10_30, record_greater_30]):\n    plt.annotate(str(count), xy = (name, count + 10)) \n    \nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:26:08.559939Z","iopub.execute_input":"2024-11-18T02:26:08.560309Z","iopub.status.idle":"2024-11-18T02:26:08.908675Z","shell.execute_reply.started":"2024-11-18T02:26:08.560272Z","shell.execute_reply":"2024-11-18T02:26:08.907198Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"##### Train_Actigraphy grouping examination\n\nLets look at the consistency of participation based on weekday for each group.","metadata":{}},{"cell_type":"code","source":"record_less_10_filter = Actigraphy_shapes[Actigraphy_shapes.record_count <= record_days_10].index\nrecord_between_10_30_filter  = Actigraphy_shapes[(Actigraphy_shapes.record_count > record_days_10) & (Actigraphy_shapes.record_count <= record_days_30)].index\nrecord_greater_30_filter = Actigraphy_shapes[Actigraphy_shapes.record_count > record_days_30].index\n\ngroup_filters = [record_less_10_filter, record_between_10_30_filter, record_greater_30_filter]","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:26:08.910590Z","iopub.execute_input":"2024-11-18T02:26:08.911124Z","iopub.status.idle":"2024-11-18T02:26:08.922608Z","shell.execute_reply.started":"2024-11-18T02:26:08.911068Z","shell.execute_reply":"2024-11-18T02:26:08.921287Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"weekday_dict = {1 : \"Monday\", 2 : \"Tuesday\", 3 : \"Wednesday\", 4 : \"Thursday\", 5 : \"Friday\", 6 : \"Saturday\", 7 : \"Sunday\"}\n\nfig, ax = plt.subplots(nrows = 3, figsize = (20, 10))\n#examination of weekdays\nfor i, group in enumerate([record_less_10_filter, record_between_10_30_filter, record_greater_30_filter]):\n    #check for frequency of weekdays. reduce df size by forcing weekday and days to be 1 to 1. \n    grouped_data = [Train_Actigraphy[c][[\"Days\", \"weekday\"]].drop_duplicates()[\"weekday\"] for c in np.arange(len(Train_Actigraphy)) if c in group]\n    for df in grouped_data:\n        #autoselect colors by default\n        sns.kdeplot(df, ax = ax[i], color=\"#69b3a2\", alpha = 0.6)\n    ax[i].set_title(\"Group \" + str(i + 1) + \" participation by weekday attendance\")\n    ax[i].set_ylabel(\"Relative Density\")\n    ax[i].set_xlabel(None)\n    ax[i].set_xlim(1, 7)\n    ax[i].set_xticklabels([weekday_dict[x] for x in np.arange(1, 8)])\n\n    ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:26:08.924018Z","iopub.execute_input":"2024-11-18T02:26:08.924475Z","iopub.status.idle":"2024-11-18T02:26:34.895773Z","shell.execute_reply.started":"2024-11-18T02:26:08.924424Z","shell.execute_reply":"2024-11-18T02:26:34.894516Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Looks like Group 2 and 3 have similar distributions of weekdays. The first group shows more variety. Some children have data for a portion of the available weekdays while others show a spread similar to that of the other groups. Lets write something up to catch any children who follow a schedule in which at least one day of the week is missed consistently. ","metadata":{}},{"cell_type":"code","source":"def give_unique_day_prop(grouping_filter, days):\n\n    grouping = [Train_Actigraphy[c][\"weekday\"].drop_duplicates().count() for c in np.arange(len(Train_Actigraphy)) if c in grouping_filter]\n    \n    return(sum([v == days for v in grouping]), len(grouping))\n\nfull_days =[]\nmissing_days = []\nfull_day_prop = []\n\nfor i, group in enumerate(group_filters):\n\n    data_prop = give_unique_day_prop(group, 7)\n\n    full_days.append(data_prop[0])\n    missing_days.append(data_prop[1] - data_prop[0])\n    full_day_prop.append(data_prop[0]/data_prop[1])\n\nweekday_consistency = pd.DataFrame({\"full_weekday_attendance\" : full_days, \"missing_at_least_one_weekday\" : missing_days, \"proportion_all_weekdays\" : [\"{:.4f}\".format(x) for x in full_day_prop]},\n                                  index = [\"Group 1 (10 or less)\", \"Group 2 (10 - 30)\", \"Group 3 (30 or more)\"])\n\n\nweekday_consistency.style.set_caption(\"Only group 1 contains children who do not have data for every weekday\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:26:34.897558Z","iopub.execute_input":"2024-11-18T02:26:34.897946Z","iopub.status.idle":"2024-11-18T02:26:40.368001Z","shell.execute_reply.started":"2024-11-18T02:26:34.897903Z","shell.execute_reply":"2024-11-18T02:26:40.366886Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Lets examine this subgroup of 28 children. \n\nPerhaps some of these children participated in the study for a period nearing the 30 day cap but only wore the watch for a few days each week? Lets also check if these strange children show up in the HBN Instruments test dataset","metadata":{}},{"cell_type":"code","source":"def day_participation_counter(weekdays):\n    \n    group_1_weekdays = [Train_Actigraphy[c] for c in np.arange(len(Train_Actigraphy)) if c in (record_less_10_filter) and \n                    (Train_Actigraphy[c][\"weekday\"].drop_duplicates().count() < weekdays)]\n\n    max_days = pd.DataFrame({\"days_participation\" : [c.Days.max() for c in group_1_weekdays]}, index  = [c.id[0] for c in group_1_weekdays])\n\n    max_days[\"Child Count\"] = 0\n    \n    return(max_days)","metadata":{"execution":{"iopub.status.busy":"2024-11-18T02:26:40.369140Z","iopub.execute_input":"2024-11-18T02:26:40.369480Z","iopub.status.idle":"2024-11-18T02:26:40.376607Z","shell.execute_reply.started":"2024-11-18T02:26:40.369446Z","shell.execute_reply":"2024-11-18T02:26:40.375281Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"group_1_weekdays = [Train_Actigraphy[c] for c in np.arange(len(Train_Actigraphy)) if c in (record_less_10_filter) and \n                    (Train_Actigraphy[c][\"weekday\"].drop_duplicates().count() < 7)]\n\nmax_days = pd.DataFrame({\"days_participation\" : [c.Days.max() for c in group_1_weekdays]}, index  = [c.id[0] for c in group_1_weekdays])\n\nmax_days[\"Child Count\"] = 0\n\n\n\n\nprint(\"None of the children within this subset of 28 show up within the HBN Test Dataset: \" + \n      str(len([c for c in max_days.index.values if c in test_HBN.id.values]) == 0))\n\n\nmax_days.groupby(\"days_participation\").count().style.set_caption(\"only three children of this subgroup show greater than 10 days of study participation\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:26:40.378367Z","iopub.execute_input":"2024-11-18T02:26:40.378849Z","iopub.status.idle":"2024-11-18T02:26:40.544536Z","shell.execute_reply.started":"2024-11-18T02:26:40.378797Z","shell.execute_reply":"2024-11-18T02:26:40.543244Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\ngreater_10_participation = [df for df in group_1_weekdays if df.Days.max() >= 10]\n\nweekday_bump = -0.25\n\n#thats a pretty style\nplt.style.use('fivethirtyeight')\n\nfig, ax = plt.subplots(figsize = (20, 10))\n\nProp_weekday_counts = greater_10_participation[0].groupby(\"weekday\").count().iloc[:, 1] / greater_10_participation[0].shape[0]\n\nax.bar(Prop_weekday_counts.index, Prop_weekday_counts.values, width = 0.25, label = greater_10_participation[0].id[0], alpha = 0.5)\n\nfor child in greater_10_participation[1:3]:\n    \n    Prop_weekday_counts = child.groupby(\"weekday\").count().iloc[:, 1] / child.shape[0]\n\n    ax.bar(Prop_weekday_counts.index + weekday_bump + 0.5, Prop_weekday_counts.values, width = 0.25, label = child.id[0], alpha = 0.5)\n    \n    print(Prop_weekday_counts.index + weekday_bump)\n    \n    weekday_bump += 0.25\n    \nax_2 = ax.twinx()\n\nfor child in greater_10_participation:\n    \n    weekday_counts = child.groupby(\"weekday\").count().iloc[:, 1]\n    \n    ax_2.scatter(weekday_counts.index, weekday_counts.values, label = child.id[0])\n    \nax.set_xticks(np.arange(1, 8))\n    \nax.set_xticklabels([weekday_dict[i] for i in [1, 2, 3, 4, 5, 6, 7]])\n\nax.set_title(\"3 of 28 with more than 10 days participation\")\n\nax.set_ylabel(\"Proportion of days (Bar)\")\nax_2.set_ylabel(\"total days (Point)\")\n\nax.legend(title = \"Child id\")\n\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:26:40.551333Z","iopub.execute_input":"2024-11-18T02:26:40.551715Z","iopub.status.idle":"2024-11-18T02:26:41.203528Z","shell.execute_reply.started":"2024-11-18T02:26:40.551678Z","shell.execute_reply":"2024-11-18T02:26:41.202233Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Lets look at the other 155 children within group 1 who did attend each weekday.","metadata":{"execution":{"iopub.status.busy":"2024-10-28T02:15:42.652892Z","iopub.execute_input":"2024-10-28T02:15:42.653402Z","iopub.status.idle":"2024-10-28T02:15:42.664359Z","shell.execute_reply.started":"2024-10-28T02:15:42.653358Z","shell.execute_reply":"2024-10-28T02:15:42.662537Z"}}},{"cell_type":"code","source":"days_participation_less_10_155 = day_participation_counter(8).groupby(\"days_participation\").count()\n\nplt.style.use('default')\n\nplt.rcParams[\"figure.figsize\"] = (10, 10)\n\nplt.plot(days_participation_less_10_155.index, days_participation_less_10_155[\"Child Count\"])\n\nplt.title(\"Children with less than 10 summed days of participation\")\nplt.xlabel(\"Non-Continuous days of study participation\")\nplt.ylabel(\"Child Count\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:26:41.205008Z","iopub.execute_input":"2024-11-18T02:26:41.205407Z","iopub.status.idle":"2024-11-18T02:26:41.726500Z","shell.execute_reply.started":"2024-11-18T02:26:41.205367Z","shell.execute_reply":"2024-11-18T02:26:41.725176Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Actigraphy Data Cleaning\n\nNow lets wittle down the actigraphy dataset. First we will remove records from the group of children with less than 10 total days of participation. \n\nThen we will examine missing records of the columns associated with HBN data. i.e. does an actigraphy id have data in the HBN Instruments table for this column? The analysis of missing data here will be relatively simple as the HBN_Instruments value was copied to all rows of its associated Actigraphy dataframe. ","metadata":{}},{"cell_type":"code","source":"Train_Actigraphy = [Train_Actigraphy[c] for c in np.arange(len(Train_Actigraphy)) if c not in record_less_10_filter]","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:26:41.728136Z","iopub.execute_input":"2024-11-18T02:26:41.728553Z","iopub.status.idle":"2024-11-18T02:26:41.749566Z","shell.execute_reply.started":"2024-11-18T02:26:41.728512Z","shell.execute_reply":"2024-11-18T02:26:41.748132Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"Actigraphy_HBN_cols_with_data = [df[df.iloc[:, 8:21].columns.values].dropna(how = 'all', axis = 1).columns.values for df in Train_Actigraphy]\n\nActigraphy_HBN_cols_with_data = [e for nested in Actigraphy_HBN_cols_with_data for e in nested]","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:26:41.753629Z","iopub.execute_input":"2024-11-18T02:26:41.754037Z","iopub.status.idle":"2024-11-18T02:26:51.208422Z","shell.execute_reply.started":"2024-11-18T02:26:41.753997Z","shell.execute_reply":"2024-11-18T02:26:51.207264Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"prop_w_data = [Actigraphy_HBN_cols_with_data.count(val) / len(Train_Actigraphy) for val in set(Actigraphy_HBN_cols_with_data)]\ncol_name = [val for val in set(Actigraphy_HBN_cols_with_data)]\n\nplt.bar(col_name, prop_w_data)\n\nplt.title(\"Actigraphy ids with associated data in HBN columns\")\nplt.ylabel(\"proportion of Actigraphy dataframes with non-null data\")\nplt.xticks(rotation = 45)\n\nplt.tight_layout()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-18T02:26:51.210123Z","iopub.execute_input":"2024-11-18T02:26:51.210512Z","iopub.status.idle":"2024-11-18T02:26:51.729881Z","shell.execute_reply.started":"2024-11-18T02:26:51.210472Z","shell.execute_reply":"2024-11-18T02:26:51.728824Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The low proportion of grip values makes sense based on our examination of the HBN_Instruments dataset. seems that the proportion of dataframes missing values is within an acceptable range, so we will not remove any Actigraphy dataframes from our set based off of this criteria. \n\nMissing Actigraphy Data Assumptions:\n\n1. The Actigraphy data will not contain any outright null columns. \n2. Missing values would be shown as skipped times","metadata":{}},{"cell_type":"code","source":"missing_val_cols = sum([df.shape[0] - df.iloc[:,0:8].dropna(how = 'any', axis = 1).shape[0] for df in Train_Actigraphy])","metadata":{"execution":{"iopub.status.busy":"2024-11-18T02:26:51.731088Z","iopub.execute_input":"2024-11-18T02:26:51.731441Z","iopub.status.idle":"2024-11-18T02:26:58.888078Z","shell.execute_reply.started":"2024-11-18T02:26:51.731404Z","shell.execute_reply":"2024-11-18T02:26:58.886892Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#iterate through each row.\n\n#row(i) - row(i - 1) = x\n\n#within day check for missing time\n#x / 5 - 1 = # of skipped events \n\n#day to day check for missing time\n#@ time = 0 Days n - 1\n#x - 1 = # of skipped days ? \n\n#max seconds in day = 86395\n#min seconds in day = 0\n\ndef missed_time_calc(df):\n    \n    test = df\n    \n    test[\"Days_lagged\"] = test[\"Days\"].shift(1)\n    test[\"times_lagged\"] = test['time_of_day'].shift(1)\n\n    test[\"day_diff\"] = test[\"Days\"] - test[\"Days_lagged\"]\n    test[\"time_diff\"] = test['time_of_day'] - test[\"times_lagged\"]\n    test[\"Day_Swap\"] = (test['time_diff'] == -86395) & (test['day_diff'] == 1)\n    \n\n    def missed_event_counter(x):\n        \n        #within the same day\n        if x.day_diff == 0:\n\n            missed_events = x.time_diff / 5 - 1\n            \n        elif x.Day_Swap:\n            \n            missed_events = 0\n            \n        #between days\n            #exactly a # of days of difference\n        elif x.day_diff > 0: \n        \n            if x.times_lagged == x.time_of_day:\n                #exactly 86340 seconds between days have passed\n                missed_events = ((x.day_diff * 86400) / 5) - 1\n            elif x.times_lagged != x.time_of_day:\n                #cut out time from first half of second date and last half of first date\n                seconds_missed = x.time_of_day +  86395 - x.times_lagged\n                #remove last day as that is caught by \"seconds missed\"\n                missed_events = (((x.day_diff - 1) * 86400) + seconds_missed) / 5 - 1\n\n        elif x.day_diff < 0:\n\n            raise Exception(\"error, negative day_diff\")\n\n        elif pd.isnull(x.day_diff):\n\n             missed_events = np.NaN\n                \n        return(missed_events)\n        \n            \n    \n    test[\"missed_events\"] = test.apply(lambda x: missed_event_counter(x), axis = 1) \n    \n    return(test)\n\n","metadata":{"execution":{"iopub.status.busy":"2024-11-18T02:26:58.889859Z","iopub.execute_input":"2024-11-18T02:26:58.890334Z","iopub.status.idle":"2024-11-18T02:26:58.902486Z","shell.execute_reply.started":"2024-11-18T02:26:58.890279Z","shell.execute_reply":"2024-11-18T02:26:58.900978Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"printProgressBar(0, len(Train_Actigraphy), \"checking for missing data\", \"donesky\")\n\nfor i, df in enumerate(Train_Actigraphy):\n\n    df = missed_time_calc(df)\n\n    Train_Actigraphy[i] = df\n\n    printProgressBar(i + 1, len(Train_Actigraphy), \"checking for missing data\", \"donesky\")\n    ","metadata":{"execution":{"iopub.status.busy":"2024-11-18T02:26:58.904070Z","iopub.execute_input":"2024-11-18T02:26:58.904547Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"dfs_missed_events = [df[df.missed_events > 0].shape[0] for df in Train_Actigraphy]","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize = (10, 5))\n\nax.hist(dfs_missed_events, bins = 50);","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Looks like there are peaks near 0 and around 5000 missed events. Given the definition of missed events, this 5000 number represents about 7 hours of time. Maybe it has something to do with removal of the device at night?","metadata":{}},{"cell_type":"code","source":"dir = \"/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet\"\nprint(\"if we select a threshold of 5 missed events and remove all other dataframes from the actigraphy set, we are left with \" \n      + \"{:.2f}\".format(sum([i < 5 for i in dfs_missed_events]) / len(os.listdir(dir))) +\n     \" of the original dataset\\n\\nWhile representing \" +\n     \"{:.2f}\".format(sum([i < 5 for i in dfs_missed_events]) / len(dfs_missed_events)) +\n     \" of the dataset less those dfs with fewer than 10 records\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"I believe that we can move forward with this subset of actigraphy dataframes. Lets redefine the Actigraphy list to remove those frames with more than 5 missing events. We will also need to re-code the Train_Actigraphy_ids list to account for the removal of these dataframes. Also, with the removal of ","metadata":{}},{"cell_type":"code","source":"Train_Actigraphy = [Train_Actigraphy[i] for i in np.arange(len(Train_Actigraphy)) if dfs_missed_events[i] < 5]\nTrain_Actigraphy_ids = [df.id[0] for df in Train_Actigraphy]\n\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Physical-Waist_Circumference    0.773232\n\r\nFitness_Endurance-Max_Stage     0.81237\n4\r\nFitness_Endurance-Time_Mins     0.8131\n31\r\nFitness_Endurance-Time_Sec      0.813\n131\r\nFGC-FGC_GSND                    0.72\n8788\r\nFGC-FGC_GSD                     0.7\n28788\r\nFGC-FGC_GSD_Zone                0.\n731566\r\nFitness_Endurance-Season \n\n#### examination of d74e4d7c\n\n\nlets checkactivity_df to see if this child has actigraphy data, then take a look at what their profile looks like. ","metadata":{}},{"cell_type":"code","source":"strange_child = train_HBN[train_HBN[[i[0] for i in adolescent_activity + child_activity]].isnull().sum(axis = 1) == 0][\"id\"].values[0]\n\nprint(\"This child has actigraphy data: \" + str(strange_child in Train_Actigraphy_ids) +\n      \"\\nThis id is found at index \" +  \n     \" of the list of parquet ids.\\nTherefore it should be located at this same index in \" +\n     \"the list of actigraphy dataframes.\\nLets take a look at their data, \" +\n     \"beginning with the HBN set, but first lets double check the id in the \" + \n      \"actigraphy dataframe to be completely sure.\")\n\nstrange_child_in_dataframe = strange_child == Train_Actigraphy[Train_Actigraphy_ids.index(strange_child)][\"id\"].iloc[0]\n\nprint(strange_child + \" is present in dataframe \" +\n      str(Train_Actigraphy_ids.index(strange_child)) + \n      \" of Train_Actigraphy: \" + \n      str(strange_child_in_dataframe))","metadata":{"_kg_hide-input":true,"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"##### d74e4d7c HBN data - what is missing?\n\nlets check null values for this child within the HBN table.","metadata":{}},{"cell_type":"code","source":"plt.rcParams[\"figure.figsize\"] = (10, 1)\n\nsns.heatmap(train_HBN[train_HBN[\"id\"] == strange_child].isnull(), cbar = False)\nplt.title(\"Null check for \" + strange_child)\nplt.show()","metadata":{"_kg_hide-input":true,"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Well, it looks like this child is lacking data for all fitness gram categories. They are also missing data for","metadata":{}},{"cell_type":"markdown","source":"## Weekly Aggregate examination\n\nLets examine some summary statistics of weekly child activity measures\n\nneed to iterate over each weekday to save space...constructive criticism welcome here\n\nthrowing **everything** but the kitchen sink at it to see if something interesting pops up. This will take time to run, but there should be at least one interesting finding that pops up, right? right?!\n\nWe'll use the measures from the describe function. With that we will take the associated aggregate measure from each child, broken down by weekday, and compare all those aggregates among the entire group of child actigraphy data. This should be ~1000 or so observations (can fluctuate a bit, see comments in code). \n\n**the output is hidden as its quite long.**","metadata":{}},{"cell_type":"code","source":"%%time\n\nmeasures = list(Train_Actigraphy[0].describe().index)\n\nfeatures = [\"enmo\", \"anglez\", 'non-wear_flag', 'light']\n\nprintProgressBar(0, len(cat_cols_better_names), \"Drawing\", \"Complete\")\n\nfor i, cat in enumerate(cat_cols_better_names):\n    \n    cat_val = list(rename_dict.keys())[list(rename_dict.values()).index(cat)]\n    \n    unique_cat_vals = train_HBN[cat_val].drop_duplicates()\n    \n    for feature in features:\n            \n        for measure in measures:\n            \n            fig, ax = plt.subplots(nrows = 7, figsize = (20, 20), sharex = True)\n            \n            ax[0].set_title(measure + \" of \" + feature + \" grouped by \" + cat)\n\n            for i in np.arange(1, 8):\n                \n                ax[i - 1].set_ylabel(\"Weekday_\" + str(i))\n                \n                \n                \n                Weekday_Feature_Cat = [df[df[\"weekday\"] == i][[feature, cat]].reset_index(drop = True) \n                                       for df in Train_Actigraphy]\n                \n\n                \n                \n                \n                #each df represents one child, who occupies 1 category type, so we can perform operation up here to save\n                #processing time\n                #in some cases a feature column may not be present...this warrants further investigation\n                weekday_measure = [df.describe().loc[measure, feature] if feature in df.columns else np.NaN for df in Weekday_Feature_Cat]\n                \n                #in some cases a child only has cuff data for one weekday, so filtering for a weekday yields\n                #an empty table, so we need to catch for that\n                #could be interesting to see what the heck these 1 weekday kiddos are up to\n                weekday_cat = [df.iloc[0, 1] if df.shape[0] > 0 else np.NaN for df in Weekday_Feature_Cat ]\n                \n                weekday = pd.DataFrame({measure : weekday_measure, cat : weekday_cat}).dropna(how = \"any\")\n                \n                for child_cat in unique_cat_vals:\n                    \n                    sns.kdeplot(weekday[weekday[cat] == child_cat][measure], label = child_cat, ax = ax[i - 1])\n                \n                \n               \n                \n            \n            ax[0].legend()    \n\n    printProgressBar(i + 1, len(cat_cols_better_names), \"Drawing\", \"Complete\")","metadata":{"execution":{"iopub.status.busy":"2024-11-17T23:08:08.698662Z","iopub.execute_input":"2024-11-17T23:08:08.699415Z"},"_kg_hide-output":true,"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Interesting results from weekly aggregate \nThere are a few insights that can be gathered from looking at the weekly data. \nit looks like the \n\n","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}