{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30775,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"1.Extract specific data:\nThe code extracts data for the 13th day from the dictionary grouped_by_day_dict and stores it in the variable tmp.\n\n2.Calculate z-angle differences:\nThe z-angle_diff column is created, which computes the absolute difference of the ‘anglez’ values between consecutive rows.\n\n3.Compute rolling median:\nThe z-angle_diff_median column is added, which calculates a rolling median with a window size of 60 (1-minute window) for the z-angle_diff values. This helps smooth the differences.\n\n4.Threshold for stationary periods:\nThe threshold is set as the 10th percentile of the z-angle_diff_median, assuming lower values indicate minimal movement or stationary periods.\n\n5.Identify stationary periods:\nThe is_stationary column is created to indicate whether the z-angle_diff_median is below the threshold, marking those periods as stationary (True or False).\n\n6.Cumulative stationary periods:\nThe stationary_periods column calculates cumulative sums for consecutive stationary periods, grouping them together to track duration.\n\n7.Sleep detection:\nA threshold of 30 minutes (or 360 seconds) of stationary time is set. The code creates an is_sleep column, which flags periods longer than 30 minutes as potential sleep periods.\n\n8.Plot stationary periods:\nFinally, the stationary_periods data is plotted to visualize periods of stillness over the course of the day.\n\nThis algorithm is inspired by the paper “Estimating Sleep Parameters Using an Accelerometer Without a Sleep Diary” (van Hees et al., 2018) ￼. In the paper, the authors developed a heuristic algorithm to detect sleep periods by analyzing the variance in the z-axis angle from accelerometer data. This method, which does not rely on a sleep diary, identifies periods of low movement (stationary periods) based on changes in arm angle and detects the Sleep Period Time (SPT) window.\n\nIn our algorithm, we adopt a similar approach by calculating the difference in the z-axis angle (z-angle_diff) from accelerometer data and using a rolling median to smooth these differences. We set a threshold based on the 10th percentile of these smoothed values to classify stationary periods. If the stationary periods exceed a certain duration (30 minutes), they are flagged as potential sleep periods. This approach allows us to estimate sleep duration and patterns without the need for subjective input, aligning with the principles of the heuristic method outlined in the paper.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport polars as pl\nimport pandas as pd\nfrom sklearn.base import clone\nfrom copy import deepcopy\nimport optuna\nfrom scipy.optimize import minimize\nimport os\nimport matplotlib.pyplot as plt\nimport re\nfrom colorama import Fore, Style\n\nfrom tqdm import tqdm\nfrom IPython.display import clear_output\nfrom concurrent.futures import ThreadPoolExecutor\n\nimport warnings\nwarnings.filterwarnings('ignore')\npd.options.display.max_columns = None\n\nimport lightgbm as lgb\nfrom catboost import CatBoostRegressor, CatBoostClassifier\nfrom xgboost import XGBRegressor\nfrom sklearn.ensemble import VotingRegressor\nfrom sklearn.model_selection import *\nfrom sklearn.metrics import *\n\nSEED = 42\nn_splits = 5","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-09-26T07:59:46.118594Z","iopub.execute_input":"2024-09-26T07:59:46.119412Z","iopub.status.idle":"2024-09-26T07:59:46.128029Z","shell.execute_reply.started":"2024-09-26T07:59:46.119365Z","shell.execute_reply":"2024-09-26T07:59:46.126728Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Preprocess","metadata":{}},{"cell_type":"code","source":"pp=pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/train.csv\")\npp.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-26T07:59:46.130033Z","iopub.execute_input":"2024-09-26T07:59:46.130462Z","iopub.status.idle":"2024-09-26T07:59:46.289688Z","shell.execute_reply.started":"2024-09-26T07:59:46.130411Z","shell.execute_reply":"2024-09-26T07:59:46.288553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"directory_path = '/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet'\nid_directories = [d[3:] for d in os.listdir(directory_path) if d.startswith('id=') and os.path.isdir(os.path.join(directory_path, d))]","metadata":{"execution":{"iopub.status.busy":"2024-09-26T07:59:46.290872Z","iopub.execute_input":"2024-09-26T07:59:46.291236Z","iopub.status.idle":"2024-09-26T07:59:46.380077Z","shell.execute_reply.started":"2024-09-26T07:59:46.291198Z","shell.execute_reply":"2024-09-26T07:59:46.379010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pp_with_parquet = pp[pp['id'].isin(id_directories)]\npp_with_parquet=pp_with_parquet.reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2024-09-26T07:59:46.382383Z","iopub.execute_input":"2024-09-26T07:59:46.382799Z","iopub.status.idle":"2024-09-26T07:59:46.399692Z","shell.execute_reply.started":"2024-09-26T07:59:46.382737Z","shell.execute_reply":"2024-09-26T07:59:46.398496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pp_with_parquet=pp_with_parquet.drop([\n'PCIAT-PCIAT_01',\n'PCIAT-PCIAT_02',\n'PCIAT-PCIAT_03',\n'PCIAT-PCIAT_04',\n'PCIAT-PCIAT_05',\n'PCIAT-PCIAT_06',\n'PCIAT-PCIAT_07',\n'PCIAT-PCIAT_08',\n'PCIAT-PCIAT_09',\n'PCIAT-PCIAT_10',\n'PCIAT-PCIAT_11',\n'PCIAT-PCIAT_12',\n'PCIAT-PCIAT_13',\n'PCIAT-PCIAT_14',\n'PCIAT-PCIAT_15',\n'PCIAT-PCIAT_16',\n'PCIAT-PCIAT_17',\n'PCIAT-PCIAT_18',\n'PCIAT-PCIAT_19',\n'PCIAT-PCIAT_20',\n'PCIAT-PCIAT_Total',\n'Basic_Demos-Enroll_Season',\n'CGAS-Season',\n'Physical-Season',\n'Fitness_Endurance-Season',\n'FGC-Season',\n'BIA-Season',\n'PAQ_A-Season',\n'PAQ_C-Season',\n'PCIAT-Season',\n'SDS-Season',\n'PreInt_EduHx-Season',\n],axis=1)","metadata":{"execution":{"iopub.status.busy":"2024-09-26T07:59:46.401423Z","iopub.execute_input":"2024-09-26T07:59:46.402106Z","iopub.status.idle":"2024-09-26T07:59:46.412654Z","shell.execute_reply.started":"2024-09-26T07:59:46.402053Z","shell.execute_reply":"2024-09-26T07:59:46.411388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pp_with_parquet_zero=pp_with_parquet[pp_with_parquet[\"sii\"]==0]\npp_with_parquet_one=pp_with_parquet[pp_with_parquet[\"sii\"]==1]\npp_with_parquet_two=pp_with_parquet[pp_with_parquet[\"sii\"]==2]\npp_with_parquet_three=pp_with_parquet[pp_with_parquet[\"sii\"]==3]","metadata":{"execution":{"iopub.status.busy":"2024-09-26T07:59:46.414165Z","iopub.execute_input":"2024-09-26T07:59:46.414639Z","iopub.status.idle":"2024-09-26T07:59:46.432741Z","shell.execute_reply.started":"2024-09-26T07:59:46.414587Z","shell.execute_reply":"2024-09-26T07:59:46.431590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"base_path=\"/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet/id={}/part-0.parquet\"\n\nid_zero=pp_with_parquet_zero['id'].sample(n=10)\nid_one=pp_with_parquet_one['id'].sample(n=10)\nid_two=pp_with_parquet_two['id'].sample(n=10)\nid_three=pp_with_parquet_three['id'].sample(n=10)\nids=[id_zero,id_one,id_two,id_three]\n\npaths_zero=[base_path.format(id)for id in id_zero]\npaths_one=[base_path.format(id)for id in id_one]\npaths_two=[base_path.format(id)for id in id_two]\npaths_three=[base_path.format(id)for id in id_three]\npaths_list = [paths_zero, paths_one, paths_two, paths_three]\nparquet_dfs = {}\n\nfor i, name in enumerate(['zero', 'one', 'two', 'three']):\n    parquet_list = []\n    \n    for id, path in zip(ids[i], paths_list[i]):\n        df = pd.read_parquet(path)\n        df['id'] = id\n        parquet_list.append(df)\n    \n    parquet_dfs[name] = pd.concat(parquet_list, ignore_index=True)","metadata":{"execution":{"iopub.status.busy":"2024-09-26T07:59:46.434225Z","iopub.execute_input":"2024-09-26T07:59:46.434589Z","iopub.status.idle":"2024-09-26T07:59:51.053209Z","shell.execute_reply.started":"2024-09-26T07:59:46.434552Z","shell.execute_reply":"2024-09-26T07:59:51.051960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"colums_keep=['id','enmo','anglez','non-wear_flag','light',\"time_of_day\",'weekday']\nzero_parquet=parquet_dfs['zero'][colums_keep]\none_parquet=parquet_dfs['one'][colums_keep]\ntwo_parquet=parquet_dfs['two'][colums_keep]\nthree_parquet=parquet_dfs['three'][colums_keep]","metadata":{"execution":{"iopub.status.busy":"2024-09-26T07:59:51.054639Z","iopub.execute_input":"2024-09-26T07:59:51.055008Z","iopub.status.idle":"2024-09-26T07:59:51.454046Z","shell.execute_reply.started":"2024-09-26T07:59:51.054971Z","shell.execute_reply":"2024-09-26T07:59:51.452796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1.Group by id: The dataset is grouped based on the id column to organize the data for each unique identifier.\n\n2.Add a day column: A new column called day is added by identifying one day as the period between two consecutive noon timestamps. This is done by detecting when the time_of_day crosses the noon mark and incrementing the day count accordingly.\n\n3.Re-group by day: After adding the day column, the data is further grouped by this new column to separate the data into distinct days for each id.\n\nThis process ensures the data is structured by id and then further organized into individual days based on time of day.","metadata":{}},{"cell_type":"code","source":"three_parquet_dict = {id_: group for id_, group in three_parquet.groupby('id')}\nnoon_ns = 43200000000000\nfor id_, df in three_parquet_dict.items():\n    a = 1\n    day_list = []\n\n    for time_of_day in df['time_of_day']:\n        if time_of_day == noon_ns:\n            a += 1\n        day_list.append(a) \n    \n    df['day'] = day_list\n    three_parquet_dict[id_] = df\ngrouped_by_day_dict={}\nfor id_, df in three_parquet_dict.items():\n    grouped_by_day_dict[id_] = {day: group for day, group in df.groupby('day')}\n    \nfor id_, day_groups in grouped_by_day_dict.items():\n    for day, df in day_groups.items():\n        grouped_by_day_dict[id_][day] = df.reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2024-09-26T08:00:07.218384Z","iopub.execute_input":"2024-09-26T08:00:07.218825Z","iopub.status.idle":"2024-09-26T08:00:09.473462Z","shell.execute_reply.started":"2024-09-26T08:00:07.218779Z","shell.execute_reply":"2024-09-26T08:00:09.472238Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ids = ['c308b134', 'df556fd2']  \nnum_days = max(len(grouped_by_day_dict[id_]) for id_ in ids) \nfig, axes = plt.subplots(num_days, len(ids), figsize=(7, num_days * 3))\n\nfor i, id_ in enumerate(ids):\n    for j, (day, group) in enumerate(grouped_by_day_dict[id_].items()):\n        ax1 = axes[j, i]  \n        ax2 = ax1.twinx()  \n\n        ax1.plot(group.index, group['enmo'], color='blue', label='Enmo')\n        ax1.tick_params(axis='y', labelcolor='blue')\n        \n        ax2.plot(group.index, group['non-wear_flag'], color='green', label='Non-Wear Flag')\n        ax2.tick_params(axis='y', labelcolor='green')\n\n        ax1.set_title(f'ID: {id_}, Day {day}')\n        ax1.set_xlabel('Index')","metadata":{"execution":{"iopub.status.busy":"2024-09-26T08:03:07.518702Z","iopub.execute_input":"2024-09-26T08:03:07.519712Z","iopub.status.idle":"2024-09-26T08:03:23.591084Z","shell.execute_reply.started":"2024-09-26T08:03:07.519659Z","shell.execute_reply":"2024-09-26T08:03:23.589818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Calculate the percentage of time the accelerometer is worn during the day and the amount of data collected for each day.","metadata":{}},{"cell_type":"code","source":"non_wear_flag_ratios = {}\n\nfor id_, day_groups in grouped_by_day_dict.items():\n    non_wear_flag_ratios[id_] = {}\n    for day, df in day_groups.items():\n        ratio = df['non-wear_flag'].mean()  \n        non_wear_flag_ratios[id_][day] = ratio\n\nfor id_, day_ratios in non_wear_flag_ratios.items():\n    for day, ratio in day_ratios.items():\n        print(f'{id_},{day},{ratio:.2f}')","metadata":{"execution":{"iopub.status.busy":"2024-09-26T08:00:22.717773Z","iopub.execute_input":"2024-09-26T08:00:22.718161Z","iopub.status.idle":"2024-09-26T08:00:22.757736Z","shell.execute_reply.started":"2024-09-26T08:00:22.718103Z","shell.execute_reply":"2024-09-26T08:00:22.756511Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"index_lengths = {}\n\nfor id_, day_groups in grouped_by_day_dict.items():\n    index_lengths[id_] = {}\n    for day, df in day_groups.items():\n        length = len(df)  \n        index_lengths[id_][day] = length\n\nfor id_, day_lengths in index_lengths.items():\n    for day, length in day_lengths.items():\n        print(f'{length}')","metadata":{"execution":{"iopub.status.busy":"2024-09-26T08:00:22.759111Z","iopub.execute_input":"2024-09-26T08:00:22.759476Z","iopub.status.idle":"2024-09-26T08:00:22.777540Z","shell.execute_reply.started":"2024-09-26T08:00:22.759438Z","shell.execute_reply":"2024-09-26T08:00:22.776298Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Sleep estimation","metadata":{}},{"cell_type":"code","source":"tmp=grouped_by_day_dict[\"df556fd2\"][13]\n\ntmp[\"z-angle_diff\"] = tmp[\"anglez\"].diff().abs()\n\ntmp[\"z-angle_diff_median\"] = tmp[\"z-angle_diff\"].rolling(window=60, min_periods=1).median()\n\nthreshold = tmp[\"z-angle_diff_median\"].quantile(0.10)\n\ntmp[\"is_stationary\"] = tmp[\"z-angle_diff_median\"] < threshold\n\ntmp[\"stationary_periods\"] = tmp[\"is_stationary\"].astype(int).groupby(tmp[\"is_stationary\"].ne(tmp[\"is_stationary\"].shift()).cumsum()).cumsum()\n\nsleep_duration_threshold = 30 * 60 // 5 \ntmp[\"is_sleep\"] = tmp[\"stationary_periods\"] >= sleep_duration_threshold\n\ntmp[\"stationary_periods\"].plot(figsize=(10,5))","metadata":{"execution":{"iopub.status.busy":"2024-09-26T08:26:35.909149Z","iopub.execute_input":"2024-09-26T08:26:35.909575Z","iopub.status.idle":"2024-09-26T08:26:36.211956Z","shell.execute_reply.started":"2024-09-26T08:26:35.909535Z","shell.execute_reply":"2024-09-26T08:26:36.210798Z"},"trusted":true},"execution_count":null,"outputs":[]}]}