{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30775,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# First attempt","metadata":{}},{"cell_type":"markdown","source":"## The goal\n> Predict the level of problematic internet usage exhibited by children and adolescents, based on their physical activity.\n## Evaluation\n> Submissions are scored based on the quadratic weighted kappa, which measures the agreement between 2 outcomes. This metric varies from 0 (random agreement) to 1 (complete agreement). When there is less agreement than expected by change, the metric may go below 0.\n## Submission file\n> For each `id` in the test set, you must predit the corresponding `sii`","metadata":{}},{"cell_type":"markdown","source":"id,sii\n000046df,0\n000089ff,1\n00012558,2\n00017ccd,3","metadata":{}},{"cell_type":"markdown","source":"## Dataset\nIs a clinical sample of about 5000 5-22 year olds.\n`sii` => Severity Impairment Index.\nThe full test set comprises about 3800 instances.\nThere are 2 sources:\n1. `parquet` files containing the accelerometer (actigraphy) series.\n2. `csv` files containing the remaining tabular data.\nMajority of measures are missing for most participants.`sii` is missing for a portion of the participants in the training set. It is present for all instances in the test set.\n**Actigraphy**: Is an objective measure of ecological physical activity through a research-grade biotracker.\n`sii` values:\n- `0` for `None`.\n- `1` for `Mild`.\n- `2` for `Moderate`.\n- `3` for `Severe`.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:58.986637Z","iopub.execute_input":"2024-09-25T13:23:58.987129Z","iopub.status.idle":"2024-09-25T13:23:58.992836Z","shell.execute_reply.started":"2024-09-25T13:23:58.987077Z","shell.execute_reply":"2024-09-25T13:23:58.991653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/train.csv\")\ntest_df = pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/test.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:58.994684Z","iopub.execute_input":"2024-09-25T13:23:58.995090Z","iopub.status.idle":"2024-09-25T13:23:59.066961Z","shell.execute_reply.started":"2024-09-25T13:23:58.995043Z","shell.execute_reply":"2024-09-25T13:23:59.065683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.069138Z","iopub.execute_input":"2024-09-25T13:23:59.069591Z","iopub.status.idle":"2024-09-25T13:23:59.105639Z","shell.execute_reply.started":"2024-09-25T13:23:59.069541Z","shell.execute_reply":"2024-09-25T13:23:59.104160Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.107272Z","iopub.execute_input":"2024-09-25T13:23:59.107752Z","iopub.status.idle":"2024-09-25T13:23:59.144608Z","shell.execute_reply.started":"2024-09-25T13:23:59.107685Z","shell.execute_reply":"2024-09-25T13:23:59.143571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It looks like the train data has 82 columns and the test data has 59. Let's check why is this and which columns are common and unique to each csv file.","metadata":{}},{"cell_type":"code","source":"train_cols = set(train_df.columns)\ntest_cols = set(test_df.columns)\ncommon_cols = train_cols.intersection(test_cols)\nunique_train_cols = train_cols.difference(test_cols)\nunique_test_cols = test_cols.difference(train_cols)","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.145828Z","iopub.execute_input":"2024-09-25T13:23:59.146122Z","iopub.status.idle":"2024-09-25T13:23:59.152261Z","shell.execute_reply.started":"2024-09-25T13:23:59.146089Z","shell.execute_reply":"2024-09-25T13:23:59.151216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"common_cols","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.153802Z","iopub.execute_input":"2024-09-25T13:23:59.154250Z","iopub.status.idle":"2024-09-25T13:23:59.163999Z","shell.execute_reply.started":"2024-09-25T13:23:59.154213Z","shell.execute_reply":"2024-09-25T13:23:59.162940Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"unique_train_cols, unique_test_cols","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.165573Z","iopub.execute_input":"2024-09-25T13:23:59.165910Z","iopub.status.idle":"2024-09-25T13:23:59.176093Z","shell.execute_reply.started":"2024-09-25T13:23:59.165863Z","shell.execute_reply":"2024-09-25T13:23:59.174992Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"All those PCIAT values are only present in the train csv file but not in the test. I believe it does not make too much sense to use it, since we cannot use them after that to predict the values of test data. Let's drop them.","metadata":{}},{"cell_type":"code","source":"pciat_cols = [col for col in train_cols if 'PCIAT' in col]\npciat_cols","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.177283Z","iopub.execute_input":"2024-09-25T13:23:59.177717Z","iopub.status.idle":"2024-09-25T13:23:59.187881Z","shell.execute_reply.started":"2024-09-25T13:23:59.177675Z","shell.execute_reply":"2024-09-25T13:23:59.186573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = train_df.drop(columns=pciat_cols)\ntrain_df","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.198111Z","iopub.execute_input":"2024-09-25T13:23:59.198436Z","iopub.status.idle":"2024-09-25T13:23:59.240186Z","shell.execute_reply.started":"2024-09-25T13:23:59.198402Z","shell.execute_reply":"2024-09-25T13:23:59.239064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So now, we have our training data without those PCIAT columns that were not in the test data. Now let's see how we are doing with the missing/null values.","metadata":{}},{"cell_type":"code","source":"train_df.shape, train_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.247857Z","iopub.execute_input":"2024-09-25T13:23:59.248185Z","iopub.status.idle":"2024-09-25T13:23:59.261930Z","shell.execute_reply.started":"2024-09-25T13:23:59.248153Z","shell.execute_reply":"2024-09-25T13:23:59.260797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Any column with a ratio greater than 70% are good to be dropped in my opinion.","metadata":{}},{"cell_type":"code","source":"cols_many_missing_values = [col for col in train_df.columns if (train_df[col].isna().sum() / len(train_df)) > 0.7]\ncols_many_missing_values","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.263211Z","iopub.execute_input":"2024-09-25T13:23:59.263520Z","iopub.status.idle":"2024-09-25T13:23:59.288883Z","shell.execute_reply.started":"2024-09-25T13:23:59.263487Z","shell.execute_reply":"2024-09-25T13:23:59.287762Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = train_df.drop(columns=cols_many_missing_values)\ntrain_df","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.290406Z","iopub.execute_input":"2024-09-25T13:23:59.290790Z","iopub.status.idle":"2024-09-25T13:23:59.333760Z","shell.execute_reply.started":"2024-09-25T13:23:59.290747Z","shell.execute_reply":"2024-09-25T13:23:59.332786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I think we should drop also all the data without sii, at least at first. Maybe later I can investigate non-supervised reinforcement learning, but right now I prefer something simple.","metadata":{}},{"cell_type":"code","source":"train_df = train_df.dropna(subset=[\"sii\"])\ntrain_df.shape, train_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.334940Z","iopub.execute_input":"2024-09-25T13:23:59.335250Z","iopub.status.idle":"2024-09-25T13:23:59.350002Z","shell.execute_reply.started":"2024-09-25T13:23:59.335218Z","shell.execute_reply":"2024-09-25T13:23:59.349189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's recalculate now the percentages and see if there is more columns to drop.","metadata":{}},{"cell_type":"code","source":"for col in train_df.columns:\n    miss_count = train_df[col].isna().sum()\n    percent_miss = (miss_count / len(train_df)) * 100\n    print(f\"{col}: {percent_miss:.2f}%\")","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.351253Z","iopub.execute_input":"2024-09-25T13:23:59.351758Z","iopub.status.idle":"2024-09-25T13:23:59.371618Z","shell.execute_reply.started":"2024-09-25T13:23:59.351703Z","shell.execute_reply":"2024-09-25T13:23:59.370630Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's drop now everything over 30% to skim more the data.","metadata":{}},{"cell_type":"code","source":"cols_many_missing_values = [col for col in train_df.columns if (train_df[col].isna().sum() / len(train_df)) > 0.3]\ncols_many_missing_values","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.373005Z","iopub.execute_input":"2024-09-25T13:23:59.373360Z","iopub.status.idle":"2024-09-25T13:23:59.391770Z","shell.execute_reply.started":"2024-09-25T13:23:59.373326Z","shell.execute_reply":"2024-09-25T13:23:59.390758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = train_df.drop(columns=cols_many_missing_values)\ntrain_df","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.393054Z","iopub.execute_input":"2024-09-25T13:23:59.393379Z","iopub.status.idle":"2024-09-25T13:23:59.436654Z","shell.execute_reply.started":"2024-09-25T13:23:59.393329Z","shell.execute_reply":"2024-09-25T13:23:59.435706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"categorical_cols = train_df.select_dtypes(include=[\"object\"]).columns\ncategorical_cols = categorical_cols[1:].to_numpy()\ncategorical_cols","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.438344Z","iopub.execute_input":"2024-09-25T13:23:59.438651Z","iopub.status.idle":"2024-09-25T13:23:59.446940Z","shell.execute_reply.started":"2024-09-25T13:23:59.438619Z","shell.execute_reply":"2024-09-25T13:23:59.445894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"non_categorical_cols = [col for col in train_df.columns if col not in categorical_cols]\nnon_categorical_cols = non_categorical_cols[1:-1]\nnon_categorical_cols","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.448934Z","iopub.execute_input":"2024-09-25T13:23:59.449371Z","iopub.status.idle":"2024-09-25T13:23:59.458447Z","shell.execute_reply.started":"2024-09-25T13:23:59.449332Z","shell.execute_reply":"2024-09-25T13:23:59.457252Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's use a TabularLearner from fastai now.","metadata":{}},{"cell_type":"code","source":"from fastai.tabular.all import *","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.460069Z","iopub.execute_input":"2024-09-25T13:23:59.460444Z","iopub.status.idle":"2024-09-25T13:23:59.466541Z","shell.execute_reply.started":"2024-09-25T13:23:59.460408Z","shell.execute_reply":"2024-09-25T13:23:59.465443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"splits = RandomSplitter(valid_pct=0.2)(range_of(train_df))","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.468069Z","iopub.execute_input":"2024-09-25T13:23:59.468433Z","iopub.status.idle":"2024-09-25T13:23:59.476789Z","shell.execute_reply.started":"2024-09-25T13:23:59.468398Z","shell.execute_reply":"2024-09-25T13:23:59.475502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"to = TabularPandas(train_df, procs=[Categorify, FillMissing, Normalize],\n                   cat_names=['Basic_Demos-Enroll_Season', 'CGAS-Season', 'Physical-Season', 'FGC-Season', 'SDS-Season', 'PreInt_EduHx-Season'],\n                   cont_names=['Basic_Demos-Age','Basic_Demos-Sex','CGAS-CGAS_Score','Physical-BMI','Physical-Height','Physical-Weight','Physical-Diastolic_BP','Physical-HeartRate','Physical-Systolic_BP','FGC-FGC_CU','FGC-FGC_TL','SDS-SDS_Total_Raw','SDS-SDS_Total_T','PreInt_EduHx-computerinternet_hoursday'],\n                   y_names='sii',\n                   y_block=CategoryBlock,\n                   splits=splits)","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.478664Z","iopub.execute_input":"2024-09-25T13:23:59.478991Z","iopub.status.idle":"2024-09-25T13:23:59.555798Z","shell.execute_reply.started":"2024-09-25T13:23:59.478959Z","shell.execute_reply":"2024-09-25T13:23:59.554625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dls = to.dataloaders(bs=64)","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.557743Z","iopub.execute_input":"2024-09-25T13:23:59.558209Z","iopub.status.idle":"2024-09-25T13:23:59.576938Z","shell.execute_reply.started":"2024-09-25T13:23:59.558161Z","shell.execute_reply":"2024-09-25T13:23:59.575697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dls.show_batch()","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.578936Z","iopub.execute_input":"2024-09-25T13:23:59.579351Z","iopub.status.idle":"2024-09-25T13:23:59.640557Z","shell.execute_reply.started":"2024-09-25T13:23:59.579294Z","shell.execute_reply":"2024-09-25T13:23:59.639499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn = tabular_learner(dls, metrics=accuracy)","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.642401Z","iopub.execute_input":"2024-09-25T13:23:59.642863Z","iopub.status.idle":"2024-09-25T13:23:59.658848Z","shell.execute_reply.started":"2024-09-25T13:23:59.642817Z","shell.execute_reply":"2024-09-25T13:23:59.657607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.fit_one_cycle(1)","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:23:59.660275Z","iopub.execute_input":"2024-09-25T13:23:59.661261Z","iopub.status.idle":"2024-09-25T13:24:00.284238Z","shell.execute_reply.started":"2024-09-25T13:23:59.661209Z","shell.execute_reply":"2024-09-25T13:24:00.283299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.show_results()","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:24:00.285308Z","iopub.execute_input":"2024-09-25T13:24:00.285616Z","iopub.status.idle":"2024-09-25T13:24:00.349769Z","shell.execute_reply.started":"2024-09-25T13:24:00.285583Z","shell.execute_reply":"2024-09-25T13:24:00.348930Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dl = learn.dls.test_dl(test_df)","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:24:00.351042Z","iopub.execute_input":"2024-09-25T13:24:00.351498Z","iopub.status.idle":"2024-09-25T13:24:00.391286Z","shell.execute_reply.started":"2024-09-25T13:24:00.351459Z","shell.execute_reply":"2024-09-25T13:24:00.390394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds = learn.get_preds(dl=dl)","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:24:00.392961Z","iopub.execute_input":"2024-09-25T13:24:00.393658Z","iopub.status.idle":"2024-09-25T13:24:00.422509Z","shell.execute_reply.started":"2024-09-25T13:24:00.393615Z","shell.execute_reply":"2024-09-25T13:24:00.421621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"results = []\ncount = 0\nfor pred in preds[0]:\n    max_idx = np.argmax(pred.numpy())\n    results.append((test_df['id'].iloc[count], str(max_idx)))\n    count += 1\n\nresults_df = pd.DataFrame(results, columns=['id', 'sii'])\nresults_df.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:24:00.424110Z","iopub.execute_input":"2024-09-25T13:24:00.424826Z","iopub.status.idle":"2024-09-25T13:24:00.434584Z","shell.execute_reply.started":"2024-09-25T13:24:00.424783Z","shell.execute_reply":"2024-09-25T13:24:00.433607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!head submission.csv","metadata":{"execution":{"iopub.status.busy":"2024-09-25T13:24:00.436890Z","iopub.execute_input":"2024-09-25T13:24:00.437904Z","iopub.status.idle":"2024-09-25T13:24:01.604785Z","shell.execute_reply.started":"2024-09-25T13:24:00.437855Z","shell.execute_reply":"2024-09-25T13:24:01.603314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}