{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🐠 Reef - CV strategy: subsequences!\n\n![](https://storage.googleapis.com/kaggle-competitions/kaggle/31703/logos/header.png)\n\n## Problem: there are 3 videos. Using videos as split units for cross-validation or train-validation splits is not optimal, as it generates a too large validation portion.\n\n\n## In this notebook we explore sequences as potential units for cross-validation, but since there are only 20 sequences and their sizes are quite disimilar, we propose an approach to split them into smaller chunks, that we name _subsequences_.\n\nA **sequence**, as stated in the [data tab of the competition](https://www.kaggle.com/c/tensorflow-great-barrier-reef/data), is:\n> sequence - ID of a gap-free subset of a given video. The sequence ids are not meaningfully ordered.\n\n**Subsequences**, as we will define them below,  are parts of a sequences where objects are continually present or are continually not present. We isolate 2 kind of subsequences: with objects and with no objects.\n\n&nbsp;\n\nLet's see an **example**. Consider the sequence `A` with the following frames:\n* `1-20` - No annotations present\n* `21-30` - Annotations present\n* `31-60` - No annotations\n* `61-80` - Annotations present\n\nIn this case, we say that the sequence `A` has `4` subsequences (`1-20`, `21-30`, `31-60`, `61-80`).\n\n&nbsp;\n\n\nA subsequence seems to me like the minimal atom for ensuring no leaks happen between train and test.\n\n\n&nbsp;\n&nbsp;\n&nbsp;\n\n---\n\nThe notebook goes as follows:\n1. Analize sequences as potential units for spliting\n2. Propose and create subsequences, and create videos for sequences and subsequences to get a feeling of them\n3. Create common train-validation splits (1%, 5%, 10%, 20%) using subsequences\n4. Create 5-fold splits and 10-fold splits using subsequences\n\n&nbsp;\n&nbsp;\n\n\n### The resulting dataframes are provided as a dataset for ease of use here: [reef-cv-strategy-subsequences-dataframes](https://www.kaggle.com/julian3833/reef-cv-strategy-subsequences-dataframes)\n\n\n# Please, _DO_ upvote if you find this useful or interesting!\n\n","metadata":{}},{"cell_type":"markdown","source":"# Analyze sequences","metadata":{}},{"cell_type":"code","source":"import os\nimport cv2\nimport subprocess\nfrom tqdm.auto import tqdm\nimport pandas as pd\nfrom IPython.display import Video, display, HTML\nimport warnings; warnings.simplefilter(\"ignore\")\n\n\nBASE_PATH = '../input/tensorflow-great-barrier-reef/train_images/'\n\ndf = pd.read_csv(\"/kaggle/input/tensorflow-great-barrier-reef/train.csv\")\ndf['annotations'] = df['annotations'].apply(eval)\ndf['n_annotations'] = df['annotations'].str.len()\ndf['has_annotations'] = df['annotations'].str.len() > 0\ndf['has_2_or_more_annotations'] = df['annotations'].str.len() >= 2\ndf['doesnt_have_annotations'] = df['annotations'].str.len() == 0\ndf['image_path'] = BASE_PATH + \"video_\" + df['video_id'].astype(str) + \"/\" + df['video_frame'].astype(str) + \".jpg\"","metadata":{"execution":{"iopub.status.busy":"2021-11-23T16:49:53.716926Z","iopub.execute_input":"2021-11-23T16:49:53.717251Z","iopub.status.idle":"2021-11-23T16:49:54.453663Z","shell.execute_reply.started":"2021-11-23T16:49:53.717169Z","shell.execute_reply":"2021-11-23T16:49:54.452887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**There are 20 sequences**:","metadata":{}},{"cell_type":"code","source":"df['sequence'].unique()","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:32:43.792759Z","iopub.execute_input":"2021-11-23T15:32:43.793023Z","iopub.status.idle":"2021-11-23T15:32:43.802394Z","shell.execute_reply.started":"2021-11-23T15:32:43.792993Z","shell.execute_reply":"2021-11-23T15:32:43.80174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['sequence'].nunique()","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:32:43.803869Z","iopub.execute_input":"2021-11-23T15:32:43.804181Z","iopub.status.idle":"2021-11-23T15:32:43.813719Z","shell.execute_reply.started":"2021-11-23T15:32:43.80415Z","shell.execute_reply":"2021-11-23T15:32:43.812863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**\"sequence\" is a global identifier for a sequence (aka it's not relative to the video_id)**","metadata":{}},{"cell_type":"code","source":"df.groupby(\"sequence\")['video_id'].nunique()","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:32:43.817528Z","iopub.execute_input":"2021-11-23T15:32:43.817763Z","iopub.status.idle":"2021-11-23T15:32:43.831953Z","shell.execute_reply.started":"2021-11-23T15:32:43.817737Z","shell.execute_reply":"2021-11-23T15:32:43.83136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Videos 0 and 1 have 8 sequences, while video 2 has 4\ndf.groupby(\"video_id\")['sequence'].nunique()","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:32:43.832904Z","iopub.execute_input":"2021-11-23T15:32:43.833242Z","iopub.status.idle":"2021-11-23T15:32:43.840014Z","shell.execute_reply.started":"2021-11-23T15:32:43.833217Z","shell.execute_reply":"2021-11-23T15:32:43.839188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_agg = df.groupby([\"video_id\", 'sequence']).agg({'sequence_frame': 'count', 'has_annotations': 'sum', 'doesnt_have_annotations': 'sum'})\\\n           .rename(columns={'sequence_frame': 'Total Frames', 'has_annotations': 'Frames with at least 1 object', 'doesnt_have_annotations': \"Frames with no object\"})\ndf_agg","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:32:43.842711Z","iopub.execute_input":"2021-11-23T15:32:43.842918Z","iopub.status.idle":"2021-11-23T15:32:43.867156Z","shell.execute_reply.started":"2021-11-23T15:32:43.842895Z","shell.execute_reply":"2021-11-23T15:32:43.866502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_agg.sort_values(\"Total Frames\")","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:32:43.868422Z","iopub.execute_input":"2021-11-23T15:32:43.869021Z","iopub.status.idle":"2021-11-23T15:32:43.880731Z","shell.execute_reply.started":"2021-11-23T15:32:43.868975Z","shell.execute_reply":"2021-11-23T15:32:43.879972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_agg.sort_values(\"Frames with at least 1 object\")","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:32:43.881741Z","iopub.execute_input":"2021-11-23T15:32:43.882364Z","iopub.status.idle":"2021-11-23T15:32:43.89954Z","shell.execute_reply.started":"2021-11-23T15:32:43.882329Z","shell.execute_reply":"2021-11-23T15:32:43.898678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The total amount of frames in each sequence varies a lot, so it might be quite difficult to use sequences as the splitting unit.","metadata":{}},{"cell_type":"code","source":"# image_id is a unique identifier for a row\ndf['image_id'].nunique() == len(df)","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:32:43.900607Z","iopub.execute_input":"2021-11-23T15:32:43.900831Z","iopub.status.idle":"2021-11-23T15:32:43.915737Z","shell.execute_reply.started":"2021-11-23T15:32:43.900804Z","shell.execute_reply":"2021-11-23T15:32:43.914977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## What do we want, ideally?\n\nThe ideal scenario would be that, if we split 80 - 20 (train - validation)\nThen the 80% of the training data has:\n* 80% sequences\n* 80% of frames \n* 80% of frames with objects\n* 80% of individuals\n\nSplitting by sequence ensures that no region of the coral appears both in training and validation data.\n","metadata":{}},{"cell_type":"markdown","source":"# Can we go into a subsequence level? \n\nSo, one idea is the following: There are large gaps of a sequence with no objects in sights. May be we can split the sequence into a smaller piece during those \"empty\" times and that might ensure that the same region of the coral, with the same individuals, doesn't appear in both training and validation data.\n\n\n### See for example, sequence 40258:","metadata":{}},{"cell_type":"code","source":"df_agg.loc[[(0, 40258)]]","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:32:43.916811Z","iopub.execute_input":"2021-11-23T15:32:43.917249Z","iopub.status.idle":"2021-11-23T15:32:43.934071Z","shell.execute_reply.started":"2021-11-23T15:32:43.917215Z","shell.execute_reply":"2021-11-23T15:32:43.93325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.set_option(\"display.max_rows\", 500)\ndf[df['sequence'] == 40258]","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:32:43.935432Z","iopub.execute_input":"2021-11-23T15:32:43.935867Z","iopub.status.idle":"2021-11-23T15:32:44.229921Z","shell.execute_reply.started":"2021-11-23T15:32:43.935826Z","shell.execute_reply":"2021-11-23T15:32:44.228975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The sequence has 4 subsequences with objects surrounded by other parts with no objects at all. We could split this sequence in these 4 subsequences and use that as units for train-validation splits.\n\n\n### Here we cut continuous subsequence with objects and without objects:","metadata":{}},{"cell_type":"code","source":"df['start_cut_here'] = df['has_annotations'] & df['doesnt_have_annotations'].shift(1)  & df['doesnt_have_annotations'].shift(2)\ndf['end_cut_here'] = df['doesnt_have_annotations'] & df['has_annotations'].shift(1)  & df['has_annotations'].shift(2)\ndf['sequence_change'] = df['sequence'] != df['sequence'].shift(1)\ndf['last_row'] =  df.index == len(df)-1\ndf['cut_here'] = df['start_cut_here'] | df['end_cut_here'] | df['sequence_change'] | df['last_row']\n","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:32:44.230986Z","iopub.execute_input":"2021-11-23T15:32:44.2312Z","iopub.status.idle":"2021-11-23T15:32:44.24948Z","shell.execute_reply.started":"2021-11-23T15:32:44.231175Z","shell.execute_reply":"2021-11-23T15:32:44.248785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"start_idx = 0\nfor subsequence_id, end_idx in enumerate(df[df['cut_here']].index):\n    df.loc[start_idx:end_idx, 'subsequence_id'] = subsequence_id\n    start_idx = end_idx","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:32:44.251758Z","iopub.execute_input":"2021-11-23T15:32:44.251989Z","iopub.status.idle":"2021-11-23T15:32:44.295902Z","shell.execute_reply.started":"2021-11-23T15:32:44.251956Z","shell.execute_reply":"2021-11-23T15:32:44.295172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['subsequence_id'] = df['subsequence_id'].astype(int)","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:48:33.05707Z","iopub.execute_input":"2021-11-23T15:48:33.057374Z","iopub.status.idle":"2021-11-23T15:48:33.064879Z","shell.execute_reply.started":"2021-11-23T15:48:33.057345Z","shell.execute_reply":"2021-11-23T15:48:33.064145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['subsequence_id'].nunique()","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:48:34.433849Z","iopub.execute_input":"2021-11-23T15:48:34.434127Z","iopub.status.idle":"2021-11-23T15:48:34.44126Z","shell.execute_reply.started":"2021-11-23T15:48:34.434097Z","shell.execute_reply":"2021-11-23T15:48:34.440438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"drop_cols = ['start_cut_here', 'end_cut_here', 'sequence_change', 'last_row', 'cut_here', 'has_2_or_more_annotations', 'doesnt_have_annotations']\ndf = df.drop(drop_cols, axis=1)\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:36:12.733685Z","iopub.execute_input":"2021-11-23T15:36:12.733948Z","iopub.status.idle":"2021-11-23T15:36:12.753564Z","shell.execute_reply.started":"2021-11-23T15:36:12.733906Z","shell.execute_reply":"2021-11-23T15:36:12.75299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The method didn't work perfectly for some subsequences, but the \"broken ones\" don't look that bad so we are cool 👍👍","metadata":{}},{"cell_type":"code","source":"df.groupby(\"subsequence_id\")['has_annotations'].mean().round(2).sort_values().value_counts()","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:48:37.380286Z","iopub.execute_input":"2021-11-23T15:48:37.380753Z","iopub.status.idle":"2021-11-23T15:48:37.390254Z","shell.execute_reply.started":"2021-11-23T15:48:37.380722Z","shell.execute_reply":"2021-11-23T15:48:37.389698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_subseq_agg = df.groupby(\"subsequence_id\")['has_annotations'].mean()\ndf_subseq_agg[~df_subseq_agg.isin([0, 1])]","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:48:38.201697Z","iopub.execute_input":"2021-11-23T15:48:38.201997Z","iopub.status.idle":"2021-11-23T15:48:38.210589Z","shell.execute_reply.started":"2021-11-23T15:48:38.201963Z","shell.execute_reply":"2021-11-23T15:48:38.20957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df[df['subsequence_id'] == 52]","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:48:38.879374Z","iopub.execute_input":"2021-11-23T15:48:38.879657Z","iopub.status.idle":"2021-11-23T15:48:38.924836Z","shell.execute_reply.started":"2021-11-23T15:48:38.879625Z","shell.execute_reply":"2021-11-23T15:48:38.924233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df[df['subsequence_id'] == 53]","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:48:39.519847Z","iopub.execute_input":"2021-11-23T15:48:39.520151Z","iopub.status.idle":"2021-11-23T15:48:39.536061Z","shell.execute_reply.started":"2021-11-23T15:48:39.520115Z","shell.execute_reply":"2021-11-23T15:48:39.535253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df[df['subsequence_id'] == 54]","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:48:39.985191Z","iopub.execute_input":"2021-11-23T15:48:39.985475Z","iopub.status.idle":"2021-11-23T15:48:40.002103Z","shell.execute_reply.started":"2021-11-23T15:48:39.985441Z","shell.execute_reply":"2021-11-23T15:48:40.001397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Let's see how a sequence and a subsequence look like as videos!!\n\n## mp4 generating code from [create annotated video](https://www.kaggle.com/bamps53/create-annotated-video)\n\n### I changed it to have sequence as parameter instead of video_id","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:32:44.407146Z","iopub.status.idle":"2021-11-23T15:32:44.407766Z","shell.execute_reply.started":"2021-11-23T15:32:44.407511Z","shell.execute_reply":"2021-11-23T15:32:44.407538Z"}}},{"cell_type":"code","source":"! mkdir videos/","metadata":{"execution":{"iopub.status.busy":"2021-11-23T16:51:30.303739Z","iopub.execute_input":"2021-11-23T16:51:30.304008Z","iopub.status.idle":"2021-11-23T16:51:31.062723Z","shell.execute_reply.started":"2021-11-23T16:51:30.303978Z","shell.execute_reply":"2021-11-23T16:51:31.061685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def load_image(img_path):\n    assert os.path.exists(img_path), f'{img_path} does not exist.'\n    img = cv2.imread(img_path)\n    return img\n\ndef load_image_with_annotations(img_path, annotations):\n    img = load_image(img_path)\n    if len(annotations) > 0:\n        for ann in annotations:\n            cv2.rectangle(img, (ann['x'], ann['y']),\n                (ann['x'] + ann['width'], ann['y'] + ann['height']),\n                (255, 255, 0), thickness=2,)\n    return img\n\ndef make_video(df, part_id, is_subsequence=False):\n    \"\"\"\n    Args:\n        - part_id: either a sequence or a subsequence id\n    \"\"\"\n    \n    if is_subsequence:\n        part_str = \"subsequence_id\"\n    else:\n        part_str = \"sequence\"\n    \n    print(f\"Creating video for part={part_id}, is_subsequence={is_subsequence} (querying by {part_str})\")\n    # partly borrowed from https://github.com/RobMulla/helmet-assignment/blob/main/helmet_assignment/video.py\n    fps = 15 # don't know exact value\n    width = 1280\n    height = 720\n    save_path = f'videos/video_{part_str}_{part_id}.mp4'\n    tmp_path = f'videos/tmp_video_{part_str}_{part_id}.mp4'\n    \n    \n    output_video = cv2.VideoWriter(tmp_path, cv2.VideoWriter_fourcc(*\"MP4V\"), fps, (width, height))\n    \n    df_part = df.query(f'{part_str} == @part_id')\n    for _, row in tqdm(df_part.iterrows(), total=len(df_part)):\n        img = load_image_with_annotations(row.image_path, row.annotations)\n        output_video.write(img)\n    \n    output_video.release()\n    # Not all browsers support the codec, we will re-load the file at tmp_output_path\n    # and convert to a codec that is more broadly readable using ffmpeg\n    if os.path.exists(save_path):\n        os.remove(save_path)\n    subprocess.run(\n        [\"ffmpeg\", \"-i\", tmp_path, \"-crf\", \"18\", \"-preset\", \"veryfast\", \"-vcodec\", \"libx264\", save_path],\n        stdout=subprocess.DEVNULL,\n        stderr=subprocess.DEVNULL\n    )\n    os.remove(tmp_path)\n    print(f\"Finished creating video for {part_id}... saved as {save_path}\")\n    return save_path","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:51:34.280643Z","iopub.status.idle":"2021-11-23T15:51:34.28106Z","shell.execute_reply.started":"2021-11-23T15:51:34.280841Z","shell.execute_reply":"2021-11-23T15:51:34.280864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Video for sequence _40258_ and its subsequences","metadata":{}},{"cell_type":"code","source":"video_path = make_video(df, 40258)","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:44:48.62454Z","iopub.execute_input":"2021-11-23T15:44:48.625286Z","iopub.status.idle":"2021-11-23T15:45:21.385756Z","shell.execute_reply.started":"2021-11-23T15:44:48.625239Z","shell.execute_reply":"2021-11-23T15:45:21.384992Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Video(video_path, width= 1280/2, height= 720/2)","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:46:06.703197Z","iopub.execute_input":"2021-11-23T15:46:06.703919Z","iopub.status.idle":"2021-11-23T15:46:06.709426Z","shell.execute_reply.started":"2021-11-23T15:46:06.703884Z","shell.execute_reply":"2021-11-23T15:46:06.708616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"subsequences = df.loc[df['sequence'] == 40258, 'subsequence_id'].unique()\nsubsequences","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:48:44.395892Z","iopub.execute_input":"2021-11-23T15:48:44.396303Z","iopub.status.idle":"2021-11-23T15:48:44.402657Z","shell.execute_reply.started":"2021-11-23T15:48:44.396273Z","shell.execute_reply":"2021-11-23T15:48:44.401861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for subsequence in subsequences:\n    video_path = make_video(df, subsequence, is_subsequence=True)\n    display(HTML(f\"<h2>Subsequence ID: {subsequence}</h2>\"))\n    display(Video(video_path, width= 1280/2, height= 720/2))","metadata":{"execution":{"iopub.status.busy":"2021-11-23T15:51:36.06963Z","iopub.execute_input":"2021-11-23T15:51:36.069892Z","iopub.status.idle":"2021-11-23T15:52:05.412125Z","shell.execute_reply.started":"2021-11-23T15:51:36.069865Z","shell.execute_reply":"2021-11-23T15:52:05.411171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This looks good 😁, let's use it for creating the splits...\n\n# Generate some common splits based on _subsequences_","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split, StratifiedKFold\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2021-11-23T16:06:57.867626Z","iopub.execute_input":"2021-11-23T16:06:57.868516Z","iopub.status.idle":"2021-11-23T16:06:57.882913Z","shell.execute_reply.started":"2021-11-23T16:06:57.868473Z","shell.execute_reply":"2021-11-23T16:06:57.881986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_split  = df.groupby(\"subsequence_id\").agg({'has_annotations': 'max', 'video_frame': 'count'}).astype(int).reset_index()\ndf_split.head()","metadata":{"execution":{"iopub.status.busy":"2021-11-23T16:07:11.536978Z","iopub.execute_input":"2021-11-23T16:07:11.537251Z","iopub.status.idle":"2021-11-23T16:07:11.550604Z","shell.execute_reply.started":"2021-11-23T16:07:11.537221Z","shell.execute_reply":"2021-11-23T16:07:11.550026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train-validation splits for 1%, 5%, 10% and 20%","metadata":{}},{"cell_type":"code","source":"!mkdir train-validation-split/","metadata":{"execution":{"iopub.status.busy":"2021-11-23T16:26:56.719747Z","iopub.execute_input":"2021-11-23T16:26:56.720483Z","iopub.status.idle":"2021-11-23T16:26:57.535511Z","shell.execute_reply.started":"2021-11-23T16:26:56.720435Z","shell.execute_reply":"2021-11-23T16:26:57.534284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def analize_split(df_train, df_val, df):\n     # Analize results\n    print(f\"   Train images                 : {len(df_train) / len(df):.3f}\")\n    print(f\"   Val   images                 : {len(df_val) / len(df):.3f}\")\n    print()\n    print(f\"   Train images with annotations: {len(df_train[df_train['has_annotations']]) / len(df[df['has_annotations']]):.3f}\")\n    print(f\"   Val   images with annotations: {len(df_val[df_val['has_annotations']]) / len(df[df['has_annotations']]):.3f}\")\n    print()\n    print(f\"   Train images w/no annotations: {len(df_train[~df_train['has_annotations']]) / len(df[~df['has_annotations']]):.3f}\")\n    print(f\"   Val   images w/no annotations: {len(df_val[~df_val['has_annotations']]) / len(df[~df['has_annotations']]):.3f}\")\n    print()\n    print(f\"   Train mean annotations       : {df_train['n_annotations'].mean():.3f}\")\n    print(f\"   Val   mean annotations       : {df_val['n_annotations'].mean():.3f}\")\n    \n    print()","metadata":{"execution":{"iopub.status.busy":"2021-11-23T16:30:38.82933Z","iopub.execute_input":"2021-11-23T16:30:38.830054Z","iopub.status.idle":"2021-11-23T16:30:38.835489Z","shell.execute_reply.started":"2021-11-23T16:30:38.830016Z","shell.execute_reply":"2021-11-23T16:30:38.834898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for test_size in [0.01, 0.05, 0.1, 0.2]:\n    print(f\"Generating train-validation split with {test_size*100}% validation\")\n    df_train_idx, df_val_idx = train_test_split(df_split['subsequence_id'], stratify=df_split[\"has_annotations\"], test_size=test_size, random_state=42)\n    df['is_train'] = df['subsequence_id'].isin(df_train_idx)\n    df_train, df_val = df[df['is_train']], df[~df['is_train']]\n    \n    # Print some statistics\n    analize_split(df_train, df_val, df)\n    \n    # Save to file\n    f_name = f\"train-validation-split/train-{test_size}.csv\"\n    print(f\"Saving file to {f_name}\")\n    df.to_csv(f_name, index=False)\n    print()\n    ","metadata":{"execution":{"iopub.status.busy":"2021-11-23T16:30:51.617501Z","iopub.execute_input":"2021-11-23T16:30:51.618195Z","iopub.status.idle":"2021-11-23T16:30:52.153441Z","shell.execute_reply.started":"2021-11-23T16:30:51.618154Z","shell.execute_reply":"2021-11-23T16:30:52.152596Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!ls -l train-validation-split/","metadata":{"execution":{"iopub.status.busy":"2021-11-23T16:28:01.769785Z","iopub.execute_input":"2021-11-23T16:28:01.770108Z","iopub.status.idle":"2021-11-23T16:28:02.576369Z","shell.execute_reply.started":"2021-11-23T16:28:01.770073Z","shell.execute_reply":"2021-11-23T16:28:02.575012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create 5-folds cross validation","metadata":{}},{"cell_type":"code","source":"df = df.drop(\"is_train\", axis=1)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_splits = 5\nkf = StratifiedKFold(n_splits=n_splits, shuffle=True, random_state=2021)\nfor fold_id, (_, val_idx) in enumerate(kf.split(df_split['subsequence_id'], y=df_split[\"has_annotations\"])):\n    subseq_val_idx = df_split['subsequence_id'].iloc[val_idx]\n    df.loc[df['subsequence_id'].isin(subseq_val_idx), 'fold'] = fold_id\n    \ndf['fold'] = df['fold'].astype(int)\ndf['fold'].value_counts(dropna=False)","metadata":{"execution":{"iopub.status.busy":"2021-11-23T16:40:01.918967Z","iopub.execute_input":"2021-11-23T16:40:01.919379Z","iopub.status.idle":"2021-11-23T16:40:01.934152Z","shell.execute_reply.started":"2021-11-23T16:40:01.919351Z","shell.execute_reply":"2021-11-23T16:40:01.933499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for fold_id in df['fold'].sort_values().unique():\n    print(\"=============================\")\n    print(f\"Analyzing fold {fold_id}\")\n    df_train, df_val = df[df['fold'] != fold_id], df[df['fold'] == fold_id]\n    analize_split(df_train, df_val, df)\n    print()","metadata":{"execution":{"iopub.status.busy":"2021-11-23T16:38:03.899089Z","iopub.execute_input":"2021-11-23T16:38:03.899527Z","iopub.status.idle":"2021-11-23T16:38:03.9773Z","shell.execute_reply.started":"2021-11-23T16:38:03.899485Z","shell.execute_reply":"2021-11-23T16:38:03.976361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!mkdir cross-validation/","metadata":{"execution":{"iopub.status.busy":"2021-11-23T16:39:16.169721Z","iopub.execute_input":"2021-11-23T16:39:16.169997Z","iopub.status.idle":"2021-11-23T16:39:16.967729Z","shell.execute_reply.started":"2021-11-23T16:39:16.169966Z","shell.execute_reply":"2021-11-23T16:39:16.96663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.to_csv(\"cross-validation/train-5folds.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2021-11-23T16:39:38.54036Z","iopub.execute_input":"2021-11-23T16:39:38.540762Z","iopub.status.idle":"2021-11-23T16:39:38.664848Z","shell.execute_reply.started":"2021-11-23T16:39:38.540734Z","shell.execute_reply":"2021-11-23T16:39:38.663991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create 10-fold cross validation","metadata":{}},{"cell_type":"code","source":"n_splits = 10\nkf = StratifiedKFold(n_splits=n_splits, shuffle=True, random_state=2021)\nfor fold_id, (_, val_idx) in enumerate(kf.split(df_split['subsequence_id'], y=df_split[\"has_annotations\"])):\n    subseq_val_idx = df_split['subsequence_id'].iloc[val_idx]\n    df.loc[df['subsequence_id'].isin(subseq_val_idx), 'fold'] = fold_id\n    \ndf['fold'] = df['fold'].astype(int)\ndf['fold'].value_counts(dropna=False)","metadata":{"execution":{"iopub.status.busy":"2021-11-23T16:40:14.014863Z","iopub.execute_input":"2021-11-23T16:40:14.015422Z","iopub.status.idle":"2021-11-23T16:40:14.037476Z","shell.execute_reply.started":"2021-11-23T16:40:14.015374Z","shell.execute_reply":"2021-11-23T16:40:14.036751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for fold_id in df['fold'].sort_values().unique():\n    print(\"=============================\")\n    print(f\"Analyzing fold {fold_id}\")\n    df_train, df_val = df[df['fold'] != fold_id], df[df['fold'] == fold_id]\n    analize_split(df_train, df_val, df)\n    print()","metadata":{"execution":{"iopub.status.busy":"2021-11-23T16:40:15.229298Z","iopub.execute_input":"2021-11-23T16:40:15.230144Z","iopub.status.idle":"2021-11-23T16:40:15.36243Z","shell.execute_reply.started":"2021-11-23T16:40:15.230091Z","shell.execute_reply":"2021-11-23T16:40:15.36167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.to_csv(\"cross-validation/train-10folds.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2021-11-23T16:40:34.410813Z","iopub.execute_input":"2021-11-23T16:40:34.411091Z","iopub.status.idle":"2021-11-23T16:40:34.535513Z","shell.execute_reply.started":"2021-11-23T16:40:34.411063Z","shell.execute_reply":"2021-11-23T16:40:34.534693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Please, _DO_ upvote if you find this useful or interesting!","metadata":{}}]}