{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"#### What are you trying to do in this notebook?\nMy goal for this competition is to accurate identify starfish in real-time by building an object detection model trained on underwater videos of coral reefs.\nMy work will help researchers to identify species that are threatening Australia's Great Barrier Reef and take well-informed action to protect the reef for future generations.\nIn this notebook we explore sequences as potential units for cross-validation, but since there are only 20 sequences and their sizes are quite disimilar, we propose an approach to split them into smaller chunks, that we name subsequences.\n\n#### Why are you trying it?\nTo detect crown-of-thorns starfish in underwater image data. In this competition, I will predict the presence and position of crown-of-thorns starfish in sequences of underwater images taken at various times and locations around the Great Barrier Reef. Predictions take the form of a bounding box together with a confidence score for each identified starfish. An image may contain zero or more starfish.\n\nIn this notebook we explore sequences as potential units for cross-validation, but since there are only 20 sequences and their sizes are quite disimilar, we propose an approach to split them into smaller chunks, that we name subsequences.\n\nA sequence, as stated in the data tab of the competition, is:\n\nsequence - ID of a gap-free subset of a given video. The sequence ids are not meaningfully ordered.\n\nSubsequences, as we will define them below, are parts of a sequences where objects are continually present or are continually not present. We isolate 2 kind of subsequences: with objects and with no objects.","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:01.004236Z","iopub.execute_input":"2021-12-08T10:05:01.004616Z","iopub.status.idle":"2021-12-08T10:05:29.813146Z","shell.execute_reply.started":"2021-12-08T10:05:01.004520Z","shell.execute_reply":"2021-12-08T10:05:29.811200Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport cv2\nimport subprocess\nfrom tqdm.auto import tqdm\nimport pandas as pd\nfrom IPython.display import Video, display, HTML\nimport warnings; warnings.simplefilter(\"ignore\")\n\n\nBASE_PATH = '../input/tensorflow-great-barrier-reef/train_images/'\n\ndf = pd.read_csv(\"/kaggle/input/tensorflow-great-barrier-reef/train.csv\")\ndf['annotations'] = df['annotations'].apply(eval)\ndf['n_annotations'] = df['annotations'].str.len()\ndf['has_annotations'] = df['annotations'].str.len() > 0\ndf['has_2_or_more_annotations'] = df['annotations'].str.len() >= 2\ndf['doesnt_have_annotations'] = df['annotations'].str.len() == 0\ndf['image_path'] = BASE_PATH + \"video_\" + df['video_id'].astype(str) + \"/\" + df['video_frame'].astype(str) + \".jpg\"","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:29.815383Z","iopub.execute_input":"2021-12-08T10:05:29.815788Z","iopub.status.idle":"2021-12-08T10:05:30.569112Z","shell.execute_reply.started":"2021-12-08T10:05:29.815745Z","shell.execute_reply":"2021-12-08T10:05:30.568208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['sequence'].unique()","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:30.570656Z","iopub.execute_input":"2021-12-08T10:05:30.571146Z","iopub.status.idle":"2021-12-08T10:05:30.584282Z","shell.execute_reply.started":"2021-12-08T10:05:30.571091Z","shell.execute_reply":"2021-12-08T10:05:30.582826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['sequence'].nunique()","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:30.587807Z","iopub.execute_input":"2021-12-08T10:05:30.588492Z","iopub.status.idle":"2021-12-08T10:05:30.597099Z","shell.execute_reply.started":"2021-12-08T10:05:30.588449Z","shell.execute_reply":"2021-12-08T10:05:30.595921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.groupby(\"sequence\")['video_id'].nunique()","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:30.598924Z","iopub.execute_input":"2021-12-08T10:05:30.600128Z","iopub.status.idle":"2021-12-08T10:05:30.616667Z","shell.execute_reply.started":"2021-12-08T10:05:30.600086Z","shell.execute_reply":"2021-12-08T10:05:30.615591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Videos 0 and 1 have 8 sequences, while video 2 has 4\ndf.groupby(\"video_id\")['sequence'].nunique()\n","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:30.618377Z","iopub.execute_input":"2021-12-08T10:05:30.618727Z","iopub.status.idle":"2021-12-08T10:05:30.630143Z","shell.execute_reply.started":"2021-12-08T10:05:30.618682Z","shell.execute_reply":"2021-12-08T10:05:30.628987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_agg = df.groupby([\"video_id\", 'sequence']).agg({'sequence_frame': 'count', 'has_annotations': 'sum', 'doesnt_have_annotations': 'sum'})\\\n           .rename(columns={'sequence_frame': 'Total Frames', 'has_annotations': 'Frames with at least 1 object', 'doesnt_have_annotations': \"Frames with no object\"})\ndf_agg","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:30.631938Z","iopub.execute_input":"2021-12-08T10:05:30.632724Z","iopub.status.idle":"2021-12-08T10:05:30.662115Z","shell.execute_reply.started":"2021-12-08T10:05:30.632681Z","shell.execute_reply":"2021-12-08T10:05:30.661259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_agg.sort_values(\"Total Frames\")","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:30.663332Z","iopub.execute_input":"2021-12-08T10:05:30.664049Z","iopub.status.idle":"2021-12-08T10:05:30.680322Z","shell.execute_reply.started":"2021-12-08T10:05:30.664008Z","shell.execute_reply":"2021-12-08T10:05:30.678989Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_agg.sort_values(\"Frames with at least 1 object\")\n","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:30.682305Z","iopub.execute_input":"2021-12-08T10:05:30.683001Z","iopub.status.idle":"2021-12-08T10:05:30.701557Z","shell.execute_reply.started":"2021-12-08T10:05:30.682941Z","shell.execute_reply":"2021-12-08T10:05:30.700593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# image_id is a unique identifier for a row\ndf['image_id'].nunique() == len(df)","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:30.705397Z","iopub.execute_input":"2021-12-08T10:05:30.705659Z","iopub.status.idle":"2021-12-08T10:05:30.720254Z","shell.execute_reply.started":"2021-12-08T10:05:30.705630Z","shell.execute_reply":"2021-12-08T10:05:30.718771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_agg.loc[[(0, 40258)]]\n","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:30.721664Z","iopub.execute_input":"2021-12-08T10:05:30.722661Z","iopub.status.idle":"2021-12-08T10:05:30.741811Z","shell.execute_reply.started":"2021-12-08T10:05:30.722618Z","shell.execute_reply":"2021-12-08T10:05:30.740602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.set_option(\"display.max_rows\", 500)\ndf[df['sequence'] == 40258]","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:30.743415Z","iopub.execute_input":"2021-12-08T10:05:30.743806Z","iopub.status.idle":"2021-12-08T10:05:31.141426Z","shell.execute_reply.started":"2021-12-08T10:05:30.743761Z","shell.execute_reply":"2021-12-08T10:05:31.140547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['start_cut_here'] = df['has_annotations'] & df['doesnt_have_annotations'].shift(1)  & df['doesnt_have_annotations'].shift(2)\ndf['end_cut_here'] = df['doesnt_have_annotations'] & df['has_annotations'].shift(1)  & df['has_annotations'].shift(2)\ndf['sequence_change'] = df['sequence'] != df['sequence'].shift(1)\ndf['last_row'] =  df.index == len(df)-1\ndf['cut_here'] = df['start_cut_here'] | df['end_cut_here'] | df['sequence_change'] | df['last_row']","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:31.143158Z","iopub.execute_input":"2021-12-08T10:05:31.144150Z","iopub.status.idle":"2021-12-08T10:05:31.177757Z","shell.execute_reply.started":"2021-12-08T10:05:31.144110Z","shell.execute_reply":"2021-12-08T10:05:31.176935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"start_idx = 0\nfor subsequence_id, end_idx in enumerate(df[df['cut_here']].index):\n    df.loc[start_idx:end_idx, 'subsequence_id'] = subsequence_id\n    start_idx = end_idx","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:31.179079Z","iopub.execute_input":"2021-12-08T10:05:31.180290Z","iopub.status.idle":"2021-12-08T10:05:31.239716Z","shell.execute_reply.started":"2021-12-08T10:05:31.180250Z","shell.execute_reply":"2021-12-08T10:05:31.238625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['subsequence_id'] = df['subsequence_id'].astype(int)\n","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:31.241278Z","iopub.execute_input":"2021-12-08T10:05:31.242026Z","iopub.status.idle":"2021-12-08T10:05:31.248349Z","shell.execute_reply.started":"2021-12-08T10:05:31.241985Z","shell.execute_reply":"2021-12-08T10:05:31.246984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['subsequence_id'].nunique()\n","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:31.250179Z","iopub.execute_input":"2021-12-08T10:05:31.251197Z","iopub.status.idle":"2021-12-08T10:05:31.263450Z","shell.execute_reply.started":"2021-12-08T10:05:31.251151Z","shell.execute_reply":"2021-12-08T10:05:31.262395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"drop_cols = ['start_cut_here', 'end_cut_here', 'sequence_change', 'last_row', 'cut_here', 'has_2_or_more_annotations', 'doesnt_have_annotations']\ndf = df.drop(drop_cols, axis=1)\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:31.264775Z","iopub.execute_input":"2021-12-08T10:05:31.265854Z","iopub.status.idle":"2021-12-08T10:05:31.294982Z","shell.execute_reply.started":"2021-12-08T10:05:31.265794Z","shell.execute_reply":"2021-12-08T10:05:31.294141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.groupby(\"subsequence_id\")['has_annotations'].mean().round(2).sort_values().value_counts()","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:31.296738Z","iopub.execute_input":"2021-12-08T10:05:31.297318Z","iopub.status.idle":"2021-12-08T10:05:31.312625Z","shell.execute_reply.started":"2021-12-08T10:05:31.297243Z","shell.execute_reply":"2021-12-08T10:05:31.311242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_subseq_agg = df.groupby(\"subsequence_id\")['has_annotations'].mean()\ndf_subseq_agg[~df_subseq_agg.isin([0, 1])]","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:31.314046Z","iopub.execute_input":"2021-12-08T10:05:31.314975Z","iopub.status.idle":"2021-12-08T10:05:31.327081Z","shell.execute_reply.started":"2021-12-08T10:05:31.314931Z","shell.execute_reply":"2021-12-08T10:05:31.325789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df[df['subsequence_id'] == 52]","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:31.328842Z","iopub.execute_input":"2021-12-08T10:05:31.330099Z","iopub.status.idle":"2021-12-08T10:05:31.388543Z","shell.execute_reply.started":"2021-12-08T10:05:31.330052Z","shell.execute_reply":"2021-12-08T10:05:31.387612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df[df['subsequence_id'] == 54]","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:31.389925Z","iopub.execute_input":"2021-12-08T10:05:31.391004Z","iopub.status.idle":"2021-12-08T10:05:31.414348Z","shell.execute_reply.started":"2021-12-08T10:05:31.390962Z","shell.execute_reply":"2021-12-08T10:05:31.413248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"! mkdir videos/","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:31.416407Z","iopub.execute_input":"2021-12-08T10:05:31.416853Z","iopub.status.idle":"2021-12-08T10:05:32.352094Z","shell.execute_reply.started":"2021-12-08T10:05:31.416809Z","shell.execute_reply":"2021-12-08T10:05:32.350838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def load_image(img_path):\n    assert os.path.exists(img_path), f'{img_path} does not exist.'\n    img = cv2.imread(img_path)\n    return img\n\ndef load_image_with_annotations(img_path, annotations):\n    img = load_image(img_path)\n    if len(annotations) > 0:\n        for ann in annotations:\n            cv2.rectangle(img, (ann['x'], ann['y']),\n                (ann['x'] + ann['width'], ann['y'] + ann['height']),\n                (255, 255, 0), thickness=2,)\n    return img\n\ndef make_video(df, part_id, is_subsequence=False):\n    \"\"\"\n    Args:\n        - part_id: either a sequence or a subsequence id\n    \"\"\"\n    \n    if is_subsequence:\n        part_str = \"subsequence_id\"\n    else:\n        part_str = \"sequence\"\n    \n    print(f\"Creating video for part={part_id}, is_subsequence={is_subsequence} (querying by {part_str})\")\n    # partly borrowed from https://github.com/RobMulla/helmet-assignment/blob/main/helmet_assignment/video.py\n    fps = 15 # don't know exact value\n    width = 1280\n    height = 720\n    save_path = f'videos/video_{part_str}_{part_id}.mp4'\n    tmp_path = f'videos/tmp_video_{part_str}_{part_id}.mp4'\n    \n    \n    output_video = cv2.VideoWriter(tmp_path, cv2.VideoWriter_fourcc(*\"MP4V\"), fps, (width, height))\n    \n    df_part = df.query(f'{part_str} == @part_id')\n    for _, row in tqdm(df_part.iterrows(), total=len(df_part)):\n        img = load_image_with_annotations(row.image_path, row.annotations)\n        output_video.write(img)\n    \n    output_video.release()\n    # Not all browsers support the codec, we will re-load the file at tmp_output_path\n    # and convert to a codec that is more broadly readable using ffmpeg\n    if os.path.exists(save_path):\n        os.remove(save_path)\n    subprocess.run(\n        [\"ffmpeg\", \"-i\", tmp_path, \"-crf\", \"18\", \"-preset\", \"veryfast\", \"-vcodec\", \"libx264\", save_path],\n        stdout=subprocess.DEVNULL,\n        stderr=subprocess.DEVNULL\n    )\n    os.remove(tmp_path)\n    print(f\"Finished creating video for {part_id}... saved as {save_path}\")\n    return save_path","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:32.354702Z","iopub.execute_input":"2021-12-08T10:05:32.355350Z","iopub.status.idle":"2021-12-08T10:05:32.369804Z","shell.execute_reply.started":"2021-12-08T10:05:32.355293Z","shell.execute_reply":"2021-12-08T10:05:32.368786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"video_path = make_video(df, 40258)","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:05:32.373688Z","iopub.execute_input":"2021-12-08T10:05:32.374501Z","iopub.status.idle":"2021-12-08T10:06:15.452033Z","shell.execute_reply.started":"2021-12-08T10:05:32.374468Z","shell.execute_reply":"2021-12-08T10:06:15.450779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Video(video_path, width= 1280/2, height= 720/2)","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:06:15.454558Z","iopub.execute_input":"2021-12-08T10:06:15.454953Z","iopub.status.idle":"2021-12-08T10:06:15.463141Z","shell.execute_reply.started":"2021-12-08T10:06:15.454905Z","shell.execute_reply":"2021-12-08T10:06:15.461953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"subsequences = df.loc[df['sequence'] == 40258, 'subsequence_id'].unique()\nsubsequences","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:06:15.465001Z","iopub.execute_input":"2021-12-08T10:06:15.465827Z","iopub.status.idle":"2021-12-08T10:06:15.482971Z","shell.execute_reply.started":"2021-12-08T10:06:15.465607Z","shell.execute_reply":"2021-12-08T10:06:15.481677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for subsequence in subsequences:\n    video_path = make_video(df, subsequence, is_subsequence=True)\n    display(HTML(f\"<h2>Subsequence ID: {subsequence}</h2>\"))\n    display(Video(video_path, width= 1280/2, height= 720/2))","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:06:15.485917Z","iopub.execute_input":"2021-12-08T10:06:15.486185Z","iopub.status.idle":"2021-12-08T10:06:56.386798Z","shell.execute_reply.started":"2021-12-08T10:06:15.486155Z","shell.execute_reply":"2021-12-08T10:06:56.385502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split, StratifiedKFold\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:06:56.392545Z","iopub.execute_input":"2021-12-08T10:06:56.392782Z","iopub.status.idle":"2021-12-08T10:06:57.392991Z","shell.execute_reply.started":"2021-12-08T10:06:56.392752Z","shell.execute_reply":"2021-12-08T10:06:57.390961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_split  = df.groupby(\"subsequence_id\").agg({'has_annotations': 'max', 'video_frame': 'count'}).astype(int).reset_index()\ndf_split.head()","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:06:57.394799Z","iopub.execute_input":"2021-12-08T10:06:57.395166Z","iopub.status.idle":"2021-12-08T10:06:57.415827Z","shell.execute_reply.started":"2021-12-08T10:06:57.395106Z","shell.execute_reply":"2021-12-08T10:06:57.414936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!mkdir train-validation-split/","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:06:57.417433Z","iopub.execute_input":"2021-12-08T10:06:57.417935Z","iopub.status.idle":"2021-12-08T10:06:58.167534Z","shell.execute_reply.started":"2021-12-08T10:06:57.417872Z","shell.execute_reply":"2021-12-08T10:06:58.166253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def analize_split(df_train, df_val, df):\n     # Analize results\n    print(f\"   Train images                 : {len(df_train) / len(df):.3f}\")\n    print(f\"   Val   images                 : {len(df_val) / len(df):.3f}\")\n    print()\n    print(f\"   Train images with annotations: {len(df_train[df_train['has_annotations']]) / len(df[df['has_annotations']]):.3f}\")\n    print(f\"   Val   images with annotations: {len(df_val[df_val['has_annotations']]) / len(df[df['has_annotations']]):.3f}\")\n    print()\n    print(f\"   Train images w/no annotations: {len(df_train[~df_train['has_annotations']]) / len(df[~df['has_annotations']]):.3f}\")\n    print(f\"   Val   images w/no annotations: {len(df_val[~df_val['has_annotations']]) / len(df[~df['has_annotations']]):.3f}\")\n    print()\n    print(f\"   Train mean annotations       : {df_train['n_annotations'].mean():.3f}\")\n    print(f\"   Val   mean annotations       : {df_val['n_annotations'].mean():.3f}\")\n    \n    print()","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:06:58.169658Z","iopub.execute_input":"2021-12-08T10:06:58.170339Z","iopub.status.idle":"2021-12-08T10:06:58.179458Z","shell.execute_reply.started":"2021-12-08T10:06:58.170269Z","shell.execute_reply":"2021-12-08T10:06:58.177723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for test_size in [0.01, 0.05, 0.1, 0.2]:\n    print(f\"Generating train-validation split with {test_size*100}% validation\")\n    df_train_idx, df_val_idx = train_test_split(df_split['subsequence_id'], stratify=df_split[\"has_annotations\"], test_size=test_size, random_state=42)\n    df['is_train'] = df['subsequence_id'].isin(df_train_idx)\n    df_train, df_val = df[df['is_train']], df[~df['is_train']]\n    \n    # Print some statistics\n    analize_split(df_train, df_val, df)\n    \n    # Save to file\n    f_name = f\"train-validation-split/train-{test_size}.csv\"\n    print(f\"Saving file to {f_name}\")\n    df.to_csv(f_name, index=False)\n    print()","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:06:58.181164Z","iopub.execute_input":"2021-12-08T10:06:58.181706Z","iopub.status.idle":"2021-12-08T10:06:59.007326Z","shell.execute_reply.started":"2021-12-08T10:06:58.181643Z","shell.execute_reply":"2021-12-08T10:06:59.006377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!ls -l train-validation-split/","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:06:59.008639Z","iopub.execute_input":"2021-12-08T10:06:59.009019Z","iopub.status.idle":"2021-12-08T10:06:59.843388Z","shell.execute_reply.started":"2021-12-08T10:06:59.008975Z","shell.execute_reply":"2021-12-08T10:06:59.842294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df.drop(\"is_train\", axis=1)","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:06:59.846910Z","iopub.execute_input":"2021-12-08T10:06:59.847272Z","iopub.status.idle":"2021-12-08T10:06:59.868058Z","shell.execute_reply.started":"2021-12-08T10:06:59.847208Z","shell.execute_reply":"2021-12-08T10:06:59.867178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_splits = 5\nkf = StratifiedKFold(n_splits=n_splits, shuffle=True, random_state=2021)\nfor fold_id, (_, val_idx) in enumerate(kf.split(df_split['subsequence_id'], y=df_split[\"has_annotations\"])):\n    subseq_val_idx = df_split['subsequence_id'].iloc[val_idx]\n    df.loc[df['subsequence_id'].isin(subseq_val_idx), 'fold'] = fold_id\n    \ndf['fold'] = df['fold'].astype(int)\ndf['fold'].value_counts(dropna=False)","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:06:59.872901Z","iopub.execute_input":"2021-12-08T10:06:59.875297Z","iopub.status.idle":"2021-12-08T10:06:59.910780Z","shell.execute_reply.started":"2021-12-08T10:06:59.875253Z","shell.execute_reply":"2021-12-08T10:06:59.909862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for fold_id in df['fold'].sort_values().unique():\n    print(\"=============================\")\n    print(f\"Analyzing fold {fold_id}\")\n    df_train, df_val = df[df['fold'] != fold_id], df[df['fold'] == fold_id]\n    analize_split(df_train, df_val, df)\n    print()","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:06:59.919445Z","iopub.execute_input":"2021-12-08T10:06:59.920020Z","iopub.status.idle":"2021-12-08T10:07:00.029224Z","shell.execute_reply.started":"2021-12-08T10:06:59.919980Z","shell.execute_reply":"2021-12-08T10:07:00.028313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!mkdir cross-validation/","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:07:00.033847Z","iopub.execute_input":"2021-12-08T10:07:00.036383Z","iopub.status.idle":"2021-12-08T10:07:00.812087Z","shell.execute_reply.started":"2021-12-08T10:07:00.036338Z","shell.execute_reply":"2021-12-08T10:07:00.810920Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.to_csv(\"cross-validation/train-5folds.csv\", index=False)\n","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:07:00.814411Z","iopub.execute_input":"2021-12-08T10:07:00.815225Z","iopub.status.idle":"2021-12-08T10:07:01.002052Z","shell.execute_reply.started":"2021-12-08T10:07:00.815177Z","shell.execute_reply":"2021-12-08T10:07:01.001108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_splits = 10\nkf = StratifiedKFold(n_splits=n_splits, shuffle=True, random_state=2021)\nfor fold_id, (_, val_idx) in enumerate(kf.split(df_split['subsequence_id'], y=df_split[\"has_annotations\"])):\n    subseq_val_idx = df_split['subsequence_id'].iloc[val_idx]\n    df.loc[df['subsequence_id'].isin(subseq_val_idx), 'fold'] = fold_id\n    \ndf['fold'] = df['fold'].astype(int)\ndf['fold'].value_counts(dropna=False)","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:07:01.003494Z","iopub.execute_input":"2021-12-08T10:07:01.004495Z","iopub.status.idle":"2021-12-08T10:07:01.032788Z","shell.execute_reply.started":"2021-12-08T10:07:01.004452Z","shell.execute_reply":"2021-12-08T10:07:01.031913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for fold_id in df['fold'].sort_values().unique():\n    print(\"=============================\")\n    print(f\"Analyzing fold {fold_id}\")\n    df_train, df_val = df[df['fold'] != fold_id], df[df['fold'] == fold_id]\n    analize_split(df_train, df_val, df)\n    print()","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:07:01.034230Z","iopub.execute_input":"2021-12-08T10:07:01.036049Z","iopub.status.idle":"2021-12-08T10:07:01.196640Z","shell.execute_reply.started":"2021-12-08T10:07:01.036006Z","shell.execute_reply":"2021-12-08T10:07:01.195599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.to_csv(\"cross-validation/train-10folds.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2021-12-08T10:07:01.198273Z","iopub.execute_input":"2021-12-08T10:07:01.198581Z","iopub.status.idle":"2021-12-08T10:07:01.392586Z","shell.execute_reply.started":"2021-12-08T10:07:01.198529Z","shell.execute_reply":"2021-12-08T10:07:01.391499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Did it work?\nThis competition uses a hidden test set that will be served by an API to ensure you evaluate the images in the same order they were recorded within each video.\nI tried doing EDA for the provided dataset. I plan to dig much dipper and understand the data in a much better way. There are alot of things to discover from this dataset.\n\n#### What did you not understand about this process?\nWell, everything provides in the competition data page. I've no problem while working on it. If you guys don't understand the thing that I'll do in this notebook then please comment on this notebook.\n\n#### What else do you think you can try as part of this approach?\nWe solve the greatest challenges through innovative science and technology to unlock a better future for everyone. We are thinkers, problem solvers, leaders. We blaze new trails of discovery. We aim to inspire the next generation. The Great Barrier Reef Foundation creates a better future for coral reefs and their marine life through innovative projects and global advocacy efforts.","metadata":{}},{"cell_type":"markdown","source":"#### PLEASE UPVOTE if you find this notebook is useful for you !","metadata":{}}]}