{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1><center>HuBMAP: Advance Data Augmentations | W&B</center></h1>\n                                                      \n<center><img src = \"https://hubmapconsortium.org/wp-content/uploads/2019/01/HuBMAP-Retina-Logo-Color.png\" width = \"750\" height = \"500\"/></center>                                                                          ","metadata":{}},{"cell_type":"markdown","source":"In the first version of the [EDA Notebook](https://www.kaggle.com/code/ishandutta/hubmap-complete-understanding-and-eda-w-b) (Presently highest upvoted for this competition), we learnt how we can combine the images with their masks and visualize everything with weights and biases.\n  \nThis notebook is a continuation, which teaches you how to perform advanced data augmentations on the image, mask and the combined image. We shall se real time comparison of the images and how every augmentation affects them.\n  \nThis is a highly detailed notebook on how you can apply any type of transformation on your dataset.","metadata":{}},{"cell_type":"markdown","source":"<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='background:orange; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>Contents</center></h2>","metadata":{}},{"cell_type":"markdown","source":"> | S.No       |                   Heading                |\n> | :------------- | :-------------------:                |                         \n> |  01 |  [**Libraries**](#libraries)                        |  \n> |  02 |  [**Global Config**](#global-config)                |\n> |  03 |  [**Weights and Biases**](#weights-and-biases)      |\n> |  04 |  [**Load Datasets**](#load-datasets)                |\n> |  05 |  [**Basic Image Augmentations**](#basic-image-augmentations)   |","metadata":{}},{"cell_type":"markdown","source":"<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h3 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='background:maroon; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>If you find this notebook useful, do give me an upvote, it helps to keep up my motivation. This notebook will be updated frequently so keep checking for furthur developments.</center></h3>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"libraries\"></a>\n<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='background:orange; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>Libraries</center></h2>","metadata":{}},{"cell_type":"code","source":"%%sh\npip install -q albumentations\npip install -q --upgrade wandb","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\nimport os\nimport glob\nimport sys\nimport cv2\nimport imageio\nimport joblib\nimport math\nimport random\nimport wandb\nimport math\n\nimport numpy as np\nimport pandas as pd\n\nfrom scipy.stats import kstest\n\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as patches\nimport matplotlib.image as mpimg\n\nfrom PIL import Image\n\nfrom statsmodels.graphics.gofplots import qqplot\n\nplt.rcParams.update({'font.size': 18})\nplt.style.use('fivethirtyeight')\n\nimport seaborn as sns\nimport matplotlib\n\nfrom termcolor import colored\n\nfrom multiprocessing import cpu_count\nfrom tqdm.notebook import tqdm\nfrom sklearn.model_selection import StratifiedKFold\nfrom scipy.stats import pearsonr\n\nimport torch\nimport transformers\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom torch.utils.data import DataLoader, Dataset\nimport torchvision\nimport torchvision.transforms as transforms\n\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.metrics import r2_score, mean_squared_error\n\nfrom albumentations import (\n    HorizontalFlip, VerticalFlip, IAAPerspective, ShiftScaleRotate, CLAHE, RandomRotate90,\n    Transpose, ShiftScaleRotate, Blur, OpticalDistortion, GridDistortion, HueSaturationValue,\n    IAAAdditiveGaussianNoise, GaussNoise, MotionBlur, MedianBlur, IAAPiecewiseAffine, RandomResizedCrop,\n    IAASharpen, IAAEmboss, RandomBrightnessContrast, Flip, OneOf, Compose, Normalize, Cutout, CoarseDropout, ShiftScaleRotate, CenterCrop, Resize\n)\nfrom albumentations.pytorch import ToTensorV2\n\nimport tifffile as tiff \nfrom tqdm.auto import tqdm\n\nimport warnings\nwarnings.simplefilter('ignore')\n\n# Activate pandas progress apply bar\ntqdm.pandas()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T18:24:18.772844Z","iopub.execute_input":"2022-07-09T18:24:18.773172Z","iopub.status.idle":"2022-07-09T18:24:32.825731Z","shell.execute_reply.started":"2022-07-09T18:24:18.773141Z","shell.execute_reply":"2022-07-09T18:24:32.824475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Wandb Login\nimport wandb\nwandb.login()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T18:24:40.820325Z","iopub.execute_input":"2022-07-09T18:24:40.821585Z","iopub.status.idle":"2022-07-09T18:24:49.588427Z","shell.execute_reply.started":"2022-07-09T18:24:40.821523Z","shell.execute_reply":"2022-07-09T18:24:49.586852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"global-config\"></a>\n<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='background:orange; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>Global Config</center></h2>","metadata":{}},{"cell_type":"code","source":"class config:\n    BASE_PATH = \"../input/hubmap-organ-segmentation/\"\n    TRAIN_PATH = os.path.join(BASE_PATH, \"train\")\n\n# wandb config\nWANDB_CONFIG = {\n     'competition': 'HuBMAP', \n              '_wandb_kernel': 'neuracort'\n    }\n\n# # Initialize W&B\n# run = wandb.init(\n#     project='hubmap-data-augmentation', \n#     config= WANDB_CONFIG\n# )","metadata":{"execution":{"iopub.status.busy":"2022-07-09T18:25:04.153453Z","iopub.execute_input":"2022-07-09T18:25:04.154589Z","iopub.status.idle":"2022-07-09T18:25:04.160775Z","shell.execute_reply.started":"2022-07-09T18:25:04.154536Z","shell.execute_reply":"2022-07-09T18:25:04.159598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"weights-and-biases\"></a>\n<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='background:orange; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>Weights and Biases</center></h2>","metadata":{}},{"cell_type":"markdown","source":"<center><img src = \"https://i.imgur.com/1sm6x8P.png\" width = \"750\" height = \"500\"/></center>        ","metadata":{}},{"cell_type":"markdown","source":"**Weights & Biases** is the machine learning platform for developers to build better models faster.\n\nYou can use W&B's lightweight, interoperable tools to\n\n- quickly track experiments,\n- version and iterate on datasets,\n- evaluate model performance,\n- reproduce models,\n- visualize results and spot regressions,\n- and share findings with colleagues.\n  \nSet up W&B in 5 minutes, then quickly iterate on your machine learning pipeline with the confidence that your datasets and models are tracked and versioned in a reliable system of record.\n\nIn this notebook I will use Weights and Biases's amazing features to perform wonderful visualizations and logging seamlessly.","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"<a id=\"load-datasets\"></a>\n<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='background:orange; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>Load Datasets</center></h2>","metadata":{}},{"cell_type":"markdown","source":"## **<span style=\"color:orange;\">Train Dataset</span>** ","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv(\n    os.path.join(config.BASE_PATH, \"train.csv\")\n)\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T18:25:11.005712Z","iopub.execute_input":"2022-07-09T18:25:11.006572Z","iopub.status.idle":"2022-07-09T18:25:11.364816Z","shell.execute_reply.started":"2022-07-09T18:25:11.006529Z","shell.execute_reply":"2022-07-09T18:25:11.363476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%capture\nwandb.log({\"df_train\": df})","metadata":{"execution":{"iopub.status.busy":"2022-07-08T19:37:45.808480Z","iopub.execute_input":"2022-07-08T19:37:45.808884Z","iopub.status.idle":"2022-07-08T19:37:46.946625Z","shell.execute_reply.started":"2022-07-08T19:37:45.808851Z","shell.execute_reply":"2022-07-08T19:37:46.945472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### [Interactive W&B Table for Dataframe $\\rightarrow$](https://wandb.ai/ishandutta/hubmap-data-augmentation?workspace=user-ishandutta)\n![Animation.gif](https://iili.io/wTrTCb.gif)\n\n\nThe above gif shows the following features you can seamlessly use on wandb -\n\n- Interactively looking at the dataframe\n- Groupby based on columns and visualize the results\n- Select and deselect columns\n- Sorting columns in Ascending and Descending","metadata":{}},{"cell_type":"markdown","source":"<a id=\"basic-image-augmentations\"></a>\n<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='background:orange; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>Basic Image Augmentations</center></h2>","metadata":{}},{"cell_type":"markdown","source":"In this section we will see what are the different **color spaces** in which we can transform the images. Based on your target you can select suitable color spaces and train model on that dataset.\n  \n**I will demonstrate 9 different variations of color spaces namely:**\n1. [Black and White](#1.1)   \n2. [Ben Graham: Greyscale + Gaussian Blur](#1.2)\n3. [Hue, Saturation, Brightness](#1.3) \n4. [LUV Color Space](#1.4) \n5. [Alpha Channel](#1.5)\n6. [XYZ Color Space](#1.6)\n7. [Luma Chroma](#1.7)\n8. [CIE Lab](#1.8)\n9. [YUV Color Space](#1.9)","metadata":{}},{"cell_type":"code","source":"image_ids = df.id\nimage_files = glob.glob(config.BASE_PATH + \"/train_images/*\")\n\n# For demonstration purposes we are taking into account 10 samples only\nimage_ids = image_ids[:10]\nimage_files = image_files[:10]","metadata":{"execution":{"iopub.status.busy":"2022-07-09T18:25:15.418560Z","iopub.execute_input":"2022-07-09T18:25:15.418992Z","iopub.status.idle":"2022-07-09T18:25:15.470203Z","shell.execute_reply.started":"2022-07-09T18:25:15.418957Z","shell.execute_reply":"2022-07-09T18:25:15.468862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# https://www.kaggle.com/paulorzp/rle-functions-run-length-encode-decode\ndef mask2rle(img):\n    '''\n    img: numpy array, 1 - mask, 0 - background\n    Returns run length as string formated\n    '''\n    pixels= img.T.flatten()\n    pixels = np.concatenate([[0], pixels, [0]])\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 1\n    runs[1::2] -= runs[::2]\n    return ' '.join(str(x) for x in runs)\n \ndef rle2mask(mask_rle, shape=(1600,256)):\n    '''\n    mask_rle: run-length as string formated (start length)\n    shape: (width,height) of array to return \n    Returns numpy array, 1 - mask, 0 - background\n\n    '''\n    s = mask_rle.split()\n    starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n    starts -= 1\n    ends = starts + lengths\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)\n    for lo, hi in zip(starts, ends):\n        img[lo:hi] = 1\n    return img.reshape(shape).T","metadata":{"execution":{"iopub.status.busy":"2022-07-09T18:25:15.794934Z","iopub.execute_input":"2022-07-09T18:25:15.795396Z","iopub.status.idle":"2022-07-09T18:25:15.807959Z","shell.execute_reply.started":"2022-07-09T18:25:15.795363Z","shell.execute_reply":"2022-07-09T18:25:15.806803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def basic_data_augmentations(table_name, color_space, title):\n    \n    # Initialize W&B\n    run = wandb.init(\n        project='hubmap-data-augmentation', \n        config= WANDB_CONFIG\n    )\n    \n    wandb.run.name = title\n\n    table = wandb.Table(\n        columns=[\n            'Id', \n            'Image', \n            'Image_'+ title, \n            'Mask', \n            'Mask_' + title,  \n            'Image_with_Mask', \n            'Image_with_Mask_' + title\n        ], \n        allow_mixed_types = True\n    )\n\n    for id, img in tqdm(zip(image_ids, image_files), total = len(image_ids)):\n\n        img = tiff.imread(img)\n        mask = rle2mask(df[df[\"id\"]==id][\"rle\"].iloc[-1], (img.shape[1], img.shape[0]))\n        \n        plt.figure(figsize=(10,10))\n        plt.axis(\"off\")\n        plt.imshow(img)\n        plt.imshow(mask, cmap='coolwarm', alpha=0.5)\n        plt.savefig(\"./image.jpg\")\n        plt.close()\n        \n        img_with_mask = cv2.cvtColor(cv2.imread(\"./image.jpg\"), cv2.COLOR_BGR2RGB)\n\n        table.add_data(\n            id, \n            wandb.Image(img), \n            wandb.Image(cv2.cvtColor(img, color_space)),\n            wandb.Image(mask),\n            wandb.Image(cv2.cvtColor(cv2.cvtColor(mask, cv2.COLOR_RGBA2RGB), color_space)),\n            wandb.Image(img_with_mask),\n            wandb.Image(cv2.cvtColor(cv2.cvtColor(img_with_mask, cv2.COLOR_RGBA2RGB), color_space))\n        )\n\n    wandb.log({table_name : table})\n    \n    run.finish()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T18:25:17.850750Z","iopub.execute_input":"2022-07-09T18:25:17.851152Z","iopub.status.idle":"2022-07-09T18:25:17.865508Z","shell.execute_reply.started":"2022-07-09T18:25:17.851121Z","shell.execute_reply":"2022-07-09T18:25:17.864515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"1.1\"></a>\n## **<span style=\"color:orange;\">01. Black and White</span>** ","metadata":{}},{"cell_type":"code","source":"basic_data_augmentations(\n    table_name = \"Black and White Color Space\", \n    color_space = cv2.COLOR_RGB2GRAY, \n    title = \"B&W\"\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### [Interactive W&B Table $\\rightarrow$](https://wandb.ai/ishandutta/hubmap-data-augmentation?workspace=user-ishandutta)\n![Animation.gif](https://iili.io/wuL3qF.gif)","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1.2\"></a>\n## **<span style=\"color:orange;\">02. Ben Graham</span>** ","metadata":{}},{"cell_type":"code","source":"basic_data_augmentations(\n    table_name = \"Ben Graham Color Space\", \n    color_space = cv2.COLOR_RGB2GRAY, \n    title = \"BenGraham\"\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### [Interactive W&B Table $\\rightarrow$](https://wandb.ai/ishandutta/hubmap-data-augmentation?workspace=user-ishandutta)\n![Animation.gif](https://iili.io/wusgR4.gif)","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1.3\"></a>\n## **<span style=\"color:orange;\">03. Hue, Saturation, Brightness </span>**","metadata":{}},{"cell_type":"code","source":"basic_data_augmentations(\n    table_name = \"Hue, Saturation, Brightness\", \n    color_space = cv2.COLOR_RGB2HLS, \n    title = \"HLS\"\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### [Interactive W&B Table $\\rightarrow$](https://wandb.ai/ishandutta/hubmap-data-augmentation?workspace=user-ishandutta)\n![Animation.gif](https://iili.io/wusUJf.gif)","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1.4\"></a>\n## **<span style=\"color:orange;\">04. LUV Color Space</span>**","metadata":{}},{"cell_type":"code","source":"basic_data_augmentations(\n    table_name = \"LUV Color Space\", \n    color_space = cv2.COLOR_RGB2LUV, \n    title = \"LUV\"\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### [Interactive W&B Table $\\rightarrow$](https://wandb.ai/ishandutta/hubmap-data-augmentation?workspace=user-ishandutta)\n![Animation.gif](https://iili.io/wus4b2.gif)","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1.5\"></a>\n## **<span style=\"color:orange;\">05. Alpha Channel</span>**","metadata":{}},{"cell_type":"code","source":"basic_data_augmentations(\n    table_name = \"Alpha Channel\", \n    color_space = cv2.COLOR_RGB2RGBA, \n    title = \"RGBA\"\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### [Interactive W&B Table $\\rightarrow$](https://wandb.ai/ishandutta/hubmap-data-augmentation?workspace=user-ishandutta)\n![Animation.gif](https://iili.io/wusrOl.gif)","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1.6\"></a>\n## **<span style=\"color:orange;\">06. XYZ Color Space</span>**","metadata":{}},{"cell_type":"code","source":"basic_data_augmentations(\n    table_name = \"XYZ Color Space\", \n    color_space = cv2.COLOR_RGB2XYZ, \n    title = \"XYZ\"\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### [Interactive W&B Table $\\rightarrow$](https://wandb.ai/ishandutta/hubmap-data-augmentation?workspace=user-ishandutta)\n![Animation.gif](https://iili.io/wusPxS.gif)","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1.7\"></a>\n## **<span style=\"color:orange;\">07. Luma Chroma</span>**","metadata":{}},{"cell_type":"code","source":"basic_data_augmentations(\n    table_name = \"Luma Chroma\", \n    color_space = cv2.COLOR_RGB2YCrCb, \n    title = \"LumaChroma\"\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### [Interactive W&B Table $\\rightarrow$](https://wandb.ai/ishandutta/hubmap-data-augmentation?workspace=user-ishandutta)\n![Animation.gif](https://iili.io/wusiW7.gif)","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1.8\"></a>\n## **<span style=\"color:orange;\">08. CIE Lab</span>**","metadata":{}},{"cell_type":"code","source":"basic_data_augmentations(\n    table_name = \"CIE Lab\", \n    color_space = cv2.COLOR_RGB2Lab, \n    title = \"CIE Lab\"\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### [Interactive W&B Table $\\rightarrow$](https://wandb.ai/ishandutta/hubmap-data-augmentation?workspace=user-ishandutta)\n![Animation.gif](https://iili.io/wusQfe.gif)","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1.9\"></a>\n## **<span style=\"color:orange;\">09. YUV Color Space</span>**","metadata":{}},{"cell_type":"code","source":"basic_data_augmentations(\n    table_name = \"YUV Color Space\", \n    color_space = cv2.COLOR_RGB2YUV, \n    title = \"YUV Color Space\"\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### [Interactive W&B Table $\\rightarrow$](https://wandb.ai/ishandutta/hubmap-data-augmentation?workspace=user-ishandutta)\n![Animation.gif](https://iili.io/wusss9.gif)","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"<h1><center>More Plots coming soon!</center></h1>\n                                                      \n<center><img src = \"https://static.wixstatic.com/media/5f8fae_7581e21a24a1483085024f88b0949a9d~mv2.jpg/v1/fill/w_934,h_379,al_c,q_90/5f8fae_7581e21a24a1483085024f88b0949a9d~mv2.jpg\" width = \"750\" height = \"500\"/></center> ","metadata":{}},{"cell_type":"markdown","source":"<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h3 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='background:maroon; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>If you find this notebook useful, do give me an upvote, it helps to keep up my motivation. This notebook will be updated frequently so keep checking for furthur developments.</center></h3>","metadata":{}}]}