{"cells":[{"metadata":{"_uuid":"5f04ea8c450d440af7fc828784e4d6a6c05abf85"},"cell_type":"markdown","source":"# Failure Analysis\n\nIn order to improve a model it could be beneficiall to understand when model works and when does it fails. In this notebook we will do an analysis of a model prediction. We will take the submission from this kernel: https://www.kaggle.com/iafoss/unet34-submission-0-89-public-lb, and compare it to the ground truth, which is available for us due to the data leak. Ussualy you would rather do this analysis either on validation or on hold-out set."},{"metadata":{"trusted":true,"_uuid":"8ba800329845d2f285651c448fbfa21af263f9bf"},"cell_type":"code","source":"# First we will just define some auxiliary functions\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sn\n\n# set the thresholds for the evaluation metrics\nthresholds = np.arange(0.5,1,0.05)\n\n# ref: https://www.kaggle.com/paulorzp/run-length-encode-and-decode\ndef rle_decode(mask_rle, shape=(768, 768)):\n    '''\n    mask_rle: run-length as string formated (start length)\n    shape: (height,width) of array to return \n    Returns numpy array, 1 - mask, 0 - background\n\n    '''\n    s = mask_rle.split()\n    starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n    starts -= 1\n    ends = starts + lengths\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)\n    for lo, hi in zip(starts, ends):\n        img[lo:hi] = 1\n    return img.reshape(shape).T  # Needed to align to RLE direction\n\n# ref : https://www.kaggle.com/stkbailey/step-by-step-explanation-of-scoring-metric\ndef iou_at_thresholds(target_mask, pred_mask, thresholds=np.arange(0.5,1,0.05)):\n    '''Returns True if IoU is greater than the thresholds.'''\n    intersection = np.logical_and(target_mask, pred_mask)\n    union = np.logical_or(target_mask, pred_mask)\n    iou = np.sum(intersection > 0) / np.sum(union > 0)\n    return iou > thresholds\n\ndef calculate_average_precision(x):\n    '''Calculates the average precision over a range of thresholds for one observation (with a single class).'''\n    \n    target_masks, pred_masks = x[\"EncodedPixels\"], x[\"EncodedPixels_pred\"]\n    \n    if type(pred_masks[0]) is not str:\n        pred_masks = []\n    \n    if type(target_masks[0]) is not str:\n        target_masks = []\n        \n    if len(pred_masks) == 0 and len(target_masks) == 0:\n        return 1\n    \n    if len(pred_masks) == 0 or len(target_masks) == 0:\n        return 0\n    \n    pred_masks = [rle_decode(mask) for mask in pred_masks]\n    target_masks = [rle_decode(mask) for mask in target_masks]\n    \n    iou_tensor = np.zeros([len(thresholds), len(pred_masks), len(target_masks)])\n    \n    for i, p_mask in enumerate(pred_masks):\n        for j, t_mask in enumerate(target_masks):\n            iou_tensor[:, i, j] = iou_at_thresholds(t_mask, p_mask, thresholds)\n\n    TP = np.sum((np.sum(iou_tensor, axis=2) == 1), axis=1)\n    FP = np.sum((np.sum(iou_tensor, axis=2) == 0), axis=1)\n    FN = np.sum((np.sum(iou_tensor, axis=1) == 0), axis=1)\n\n    precision = 5*TP / (5*TP + 4*FN + FP)\n    return TP, FP, FN, precision","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"dddac5a6ba0efe5d8a32d15638c0103dfb9ddce5"},"cell_type":"markdown","source":"## Evaluation of solution"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"# Change these pathes to do other comparisons\nGROUND_TRUTH_FILE = \"../input/airbus-ship-detection/test_ship_segmentations.csv\"\nPREDICTION_FILE = \"../input/unet34-submission-0-89-public-lb/submission.csv\"\n#PREDICTION_FILE = \"../input/binary-classifier-submission/submission.csv\"\n\ndf_gt = pd.read_csv(GROUND_TRUTH_FILE)\ndf_pred = pd.read_csv(PREDICTION_FILE)\n\n# Combining all schips to one list\ndf_gt = pd.DataFrame(df_gt.groupby('ImageId')['EncodedPixels'].apply(list))\ndf_pred = pd.DataFrame(df_pred.groupby('ImageId')['EncodedPixels'].apply(list))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"672546ecfffcfe8dd1a8bcdb4dfa32074d157683"},"cell_type":"markdown","source":"Here we calculate F2 score, as well as True Positives(TP), False Positives (FP) and False Negatives (FN) according to the explanation of evaluation metric: https://www.kaggle.com/c/airbus-ship-detection#evaluation"},{"metadata":{"trusted":true,"_uuid":"1492d4a721df6f5eddd9d4b9c1ba5bd27b565bca"},"cell_type":"code","source":"df = df_pred.join(df_gt, lsuffix=\"_pred\")\nprecisions = df.apply(calculate_average_precision, axis=1)\n\ndf[\"TP\"] = precisions.apply(lambda x : x[0] if (type(x) is not int) else np.nan)\ndf[\"FP\"] = precisions.apply(lambda x : x[1] if (type(x) is not int) else np.nan)\ndf[\"FN\"] = precisions.apply(lambda x : x[2] if (type(x) is not int) else np.nan)\ndf[\"IoU\"] = precisions.apply(lambda x : x[3] if (type(x) is not int) else x)\ndf[\"IoU_mean\"] = df[\"IoU\"].apply(np.mean)\ndf[\"TP_mean\"] = df[\"TP\"].apply(np.mean)\ndf[\"FP_mean\"] = df[\"FP\"].apply(np.mean)\ndf[\"FN_mean\"] = df[\"FN\"].apply(np.mean)\n\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"22ccde6aedd3d4ad849f7ad2e3ebb45a7fd53004"},"cell_type":"markdown","source":"Now we will calculate the final score, on the leaderboard the score is ~0.892, it is a bit higher, probably due to the fact, that for the leaderboard only 15% of data are used."},{"metadata":{"trusted":true,"_uuid":"0b2aabec25e13c59fe0d45cf6fa459c06c450234"},"cell_type":"code","source":"print (\"Solution score:\", np.mean(df[\"IoU_mean\"]))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4283c5de07302a85de06c71bd09630afaf248392"},"cell_type":"markdown","source":"## Confusion matrix\n\nLet's also calculate confusion matrix. Here we count one ship as an instance. The confusion matrix here is the average over all thresholds."},{"metadata":{"trusted":true,"_uuid":"93e737bdda78b9e76c6d33f1858211c71e290cf6"},"cell_type":"code","source":"cm = [[np.sum(~pd.notna(df[\"TP_mean\"])), np.nansum(df[\"FP_mean\"])],\n      [np.nansum(df[\"FN_mean\"]), np.nansum(df[\"TP_mean\"])]]\n\ndf_cm = pd.DataFrame(cm, index = [i for i in [\"No Ship\", \"Ship\"]],\n                          columns = [i for i in  [\"No Ship\", \"Ship\"]])\nplt.figure(figsize = (10,7))\nax = sn.heatmap(df_cm, annot=True, cmap=\"YlGnBu\")\nax.set(xlabel='Prediction', ylabel='Ground Truth')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8058b535e7d1c8b15b2400567b6a2b23b39590eb"},"cell_type":"markdown","source":"## Intersection over Union\n\nNow we exctract the matrix with IoU for each image in every threshold and plot some statistic"},{"metadata":{"trusted":true,"_uuid":"4956c7496fe6624f74682a8ae11e7ccac128a888"},"cell_type":"code","source":"iou_on_thresholds_matrix = np.asarray([i for i in df[\"IoU\"] if not type(i) is int])\n\nf, ax = plt.subplots(2, 5, figsize=(30,10))\n\nfor i in range(2):\n    for j in range(5):\n        ax[i,j].hist(iou_on_thresholds_matrix[:,i*5 + j])\n        ax[i,j].set_title(\"Threshold: %.2f\" % thresholds[i*5 + j])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1104bb9b099507c6ddbd3839209bad05fb7918f4"},"cell_type":"markdown","source":"## Ship number and Pixel Occupation\n\nHere we will excract some additional information from the mask, like the number of ships per image and number of pixels occupied with ships per image"},{"metadata":{"trusted":true,"_uuid":"b3a3c52dacc4e2e2d053fe0b8576139145596689"},"cell_type":"code","source":"def rle_get_mask_size(masks, shape=(768, 768)):\n    res = 0\n    \n    for mask_rle in masks:\n        if type(mask_rle) is float:\n            return 0\n        s = mask_rle.split()\n        starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n        starts -= 1\n        ends = starts + lengths\n    \n        for lo, hi in zip(starts, ends):\n            res += hi - lo\n            \n    return res\n\ndf[\"pred_ship_num\"] = df[\"EncodedPixels_pred\"].apply(lambda x : len(x) if type(x[0]) is str else 0)\ndf[\"true_ship_num\"] = df[\"EncodedPixels\"].apply(lambda x : len(x) if type(x[0]) is str else 0)\ndf[\"pred_ship_occupation\"] = df[\"EncodedPixels_pred\"].apply(rle_get_mask_size)\ndf[\"true_ship_occupation\"] = df[\"EncodedPixels\"].apply(rle_get_mask_size)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f23765604804a2e2f9af3a4899e12dba09df071b"},"cell_type":"markdown","source":"Let's plot the real versus predicted occupation of the image."},{"metadata":{"trusted":true,"_uuid":"98d962855adfae45a5b2643920a4f0485a7f7e14"},"cell_type":"code","source":"fig, (ax1, ax2) = plt.subplots(ncols=2, figsize=(14, 7))\nfig.subplots_adjust(hspace=0.5, left=0.07, right=0.93)\nhb = ax1.hexbin(df['pred_ship_occupation'], df['true_ship_occupation'], gridsize=50, cmap='BuGn', bins='log')\nxmin, xmax = 0, 20000\nax1.plot([xmin, xmax], [xmin, xmax])\nax1.axis([xmin, xmax, xmin, xmax])\n\nax1.set_ylabel('Prediction')\nax1.set_xlabel('Ground Truth')\nax1.set_title(\"Pixel Occupation\")\n\nax2.hist(df['pred_ship_occupation'] - df['true_ship_occupation'],  log=True, bins=40)\nax2.axis([-30000, 30000, 1, 100000])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b49e727e9c6c533295ac3af31ba4768ee75d975b"},"cell_type":"markdown","source":"Here is the plot that shows the correlation beetween how many ships are present in an image and average IoU that we got. It actually make sense that we got the highest IoU when only one ship is present. We also see that it is the most popular class, so it actually makes sense to optimize the model for the case with only few ships."},{"metadata":{"trusted":true,"_uuid":"c915d2d4b341574b74845dd2daa63aea2aa2f9aa"},"cell_type":"code","source":"fig, (ax) = plt.subplots(ncols=1, figsize=(14, 7))\n\nmeans = df.groupby(\"true_ship_num\")[\"IoU_mean\"].mean()\nvars = df.groupby(\"true_ship_num\")[\"IoU_mean\"].var()\ncount = df.groupby(\"true_ship_num\")[\"IoU_mean\"].count()\n\nax.bar(np.arange(1,16), count[1:]/10000.)\nax.errorbar(np.arange(1,16), means[1:], yerr=vars[1:,], fmt='-or')\nax.legend(labels=(\"Class count\", \"IoU\"))\nax.set_xlabel(\"Number of ships\")\nax.set_ylabel(\"IoU / Class count\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2093664f7f2e4eb2180f80146b4c92f8043e5362"},"cell_type":"markdown","source":"## Visualization of images\n\nHere we will plot some of FN, FP and TP images to better understand when it works and when not."},{"metadata":{"trusted":true,"_uuid":"52ac8274e8f79d3d3487d4291a0707c609cc7863"},"cell_type":"code","source":"from skimage.io import imread\nfrom skimage.util import montage\n\nmontage_rgb = lambda x: np.stack([montage(x[:, :, :, i]) for i in range(x.shape[3])], -1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ff5c9fe93e2ff579d817e5faeeb764947581108b"},"cell_type":"code","source":"img_names = df[df[\"FN_mean\"] > 10].index[:49].values.tolist()\nimg = np.stack([imread(\"../input/airbus-ship-detection/test/\" + img_name) for img_name in img_names])\n\nfig, (ax1) = plt.subplots(1, 1, figsize = (40, 20))\nax1.imshow(montage_rgb(img))\nax1.set_title('False negatives')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6fe0dee4f71ce32b4689a326e9f9b4f50dcc22ad"},"cell_type":"markdown","source":"Now we will plot some of the images with high False Positive rate."},{"metadata":{"trusted":true,"_uuid":"84bc25a1016af041c922124465fe78e4b4f83759"},"cell_type":"code","source":"img_names = df[df[\"FP_mean\"] > 10].index[:49].values.tolist()\nimg = np.stack([imread(\"../input/airbus-ship-detection/test/\" + img_name) for img_name in img_names])\n\nfig, (ax1) = plt.subplots(1, 1, figsize = (40, 20))\nax1.imshow(montage_rgb(img))\nax1.set_title('False positives')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"632bd48a625f87fe51cb27f665ee7f3ed25fa178"},"cell_type":"markdown","source":"And at the end, let's plot some images where we got high IoU."},{"metadata":{"trusted":true,"_uuid":"124862a974bff01bfe0a0de5c06d1863c5b164d3"},"cell_type":"code","source":"img_names = df[df[\"TP_mean\"] > 0.9].index[:49].values.tolist()\nimg = np.stack([imread(\"../input/airbus-ship-detection/test/\" + img_name) for img_name in img_names])\n\nfig, (ax1) = plt.subplots(1, 1, figsize = (40, 20))\nax1.imshow(montage_rgb(img))\nax1.set_title('True positives')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cb8be125071fbae07f1041ea385f5cc475e845ee"},"cell_type":"markdown","source":"## Understanding False Positives and False Negatives\n\nLet's start with false positives. In order to understand why we have high false positive rate even if it seems, that ships are present in the image, let's plot the segmentation prediction and ground truth. From the image below it is seen, that models predicts a lot of small objects that are not in ground truth."},{"metadata":{"trusted":true,"_uuid":"1734bfd531bd488987d371ef20a8610cf9262f45"},"cell_type":"code","source":"img_name = df[df[\"FP_mean\"] > 10].index[1]\nimg = imread(\"../input/airbus-ship-detection/test/\" + img_name)\n\nfig, (ax1, ax2, ax3) = plt.subplots(1, 3, figsize = (22, 12))\nax1.imshow(img)      \nax1.set_title(\"Original Image\")\n\nmask = np.sum([rle_decode(df.loc[img_name][\"EncodedPixels\"][i]) for i in range(len(df.loc[img_name][\"EncodedPixels\"]))], axis=0)\nax2.imshow(mask)\nax2.set_title(\"Ground gruth\")\n\nmask = np.sum([rle_decode(df.loc[img_name][\"EncodedPixels_pred\"][i]) for i in range(len(df.loc[img_name][\"EncodedPixels_pred\"]))], axis=0)\nax3.imshow(mask)\nax3.set_title(\"Prediction\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8a0f2a068ae0758e91d0d3324988f70b08a7cb67"},"cell_type":"markdown","source":"Let's also plot segmentation masks for one of the false negative example. Here it is clearly seen that objects are very close to each other and model merge them into one."},{"metadata":{"trusted":true,"_uuid":"83dfd489f8b963d647bc36a73e054fe056808508"},"cell_type":"code","source":"img_name = df[df[\"FN_mean\"] > 10].index[:49][5]\n\nimg = imread(\"../input/airbus-ship-detection/test/\" + img_name)\n\nfig, (ax1, ax2, ax3) = plt.subplots(1, 3, figsize = (22, 12))\nax1.imshow(img)      \nax1.set_title(\"Original Image\")\n\nmask = np.sum([rle_decode(df.loc[img_name][\"EncodedPixels\"][i]) for i in range(len(df.loc[img_name][\"EncodedPixels\"]))], axis=0)\nax2.imshow(mask)\nax2.set_title(\"Ground gruth\")\n\nmask = np.sum([rle_decode(df.loc[img_name][\"EncodedPixels_pred\"][i]) for i in range(len(df.loc[img_name][\"EncodedPixels_pred\"]))], axis=0)\nax3.imshow(mask)\nax3.set_title(\"Prediction\")","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}