{
  "id": 112807,
  "title": "Top 9th Solution: Simple but complete approach.",
  "url": "/competitions/kuzushiji-recognition/writeups/ollie-nanashi-and-tom-top-9th-solution-simple-but-",
  "author_name": "",
  "post_date": "2019-10-19T20:48:47.887Z",
  "votes": 25,
  "comment_count": 9,
  "views": 0,
  "content": "<blockquote>\n  <p><strong>GitHub:</strong> <a href=\"https://github.com/mv-lab/kuzushiji-recognition\">https://github.com/mv-lab/kuzushiji-recognition</a></p>\n</blockquote>\n\n<p>First of all I'd like to thank my mates <a href=\"/ollieperree\">@ollieperree</a> and <a href=\"/tikutiku\">@tikutiku</a>, we couldn't work a lot of time together but the experience was amazing as the result!</p>\n\n<p>Also congrats to the winners <a href=\"/tascj0\">@tascj0</a> <a href=\"/lopuhin\">@lopuhin</a> <a href=\"/knjcode\">@knjcode</a> <a href=\"/linhui\">@linhui</a> <a href=\"/seesee\">@seesee</a> \n<a href=\"/kmat2019\">@kmat2019</a> your kernel is amazing!</p>\n\n<p>And obviously the organization! <a href=\"/tkasasagi\">@tkasasagi</a>  <a href=\"/sohier\">@sohier</a> and <a href=\"/anokas\">@anokas</a> this competition was perfect! I think kaggle should reconsider medals.\nWe don't receive medals, and we are only 293 teams... I think the people here competed because of the big prize and the challenge itself, not medals :) In our case, the biggest impediment was the GPU quota...</p>\n\n<p>Anyway, I'll explain our approach. Please check the notebook: <strong><a href=\"https://www.kaggle.com/jesucristo/kuzushiji-recognition-starter\">Kuzushiji Recognition Starter</a></strong> where we are going to update the inference and the public dataset with 256x256 images for training classifier.</p>\n\n<hr>\n\n<p>From the beginning <a href=\"/ollieperree\">@ollieperree</a> was using a <strong>2-stage approach.</strong> \nOur approach to detection was directly inspired by <a href=\"https://www.kaggle.com/kmat2019/centernet-keypoint-detector\">K_mat's kernel</a>, with the main takeaway being the idea of predicting a heatmap showing the centers of characters. Initially, we used a U-Net with a resnet18 backbone to predict a heatmap consisting of ellipses placed at the centers of characters, with the radii proportional to the width and height of the bounding box, with the input to the model being a 1024x1024 pixel crop of the page resized to 256x256 pixels. \nPredictions for the centers were then obtained by picking the local maxima (note that the width and height of the bounding box were not predicted). Performance was improved by changing the ellipses to circles of constant radius.</p>\n\n<p>![](000_002.png?generation=1571161817145901&amp;alt=media\"&gt;https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2Fa81b20efc933fcdfd877b0636e229da6%2Fumgy012-042_000_002.png?generation=1571161817145901&amp;alt=media =200x*)</p>\n\n<p>We tried using <em>focal loss</em> and binary <em>cross-entropy</em> as loss functions, but using mean squared error resulted in the cleanest predictions for us (though more epochs were needed to get sensible-looking predictions).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F97924b78f0393a6e162cb477ba51c5e1%2Fimage(1\" alt=\"\">.png?generation=1571161817778557&amp;alt=media =500x*)</p>\n\n<p>One issue with using 1024x1024 crops of the page as the input were <strong>\"artifacts\"</strong> around the edges of the input. We tried a few things to try to counteract this, such as moving the sliding window over the page with a stride less than 1024x1024, then removing duplicate predictions by detecting when two predicted points of the same class were within a certain distance of each other. However, these did not give an improvement on the LB - we think that tuning parameters for these methods on the validation set, as well as the parameters for selecting maxima in the heatmap, might have caused us to \"overfit\".</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F2cd1d021e43de8c00e5f641c6e91586c%2FScreenshot%20from%202019-10-15%2019-53-12.png?generation=1571162065651936&amp;alt=media\" alt=\"\"></p>\n\n<p>These artifacts were related with the drawings and annotations!\n(See carefully the <em>red</em> dots at the images)\nHow did we fix this? Check <strong>ensemble</strong> and <strong>pseudolabels</strong></p>\n\n<hr>\n\n<p>We have a <strong>2 stage</strong> model: detection and classification.</p>\n\n<h1>Detection</h1>\n\n<p>We used as starter code the great kernel: <a href=\"https://www.kaggle.com/kmat2019/centernet-keypoint-detector\">CenterNet -Keypoint Detector-</a> by <a href=\"/kmat2019\">@kmat2019</a>\nThen I realized that <a href=\"/seesee\">@seesee</a> had his own <a href=\"https://github.com/see--/keras-centernet\">keras-centernet</a>.\nAt the end we used Hourglass and the output are boxes instead of only the centers (like the original paper).</p>\n\n<p><strong>Model</strong>\n- Detection by hourglassnet\n- Output heatmaps + width/height maps\n- generate_heatmap_crops_circular(crop1024,resize256)\n- Validation: w/o outliers GroupKFold\n- resnet34\n- MSELoss, dice and IOU (oss=0.00179, dice=0.6270, F1=0.9856, iou=0.8142)\n- Augmentations: aug(randombrightness0.2,scale0.5)\n- Learning rate: (1e-4:1e-6)\n- 20epochs</p>\n\n<p>![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F4d84f42553ac59551b480d883ad8f33b%2Fvalid_pred.png?generation=1571146402846083&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F4d84f42553ac59551b480d883ad8f33b%2Fvalid_pred.png?generation=1571146402846083&amp;alt=media</a> =300x*)</p>\n\n<p>![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F9c032decbf330957fd2b25c6a34a73bd%2FScreenshot%20from%202019-10-15%2011-05-12.png?generation=1571146515196023&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F9c032decbf330957fd2b25c6a34a73bd%2FScreenshot%20from%202019-10-15%2011-05-12.png?generation=1571146515196023&amp;alt=media</a> =500x*)</p>\n\n<h1>Classification</h1>\n\n<p>The classification model was a <strong>resnet18</strong>, pretrained on ImageNet, with the input being a fixed <em>256x256</em> pixel area, scaled down to <em>128x128</em>, centered at the (predicted) center of the character. \nThe training data included a <em>background class</em>, whose training examples were random 256x256 crops of the pages with no labelled characters. \nTraining was done using the <strong>fastai</strong> library, with <a href=\"https://docs.fast.ai/vision.transform.html#get_transforms\">standard fastai transforms</a> and MixUp. \nThis model achieved a Classification accuracy of <strong>94.6%</strong> on a validation set (20% of the train data, group split by book).</p>\n\n<h1>Augmentations</h1>\n\n<p>Like <a href=\"/lopuhin\">@lopuhin</a>  said: <em>I'm using standard augmentations from <a href=\"https://github.com/albu/albumentations/\">https://github.com/albu/albumentations/</a> library, including adjusting colors that helps simulate different paper styles</em></p>\n\n<p>&gt; Can you tell me wich one is the real one?</p>\n\n<p>![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F89c2cd2e945352655b1932b8d6ce1378%2Fimage.png?generation=1571146446903717&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F89c2cd2e945352655b1932b8d6ce1378%2Fimage.png?generation=1571146446903717&amp;alt=media</a> =500x*)</p>\n\n<p><strong>Code</strong></p>\n\n<p>```\nimport albumentations</p>\n\n<p>colorize = albumentations.RGBShift(r_shift_limit=0, g_shift_limit=0, b_shift_limit=[-80,0])</p>\n\n<p>def color_get_params():\n    a = random.uniform(-40, 0)\n    b = random.uniform(-80, -30)\n    return {\"r_shift\": a,\n            \"g_shift\": a,\n            \"b_shift\": b}</p>\n\n<p>colorize.get_params = color_get_params</p>\n\n<p>aug = albumentations.Compose([albumentations.RandomBrightnessContrast(contrast_limit=0.2, brightness_limit=0.2),\n                              albumentations.ToGray(),\n                              albumentations.Blur(),\n                              albumentations.Rotate(limit=5),\n                              colorize\n                             ])\n```</p>\n\n<h1>Ensemble</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F3d45fceb5c94db4071d6ef4e53990e1b%2Fensemble.jpg?generation=1571162716740190&amp;alt=media\" alt=\"\"></p>\n\n<p>We had a problem... our centers weren't ordered! so in order to improve the accuracy and delete <em>false positives</em> we thought in the following ensemble method.\nFor the image IMG we take the most external centers from 3 predictions: \n<code>\n(xmin, ymin), (xmin, ymax), (xmax,ymin), (xmax,ymax)\n</code> \nAt the picture this boxes are represented by 3 different colours (yellow, blue, red). \nFinally, we take the <strong>intersection</strong> of those 3 boxes, the black rectangle defined as (X,Y,Z,W), and we drop all the centers out of the black box! with this technique we could eliminate artifacts like predictions at the edges.\nI have to say that the idea is cool, but I found 2 bugs and this technique didn't work properly :(</p>\n\n<h1>Pseudolabels</h1>\n\n<p>We realized that train had only <strong>276</strong> with no-chars and we thought that was the reason why our model had so many problems with drawings. \nThe solution? we set the as threshold= 5 chars. All the predictions with &lt;=5 chars where used as pseudolabels with labels= <em>NaN</em>.</p>\n\n<h1>Bonus</h1>\n\n<p>A samurai with 3 swords... that emblem... this reminds me to something :)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2Fd7568f09270098f7a80994266db642e0%2Fkozuki.jpg?generation=1571164389635754&amp;alt=media\" alt=\"\"></p>\n\n<p>![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F7989d65dc8ad606d2cfa51dc08fd87d5%2FFamiglia_Kozuki_simbolo.webp?generation=1571164245190855&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F7989d65dc8ad606d2cfa51dc08fd87d5%2FFamiglia_Kozuki_simbolo.webp?generation=1571164245190855&amp;alt=media</a> =100x*)</p>",
  "messages": [
    {
      "id": "649533",
      "postDate": "10/15/2019 13:57:19",
      "content": "<blockquote>\n  <p><strong>GitHub:</strong> <a href=\"https://github.com/mv-lab/kuzushiji-recognition\">https://github.com/mv-lab/kuzushiji-recognition</a></p>\n</blockquote>\n\n<p>First of all I'd like to thank my mates <a href=\"/ollieperree\">@ollieperree</a> and <a href=\"/tikutiku\">@tikutiku</a>, we couldn't work a lot of time together but the experience was amazing as the result!</p>\n\n<p>Also congrats to the winners <a href=\"/tascj0\">@tascj0</a> <a href=\"/lopuhin\">@lopuhin</a> <a href=\"/knjcode\">@knjcode</a> <a href=\"/linhui\">@linhui</a> <a href=\"/seesee\">@seesee</a> \n<a href=\"/kmat2019\">@kmat2019</a> your kernel is amazing!</p>\n\n<p>And obviously the organization! <a href=\"/tkasasagi\">@tkasasagi</a>  <a href=\"/sohier\">@sohier</a> and <a href=\"/anokas\">@anokas</a> this competition was perfect! I think kaggle should reconsider medals.\nWe don't receive medals, and we are only 293 teams... I think the people here competed because of the big prize and the challenge itself, not medals :) In our case, the biggest impediment was the GPU quota...</p>\n\n<p>Anyway, I'll explain our approach. Please check the notebook: <strong><a href=\"https://www.kaggle.com/jesucristo/kuzushiji-recognition-starter\">Kuzushiji Recognition Starter</a></strong> where we are going to update the inference and the public dataset with 256x256 images for training classifier.</p>\n\n<hr>\n\n<p>From the beginning <a href=\"/ollieperree\">@ollieperree</a> was using a <strong>2-stage approach.</strong> \nOur approach to detection was directly inspired by <a href=\"https://www.kaggle.com/kmat2019/centernet-keypoint-detector\">K_mat's kernel</a>, with the main takeaway being the idea of predicting a heatmap showing the centers of characters. Initially, we used a U-Net with a resnet18 backbone to predict a heatmap consisting of ellipses placed at the centers of characters, with the radii proportional to the width and height of the bounding box, with the input to the model being a 1024x1024 pixel crop of the page resized to 256x256 pixels. \nPredictions for the centers were then obtained by picking the local maxima (note that the width and height of the bounding box were not predicted). Performance was improved by changing the ellipses to circles of constant radius.</p>\n\n<p>![](000_002.png?generation=1571161817145901&amp;alt=media\"&gt;https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2Fa81b20efc933fcdfd877b0636e229da6%2Fumgy012-042_000_002.png?generation=1571161817145901&amp;alt=media =200x*)</p>\n\n<p>We tried using <em>focal loss</em> and binary <em>cross-entropy</em> as loss functions, but using mean squared error resulted in the cleanest predictions for us (though more epochs were needed to get sensible-looking predictions).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F97924b78f0393a6e162cb477ba51c5e1%2Fimage(1\" alt=\"\">.png?generation=1571161817778557&amp;alt=media =500x*)</p>\n\n<p>One issue with using 1024x1024 crops of the page as the input were <strong>\"artifacts\"</strong> around the edges of the input. We tried a few things to try to counteract this, such as moving the sliding window over the page with a stride less than 1024x1024, then removing duplicate predictions by detecting when two predicted points of the same class were within a certain distance of each other. However, these did not give an improvement on the LB - we think that tuning parameters for these methods on the validation set, as well as the parameters for selecting maxima in the heatmap, might have caused us to \"overfit\".</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F2cd1d021e43de8c00e5f641c6e91586c%2FScreenshot%20from%202019-10-15%2019-53-12.png?generation=1571162065651936&amp;alt=media\" alt=\"\"></p>\n\n<p>These artifacts were related with the drawings and annotations!\n(See carefully the <em>red</em> dots at the images)\nHow did we fix this? Check <strong>ensemble</strong> and <strong>pseudolabels</strong></p>\n\n<hr>\n\n<p>We have a <strong>2 stage</strong> model: detection and classification.</p>\n\n<h1>Detection</h1>\n\n<p>We used as starter code the great kernel: <a href=\"https://www.kaggle.com/kmat2019/centernet-keypoint-detector\">CenterNet -Keypoint Detector-</a> by <a href=\"/kmat2019\">@kmat2019</a>\nThen I realized that <a href=\"/seesee\">@seesee</a> had his own <a href=\"https://github.com/see--/keras-centernet\">keras-centernet</a>.\nAt the end we used Hourglass and the output are boxes instead of only the centers (like the original paper).</p>\n\n<p><strong>Model</strong>\n- Detection by hourglassnet\n- Output heatmaps + width/height maps\n- generate_heatmap_crops_circular(crop1024,resize256)\n- Validation: w/o outliers GroupKFold\n- resnet34\n- MSELoss, dice and IOU (oss=0.00179, dice=0.6270, F1=0.9856, iou=0.8142)\n- Augmentations: aug(randombrightness0.2,scale0.5)\n- Learning rate: (1e-4:1e-6)\n- 20epochs</p>\n\n<p>![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F4d84f42553ac59551b480d883ad8f33b%2Fvalid_pred.png?generation=1571146402846083&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F4d84f42553ac59551b480d883ad8f33b%2Fvalid_pred.png?generation=1571146402846083&amp;alt=media</a> =300x*)</p>\n\n<p>![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F9c032decbf330957fd2b25c6a34a73bd%2FScreenshot%20from%202019-10-15%2011-05-12.png?generation=1571146515196023&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F9c032decbf330957fd2b25c6a34a73bd%2FScreenshot%20from%202019-10-15%2011-05-12.png?generation=1571146515196023&amp;alt=media</a> =500x*)</p>\n\n<h1>Classification</h1>\n\n<p>The classification model was a <strong>resnet18</strong>, pretrained on ImageNet, with the input being a fixed <em>256x256</em> pixel area, scaled down to <em>128x128</em>, centered at the (predicted) center of the character. \nThe training data included a <em>background class</em>, whose training examples were random 256x256 crops of the pages with no labelled characters. \nTraining was done using the <strong>fastai</strong> library, with <a href=\"https://docs.fast.ai/vision.transform.html#get_transforms\">standard fastai transforms</a> and MixUp. \nThis model achieved a Classification accuracy of <strong>94.6%</strong> on a validation set (20% of the train data, group split by book).</p>\n\n<h1>Augmentations</h1>\n\n<p>Like <a href=\"/lopuhin\">@lopuhin</a>  said: <em>I'm using standard augmentations from <a href=\"https://github.com/albu/albumentations/\">https://github.com/albu/albumentations/</a> library, including adjusting colors that helps simulate different paper styles</em></p>\n\n<p>&gt; Can you tell me wich one is the real one?</p>\n\n<p>![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F89c2cd2e945352655b1932b8d6ce1378%2Fimage.png?generation=1571146446903717&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F89c2cd2e945352655b1932b8d6ce1378%2Fimage.png?generation=1571146446903717&amp;alt=media</a> =500x*)</p>\n\n<p><strong>Code</strong></p>\n\n<p>```\nimport albumentations</p>\n\n<p>colorize = albumentations.RGBShift(r_shift_limit=0, g_shift_limit=0, b_shift_limit=[-80,0])</p>\n\n<p>def color_get_params():\n    a = random.uniform(-40, 0)\n    b = random.uniform(-80, -30)\n    return {\"r_shift\": a,\n            \"g_shift\": a,\n            \"b_shift\": b}</p>\n\n<p>colorize.get_params = color_get_params</p>\n\n<p>aug = albumentations.Compose([albumentations.RandomBrightnessContrast(contrast_limit=0.2, brightness_limit=0.2),\n                              albumentations.ToGray(),\n                              albumentations.Blur(),\n                              albumentations.Rotate(limit=5),\n                              colorize\n                             ])\n```</p>\n\n<h1>Ensemble</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F3d45fceb5c94db4071d6ef4e53990e1b%2Fensemble.jpg?generation=1571162716740190&amp;alt=media\" alt=\"\"></p>\n\n<p>We had a problem... our centers weren't ordered! so in order to improve the accuracy and delete <em>false positives</em> we thought in the following ensemble method.\nFor the image IMG we take the most external centers from 3 predictions: \n<code>\n(xmin, ymin), (xmin, ymax), (xmax,ymin), (xmax,ymax)\n</code> \nAt the picture this boxes are represented by 3 different colours (yellow, blue, red). \nFinally, we take the <strong>intersection</strong> of those 3 boxes, the black rectangle defined as (X,Y,Z,W), and we drop all the centers out of the black box! with this technique we could eliminate artifacts like predictions at the edges.\nI have to say that the idea is cool, but I found 2 bugs and this technique didn't work properly :(</p>\n\n<h1>Pseudolabels</h1>\n\n<p>We realized that train had only <strong>276</strong> with no-chars and we thought that was the reason why our model had so many problems with drawings. \nThe solution? we set the as threshold= 5 chars. All the predictions with &lt;=5 chars where used as pseudolabels with labels= <em>NaN</em>.</p>\n\n<h1>Bonus</h1>\n\n<p>A samurai with 3 swords... that emblem... this reminds me to something :)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2Fd7568f09270098f7a80994266db642e0%2Fkozuki.jpg?generation=1571164389635754&amp;alt=media\" alt=\"\"></p>\n\n<p>![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F7989d65dc8ad606d2cfa51dc08fd87d5%2FFamiglia_Kozuki_simbolo.webp?generation=1571164245190855&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F7989d65dc8ad606d2cfa51dc08fd87d5%2FFamiglia_Kozuki_simbolo.webp?generation=1571164245190855&amp;alt=media</a> =100x*)</p>",
      "rawMarkdown": "&gt; **GitHub:** https://github.com/mv-lab/kuzushiji-recognition\n\nFirst of all I'd like to thank my mates @ollieperree and @tikutiku, we couldn't work a lot of time together but the experience was amazing as the result!\n\nAlso congrats to the winners @tascj0 @lopuhin @knjcode @linhui @seesee \n@kmat2019 your kernel is amazing!\n\nAnd obviously the organization! @tkasasagi  @sohier and @anokas this competition was perfect! I think kaggle should reconsider medals.\nWe don't receive medals, and we are only 293 teams... I think the people here competed because of the big prize and the challenge itself, not medals :) In our case, the biggest impediment was the GPU quota...\n\nAnyway, I'll explain our approach. Please check the notebook: **[Kuzushiji Recognition Starter](https://www.kaggle.com/jesucristo/kuzushiji-recognition-starter)** where we are going to update the inference and the public dataset with 256x256 images for training classifier.\n\n---\nFrom the beginning @ollieperree was using a **2-stage approach.** \nOur approach to detection was directly inspired by [K_mat's kernel](https://www.kaggle.com/kmat2019/centernet-keypoint-detector), with the main takeaway being the idea of predicting a heatmap showing the centers of characters. Initially, we used a U-Net with a resnet18 backbone to predict a heatmap consisting of ellipses placed at the centers of characters, with the radii proportional to the width and height of the bounding box, with the input to the model being a 1024x1024 pixel crop of the page resized to 256x256 pixels. \nPredictions for the centers were then obtained by picking the local maxima (note that the width and height of the bounding box were not predicted). Performance was improved by changing the ellipses to circles of constant radius.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2Fa81b20efc933fcdfd877b0636e229da6%2Fumgy012-042___000_002.png?generation=1571161817145901&amp;alt=media =200x*)\n\n\nWe tried using *focal loss* and binary *cross-entropy* as loss functions, but using mean squared error resulted in the cleanest predictions for us (though more epochs were needed to get sensible-looking predictions).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F97924b78f0393a6e162cb477ba51c5e1%2Fimage(1).png?generation=1571161817778557&amp;alt=media =500x*)\n\nOne issue with using 1024x1024 crops of the page as the input were **\"artifacts\"** around the edges of the input. We tried a few things to try to counteract this, such as moving the sliding window over the page with a stride less than 1024x1024, then removing duplicate predictions by detecting when two predicted points of the same class were within a certain distance of each other. However, these did not give an improvement on the LB - we think that tuning parameters for these methods on the validation set, as well as the parameters for selecting maxima in the heatmap, might have caused us to \"overfit\".\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F2cd1d021e43de8c00e5f641c6e91586c%2FScreenshot%20from%202019-10-15%2019-53-12.png?generation=1571162065651936&amp;alt=media)\n\nThese artifacts were related with the drawings and annotations!\n(See carefully the *red* dots at the images)\nHow did we fix this? Check **ensemble** and **pseudolabels**\n\n---\n\nWe have a **2 stage** model: detection and classification.\n\n# Detection\n\nWe used as starter code the great kernel: [CenterNet -Keypoint Detector-](https://www.kaggle.com/kmat2019/centernet-keypoint-detector) by @kmat2019\nThen I realized that @seesee had his own [keras-centernet](https://github.com/see--/keras-centernet).\nAt the end we used Hourglass and the output are boxes instead of only the centers (like the original paper).\n\n**Model**\n- Detection by hourglassnet\n- Output heatmaps + width/height maps\n- generate_heatmap_crops_circular(crop1024,resize256)\n- Validation: w/o outliers GroupKFold\n- resnet34\n- MSELoss, dice and IOU (oss=0.00179, dice=0.6270, F1=0.9856, iou=0.8142)\n- Augmentations: aug(randombrightness0.2,scale0.5)\n- Learning rate: (1e-4:1e-6)\n- 20epochs\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F4d84f42553ac59551b480d883ad8f33b%2Fvalid_pred.png?generation=1571146402846083&amp;alt=media =300x*)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F9c032decbf330957fd2b25c6a34a73bd%2FScreenshot%20from%202019-10-15%2011-05-12.png?generation=1571146515196023&amp;alt=media =500x*)\n\n\n# Classification\n\nThe classification model was a **resnet18**, pretrained on ImageNet, with the input being a fixed *256x256* pixel area, scaled down to *128x128*, centered at the (predicted) center of the character. \nThe training data included a *background class*, whose training examples were random 256x256 crops of the pages with no labelled characters. \nTraining was done using the **fastai** library, with [standard fastai transforms](https://docs.fast.ai/vision.transform.html#get_transforms) and MixUp. \nThis model achieved a Classification accuracy of **94.6%** on a validation set (20% of the train data, group split by book).\n\n# Augmentations\n\nLike @lopuhin  said: *I'm using standard augmentations from https://github.com/albu/albumentations/ library, including adjusting colors that helps simulate different paper styles*\n\n&gt; Can you tell me wich one is the real one?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F89c2cd2e945352655b1932b8d6ce1378%2Fimage.png?generation=1571146446903717&amp;alt=media =500x*)\n\n**Code**\n\n```\nimport albumentations\n\ncolorize = albumentations.RGBShift(r_shift_limit=0, g_shift_limit=0, b_shift_limit=[-80,0])\n\ndef color_get_params():\n    a = random.uniform(-40, 0)\n    b = random.uniform(-80, -30)\n    return {\"r_shift\": a,\n            \"g_shift\": a,\n            \"b_shift\": b}\n\ncolorize.get_params = color_get_params\n\naug = albumentations.Compose([albumentations.RandomBrightnessContrast(contrast_limit=0.2, brightness_limit=0.2),\n                              albumentations.ToGray(),\n                              albumentations.Blur(),\n                              albumentations.Rotate(limit=5),\n                              colorize\n                             ])\n```\n\n# Ensemble\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F3d45fceb5c94db4071d6ef4e53990e1b%2Fensemble.jpg?generation=1571162716740190&amp;alt=media)\n\nWe had a problem... our centers weren't ordered! so in order to improve the accuracy and delete *false positives* we thought in the following ensemble method.\nFor the image IMG we take the most external centers from 3 predictions: \n```\n(xmin, ymin), (xmin, ymax), (xmax,ymin), (xmax,ymax)\n``` \nAt the picture this boxes are represented by 3 different colours (yellow, blue, red). \nFinally, we take the **intersection** of those 3 boxes, the black rectangle defined as (X,Y,Z,W), and we drop all the centers out of the black box! with this technique we could eliminate artifacts like predictions at the edges.\nI have to say that the idea is cool, but I found 2 bugs and this technique didn't work properly :(\n\n# Pseudolabels\n\nWe realized that train had only **276** with no-chars and we thought that was the reason why our model had so many problems with drawings. \nThe solution? we set the as threshold= 5 chars. All the predictions with &lt;=5 chars where used as pseudolabels with labels= *NaN*.\n\n# Bonus\n\nA samurai with 3 swords... that emblem... this reminds me to something :)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2Fd7568f09270098f7a80994266db642e0%2Fkozuki.jpg?generation=1571164389635754&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F7989d65dc8ad606d2cfa51dc08fd87d5%2FFamiglia_Kozuki_simbolo.webp?generation=1571164245190855&amp;alt=media =100x*)",
      "votes": null
    },
    {
      "id": "649592",
      "postDate": "10/15/2019 14:56:04",
      "content": "<p>Congratulations\nGreat Write-Up\nThanks for Sharing your Approach &amp; Insights.. <a href=\"/jesucristo\">@jesucristo</a> </p>",
      "rawMarkdown": "Congratulations\nGreat Write-Up\nThanks for Sharing your Approach &amp; Insights.. @jesucristo",
      "votes": null
    },
    {
      "id": "649758",
      "postDate": "10/15/2019 18:20:56",
      "content": "<p>Nice writeup, thank you!</p>",
      "rawMarkdown": "Nice writeup, thank you!",
      "votes": null
    },
    {
      "id": "649818",
      "postDate": "10/15/2019 19:34:14",
      "content": "<p>Thanks for sharing! great details. 👍 </p>",
      "rawMarkdown": "Thanks for sharing! great details. 👍",
      "votes": null
    },
    {
      "id": "649861",
      "postDate": "10/15/2019 20:52:43",
      "content": "<p>Congratulations for it bro!! Very interesting and useful solutions</p>",
      "rawMarkdown": "Congratulations for it bro!! Very interesting and useful solutions",
      "votes": null
    },
    {
      "id": "650099",
      "postDate": "10/16/2019 05:18:28",
      "content": "<p>Thank you so much!! We did our best to design this competition based on Kaggle user experience. :)</p>",
      "rawMarkdown": "Thank you so much!! We did our best to design this competition based on Kaggle user experience. :)",
      "votes": null
    },
    {
      "id": "650234",
      "postDate": "10/16/2019 07:55:42",
      "content": "<p>Congratulations! Amazing ideas! Thanks for sharing.... <a href=\"/jesucristo\">@jesucristo</a> </p>",
      "rawMarkdown": "Congratulations! Amazing ideas! Thanks for sharing.... @jesucristo",
      "votes": null
    },
    {
      "id": "651111",
      "postDate": "10/17/2019 03:55:02",
      "content": "<p>Thank you so much for sharing your solution. It's very informative!</p>",
      "rawMarkdown": "Thank you so much for sharing your solution. It's very informative!",
      "votes": null
    },
    {
      "id": "657381",
      "postDate": "10/25/2019 05:07:11",
      "content": "<p>Thank you for sharing. </p>\n\n<p>I tried to run <a href=\"https://github.com/mv-lab/kuzushiji-recognition/blob/master/Centernet%20experiment.ipynb\">Centernet experiment.ipynb</a>\nand got AttributeError 'Image' object has no attribute 'obj' at Evaluate on validation set.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F468287%2F961ce321b0c02f45d1b6860c583b9358%2F201910251403.PNG?generation=1571980121860806&amp;alt=media\" alt=\"\"></p>\n\n<p>Would you be able to tell me how to solve this error ?</p>",
      "rawMarkdown": "Thank you for sharing. \n\nI tried to run [Centernet experiment.ipynb](https://github.com/mv-lab/kuzushiji-recognition/blob/master/Centernet%20experiment.ipynb)\nand got AttributeError 'Image' object has no attribute 'obj' at Evaluate on validation set.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F468287%2F961ce321b0c02f45d1b6860c583b9358%2F201910251403.PNG?generation=1571980121860806&amp;alt=media)\n\nWould you be able to tell me how to solve this error ?",
      "votes": null
    },
    {
      "id": "675399",
      "postDate": "11/18/2019 03:09:04",
      "content": "<p>Hello, </p>\n\n<p>Organizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  </p>\n\n<p>Best, </p>\n\n<p>Alex.  </p>",
      "rawMarkdown": "Hello, \n\nOrganizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  \n\nBest, \n\nAlex.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 649592,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "10/15/2019 14:56:04",
      "content": "<p>Congratulations\nGreat Write-Up\nThanks for Sharing your Approach &amp; Insights.. <a href=\"/jesucristo\">@jesucristo</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 649758,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "10/15/2019 18:20:56",
      "content": "<p>Nice writeup, thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 649818,
      "author_name": "zinovadr",
      "author_url": "",
      "post_date": "10/15/2019 19:34:14",
      "content": "<p>Thanks for sharing! great details. 👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 649861,
      "author_name": "kabure",
      "author_url": "",
      "post_date": "10/15/2019 20:52:43",
      "content": "<p>Congratulations for it bro!! Very interesting and useful solutions</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 650099,
      "author_name": "tkasasagi",
      "author_url": "",
      "post_date": "10/16/2019 05:18:28",
      "content": "<p>Thank you so much!! We did our best to design this competition based on Kaggle user experience. :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 650234,
      "author_name": "evgenyshtepin",
      "author_url": "",
      "post_date": "10/16/2019 07:55:42",
      "content": "<p>Congratulations! Amazing ideas! Thanks for sharing.... <a href=\"/jesucristo\">@jesucristo</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 651111,
      "author_name": "kmat2019",
      "author_url": "",
      "post_date": "10/17/2019 03:55:02",
      "content": "<p>Thank you so much for sharing your solution. It's very informative!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 657381,
      "author_name": "nakajima",
      "author_url": "",
      "post_date": "10/25/2019 05:07:11",
      "content": "<p>Thank you for sharing. </p>\n\n<p>I tried to run <a href=\"https://github.com/mv-lab/kuzushiji-recognition/blob/master/Centernet%20experiment.ipynb\">Centernet experiment.ipynb</a>\nand got AttributeError 'Image' object has no attribute 'obj' at Evaluate on validation set.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F468287%2F961ce321b0c02f45d1b6860c583b9358%2F201910251403.PNG?generation=1571980121860806&amp;alt=media\" alt=\"\"></p>\n\n<p>Would you be able to tell me how to solve this error ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 675399,
      "author_name": "thenuttynetter",
      "author_url": "",
      "post_date": "11/18/2019 03:09:04",
      "content": "<p>Hello, </p>\n\n<p>Organizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  </p>\n\n<p>Best, </p>\n\n<p>Alex.  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "649533": "&gt; **GitHub:** https://github.com/mv-lab/kuzushiji-recognition\n\nFirst of all I'd like to thank my mates @ollieperree and @tikutiku, we couldn't work a lot of time together but the experience was amazing as the result!\n\nAlso congrats to the winners @tascj0 @lopuhin @knjcode @linhui @seesee \n@kmat2019 your kernel is amazing!\n\nAnd obviously the organization! @tkasasagi  @sohier and @anokas this competition was perfect! I think kaggle should reconsider medals.\nWe don't receive medals, and we are only 293 teams... I think the people here competed because of the big prize and the challenge itself, not medals :) In our case, the biggest impediment was the GPU quota...\n\nAnyway, I'll explain our approach. Please check the notebook: **[Kuzushiji Recognition Starter](https://www.kaggle.com/jesucristo/kuzushiji-recognition-starter)** where we are going to update the inference and the public dataset with 256x256 images for training classifier.\n\n---\nFrom the beginning @ollieperree was using a **2-stage approach.** \nOur approach to detection was directly inspired by [K_mat's kernel](https://www.kaggle.com/kmat2019/centernet-keypoint-detector), with the main takeaway being the idea of predicting a heatmap showing the centers of characters. Initially, we used a U-Net with a resnet18 backbone to predict a heatmap consisting of ellipses placed at the centers of characters, with the radii proportional to the width and height of the bounding box, with the input to the model being a 1024x1024 pixel crop of the page resized to 256x256 pixels. \nPredictions for the centers were then obtained by picking the local maxima (note that the width and height of the bounding box were not predicted). Performance was improved by changing the ellipses to circles of constant radius.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2Fa81b20efc933fcdfd877b0636e229da6%2Fumgy012-042___000_002.png?generation=1571161817145901&amp;alt=media =200x*)\n\n\nWe tried using *focal loss* and binary *cross-entropy* as loss functions, but using mean squared error resulted in the cleanest predictions for us (though more epochs were needed to get sensible-looking predictions).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F97924b78f0393a6e162cb477ba51c5e1%2Fimage(1).png?generation=1571161817778557&amp;alt=media =500x*)\n\nOne issue with using 1024x1024 crops of the page as the input were **\"artifacts\"** around the edges of the input. We tried a few things to try to counteract this, such as moving the sliding window over the page with a stride less than 1024x1024, then removing duplicate predictions by detecting when two predicted points of the same class were within a certain distance of each other. However, these did not give an improvement on the LB - we think that tuning parameters for these methods on the validation set, as well as the parameters for selecting maxima in the heatmap, might have caused us to \"overfit\".\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F2cd1d021e43de8c00e5f641c6e91586c%2FScreenshot%20from%202019-10-15%2019-53-12.png?generation=1571162065651936&amp;alt=media)\n\nThese artifacts were related with the drawings and annotations!\n(See carefully the *red* dots at the images)\nHow did we fix this? Check **ensemble** and **pseudolabels**\n\n---\n\nWe have a **2 stage** model: detection and classification.\n\n# Detection\n\nWe used as starter code the great kernel: [CenterNet -Keypoint Detector-](https://www.kaggle.com/kmat2019/centernet-keypoint-detector) by @kmat2019\nThen I realized that @seesee had his own [keras-centernet](https://github.com/see--/keras-centernet).\nAt the end we used Hourglass and the output are boxes instead of only the centers (like the original paper).\n\n**Model**\n- Detection by hourglassnet\n- Output heatmaps + width/height maps\n- generate_heatmap_crops_circular(crop1024,resize256)\n- Validation: w/o outliers GroupKFold\n- resnet34\n- MSELoss, dice and IOU (oss=0.00179, dice=0.6270, F1=0.9856, iou=0.8142)\n- Augmentations: aug(randombrightness0.2,scale0.5)\n- Learning rate: (1e-4:1e-6)\n- 20epochs\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F4d84f42553ac59551b480d883ad8f33b%2Fvalid_pred.png?generation=1571146402846083&amp;alt=media =300x*)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F9c032decbf330957fd2b25c6a34a73bd%2FScreenshot%20from%202019-10-15%2011-05-12.png?generation=1571146515196023&amp;alt=media =500x*)\n\n\n# Classification\n\nThe classification model was a **resnet18**, pretrained on ImageNet, with the input being a fixed *256x256* pixel area, scaled down to *128x128*, centered at the (predicted) center of the character. \nThe training data included a *background class*, whose training examples were random 256x256 crops of the pages with no labelled characters. \nTraining was done using the **fastai** library, with [standard fastai transforms](https://docs.fast.ai/vision.transform.html#get_transforms) and MixUp. \nThis model achieved a Classification accuracy of **94.6%** on a validation set (20% of the train data, group split by book).\n\n# Augmentations\n\nLike @lopuhin  said: *I'm using standard augmentations from https://github.com/albu/albumentations/ library, including adjusting colors that helps simulate different paper styles*\n\n&gt; Can you tell me wich one is the real one?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F89c2cd2e945352655b1932b8d6ce1378%2Fimage.png?generation=1571146446903717&amp;alt=media =500x*)\n\n**Code**\n\n```\nimport albumentations\n\ncolorize = albumentations.RGBShift(r_shift_limit=0, g_shift_limit=0, b_shift_limit=[-80,0])\n\ndef color_get_params():\n    a = random.uniform(-40, 0)\n    b = random.uniform(-80, -30)\n    return {\"r_shift\": a,\n            \"g_shift\": a,\n            \"b_shift\": b}\n\ncolorize.get_params = color_get_params\n\naug = albumentations.Compose([albumentations.RandomBrightnessContrast(contrast_limit=0.2, brightness_limit=0.2),\n                              albumentations.ToGray(),\n                              albumentations.Blur(),\n                              albumentations.Rotate(limit=5),\n                              colorize\n                             ])\n```\n\n# Ensemble\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F3d45fceb5c94db4071d6ef4e53990e1b%2Fensemble.jpg?generation=1571162716740190&amp;alt=media)\n\nWe had a problem... our centers weren't ordered! so in order to improve the accuracy and delete *false positives* we thought in the following ensemble method.\nFor the image IMG we take the most external centers from 3 predictions: \n```\n(xmin, ymin), (xmin, ymax), (xmax,ymin), (xmax,ymax)\n``` \nAt the picture this boxes are represented by 3 different colours (yellow, blue, red). \nFinally, we take the **intersection** of those 3 boxes, the black rectangle defined as (X,Y,Z,W), and we drop all the centers out of the black box! with this technique we could eliminate artifacts like predictions at the edges.\nI have to say that the idea is cool, but I found 2 bugs and this technique didn't work properly :(\n\n# Pseudolabels\n\nWe realized that train had only **276** with no-chars and we thought that was the reason why our model had so many problems with drawings. \nThe solution? we set the as threshold= 5 chars. All the predictions with &lt;=5 chars where used as pseudolabels with labels= *NaN*.\n\n# Bonus\n\nA samurai with 3 swords... that emblem... this reminds me to something :)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2Fd7568f09270098f7a80994266db642e0%2Fkozuki.jpg?generation=1571164389635754&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F7989d65dc8ad606d2cfa51dc08fd87d5%2FFamiglia_Kozuki_simbolo.webp?generation=1571164245190855&amp;alt=media =100x*)",
    "649592": "Congratulations\nGreat Write-Up\nThanks for Sharing your Approach &amp; Insights.. @jesucristo",
    "649758": "Nice writeup, thank you!",
    "649818": "Thanks for sharing! great details. 👍",
    "649861": "Congratulations for it bro!! Very interesting and useful solutions",
    "650099": "Thank you so much!! We did our best to design this competition based on Kaggle user experience. :)",
    "650234": "Congratulations! Amazing ideas! Thanks for sharing.... @jesucristo",
    "651111": "Thank you so much for sharing your solution. It's very informative!",
    "657381": "Thank you for sharing. \n\nI tried to run [Centernet experiment.ipynb](https://github.com/mv-lab/kuzushiji-recognition/blob/master/Centernet%20experiment.ipynb)\nand got AttributeError 'Image' object has no attribute 'obj' at Evaluate on validation set.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F468287%2F961ce321b0c02f45d1b6860c583b9358%2F201910251403.PNG?generation=1571980121860806&amp;alt=media)\n\nWould you be able to tell me how to solve this error ?",
    "675399": "Hello, \n\nOrganizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  \n\nBest, \n\nAlex."
  },
  "source": "meta"
}