{
  "id": 112899,
  "title": "7th place solution",
  "url": "/competitions/kuzushiji-recognition/writeups/k-mat-7th-place-solution",
  "author_name": "",
  "post_date": "2019-10-16T03:52:56.096698Z",
  "votes": 26,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Thanks to organizers for this interesting challenge and congrats everyone who enjoyed this challenge!\nI joined this competition to learn CNNs. That’s why I don’t have a personal GPU(so my work is conducted on google colab) and didn’t use any pretrained/predefined model, but that gave me a lot of knowledge and experience.</p>\n\n<h2>Approach</h2>\n\n<p>Most of my solution is written in <a href=\"https://www.kaggle.com/kmat2019/centernet-keypoint-detector\">my kernel</a>. The overview is drawn in the attached figure.</p>\n\n<ul>\n<li>Two stage (Detection and Classification)</li>\n<li>Validation: Group Split by book titles.\n-&gt; This results gave me the idea of domain adaptation, but in vain except pseudo labeling.</li>\n</ul>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2938236%2Fc91592a8f1bed77658962edb705723d0%2Ffig.jpg?generation=1571197209830549&amp;alt=media\" alt=\"\"></p>\n\n<h2>Detection</h2>\n\n<ul>\n<li>Architecture: CenterNet (input:512x512, output:128x128)</li>\n<li>Augmentation: Cropping, Brightness, Contrast, Horizontal flip</li>\n<li>TTA(flip) &amp; Ensemble</li>\n</ul>\n\n<p>As I wanted to use a shallow model(smaller than Resnet34) due to the lack of machine resource, I focused on the specific resolution of CNN layers. Considering the size of the objects, 64x64 or 32x32 is important and smaller than 16x16 is useless for my model. Finally, I got a very accurate model with IoU~0.90, F1(IoU&gt;0.5)~0.98 and F1(center)~0.99.</p>\n\n<h2>Classification</h2>\n\n<ul>\n<li>Architecture: Resnet base (input:64x64), aspect ratio and size of object are concatenated at FC layer</li>\n<li>Augmentation: Cropping, Erasing, Brightness, Contrast, </li>\n<li>TTA(Size, Brightness) &amp; Ensemble</li>\n<li>Pseudo labeling</li>\n</ul>\n\n<p>Classification model is not unique at all. I wanted to try other models and augmentations.</p>\n\n<h2>I also tried...</h2>\n\n<ul>\n<li><p>Language model with LSTM\n-&gt; very small improvement</p></li>\n<li><p>Domain Adaptation like DANN\n-&gt; Pseudo labeling worked better.</p></li>\n</ul>\n\n<p>Thanks again to everyone!</p>",
  "messages": [
    {
      "id": "650058",
      "postDate": "10/16/2019 03:52:56",
      "content": "<p>Thanks to organizers for this interesting challenge and congrats everyone who enjoyed this challenge!\nI joined this competition to learn CNNs. That’s why I don’t have a personal GPU(so my work is conducted on google colab) and didn’t use any pretrained/predefined model, but that gave me a lot of knowledge and experience.</p>\n\n<h2>Approach</h2>\n\n<p>Most of my solution is written in <a href=\"https://www.kaggle.com/kmat2019/centernet-keypoint-detector\">my kernel</a>. The overview is drawn in the attached figure.</p>\n\n<ul>\n<li>Two stage (Detection and Classification)</li>\n<li>Validation: Group Split by book titles.\n-&gt; This results gave me the idea of domain adaptation, but in vain except pseudo labeling.</li>\n</ul>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2938236%2Fc91592a8f1bed77658962edb705723d0%2Ffig.jpg?generation=1571197209830549&amp;alt=media\" alt=\"\"></p>\n\n<h2>Detection</h2>\n\n<ul>\n<li>Architecture: CenterNet (input:512x512, output:128x128)</li>\n<li>Augmentation: Cropping, Brightness, Contrast, Horizontal flip</li>\n<li>TTA(flip) &amp; Ensemble</li>\n</ul>\n\n<p>As I wanted to use a shallow model(smaller than Resnet34) due to the lack of machine resource, I focused on the specific resolution of CNN layers. Considering the size of the objects, 64x64 or 32x32 is important and smaller than 16x16 is useless for my model. Finally, I got a very accurate model with IoU~0.90, F1(IoU&gt;0.5)~0.98 and F1(center)~0.99.</p>\n\n<h2>Classification</h2>\n\n<ul>\n<li>Architecture: Resnet base (input:64x64), aspect ratio and size of object are concatenated at FC layer</li>\n<li>Augmentation: Cropping, Erasing, Brightness, Contrast, </li>\n<li>TTA(Size, Brightness) &amp; Ensemble</li>\n<li>Pseudo labeling</li>\n</ul>\n\n<p>Classification model is not unique at all. I wanted to try other models and augmentations.</p>\n\n<h2>I also tried...</h2>\n\n<ul>\n<li><p>Language model with LSTM\n-&gt; very small improvement</p></li>\n<li><p>Domain Adaptation like DANN\n-&gt; Pseudo labeling worked better.</p></li>\n</ul>\n\n<p>Thanks again to everyone!</p>",
      "rawMarkdown": "Thanks to organizers for this interesting challenge and congrats everyone who enjoyed this challenge!\nI joined this competition to learn CNNs. That’s why I don’t have a personal GPU(so my work is conducted on google colab) and didn’t use any pretrained/predefined model, but that gave me a lot of knowledge and experience.\n\n## Approach\nMost of my solution is written in [my kernel](https://www.kaggle.com/kmat2019/centernet-keypoint-detector). The overview is drawn in the attached figure.\n\n- Two stage (Detection and Classification)\n- Validation: Group Split by book titles.\n  -&gt; This results gave me the idea of domain adaptation, but in vain except pseudo labeling.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2938236%2Fc91592a8f1bed77658962edb705723d0%2Ffig.jpg?generation=1571197209830549&amp;alt=media)\n\n\n\n## Detection\n- Architecture: CenterNet (input:512x512, output:128x128)\n- Augmentation: Cropping, Brightness, Contrast, Horizontal flip\n- TTA(flip) &amp; Ensemble\n\nAs I wanted to use a shallow model(smaller than Resnet34) due to the lack of machine resource, I focused on the specific resolution of CNN layers. Considering the size of the objects, 64x64 or 32x32 is important and smaller than 16x16 is useless for my model. Finally, I got a very accurate model with IoU~0.90, F1(IoU&gt;0.5)~0.98 and F1(center)~0.99.\n\n## Classification\n- Architecture: Resnet base (input:64x64), aspect ratio and size of object are concatenated at FC layer\n- Augmentation: Cropping, Erasing, Brightness, Contrast, \n- TTA(Size, Brightness) &amp; Ensemble\n- Pseudo labeling\n\nClassification model is not unique at all. I wanted to try other models and augmentations.\n\n## I also tried...\n- Language model with LSTM\n-&gt; very small improvement\n\n- Domain Adaptation like DANN\n-&gt; Pseudo labeling worked better.\n\nThanks again to everyone!",
      "votes": null
    },
    {
      "id": "650068",
      "postDate": "10/16/2019 04:11:26",
      "content": "<p>Cool and clear presentation. In real life this could be solution number 1. thx</p>",
      "rawMarkdown": "Cool and clear presentation. In real life this could be solution number 1. thx",
      "votes": null
    },
    {
      "id": "650095",
      "postDate": "10/16/2019 05:12:12",
      "content": "<p>Congrats\nI started this competition with your amazing kernel!\nThank You for Sharing your Approach &amp; Insights.... <a href=\"/kmat2019\">@kmat2019</a> </p>",
      "rawMarkdown": "Congrats\nI started this competition with your amazing kernel!\nThank You for Sharing your Approach &amp; Insights.... @kmat2019",
      "votes": null
    },
    {
      "id": "650180",
      "postDate": "10/16/2019 06:59:47",
      "content": "<p>Congrats\nThank You for Sharing your Approach &amp; Insights</p>",
      "rawMarkdown": "Congrats\nThank You for Sharing your Approach &amp; Insights",
      "votes": null
    },
    {
      "id": "651003",
      "postDate": "10/17/2019 00:32:14",
      "content": "<p>\"Validation: Group Split by book titles.\n-&gt; This results gave me the idea of domain adaptation\"</p>\n\n<p>Interesting.  For Domain Adaptation, you mean classifying the book identity from the hidden states and then adversarially trying to make that classifier less accurate?  </p>\n\n<p>It seems like an interesting idea, since we want the model to learn less book-specific features.  I definitely think the book id should be used more, however one thing I wonder about is if adversarial domain adaptation might be too strict if the books simply have different types of content (for example some characters are in one book but not another).  </p>\n\n<p>I've been pretty interested in \"Invariant Risk Minimization\" which discusses this issue of generalizing across multiple \"environments\".  Appendix C is especially relevant to this discussion: </p>\n\n<p><a href=\"https://arxiv.org/pdf/1907.02893.pdf\">https://arxiv.org/pdf/1907.02893.pdf</a></p>\n\n<p>\"but in vain except pseudo labeling.\"</p>\n\n<p>What do you mean by pseudolabeling in this context?  </p>",
      "rawMarkdown": "\"Validation: Group Split by book titles.\n-&gt; This results gave me the idea of domain adaptation\"\n\nInteresting.  For Domain Adaptation, you mean classifying the book identity from the hidden states and then adversarially trying to make that classifier less accurate?  \n\nIt seems like an interesting idea, since we want the model to learn less book-specific features.  I definitely think the book id should be used more, however one thing I wonder about is if adversarial domain adaptation might be too strict if the books simply have different types of content (for example some characters are in one book but not another).  \n\nI've been pretty interested in \"Invariant Risk Minimization\" which discusses this issue of generalizing across multiple \"environments\".  Appendix C is especially relevant to this discussion: \n\nhttps://arxiv.org/pdf/1907.02893.pdf\n\n\"but in vain except pseudo labeling.\"\n\nWhat do you mean by pseudolabeling in this context?",
      "votes": null
    },
    {
      "id": "651144",
      "postDate": "10/17/2019 05:07:58",
      "content": "<p>Congrats</p>",
      "rawMarkdown": "Congrats",
      "votes": null
    },
    {
      "id": "651270",
      "postDate": "10/17/2019 09:13:11",
      "content": "<p>Thank you <a href=\"/thenuttynetter\">@thenuttynetter</a> for sharing interesting knowledge. </p>\n\n<p>Yes. I tried adversarial domain adaptation technique based on <a href=\"https://arxiv.org/abs/1505.07818\">DANN</a> to learn less book-specific features. But it looks that the increase in the book title loss and decrease in character classification loss slightly conflict. What you mentioned would be one of the reasons.\nI also made the other model like the following figure. I thought this model can generate the author-related features. But I didn’t have the time and machine resource to do it :(</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2938236%2F3d837360859405258a1ed90e3d75ecfc%2Ffig2.png?generation=1571301302463252&amp;alt=media\" alt=\"\"></p>\n\n<p>I think pseudo labeling worked as a kind of domain adaptation. Even shifting the parameters of batch normalization works as domain adaptation. A fine tuning by pseudo labeling as <a href=\"/lopuhin\">@lopuhin</a> explained is more effective in this competition.</p>",
      "rawMarkdown": "Thank you @thenuttynetter for sharing interesting knowledge. \n\nYes. I tried adversarial domain adaptation technique based on [DANN](https://arxiv.org/abs/1505.07818) to learn less book-specific features. But it looks that the increase in the book title loss and decrease in character classification loss slightly conflict. What you mentioned would be one of the reasons.\nI also made the other model like the following figure. I thought this model can generate the author-related features. But I didn’t have the time and machine resource to do it :(\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2938236%2F3d837360859405258a1ed90e3d75ecfc%2Ffig2.png?generation=1571301302463252&amp;alt=media)\n\n\nI think pseudo labeling worked as a kind of domain adaptation. Even shifting the parameters of batch normalization works as domain adaptation. A fine tuning by pseudo labeling as @lopuhin explained is more effective in this competition.",
      "votes": null
    },
    {
      "id": "651499",
      "postDate": "10/17/2019 14:40:52",
      "content": "<p>Hey, thanks for the writeup and the awesome notebook! I used a similar approach to yours, but I don't understand what you mean by 64x64, 32x32 object size. Could you please clarify on that?</p>",
      "rawMarkdown": "Hey, thanks for the writeup and the awesome notebook! I used a similar approach to yours, but I don't understand what you mean by 64x64, 32x32 object size. Could you please clarify on that?",
      "votes": null
    },
    {
      "id": "651905",
      "postDate": "10/18/2019 04:14:43",
      "content": "<p>Sorry for my poor explanation.</p>\n\n<p>First, I implemented the model [A] for the detection, but model [B] performed better with the same number of convolution layers. This is because the resolution of 8x8 is too low to find the objects. Such broad view is not necessary to detect small objects. </p>\n\n<p><strong>Model [A]</strong>\ninput layer [512x512] -&gt; Encoder -&gt; down-sampling to [8x8] -&gt;Decoder -&gt; up-sampling to [128x128] (output layer)</p>\n\n<p><strong>Model [B]</strong>\ninput layer [512x512] -&gt; Encoder -&gt; down-sampling to [32x32] -&gt;Decoder -&gt; up-sampling to [128x128] (output layer)</p>",
      "rawMarkdown": "Sorry for my poor explanation.\n\nFirst, I implemented the model [A] for the detection, but model [B] performed better with the same number of convolution layers. This is because the resolution of 8x8 is too low to find the objects. Such broad view is not necessary to detect small objects. \n\n**Model [A]**\ninput layer [512x512] -&gt; Encoder -&gt; down-sampling to [8x8] -&gt;Decoder -&gt; up-sampling to [128x128] (output layer)\n\n**Model [B]**\ninput layer [512x512] -&gt; Encoder -&gt; down-sampling to [32x32] -&gt;Decoder -&gt; up-sampling to [128x128] (output layer)",
      "votes": null
    },
    {
      "id": "653960",
      "postDate": "10/21/2019 08:19:25",
      "content": "<p>Thanks for share. I learned more from your kernel.\nCould you please also share your scoring curve ?\nex:   base model          lb 0.80\n        .....\n        +TTA                      lb 0.90\n        +pseudle labeling lb 0.92</p>",
      "rawMarkdown": "Thanks for share. I learned more from your kernel.\nCould you please also share your scoring curve ?\nex:   base model          lb 0.80\n        .....\n        +TTA                      lb 0.90\n        +pseudle labeling lb 0.92",
      "votes": null
    },
    {
      "id": "653993",
      "postDate": "10/21/2019 09:19:50",
      "content": "<p>I see. Thanks!</p>",
      "rawMarkdown": "I see. Thanks!",
      "votes": null
    },
    {
      "id": "654877",
      "postDate": "10/22/2019 12:36:33",
      "content": "<p>Thank you <a href=\"/atom1231\">@atom1231</a> . Scores of public learderboard are as follows.</p>\n\n<p>Baseline: approx. 0.85 by modifying my kernel.\n- Changing NMS calculation. (Please refer to my comment in my kernel.)\n- Unifying resizing method. In my kernel, I used both resizing by pillow(PIL) and open cv(cv2). I didn’t know that the default interpolation methods between them are different. I must apologize.\n- Add post processing of the duplicated area after detection as shown in the figure.\n- Training classifier with all datasets (no validation split).</p>\n\n<p>+Improving detector as shown above: 0.90\n+Ensemble of 2 detectors: 0.905\n+Improving classifier model and adding some augmentations: 0.915\n+Ensemble of 2 classifiers: 0.918\n+TTA &amp; pseudo labeling: 0.93</p>",
      "rawMarkdown": "Thank you @atom1231 . Scores of public learderboard are as follows.\n\nBaseline: approx. 0.85 by modifying my kernel.\n- Changing NMS calculation. (Please refer to my comment in my kernel.)\n- Unifying resizing method. In my kernel, I used both resizing by pillow(PIL) and open cv(cv2). I didn’t know that the default interpolation methods between them are different. I must apologize.\n- Add post processing of the duplicated area after detection as shown in the figure.\n- Training classifier with all datasets (no validation split).\n\n+Improving detector as shown above: 0.90\n+Ensemble of 2 detectors: 0.905\n+Improving classifier model and adding some augmentations: 0.915\n+Ensemble of 2 classifiers: 0.918\n+TTA &amp; pseudo labeling: 0.93",
      "votes": null
    },
    {
      "id": "675401",
      "postDate": "11/18/2019 03:09:21",
      "content": "<p>Hello, </p>\n\n<p>Organizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  </p>\n\n<p>Best, </p>\n\n<p>Alex.  </p>",
      "rawMarkdown": "Hello, \n\nOrganizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  \n\nBest, \n\nAlex.",
      "votes": null
    },
    {
      "id": "684151",
      "postDate": "11/29/2019 09:51:02",
      "content": "<p>I remember it was approx. 1 to 2 sec with Tesla K80, without TTA and ensemble. The accuracy of single model is about 92%.\nIt takes 10-20 times longer with TTA and ensemble.</p>",
      "rawMarkdown": "I remember it was approx. 1 to 2 sec with Tesla K80, without TTA and ensemble. The accuracy of single model is about 92%.\nIt takes 10-20 times longer with TTA and ensemble.",
      "votes": null
    },
    {
      "id": "688270",
      "postDate": "12/05/2019 12:12:49",
      "content": "<p>Thanks, I just want to use Centernet in another competition. <a href=\"/kmat2019\">@kmat2019</a> </p>",
      "rawMarkdown": "Thanks, I just want to use Centernet in another competition. @kmat2019",
      "votes": null
    },
    {
      "id": "725523",
      "postDate": "01/22/2020 07:59:34",
      "content": "<p>Thanks for sharing your valuable knowledge </p>",
      "rawMarkdown": "Thanks for sharing your valuable knowledge",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 650068,
      "author_name": "zinovadr",
      "author_url": "",
      "post_date": "10/16/2019 04:11:26",
      "content": "<p>Cool and clear presentation. In real life this could be solution number 1. thx</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 650095,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "10/16/2019 05:12:12",
      "content": "<p>Congrats\nI started this competition with your amazing kernel!\nThank You for Sharing your Approach &amp; Insights.... <a href=\"/kmat2019\">@kmat2019</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 650180,
      "author_name": "marek3000",
      "author_url": "",
      "post_date": "10/16/2019 06:59:47",
      "content": "<p>Congrats\nThank You for Sharing your Approach &amp; Insights</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 651003,
      "author_name": "thenuttynetter",
      "author_url": "",
      "post_date": "10/17/2019 00:32:14",
      "content": "<p>\"Validation: Group Split by book titles.\n-&gt; This results gave me the idea of domain adaptation\"</p>\n\n<p>Interesting.  For Domain Adaptation, you mean classifying the book identity from the hidden states and then adversarially trying to make that classifier less accurate?  </p>\n\n<p>It seems like an interesting idea, since we want the model to learn less book-specific features.  I definitely think the book id should be used more, however one thing I wonder about is if adversarial domain adaptation might be too strict if the books simply have different types of content (for example some characters are in one book but not another).  </p>\n\n<p>I've been pretty interested in \"Invariant Risk Minimization\" which discusses this issue of generalizing across multiple \"environments\".  Appendix C is especially relevant to this discussion: </p>\n\n<p><a href=\"https://arxiv.org/pdf/1907.02893.pdf\">https://arxiv.org/pdf/1907.02893.pdf</a></p>\n\n<p>\"but in vain except pseudo labeling.\"</p>\n\n<p>What do you mean by pseudolabeling in this context?  </p>",
      "votes": null,
      "replies": [
        {
          "id": 651270,
          "author_name": "kmat2019",
          "author_url": "",
          "post_date": "10/17/2019 09:13:11",
          "content": "<p>Thank you <a href=\"/thenuttynetter\">@thenuttynetter</a> for sharing interesting knowledge. </p>\n\n<p>Yes. I tried adversarial domain adaptation technique based on <a href=\"https://arxiv.org/abs/1505.07818\">DANN</a> to learn less book-specific features. But it looks that the increase in the book title loss and decrease in character classification loss slightly conflict. What you mentioned would be one of the reasons.\nI also made the other model like the following figure. I thought this model can generate the author-related features. But I didn’t have the time and machine resource to do it :(</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2938236%2F3d837360859405258a1ed90e3d75ecfc%2Ffig2.png?generation=1571301302463252&amp;alt=media\" alt=\"\"></p>\n\n<p>I think pseudo labeling worked as a kind of domain adaptation. Even shifting the parameters of batch normalization works as domain adaptation. A fine tuning by pseudo labeling as <a href=\"/lopuhin\">@lopuhin</a> explained is more effective in this competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 651144,
      "author_name": "erdeneochir",
      "author_url": "",
      "post_date": "10/17/2019 05:07:58",
      "content": "<p>Congrats</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 651499,
      "author_name": "ananschuett",
      "author_url": "",
      "post_date": "10/17/2019 14:40:52",
      "content": "<p>Hey, thanks for the writeup and the awesome notebook! I used a similar approach to yours, but I don't understand what you mean by 64x64, 32x32 object size. Could you please clarify on that?</p>",
      "votes": null,
      "replies": [
        {
          "id": 651905,
          "author_name": "kmat2019",
          "author_url": "",
          "post_date": "10/18/2019 04:14:43",
          "content": "<p>Sorry for my poor explanation.</p>\n\n<p>First, I implemented the model [A] for the detection, but model [B] performed better with the same number of convolution layers. This is because the resolution of 8x8 is too low to find the objects. Such broad view is not necessary to detect small objects. </p>\n\n<p><strong>Model [A]</strong>\ninput layer [512x512] -&gt; Encoder -&gt; down-sampling to [8x8] -&gt;Decoder -&gt; up-sampling to [128x128] (output layer)</p>\n\n<p><strong>Model [B]</strong>\ninput layer [512x512] -&gt; Encoder -&gt; down-sampling to [32x32] -&gt;Decoder -&gt; up-sampling to [128x128] (output layer)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 653993,
          "author_name": "ananschuett",
          "author_url": "",
          "post_date": "10/21/2019 09:19:50",
          "content": "<p>I see. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 653960,
      "author_name": "atom1231",
      "author_url": "",
      "post_date": "10/21/2019 08:19:25",
      "content": "<p>Thanks for share. I learned more from your kernel.\nCould you please also share your scoring curve ?\nex:   base model          lb 0.80\n        .....\n        +TTA                      lb 0.90\n        +pseudle labeling lb 0.92</p>",
      "votes": null,
      "replies": [
        {
          "id": 654877,
          "author_name": "kmat2019",
          "author_url": "",
          "post_date": "10/22/2019 12:36:33",
          "content": "<p>Thank you <a href=\"/atom1231\">@atom1231</a> . Scores of public learderboard are as follows.</p>\n\n<p>Baseline: approx. 0.85 by modifying my kernel.\n- Changing NMS calculation. (Please refer to my comment in my kernel.)\n- Unifying resizing method. In my kernel, I used both resizing by pillow(PIL) and open cv(cv2). I didn’t know that the default interpolation methods between them are different. I must apologize.\n- Add post processing of the duplicated area after detection as shown in the figure.\n- Training classifier with all datasets (no validation split).</p>\n\n<p>+Improving detector as shown above: 0.90\n+Ensemble of 2 detectors: 0.905\n+Improving classifier model and adding some augmentations: 0.915\n+Ensemble of 2 classifiers: 0.918\n+TTA &amp; pseudo labeling: 0.93</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 675401,
      "author_name": "thenuttynetter",
      "author_url": "",
      "post_date": "11/18/2019 03:09:21",
      "content": "<p>Hello, </p>\n\n<p>Organizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  </p>\n\n<p>Best, </p>\n\n<p>Alex.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 684151,
          "author_name": "kmat2019",
          "author_url": "",
          "post_date": "11/29/2019 09:51:02",
          "content": "<p>I remember it was approx. 1 to 2 sec with Tesla K80, without TTA and ensemble. The accuracy of single model is about 92%.\nIt takes 10-20 times longer with TTA and ensemble.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 688270,
      "author_name": "diegojohnson",
      "author_url": "",
      "post_date": "12/05/2019 12:12:49",
      "content": "<p>Thanks, I just want to use Centernet in another competition. <a href=\"/kmat2019\">@kmat2019</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 725523,
      "author_name": "isaadansari",
      "author_url": "",
      "post_date": "01/22/2020 07:59:34",
      "content": "<p>Thanks for sharing your valuable knowledge </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "650058": "Thanks to organizers for this interesting challenge and congrats everyone who enjoyed this challenge!\nI joined this competition to learn CNNs. That’s why I don’t have a personal GPU(so my work is conducted on google colab) and didn’t use any pretrained/predefined model, but that gave me a lot of knowledge and experience.\n\n## Approach\nMost of my solution is written in [my kernel](https://www.kaggle.com/kmat2019/centernet-keypoint-detector). The overview is drawn in the attached figure.\n\n- Two stage (Detection and Classification)\n- Validation: Group Split by book titles.\n  -&gt; This results gave me the idea of domain adaptation, but in vain except pseudo labeling.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2938236%2Fc91592a8f1bed77658962edb705723d0%2Ffig.jpg?generation=1571197209830549&amp;alt=media)\n\n\n\n## Detection\n- Architecture: CenterNet (input:512x512, output:128x128)\n- Augmentation: Cropping, Brightness, Contrast, Horizontal flip\n- TTA(flip) &amp; Ensemble\n\nAs I wanted to use a shallow model(smaller than Resnet34) due to the lack of machine resource, I focused on the specific resolution of CNN layers. Considering the size of the objects, 64x64 or 32x32 is important and smaller than 16x16 is useless for my model. Finally, I got a very accurate model with IoU~0.90, F1(IoU&gt;0.5)~0.98 and F1(center)~0.99.\n\n## Classification\n- Architecture: Resnet base (input:64x64), aspect ratio and size of object are concatenated at FC layer\n- Augmentation: Cropping, Erasing, Brightness, Contrast, \n- TTA(Size, Brightness) &amp; Ensemble\n- Pseudo labeling\n\nClassification model is not unique at all. I wanted to try other models and augmentations.\n\n## I also tried...\n- Language model with LSTM\n-&gt; very small improvement\n\n- Domain Adaptation like DANN\n-&gt; Pseudo labeling worked better.\n\nThanks again to everyone!",
    "650068": "Cool and clear presentation. In real life this could be solution number 1. thx",
    "650095": "Congrats\nI started this competition with your amazing kernel!\nThank You for Sharing your Approach &amp; Insights.... @kmat2019",
    "650180": "Congrats\nThank You for Sharing your Approach &amp; Insights",
    "651003": "\"Validation: Group Split by book titles.\n-&gt; This results gave me the idea of domain adaptation\"\n\nInteresting.  For Domain Adaptation, you mean classifying the book identity from the hidden states and then adversarially trying to make that classifier less accurate?  \n\nIt seems like an interesting idea, since we want the model to learn less book-specific features.  I definitely think the book id should be used more, however one thing I wonder about is if adversarial domain adaptation might be too strict if the books simply have different types of content (for example some characters are in one book but not another).  \n\nI've been pretty interested in \"Invariant Risk Minimization\" which discusses this issue of generalizing across multiple \"environments\".  Appendix C is especially relevant to this discussion: \n\nhttps://arxiv.org/pdf/1907.02893.pdf\n\n\"but in vain except pseudo labeling.\"\n\nWhat do you mean by pseudolabeling in this context?",
    "651144": "Congrats",
    "651270": "Thank you @thenuttynetter for sharing interesting knowledge. \n\nYes. I tried adversarial domain adaptation technique based on [DANN](https://arxiv.org/abs/1505.07818) to learn less book-specific features. But it looks that the increase in the book title loss and decrease in character classification loss slightly conflict. What you mentioned would be one of the reasons.\nI also made the other model like the following figure. I thought this model can generate the author-related features. But I didn’t have the time and machine resource to do it :(\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2938236%2F3d837360859405258a1ed90e3d75ecfc%2Ffig2.png?generation=1571301302463252&amp;alt=media)\n\n\nI think pseudo labeling worked as a kind of domain adaptation. Even shifting the parameters of batch normalization works as domain adaptation. A fine tuning by pseudo labeling as @lopuhin explained is more effective in this competition.",
    "651499": "Hey, thanks for the writeup and the awesome notebook! I used a similar approach to yours, but I don't understand what you mean by 64x64, 32x32 object size. Could you please clarify on that?",
    "651905": "Sorry for my poor explanation.\n\nFirst, I implemented the model [A] for the detection, but model [B] performed better with the same number of convolution layers. This is because the resolution of 8x8 is too low to find the objects. Such broad view is not necessary to detect small objects. \n\n**Model [A]**\ninput layer [512x512] -&gt; Encoder -&gt; down-sampling to [8x8] -&gt;Decoder -&gt; up-sampling to [128x128] (output layer)\n\n**Model [B]**\ninput layer [512x512] -&gt; Encoder -&gt; down-sampling to [32x32] -&gt;Decoder -&gt; up-sampling to [128x128] (output layer)",
    "653960": "Thanks for share. I learned more from your kernel.\nCould you please also share your scoring curve ?\nex:   base model          lb 0.80\n        .....\n        +TTA                      lb 0.90\n        +pseudle labeling lb 0.92",
    "653993": "I see. Thanks!",
    "654877": "Thank you @atom1231 . Scores of public learderboard are as follows.\n\nBaseline: approx. 0.85 by modifying my kernel.\n- Changing NMS calculation. (Please refer to my comment in my kernel.)\n- Unifying resizing method. In my kernel, I used both resizing by pillow(PIL) and open cv(cv2). I didn’t know that the default interpolation methods between them are different. I must apologize.\n- Add post processing of the duplicated area after detection as shown in the figure.\n- Training classifier with all datasets (no validation split).\n\n+Improving detector as shown above: 0.90\n+Ensemble of 2 detectors: 0.905\n+Improving classifier model and adding some augmentations: 0.915\n+Ensemble of 2 classifiers: 0.918\n+TTA &amp; pseudo labeling: 0.93",
    "675401": "Hello, \n\nOrganizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  \n\nBest, \n\nAlex.",
    "684151": "I remember it was approx. 1 to 2 sec with Tesla K80, without TTA and ensemble. The accuracy of single model is about 92%.\nIt takes 10-20 times longer with TTA and ensemble.",
    "688270": "Thanks, I just want to use Centernet in another competition. @kmat2019",
    "725523": "Thanks for sharing your valuable knowledge"
  },
  "source": "meta"
}