{
  "id": 113518,
  "title": "13th place solution + unsuccessful language model",
  "url": "/competitions/kuzushiji-recognition/writeups/james-day-13th-place-solution-unsuccessful-languag",
  "author_name": "",
  "post_date": "2019-10-20T03:52:47.993Z",
  "votes": 11,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Thanks to Kaggle and the organizers for creating such an interesting competition. Congrats to everyone who finished.</p>\n\n<h1><strong>Data preprocessing</strong></h1>\n\n<p>I resized all images to 512x512. Some of my images were converted to greyscale.</p>\n\n<h1><strong>Models</strong></h1>\n\n<p>My solution was based on Faster-RCNN, but with several small modifications.\n1. No layers are shared between the network which proposes regions of interest and the network which classifies them.  This allowed for easier experimentation because it allowed the networks to be modified independently of one another. It also simplified the training procedure. However, the inference efficiency of my solution could likely be improved by sharing all of the residual blocks.\n2.  Instead of choosing anchor sizes arbitrarily, I selected them by running k-means clustering on the widths and heights of the ground truth bounding boxes. I used the cluster centers as my anchor box sizes.\n3. I used ROI Align pooling instead of standard ROI pooling. I empirically found that ROI Align pooling yielded <em>slightly</em> better results.</p>\n\n<p>My region proposal network and my character classification network both used wide ResNet-34 backbones. My region proposal network used twice as many convolutional filters as the ResNet-34 configuration described in <a href=\"https://arxiv.org/pdf/1512.03385.pdf\">the original ResNet paper</a>. My character classification network used roughly three times as many convolutional filters as the original ResNet-34 configuration. I experimented with deeper ResNet-50 and ResNet-101 based networks, but I found wider, shallower networks worked better.</p>\n\n<p>To combat overfitting, I used dropout quite extensively. 2d dropout was used within all of my residual blocks. Ordinary 1d dropout was used outside the residual blocks near the final output layers.</p>\n\n<p>My final region proposal network uses color images and my final classification network uses greyscale images. I initially used greyscale images for both the region proposal network and the classification network because I didn't think the additional channels would be useful and wanted to conserve space. Near the end of the competition, I tried using color images for both networks. This slightly improved the region proposal network's output but worsened the classification network's accuracy due to overfitting.</p>\n\n<h1><strong>Failed attempt at using a language model for post-processing the detection results</strong></h1>\n\n<p>I was very interested in using a language model to correct my image processing code's incorrect labels. I spent over a month trying different approaches for this but was unable to get it working well.</p>\n\n<p>To get the characters into approximately the order in which a human would read them, I used DBSCAN to group the characters into columns, sorted the columns by the mean horizontal coordinates of their characters, and then sorted the characters within each column by their vertical coordinate. This worked well for most images.</p>\n\n<p>To train my correction networks, I used the ground truth labels to generate a large number of synthetic submission files that contained known errors. These errors were fairly realistic and were randomly added based on statistics from how my Faster-RCNN based model performs on a cross-validation dataset.</p>\n\n<p>To perform the character label corrections, I initially tried network architectures that are meant for performing translations (e.g. from English to Japanese), such as an encoder LSTM followed by a decoder LSTM with an attention mechanism, or a transformer network. I favored these network architectures over passing characters into a single recurrent network because in theory they should be able to elegantly handle cases in which the number of ground truth characters is different from the number of characters which was detected. Unfortunately, these networks introduced slightly more errors than they corrected. I eventually abandoned the translation based network architectures and switched over to using a simple bidirectional GRU network. It was able to improve my score by around .01, but only on submission files which were generated by weak image processing models. If I run it on my best submission file, then it will lower my score by ~.003.</p>\n\n<p>I think the reason I had so much difficulty getting a language model working well is that my Faster-RCNN based models do a better job considering a character's context than I initially expected. I think in order to get the correction network working well I would have needed to pull in an outside text corpus.</p>\n\n<h1><strong>Potential improvements</strong></h1>\n\n<p>There are several potential avenues for improvement that I did not implement or test.\n1. Data augmentation for the image processing models\n2. Pseudolabeling\n3. Using an ensemble of multiple classification networks\n4. Using an outside text corpus for the correction network</p>\n\n<p><strong>My code is available <a href=\"https://github.com/jday96314/Kuzushiji\">here</a>.</strong></p>",
  "messages": [
    {
      "id": "653194",
      "postDate": "10/20/2019 03:21:54",
      "content": "<p>Thanks to Kaggle and the organizers for creating such an interesting competition. Congrats to everyone who finished.</p>\n\n<h1><strong>Data preprocessing</strong></h1>\n\n<p>I resized all images to 512x512. Some of my images were converted to greyscale.</p>\n\n<h1><strong>Models</strong></h1>\n\n<p>My solution was based on Faster-RCNN, but with several small modifications.\n1. No layers are shared between the network which proposes regions of interest and the network which classifies them.  This allowed for easier experimentation because it allowed the networks to be modified independently of one another. It also simplified the training procedure. However, the inference efficiency of my solution could likely be improved by sharing all of the residual blocks.\n2.  Instead of choosing anchor sizes arbitrarily, I selected them by running k-means clustering on the widths and heights of the ground truth bounding boxes. I used the cluster centers as my anchor box sizes.\n3. I used ROI Align pooling instead of standard ROI pooling. I empirically found that ROI Align pooling yielded <em>slightly</em> better results.</p>\n\n<p>My region proposal network and my character classification network both used wide ResNet-34 backbones. My region proposal network used twice as many convolutional filters as the ResNet-34 configuration described in <a href=\"https://arxiv.org/pdf/1512.03385.pdf\">the original ResNet paper</a>. My character classification network used roughly three times as many convolutional filters as the original ResNet-34 configuration. I experimented with deeper ResNet-50 and ResNet-101 based networks, but I found wider, shallower networks worked better.</p>\n\n<p>To combat overfitting, I used dropout quite extensively. 2d dropout was used within all of my residual blocks. Ordinary 1d dropout was used outside the residual blocks near the final output layers.</p>\n\n<p>My final region proposal network uses color images and my final classification network uses greyscale images. I initially used greyscale images for both the region proposal network and the classification network because I didn't think the additional channels would be useful and wanted to conserve space. Near the end of the competition, I tried using color images for both networks. This slightly improved the region proposal network's output but worsened the classification network's accuracy due to overfitting.</p>\n\n<h1><strong>Failed attempt at using a language model for post-processing the detection results</strong></h1>\n\n<p>I was very interested in using a language model to correct my image processing code's incorrect labels. I spent over a month trying different approaches for this but was unable to get it working well.</p>\n\n<p>To get the characters into approximately the order in which a human would read them, I used DBSCAN to group the characters into columns, sorted the columns by the mean horizontal coordinates of their characters, and then sorted the characters within each column by their vertical coordinate. This worked well for most images.</p>\n\n<p>To train my correction networks, I used the ground truth labels to generate a large number of synthetic submission files that contained known errors. These errors were fairly realistic and were randomly added based on statistics from how my Faster-RCNN based model performs on a cross-validation dataset.</p>\n\n<p>To perform the character label corrections, I initially tried network architectures that are meant for performing translations (e.g. from English to Japanese), such as an encoder LSTM followed by a decoder LSTM with an attention mechanism, or a transformer network. I favored these network architectures over passing characters into a single recurrent network because in theory they should be able to elegantly handle cases in which the number of ground truth characters is different from the number of characters which was detected. Unfortunately, these networks introduced slightly more errors than they corrected. I eventually abandoned the translation based network architectures and switched over to using a simple bidirectional GRU network. It was able to improve my score by around .01, but only on submission files which were generated by weak image processing models. If I run it on my best submission file, then it will lower my score by ~.003.</p>\n\n<p>I think the reason I had so much difficulty getting a language model working well is that my Faster-RCNN based models do a better job considering a character's context than I initially expected. I think in order to get the correction network working well I would have needed to pull in an outside text corpus.</p>\n\n<h1><strong>Potential improvements</strong></h1>\n\n<p>There are several potential avenues for improvement that I did not implement or test.\n1. Data augmentation for the image processing models\n2. Pseudolabeling\n3. Using an ensemble of multiple classification networks\n4. Using an outside text corpus for the correction network</p>\n\n<p><strong>My code is available <a href=\"https://github.com/jday96314/Kuzushiji\">here</a>.</strong></p>",
      "rawMarkdown": "Thanks to Kaggle and the organizers for creating such an interesting competition. Congrats to everyone who finished.\n\n#**Data preprocessing**\n\nI resized all images to 512x512. Some of my images were converted to greyscale.\n\n#**Models**\n\nMy solution was based on Faster-RCNN, but with several small modifications.\n1. No layers are shared between the network which proposes regions of interest and the network which classifies them.  This allowed for easier experimentation because it allowed the networks to be modified independently of one another. It also simplified the training procedure. However, the inference efficiency of my solution could likely be improved by sharing all of the residual blocks.\n2.  Instead of choosing anchor sizes arbitrarily, I selected them by running k-means clustering on the widths and heights of the ground truth bounding boxes. I used the cluster centers as my anchor box sizes.\n3. I used ROI Align pooling instead of standard ROI pooling. I empirically found that ROI Align pooling yielded *slightly* better results.\n\nMy region proposal network and my character classification network both used wide ResNet-34 backbones. My region proposal network used twice as many convolutional filters as the ResNet-34 configuration described in [the original ResNet paper](https://arxiv.org/pdf/1512.03385.pdf). My character classification network used roughly three times as many convolutional filters as the original ResNet-34 configuration. I experimented with deeper ResNet-50 and ResNet-101 based networks, but I found wider, shallower networks worked better.\n\nTo combat overfitting, I used dropout quite extensively. 2d dropout was used within all of my residual blocks. Ordinary 1d dropout was used outside the residual blocks near the final output layers.\n\nMy final region proposal network uses color images and my final classification network uses greyscale images. I initially used greyscale images for both the region proposal network and the classification network because I didn't think the additional channels would be useful and wanted to conserve space. Near the end of the competition, I tried using color images for both networks. This slightly improved the region proposal network's output but worsened the classification network's accuracy due to overfitting.\n\n#**Failed attempt at using a language model for post-processing the detection results**\n\nI was very interested in using a language model to correct my image processing code's incorrect labels. I spent over a month trying different approaches for this but was unable to get it working well.\n\nTo get the characters into approximately the order in which a human would read them, I used DBSCAN to group the characters into columns, sorted the columns by the mean horizontal coordinates of their characters, and then sorted the characters within each column by their vertical coordinate. This worked well for most images.\n\nTo train my correction networks, I used the ground truth labels to generate a large number of synthetic submission files that contained known errors. These errors were fairly realistic and were randomly added based on statistics from how my Faster-RCNN based model performs on a cross-validation dataset.\n\nTo perform the character label corrections, I initially tried network architectures that are meant for performing translations (e.g. from English to Japanese), such as an encoder LSTM followed by a decoder LSTM with an attention mechanism, or a transformer network. I favored these network architectures over passing characters into a single recurrent network because in theory they should be able to elegantly handle cases in which the number of ground truth characters is different from the number of characters which was detected. Unfortunately, these networks introduced slightly more errors than they corrected. I eventually abandoned the translation based network architectures and switched over to using a simple bidirectional GRU network. It was able to improve my score by around .01, but only on submission files which were generated by weak image processing models. If I run it on my best submission file, then it will lower my score by ~.003.\n\nI think the reason I had so much difficulty getting a language model working well is that my Faster-RCNN based models do a better job considering a character's context than I initially expected. I think in order to get the correction network working well I would have needed to pull in an outside text corpus.\n\n#**Potential improvements**\n\nThere are several potential avenues for improvement that I did not implement or test.\n1. Data augmentation for the image processing models\n2. Pseudolabeling\n3. Using an ensemble of multiple classification networks\n4. Using an outside text corpus for the correction network\n\n**My code is available [here](https://github.com/jday96314/Kuzushiji).**",
      "votes": null
    },
    {
      "id": "653314",
      "postDate": "10/20/2019 08:25:23",
      "content": "<p>Your idea to create language model is very interesting. Though it didn't work out as you wanted,  i guess you learnt a lot  and also us too by your sharing of this experience.</p>",
      "rawMarkdown": "Your idea to create language model is very interesting. Though it didn't work out as you wanted,  i guess you learnt a lot  and also us too by your sharing of this experience.",
      "votes": null
    },
    {
      "id": "654372",
      "postDate": "10/21/2019 19:47:11",
      "content": "<p>Thanks for sharing your knowledge. I think you should think of yourself as a winner. Improvements should come in time. thx</p>",
      "rawMarkdown": "Thanks for sharing your knowledge. I think you should think of yourself as a winner. Improvements should come in time. thx",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 653314,
      "author_name": "khairulislam",
      "author_url": "",
      "post_date": "10/20/2019 08:25:23",
      "content": "<p>Your idea to create language model is very interesting. Though it didn't work out as you wanted,  i guess you learnt a lot  and also us too by your sharing of this experience.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 654372,
      "author_name": "zinovadr",
      "author_url": "",
      "post_date": "10/21/2019 19:47:11",
      "content": "<p>Thanks for sharing your knowledge. I think you should think of yourself as a winner. Improvements should come in time. thx</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "653194": "Thanks to Kaggle and the organizers for creating such an interesting competition. Congrats to everyone who finished.\n\n#**Data preprocessing**\n\nI resized all images to 512x512. Some of my images were converted to greyscale.\n\n#**Models**\n\nMy solution was based on Faster-RCNN, but with several small modifications.\n1. No layers are shared between the network which proposes regions of interest and the network which classifies them.  This allowed for easier experimentation because it allowed the networks to be modified independently of one another. It also simplified the training procedure. However, the inference efficiency of my solution could likely be improved by sharing all of the residual blocks.\n2.  Instead of choosing anchor sizes arbitrarily, I selected them by running k-means clustering on the widths and heights of the ground truth bounding boxes. I used the cluster centers as my anchor box sizes.\n3. I used ROI Align pooling instead of standard ROI pooling. I empirically found that ROI Align pooling yielded *slightly* better results.\n\nMy region proposal network and my character classification network both used wide ResNet-34 backbones. My region proposal network used twice as many convolutional filters as the ResNet-34 configuration described in [the original ResNet paper](https://arxiv.org/pdf/1512.03385.pdf). My character classification network used roughly three times as many convolutional filters as the original ResNet-34 configuration. I experimented with deeper ResNet-50 and ResNet-101 based networks, but I found wider, shallower networks worked better.\n\nTo combat overfitting, I used dropout quite extensively. 2d dropout was used within all of my residual blocks. Ordinary 1d dropout was used outside the residual blocks near the final output layers.\n\nMy final region proposal network uses color images and my final classification network uses greyscale images. I initially used greyscale images for both the region proposal network and the classification network because I didn't think the additional channels would be useful and wanted to conserve space. Near the end of the competition, I tried using color images for both networks. This slightly improved the region proposal network's output but worsened the classification network's accuracy due to overfitting.\n\n#**Failed attempt at using a language model for post-processing the detection results**\n\nI was very interested in using a language model to correct my image processing code's incorrect labels. I spent over a month trying different approaches for this but was unable to get it working well.\n\nTo get the characters into approximately the order in which a human would read them, I used DBSCAN to group the characters into columns, sorted the columns by the mean horizontal coordinates of their characters, and then sorted the characters within each column by their vertical coordinate. This worked well for most images.\n\nTo train my correction networks, I used the ground truth labels to generate a large number of synthetic submission files that contained known errors. These errors were fairly realistic and were randomly added based on statistics from how my Faster-RCNN based model performs on a cross-validation dataset.\n\nTo perform the character label corrections, I initially tried network architectures that are meant for performing translations (e.g. from English to Japanese), such as an encoder LSTM followed by a decoder LSTM with an attention mechanism, or a transformer network. I favored these network architectures over passing characters into a single recurrent network because in theory they should be able to elegantly handle cases in which the number of ground truth characters is different from the number of characters which was detected. Unfortunately, these networks introduced slightly more errors than they corrected. I eventually abandoned the translation based network architectures and switched over to using a simple bidirectional GRU network. It was able to improve my score by around .01, but only on submission files which were generated by weak image processing models. If I run it on my best submission file, then it will lower my score by ~.003.\n\nI think the reason I had so much difficulty getting a language model working well is that my Faster-RCNN based models do a better job considering a character's context than I initially expected. I think in order to get the correction network working well I would have needed to pull in an outside text corpus.\n\n#**Potential improvements**\n\nThere are several potential avenues for improvement that I did not implement or test.\n1. Data augmentation for the image processing models\n2. Pseudolabeling\n3. Using an ensemble of multiple classification networks\n4. Using an outside text corpus for the correction network\n\n**My code is available [here](https://github.com/jday96314/Kuzushiji).**",
    "653314": "Your idea to create language model is very interesting. Though it didn't work out as you wanted,  i guess you learnt a lot  and also us too by your sharing of this experience.",
    "654372": "Thanks for sharing your knowledge. I think you should think of yourself as a winner. Improvements should come in time. thx"
  },
  "source": "meta"
}