{
  "id": 113049,
  "title": "3rd place solution overview: 2-stage + FalsePositive Predictor",
  "url": "/competitions/kuzushiji-recognition/discussion/113049",
  "author_name": "kenji",
  "post_date": "2019-10-16T17:21:32.697000",
  "votes": 24,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Thanks to all for organizers this very interesting competition.\nCongratulations to all who finished the competition and to the winners.</p>\n\n<p>Below is an overview of my solution.</p>\n\n<p>Code: <a href=\"https://github.com/knjcode/kaggle-kuzushiji-recognition-2019\">https://github.com/knjcode/kaggle-kuzushiji-recognition-2019</a> <br>\nTrained weights: <a href=\"https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/releases/tag/0.0.1\">https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/releases/tag/0.0.1</a></p>\n\n<p>2-stage approach + FalsePositive Predictor.</p>\n\n<ul>\n<li>Detection with Faster R-CNN ResNet101 backbone</li>\n<li>Classification with 5 models with L2-constrained Softmax loss (EfficientNet, SE-ResNeXt101, ResNet152)</li>\n<li>Postprocessing (LightGBM FalsePositive Predictor)</li>\n</ul>\n\n<h2>Preprocess</h2>\n\n<p>Denoising and Ben's preprocessing for train and test images.</p>\n\n<p>Thanks for the following notebook <br>\n<a href=\"https://www.kaggle.com/hanmingliu/denoising-ben-s-preprocessing-better-clarity\">Denoising + Ben's Preprocessing = Better Clarity</a></p>\n\n<h2>Detection</h2>\n\n<p>Detection model use Faster R-CNN with:\n- ResNet101 backbone\n- Multi-scale train&amp;test\n- data augmentation (brightness, contrast, saturation, hue, random grayscale)\n- no vertical and horizontal flip</p>\n\n<p>use customized maskrcnn_benchmark</p>\n\n<p>Training Detection model with all train images. And validation with public Leaderboard score.</p>\n\n<p>See config for details. <a href=\"https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/blob/master/configs/kuzushiji/e2e_faster_rcnn_R_101_C4_1x_2_gpu_voc.yaml\">e2e_faster_rcnn_R_101_C4_1x_2_gpu_voc.yaml</a></p>\n\n<h3>Book title classifier</h3>\n\n<p>To know the trend of test data image distribution,\nI have trained a model to estimate the title of a book using train data.</p>\n\n<p>See training script for details train_book_title_classifier.sh</p>\n\n<p>The accuracy of trained model is about 99%.</p>\n\n<p>Then I estimated the book title for each image in the test data.</p>\n\n<p>The following is the estimation result of the book title of each image of the test data.\n<code>\n  count  book title\n      4  100241706\n      0  100249371\n     46  100249376\n     60  100249416\n      0  100249476\n      9  100249537\n    209  200003076\n    550  200003967\n      1  200004148\n      0  200005598\n    497  200006663\n    178  200014685\n    818  200014740\n     95  200015779\n      0  200021637\n      9  200021644\n     18  200021660\n      2  200021712\n    177  200021763\n      0  200021802\n     21  200021851\n      0  200021853\n    157  200021869\n      0  200021925\n     88  200022050\n   1210       brsk\n      9       hnsd\n      1       umgy\n</code></p>\n\n<p>Based on this result, I used the top 5 book tile above (brsk, 200014740, 200003967, 200006663, 200003076) for local validation.</p>\n\n<p>At the same time, I confirmed that there was almost no bias in public and private test data by submitting the recognition results with the specific book title omitted.</p>\n\n<h2>Classification</h2>\n\n<p>ensemble 5 classification models (hard voting)</p>\n\n<h3>validation strategy</h3>\n\n<p>Dataset is split train and validation by book titles.</p>\n\n<p>generate 2 patterns of training datasets.\n- validation: book_title=200015779 train: others\n- validation: book_title=200003076 train: others</p>\n\n<p>Characters with few occurrences oversampling at this time.</p>\n\n<p>For details, see <a href=\"https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/blob/master/scripts/gen_csv_denoised_pad_train_val.py\">gen_csv_denoised_pad_train_val.py</a></p>\n\n<h3>preprocessing and data augmentation</h3>\n\n<ul>\n<li>Preprocessing\n<ul><li>Denoising and Ben's preprocessing for train and test images, then crop characters</li>\n<li>When cropping each character, enlarge the area by 5% vertically and horizontally</li>\n<li>Resize each character to square ignoring aspect ratio</li>\n<li>To reduce computational resource, some models training with grayscale image</li>\n<li>To reduce computational resource, undersampling characters that appeared more than 2000 times have been undersampled to 2000 times</li></ul></li>\n<li>Data augmentation\n<ul><li>brightness, contrast, saturation, hue, random grayscale, rotate, random resize crop</li>\n<li>mixup + RandomErasing or ICAP + RandomErasing (details of ICAP will be described later)</li>\n<li>no vertical and horizontal flip</li></ul></li>\n<li>Others\n<ul><li>Use L2-constrained Softmax loss for all models</li>\n<li>I also tried AdaCos, ArcFace, CosFace, but L2-constrained Softmax was better.</li>\n<li>Warmup learning rate (5epochs)</li>\n<li>SGD + momentum Optimizer</li>\n<li>MultiStep LR or Cosine Annealing LR</li>\n<li>Test-Time-Augmentation 7 crop</li></ul></li>\n</ul>\n\n<p>For more details, see options of <a href=\"https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/blob/master/train_scripts/01_efficientnet_b4_val15779_l2softmax_mixup_re_normalize_gray190.sh\">model training scripts</a></p>\n\nmodels\n\n<p>|model arch     |channel  |input size|data augmentation    |validation|\n|:--------------|:--------|:---------|:--------------------|:---------|\n|EfficientNet-B4|Grayscale|190x190   |mixup + RandomErasing|200015779 |\n|ResNet152      |Grayscale|112x112   |mixup + RandomErasing|200015779 |\n|SE-ResNeXt101  |RGB      |112x112   |mixup + RandomErasing|200015779 |\n|SE-ResNeXt101  |RGB      |112x112   |ICAP + RandomErasing |200003076 |\n|ResNet152      |RGB      |112x112   |ICAP + RandomErasing |200003076 |</p>\n\nexample of mixup + RandomErasing\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F998412%2F0ec7eec68b6681e85d213ca1ebbf2086%2Fmixup_re.jpg?generation=1572017422079231&amp;alt=media\" alt=\"mixup_re\"></p>\n\nexample of ICAP + RandomErasing\n\n<p>ICAP is the data augmentation method I have implemented.\nImages cut from the four images are pasted while keeping the original image position.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F998412%2F10e723531d90441d54afa001a1961c88%2Ficap_re.jpg?generation=1572017457420689&amp;alt=media\" alt=\"icap_re\"></p>\n\n<h2>Pseudo labeling</h2>\n\n<p>Use the predicted result using the model trained in the previous steps as the pseudo label.</p>\n\n<h3>Retrian classification model</h3>\n\n<ul>\n<li>train with pseudo label</li>\n<li>train final 3 epoch without data augmentation</li>\n</ul>\n\n<p>|model arch     |local val acc|Private Leadearbord|\n|:--------------|:------------|:------------------|\n|EfficientNet-B4|0.9666       |0.941              |\n|ResNet152      |0.9635       |-                  |\n|SE-ResNeXt101  |0.9610       |0.934              |\n|SE-ResNeXt101  |0.9629       |-                  |\n|ResNet152      |0.9644       |-                  |</p>\n\n<h3>NMS with 2 ensemble results</h3>\n\n<p>NMS 2 snapshot result of Faster R-CNN object detetor (iter=60000, 100000)</p>\n\n<h2>Postprocessing (FalsePositive Predictor)</h2>\n\n<p>Train a classifier that predicts false positive results from validation results.</p>\n\n<p>Use classification and detection score(probability) and geometric relationship between bounding boxes as a feature.</p>\n\n<p>For details, see scripts/optuna_search_for_false_positive_detector.py and scripts/gen_false_positive_detector.py</p>\n\n<p>Using this predictor increased the score by about 0.002 ~ 0.004</p>\n\n<p>The following is a sample of false positive prediction result (The box predicted to be false positive is drawn with a red line).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F998412%2Ff164b6805292396a9ca159a79e34e32e%2Ffalse_positive_prediction_sample.jpg?generation=1571246958731731&amp;alt=media\" alt=\"postprocessing_sample\"></p>\n\n<h2>Ideas not yet tried</h2>\n\n<ul>\n<li>Using SoftTriple loss (<a href=\"https://arxiv.org/abs/1909.05235\">SoftTriple Loss: Deep Metric Learning Without Triplet Sampling</a>)\n<ul><li>SoftTriple loss to extend the SoftMax loss with multiple centers for each class.</li>\n<li>I think SoftTriple loss is especially effective in classifying the 変体仮名 (hentaigana).</li></ul></li>\n<li>Use the aspect ratio of characters for classification\n<ul><li>When recognizing characters, they are resized to a fixed size without considering the aspect ratio of each character. Aspect ratio also seems to be effective for character recognition, so it should be used as an input feature of the classification model.</li></ul></li>\n</ul>\n\n<h2>Hardware and libraries</h2>\n\n<p>GCP (V100x2) 128or256GB Memory 32Core</p>",
  "messages": [
    {
      "id": 650789,
      "postDate": "2019-10-16T17:21:32.697Z",
      "content": "<p>Thanks to all for organizers this very interesting competition.\nCongratulations to all who finished the competition and to the winners.</p>\n\n<p>Below is an overview of my solution.</p>\n\n<p>Code: <a href=\"https://github.com/knjcode/kaggle-kuzushiji-recognition-2019\">https://github.com/knjcode/kaggle-kuzushiji-recognition-2019</a> <br>\nTrained weights: <a href=\"https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/releases/tag/0.0.1\">https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/releases/tag/0.0.1</a></p>\n\n<p>2-stage approach + FalsePositive Predictor.</p>\n\n<ul>\n<li>Detection with Faster R-CNN ResNet101 backbone</li>\n<li>Classification with 5 models with L2-constrained Softmax loss (EfficientNet, SE-ResNeXt101, ResNet152)</li>\n<li>Postprocessing (LightGBM FalsePositive Predictor)</li>\n</ul>\n\n<h2>Preprocess</h2>\n\n<p>Denoising and Ben's preprocessing for train and test images.</p>\n\n<p>Thanks for the following notebook <br>\n<a href=\"https://www.kaggle.com/hanmingliu/denoising-ben-s-preprocessing-better-clarity\">Denoising + Ben's Preprocessing = Better Clarity</a></p>\n\n<h2>Detection</h2>\n\n<p>Detection model use Faster R-CNN with:\n- ResNet101 backbone\n- Multi-scale train&amp;test\n- data augmentation (brightness, contrast, saturation, hue, random grayscale)\n- no vertical and horizontal flip</p>\n\n<p>use customized maskrcnn_benchmark</p>\n\n<p>Training Detection model with all train images. And validation with public Leaderboard score.</p>\n\n<p>See config for details. <a href=\"https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/blob/master/configs/kuzushiji/e2e_faster_rcnn_R_101_C4_1x_2_gpu_voc.yaml\">e2e_faster_rcnn_R_101_C4_1x_2_gpu_voc.yaml</a></p>\n\n<h3>Book title classifier</h3>\n\n<p>To know the trend of test data image distribution,\nI have trained a model to estimate the title of a book using train data.</p>\n\n<p>See training script for details train_book_title_classifier.sh</p>\n\n<p>The accuracy of trained model is about 99%.</p>\n\n<p>Then I estimated the book title for each image in the test data.</p>\n\n<p>The following is the estimation result of the book title of each image of the test data.\n<code>\n  count  book title\n      4  100241706\n      0  100249371\n     46  100249376\n     60  100249416\n      0  100249476\n      9  100249537\n    209  200003076\n    550  200003967\n      1  200004148\n      0  200005598\n    497  200006663\n    178  200014685\n    818  200014740\n     95  200015779\n      0  200021637\n      9  200021644\n     18  200021660\n      2  200021712\n    177  200021763\n      0  200021802\n     21  200021851\n      0  200021853\n    157  200021869\n      0  200021925\n     88  200022050\n   1210       brsk\n      9       hnsd\n      1       umgy\n</code></p>\n\n<p>Based on this result, I used the top 5 book tile above (brsk, 200014740, 200003967, 200006663, 200003076) for local validation.</p>\n\n<p>At the same time, I confirmed that there was almost no bias in public and private test data by submitting the recognition results with the specific book title omitted.</p>\n\n<h2>Classification</h2>\n\n<p>ensemble 5 classification models (hard voting)</p>\n\n<h3>validation strategy</h3>\n\n<p>Dataset is split train and validation by book titles.</p>\n\n<p>generate 2 patterns of training datasets.\n- validation: book_title=200015779 train: others\n- validation: book_title=200003076 train: others</p>\n\n<p>Characters with few occurrences oversampling at this time.</p>\n\n<p>For details, see <a href=\"https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/blob/master/scripts/gen_csv_denoised_pad_train_val.py\">gen_csv_denoised_pad_train_val.py</a></p>\n\n<h3>preprocessing and data augmentation</h3>\n\n<ul>\n<li>Preprocessing\n<ul><li>Denoising and Ben's preprocessing for train and test images, then crop characters</li>\n<li>When cropping each character, enlarge the area by 5% vertically and horizontally</li>\n<li>Resize each character to square ignoring aspect ratio</li>\n<li>To reduce computational resource, some models training with grayscale image</li>\n<li>To reduce computational resource, undersampling characters that appeared more than 2000 times have been undersampled to 2000 times</li></ul></li>\n<li>Data augmentation\n<ul><li>brightness, contrast, saturation, hue, random grayscale, rotate, random resize crop</li>\n<li>mixup + RandomErasing or ICAP + RandomErasing (details of ICAP will be described later)</li>\n<li>no vertical and horizontal flip</li></ul></li>\n<li>Others\n<ul><li>Use L2-constrained Softmax loss for all models</li>\n<li>I also tried AdaCos, ArcFace, CosFace, but L2-constrained Softmax was better.</li>\n<li>Warmup learning rate (5epochs)</li>\n<li>SGD + momentum Optimizer</li>\n<li>MultiStep LR or Cosine Annealing LR</li>\n<li>Test-Time-Augmentation 7 crop</li></ul></li>\n</ul>\n\n<p>For more details, see options of <a href=\"https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/blob/master/train_scripts/01_efficientnet_b4_val15779_l2softmax_mixup_re_normalize_gray190.sh\">model training scripts</a></p>\n\nmodels\n\n<p>|model arch     |channel  |input size|data augmentation    |validation|\n|:--------------|:--------|:---------|:--------------------|:---------|\n|EfficientNet-B4|Grayscale|190x190   |mixup + RandomErasing|200015779 |\n|ResNet152      |Grayscale|112x112   |mixup + RandomErasing|200015779 |\n|SE-ResNeXt101  |RGB      |112x112   |mixup + RandomErasing|200015779 |\n|SE-ResNeXt101  |RGB      |112x112   |ICAP + RandomErasing |200003076 |\n|ResNet152      |RGB      |112x112   |ICAP + RandomErasing |200003076 |</p>\n\nexample of mixup + RandomErasing\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F998412%2F0ec7eec68b6681e85d213ca1ebbf2086%2Fmixup_re.jpg?generation=1572017422079231&amp;alt=media\" alt=\"mixup_re\"></p>\n\nexample of ICAP + RandomErasing\n\n<p>ICAP is the data augmentation method I have implemented.\nImages cut from the four images are pasted while keeping the original image position.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F998412%2F10e723531d90441d54afa001a1961c88%2Ficap_re.jpg?generation=1572017457420689&amp;alt=media\" alt=\"icap_re\"></p>\n\n<h2>Pseudo labeling</h2>\n\n<p>Use the predicted result using the model trained in the previous steps as the pseudo label.</p>\n\n<h3>Retrian classification model</h3>\n\n<ul>\n<li>train with pseudo label</li>\n<li>train final 3 epoch without data augmentation</li>\n</ul>\n\n<p>|model arch     |local val acc|Private Leadearbord|\n|:--------------|:------------|:------------------|\n|EfficientNet-B4|0.9666       |0.941              |\n|ResNet152      |0.9635       |-                  |\n|SE-ResNeXt101  |0.9610       |0.934              |\n|SE-ResNeXt101  |0.9629       |-                  |\n|ResNet152      |0.9644       |-                  |</p>\n\n<h3>NMS with 2 ensemble results</h3>\n\n<p>NMS 2 snapshot result of Faster R-CNN object detetor (iter=60000, 100000)</p>\n\n<h2>Postprocessing (FalsePositive Predictor)</h2>\n\n<p>Train a classifier that predicts false positive results from validation results.</p>\n\n<p>Use classification and detection score(probability) and geometric relationship between bounding boxes as a feature.</p>\n\n<p>For details, see scripts/optuna_search_for_false_positive_detector.py and scripts/gen_false_positive_detector.py</p>\n\n<p>Using this predictor increased the score by about 0.002 ~ 0.004</p>\n\n<p>The following is a sample of false positive prediction result (The box predicted to be false positive is drawn with a red line).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F998412%2Ff164b6805292396a9ca159a79e34e32e%2Ffalse_positive_prediction_sample.jpg?generation=1571246958731731&amp;alt=media\" alt=\"postprocessing_sample\"></p>\n\n<h2>Ideas not yet tried</h2>\n\n<ul>\n<li>Using SoftTriple loss (<a href=\"https://arxiv.org/abs/1909.05235\">SoftTriple Loss: Deep Metric Learning Without Triplet Sampling</a>)\n<ul><li>SoftTriple loss to extend the SoftMax loss with multiple centers for each class.</li>\n<li>I think SoftTriple loss is especially effective in classifying the 変体仮名 (hentaigana).</li></ul></li>\n<li>Use the aspect ratio of characters for classification\n<ul><li>When recognizing characters, they are resized to a fixed size without considering the aspect ratio of each character. Aspect ratio also seems to be effective for character recognition, so it should be used as an input feature of the classification model.</li></ul></li>\n</ul>\n\n<h2>Hardware and libraries</h2>\n\n<p>GCP (V100x2) 128or256GB Memory 32Core</p>",
      "rawMarkdown": "Thanks to all for organizers this very interesting competition.\nCongratulations to all who finished the competition and to the winners.\n\nBelow is an overview of my solution.\n\nCode: [https://github.com/knjcode/kaggle-kuzushiji-recognition-2019](https://github.com/knjcode/kaggle-kuzushiji-recognition-2019)  \nTrained weights: [https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/releases/tag/0.0.1](https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/releases/tag/0.0.1)\n\n2-stage approach + FalsePositive Predictor.\n\n- Detection with Faster R-CNN ResNet101 backbone\n- Classification with 5 models with L2-constrained Softmax loss (EfficientNet, SE-ResNeXt101, ResNet152)\n- Postprocessing (LightGBM FalsePositive Predictor)\n\n## Preprocess\n\nDenoising and Ben's preprocessing for train and test images.\n\nThanks for the following notebook  \n[Denoising + Ben's Preprocessing = Better Clarity](https://www.kaggle.com/hanmingliu/denoising-ben-s-preprocessing-better-clarity)\n\n\n## Detection\n\nDetection model use Faster R-CNN with:\n- ResNet101 backbone\n- Multi-scale train&amp;test\n- data augmentation (brightness, contrast, saturation, hue, random grayscale)\n- no vertical and horizontal flip\n\nuse customized [maskrcnn_benchmark](maskrcnn_benchmark/)\n\nTraining Detection model with all train images. And validation with public Leaderboard score.\n\nSee config for details. [e2e_faster_rcnn_R_101_C4_1x_2_gpu_voc.yaml](https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/blob/master/configs/kuzushiji/e2e_faster_rcnn_R_101_C4_1x_2_gpu_voc.yaml)\n\n### Book title classifier\n\nTo know the trend of test data image distribution,\nI have trained a model to estimate the title of a book using train data.\n\nSee training script for details [train_book_title_classifier.sh](train_scripts/train_book_title_classifier.sh)\n\nThe accuracy of trained model is about 99%.\n\nThen I estimated the book title for each image in the test data.\n\nThe following is the estimation result of the book title of each image of the test data.\n```\n  count  book title\n      4  100241706\n      0  100249371\n     46  100249376\n     60  100249416\n      0  100249476\n      9  100249537\n    209  200003076\n    550  200003967\n      1  200004148\n      0  200005598\n    497  200006663\n    178  200014685\n    818  200014740\n     95  200015779\n      0  200021637\n      9  200021644\n     18  200021660\n      2  200021712\n    177  200021763\n      0  200021802\n     21  200021851\n      0  200021853\n    157  200021869\n      0  200021925\n     88  200022050\n   1210       brsk\n      9       hnsd\n      1       umgy\n```\n\nBased on this result, I used the top 5 book tile above (brsk, 200014740, 200003967, 200006663, 200003076) for local validation.\n\nAt the same time, I confirmed that there was almost no bias in public and private test data by submitting the recognition results with the specific book title omitted.\n\n\n## Classification\n\nensemble 5 classification models (hard voting)\n\n### validation strategy\n\nDataset is split train and validation by book titles.\n\ngenerate 2 patterns of training datasets.\n- validation: book_title=200015779 train: others\n- validation: book_title=200003076 train: others\n\nCharacters with few occurrences oversampling at this time.\n\nFor details, see [gen_csv_denoised_pad_train_val.py](https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/blob/master/scripts/gen_csv_denoised_pad_train_val.py)\n\n\n### preprocessing and data augmentation\n\n- Preprocessing\n  - Denoising and Ben's preprocessing for train and test images, then crop characters\n  - When cropping each character, enlarge the area by 5% vertically and horizontally\n  - Resize each character to square ignoring aspect ratio\n  - To reduce computational resource, some models training with grayscale image\n  - To reduce computational resource, undersampling characters that appeared more than 2000 times have been undersampled to 2000 times\n- Data augmentation\n  - brightness, contrast, saturation, hue, random grayscale, rotate, random resize crop\n  - mixup + RandomErasing or ICAP + RandomErasing (details of ICAP will be described later)\n  - no vertical and horizontal flip\n- Others\n  - Use L2-constrained Softmax loss for all models\n    - I also tried AdaCos, ArcFace, CosFace, but L2-constrained Softmax was better.\n  - Warmup learning rate (5epochs)\n  - SGD + momentum Optimizer\n  - MultiStep LR or Cosine Annealing LR\n  - Test-Time-Augmentation 7 crop\n\nFor more details, see options of [model training scripts](https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/blob/master/train_scripts/01_efficientnet_b4_val15779_l2softmax_mixup_re_normalize_gray190.sh)\n\n\n#### models\n|model arch     |channel  |input size|data augmentation    |validation|\n|:--------------|:--------|:---------|:--------------------|:---------|\n|EfficientNet-B4|Grayscale|190x190   |mixup + RandomErasing|200015779 |\n|ResNet152      |Grayscale|112x112   |mixup + RandomErasing|200015779 |\n|SE-ResNeXt101  |RGB      |112x112   |mixup + RandomErasing|200015779 |\n|SE-ResNeXt101  |RGB      |112x112   |ICAP + RandomErasing |200003076 |\n|ResNet152      |RGB      |112x112   |ICAP + RandomErasing |200003076 |\n\n\n#### example of mixup + RandomErasing\n\n![mixup_re](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F998412%2F0ec7eec68b6681e85d213ca1ebbf2086%2Fmixup_re.jpg?generation=1572017422079231&amp;alt=media)\n\n\n#### example of ICAP + RandomErasing\n\nICAP is the data augmentation method I have implemented.\nImages cut from the four images are pasted while keeping the original image position.\n\n![icap_re](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F998412%2F10e723531d90441d54afa001a1961c88%2Ficap_re.jpg?generation=1572017457420689&amp;alt=media)\n\n\n\n## Pseudo labeling\n\nUse the predicted result using the model trained in the previous steps as the pseudo label.\n\n### Retrian classification model\n\n- train with pseudo label\n- train final 3 epoch without data augmentation\n\n|model arch     |local val acc|Private Leadearbord|\n|:--------------|:------------|:------------------|\n|EfficientNet-B4|0.9666       |0.941              |\n|ResNet152      |0.9635       |-                  |\n|SE-ResNeXt101  |0.9610       |0.934              |\n|SE-ResNeXt101  |0.9629       |-                  |\n|ResNet152      |0.9644       |-                  |\n\n\n### NMS with 2 ensemble results\n\nNMS 2 snapshot result of Faster R-CNN object detetor (iter=60000, 100000)\n\n\n## Postprocessing (FalsePositive Predictor)\n\nTrain a classifier that predicts false positive results from validation results.\n\nUse classification and detection score(probability) and geometric relationship between bounding boxes as a feature.\n\nFor details, see [scripts/optuna_search_for_false_positive_detector.py](scripts/optuna_search_for_false_positive_detector.py) and [scripts/gen_false_positive_detector.py](scripts/gen_false_positive_detector.py)\n\nUsing this predictor increased the score by about 0.002 ~ 0.004\n\nThe following is a sample of false positive prediction result (The box predicted to be false positive is drawn with a red line).\n\n![postprocessing_sample](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F998412%2Ff164b6805292396a9ca159a79e34e32e%2Ffalse_positive_prediction_sample.jpg?generation=1571246958731731&amp;alt=media)\n\n## Ideas not yet tried\n\n- Using SoftTriple loss ([SoftTriple Loss: Deep Metric Learning Without Triplet Sampling](https://arxiv.org/abs/1909.05235))\n  - SoftTriple loss to extend the SoftMax loss with multiple centers for each class.\n  - I think SoftTriple loss is especially effective in classifying the 変体仮名 (hentaigana).\n- Use the aspect ratio of characters for classification\n  - When recognizing characters, they are resized to a fixed size without considering the aspect ratio of each character. Aspect ratio also seems to be effective for character recognition, so it should be used as an input feature of the classification model.\n\n\n## Hardware and libraries\n\nGCP (V100x2) 128or256GB Memory 32Core\n",
      "votes": 24
    },
    {
      "id": 658023,
      "postDate": "2019-10-25T15:43:18.890Z",
      "content": "<p>Published implementation and trained weights.\nCode: <a href=\"https://github.com/knjcode/kaggle-kuzushiji-recognition-2019\">https://github.com/knjcode/kaggle-kuzushiji-recognition-2019</a>\nTrained weights: <a href=\"https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/releases\">https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/releases</a></p>",
      "rawMarkdown": "Published implementation and trained weights.\nCode: [https://github.com/knjcode/kaggle-kuzushiji-recognition-2019](https://github.com/knjcode/kaggle-kuzushiji-recognition-2019)\nTrained weights: [https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/releases](https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/releases)",
      "votes": 1
    },
    {
      "id": 675402,
      "postDate": "2019-11-18T03:09:47.170Z",
      "content": "<p>Hello, </p>\n\n<p>Organizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  </p>\n\n<p>Best, </p>\n\n<p>Alex.  </p>",
      "rawMarkdown": "Hello, \n\nOrganizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  \n\nBest, \n\nAlex.  "
    },
    {
      "id": 650983,
      "postDate": "2019-10-16T22:57:13.877Z",
      "content": "<p>Very cool!  Does the classification model use the whole page image (and I guess the RCNN features then) or just a crop of each image?  </p>",
      "rawMarkdown": "Very cool!  Does the classification model use the whole page image (and I guess the RCNN features then) or just a crop of each image?  ",
      "replies": [
        {
          "id": 650988,
          "postDate": "2019-10-16T23:23:08.200Z",
          "content": "<p>Thanks! Classification model uses just a crop of each image.</p>",
          "rawMarkdown": "Thanks! Classification model uses just a crop of each image."
        }
      ]
    },
    {
      "id": 651138,
      "postDate": "2019-10-17T04:55:55.533Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 658023,
      "author_name": "kenji",
      "author_url": "",
      "post_date": "2019-10-25T15:43:18.890000",
      "content": "<p>Published implementation and trained weights.\nCode: <a href=\"https://github.com/knjcode/kaggle-kuzushiji-recognition-2019\">https://github.com/knjcode/kaggle-kuzushiji-recognition-2019</a>\nTrained weights: <a href=\"https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/releases\">https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/releases</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 675402,
      "author_name": "TheNuttyNetter",
      "author_url": "",
      "post_date": "2019-11-18T03:09:47.170000",
      "content": "<p>Hello, </p>\n\n<p>Organizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  </p>\n\n<p>Best, </p>\n\n<p>Alex.  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 650983,
      "author_name": "TheNuttyNetter",
      "author_url": "",
      "post_date": "2019-10-16T22:57:13.877000",
      "content": "<p>Very cool!  Does the classification model use the whole page image (and I guess the RCNN features then) or just a crop of each image?  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 650988,
          "author_name": "kenji",
          "author_url": "",
          "post_date": "2019-10-16T23:23:08.200000",
          "content": "<p>Thanks! Classification model uses just a crop of each image.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 651138,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-10-17T04:55:55.533000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "650789": "Thanks to all for organizers this very interesting competition.\nCongratulations to all who finished the competition and to the winners.\n\nBelow is an overview of my solution.\n\nCode: [https://github.com/knjcode/kaggle-kuzushiji-recognition-2019](https://github.com/knjcode/kaggle-kuzushiji-recognition-2019)  \nTrained weights: [https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/releases/tag/0.0.1](https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/releases/tag/0.0.1)\n\n2-stage approach + FalsePositive Predictor.\n\n- Detection with Faster R-CNN ResNet101 backbone\n- Classification with 5 models with L2-constrained Softmax loss (EfficientNet, SE-ResNeXt101, ResNet152)\n- Postprocessing (LightGBM FalsePositive Predictor)\n\n## Preprocess\n\nDenoising and Ben's preprocessing for train and test images.\n\nThanks for the following notebook  \n[Denoising + Ben's Preprocessing = Better Clarity](https://www.kaggle.com/hanmingliu/denoising-ben-s-preprocessing-better-clarity)\n\n\n## Detection\n\nDetection model use Faster R-CNN with:\n- ResNet101 backbone\n- Multi-scale train&amp;test\n- data augmentation (brightness, contrast, saturation, hue, random grayscale)\n- no vertical and horizontal flip\n\nuse customized [maskrcnn_benchmark](maskrcnn_benchmark/)\n\nTraining Detection model with all train images. And validation with public Leaderboard score.\n\nSee config for details. [e2e_faster_rcnn_R_101_C4_1x_2_gpu_voc.yaml](https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/blob/master/configs/kuzushiji/e2e_faster_rcnn_R_101_C4_1x_2_gpu_voc.yaml)\n\n### Book title classifier\n\nTo know the trend of test data image distribution,\nI have trained a model to estimate the title of a book using train data.\n\nSee training script for details [train_book_title_classifier.sh](train_scripts/train_book_title_classifier.sh)\n\nThe accuracy of trained model is about 99%.\n\nThen I estimated the book title for each image in the test data.\n\nThe following is the estimation result of the book title of each image of the test data.\n```\n  count  book title\n      4  100241706\n      0  100249371\n     46  100249376\n     60  100249416\n      0  100249476\n      9  100249537\n    209  200003076\n    550  200003967\n      1  200004148\n      0  200005598\n    497  200006663\n    178  200014685\n    818  200014740\n     95  200015779\n      0  200021637\n      9  200021644\n     18  200021660\n      2  200021712\n    177  200021763\n      0  200021802\n     21  200021851\n      0  200021853\n    157  200021869\n      0  200021925\n     88  200022050\n   1210       brsk\n      9       hnsd\n      1       umgy\n```\n\nBased on this result, I used the top 5 book tile above (brsk, 200014740, 200003967, 200006663, 200003076) for local validation.\n\nAt the same time, I confirmed that there was almost no bias in public and private test data by submitting the recognition results with the specific book title omitted.\n\n\n## Classification\n\nensemble 5 classification models (hard voting)\n\n### validation strategy\n\nDataset is split train and validation by book titles.\n\ngenerate 2 patterns of training datasets.\n- validation: book_title=200015779 train: others\n- validation: book_title=200003076 train: others\n\nCharacters with few occurrences oversampling at this time.\n\nFor details, see [gen_csv_denoised_pad_train_val.py](https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/blob/master/scripts/gen_csv_denoised_pad_train_val.py)\n\n\n### preprocessing and data augmentation\n\n- Preprocessing\n  - Denoising and Ben's preprocessing for train and test images, then crop characters\n  - When cropping each character, enlarge the area by 5% vertically and horizontally\n  - Resize each character to square ignoring aspect ratio\n  - To reduce computational resource, some models training with grayscale image\n  - To reduce computational resource, undersampling characters that appeared more than 2000 times have been undersampled to 2000 times\n- Data augmentation\n  - brightness, contrast, saturation, hue, random grayscale, rotate, random resize crop\n  - mixup + RandomErasing or ICAP + RandomErasing (details of ICAP will be described later)\n  - no vertical and horizontal flip\n- Others\n  - Use L2-constrained Softmax loss for all models\n    - I also tried AdaCos, ArcFace, CosFace, but L2-constrained Softmax was better.\n  - Warmup learning rate (5epochs)\n  - SGD + momentum Optimizer\n  - MultiStep LR or Cosine Annealing LR\n  - Test-Time-Augmentation 7 crop\n\nFor more details, see options of [model training scripts](https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/blob/master/train_scripts/01_efficientnet_b4_val15779_l2softmax_mixup_re_normalize_gray190.sh)\n\n\n#### models\n|model arch     |channel  |input size|data augmentation    |validation|\n|:--------------|:--------|:---------|:--------------------|:---------|\n|EfficientNet-B4|Grayscale|190x190   |mixup + RandomErasing|200015779 |\n|ResNet152      |Grayscale|112x112   |mixup + RandomErasing|200015779 |\n|SE-ResNeXt101  |RGB      |112x112   |mixup + RandomErasing|200015779 |\n|SE-ResNeXt101  |RGB      |112x112   |ICAP + RandomErasing |200003076 |\n|ResNet152      |RGB      |112x112   |ICAP + RandomErasing |200003076 |\n\n\n#### example of mixup + RandomErasing\n\n![mixup_re](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F998412%2F0ec7eec68b6681e85d213ca1ebbf2086%2Fmixup_re.jpg?generation=1572017422079231&amp;alt=media)\n\n\n#### example of ICAP + RandomErasing\n\nICAP is the data augmentation method I have implemented.\nImages cut from the four images are pasted while keeping the original image position.\n\n![icap_re](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F998412%2F10e723531d90441d54afa001a1961c88%2Ficap_re.jpg?generation=1572017457420689&amp;alt=media)\n\n\n\n## Pseudo labeling\n\nUse the predicted result using the model trained in the previous steps as the pseudo label.\n\n### Retrian classification model\n\n- train with pseudo label\n- train final 3 epoch without data augmentation\n\n|model arch     |local val acc|Private Leadearbord|\n|:--------------|:------------|:------------------|\n|EfficientNet-B4|0.9666       |0.941              |\n|ResNet152      |0.9635       |-                  |\n|SE-ResNeXt101  |0.9610       |0.934              |\n|SE-ResNeXt101  |0.9629       |-                  |\n|ResNet152      |0.9644       |-                  |\n\n\n### NMS with 2 ensemble results\n\nNMS 2 snapshot result of Faster R-CNN object detetor (iter=60000, 100000)\n\n\n## Postprocessing (FalsePositive Predictor)\n\nTrain a classifier that predicts false positive results from validation results.\n\nUse classification and detection score(probability) and geometric relationship between bounding boxes as a feature.\n\nFor details, see [scripts/optuna_search_for_false_positive_detector.py](scripts/optuna_search_for_false_positive_detector.py) and [scripts/gen_false_positive_detector.py](scripts/gen_false_positive_detector.py)\n\nUsing this predictor increased the score by about 0.002 ~ 0.004\n\nThe following is a sample of false positive prediction result (The box predicted to be false positive is drawn with a red line).\n\n![postprocessing_sample](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F998412%2Ff164b6805292396a9ca159a79e34e32e%2Ffalse_positive_prediction_sample.jpg?generation=1571246958731731&amp;alt=media)\n\n## Ideas not yet tried\n\n- Using SoftTriple loss ([SoftTriple Loss: Deep Metric Learning Without Triplet Sampling](https://arxiv.org/abs/1909.05235))\n  - SoftTriple loss to extend the SoftMax loss with multiple centers for each class.\n  - I think SoftTriple loss is especially effective in classifying the 変体仮名 (hentaigana).\n- Use the aspect ratio of characters for classification\n  - When recognizing characters, they are resized to a fixed size without considering the aspect ratio of each character. Aspect ratio also seems to be effective for character recognition, so it should be used as an input feature of the classification model.\n\n\n## Hardware and libraries\n\nGCP (V100x2) 128or256GB Memory 32Core\n",
    "658023": "Published implementation and trained weights.\nCode: [https://github.com/knjcode/kaggle-kuzushiji-recognition-2019](https://github.com/knjcode/kaggle-kuzushiji-recognition-2019)\nTrained weights: [https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/releases](https://github.com/knjcode/kaggle-kuzushiji-recognition-2019/releases)",
    "675402": "Hello, \n\nOrganizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  \n\nBest, \n\nAlex.  ",
    "650983": "Very cool!  Does the classification model use the whole page image (and I guess the RCNN features then) or just a crop of each image?  ",
    "651138": ""
  }
}