{
  "id": 112712,
  "title": "2nd place solution overview: detection + full-page classification",
  "url": "/competitions/kuzushiji-recognition/discussion/112712",
  "author_name": "Konstantin Lopukhin",
  "post_date": "2019-10-14T22:55:12.148000",
  "votes": 75,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Thanks to all organizers and Kaggle for such an interesting dataset and competition, and congrats to everyone who finished 🎉 </p>\n\n<p>Below is an overview of my solution. Code is at <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019\">https://github.com/lopuhin/kaggle-kuzushiji-2019</a>, model weights for the best model are at <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/releases/tag/v1.0\">https://github.com/lopuhin/kaggle-kuzushiji-2019/releases/tag/v1.0</a></p>\n\n<p>General approach is as follows:</p>\n\n<ul>\n<li>Dataset is split into 5 folds by book.</li>\n<li>Class-agnostic bounding boxes are predicted for all characters using an object detection network (with <code>resnet152</code> backbone pretrained on ImageNet). Out-of-fold predictions are obtained for all 5 folds.</li>\n<li>A \"classification\" model is trained using OOF detection predictions. An extra class <code>seg_fp</code> (segmentation false-positive) is added for bounding boxes which have low overlap with ground truth boxes, so classification model can correct errors of segmentation model. Classification model is trained on all folds. Models with <code>resnet152</code> and <code>resnext101_32x8d_wsl</code> backbones are used, they are trained on large crops containing multiple symbols, using FPN and roi align with a classification head.</li>\n<li>Pseudolabelling is performed, in OCR terms this is similar to \"writer adaptation\", although here it is applied to the whole test for simplicity.</li>\n<li>A second level model is trained on classification predictions,  which creates the final submission.</li>\n</ul>\n\n<p>Why such approach was chosen? There are two other candidate approaches:</p>\n\n<ul>\n<li>End-to-end model which does detection and classification (e.g. Faster-RCNN). This may be possible with some effort, but here it seems that segmentation is quite easy, while classification is hard, and it's more convenient to tune a classification model alone without worrying about detection, also pipeline is easier and more flexible.</li>\n<li>A separate detection model, and then a classifier on single-character crops. This is probably the easiest approach to get a reasonable result, and makes it very easy to improve a classification model. Still I felt that using larger crops as inputs should provide better context for the model, so that it can see nearby symbols and would not suffer from not ideal crops. But it could be that classification on character crops can be better.</li>\n</ul>\n\n<p>Next come more details on each stage.</p>\n\n<h2>Segmentation</h2>\n\n<p>Segmentation into characters is done with a Faster-RCNN model with <code>resnet152</code> backbone trained with torchvision. Only one class is used, so it does not try to predict the character class. This model trains very fast and gives high quality boxes. Competition F1 metric (assuming\nperfect prediction for the classes) was around ~0.99 on validation.</p>\n\n<p>Some details:</p>\n\n<ul>\n<li>torchvision detection pipeline was adapted,</li>\n<li><code>resnet152</code> backbone worked a bit better than default <code>resnet50</code> (even though it was not pre-trained on COCO, doing this would offer another small boost),</li>\n<li>pipeline was modified to accept empty crops (crops without ground truth objects) to reduce amount of false positives,</li>\n<li>it was trained on 512x384 crops, with page height around 1500 px, and full pages were used for inference,</li>\n<li>augmentations used: scale, minor color augmentations (hue/saturation/value), Albumentations library was used.</li>\n</ul>\n\n<p>Overall many more improvements are possible here: using mmdetection, better models, pre-training on COCO, blending predictions from different folds for submission, TTA, separate model to discard out-of-page symbols, etc. Still it seemed that classification was more important.</p>\n\n<p>This is implemented in <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/tree/master/kuzushiji/segment\">https://github.com/lopuhin/kaggle-kuzushiji-2019/tree/master/kuzushiji/segment</a> (which is based on reference torchvision detection code), dataset is defined in <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/segment/dataset.py\">https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/segment/dataset.py</a></p>\n\n<p>Here are validation F1 scores (assuming perfect class prediction) for resnet50 pretrained on COCO (orange), resnet152 pretrained on ImageNet (green), and same but with a better training schedule (blue, used for the submission).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F3fdbb6b475cd1c5efa06572460c2b783%2Fsegment-charts.jpg?generation=1572187112654283&amp;alt=media\" alt=\"\"></p>\n\n<h2>Classification</h2>\n\n<p>Classification is performed by a model which gets as input a large crop from the image (512x768 from 2500-3000px high image) which contains multiple characters. It also receives as input boudning boxes predicted by segmentation model (these are out-of-fold predictions). This is similar to multi-class detection, but with frozen bounding boxes. ResNet base is used, <code>layer4</code> is discarded, features are extracted for each bounding box with <code>roi_align</code> from <code>layer2</code> and <code>layer3</code> and concatenated, and then passed into a classification head.</p>\n\n<p>Some details:</p>\n\n<ul>\n<li>surprisingly, details such as architecture, backbone and learning regime made a lot of difference, much more than usual.</li>\n<li>head with two fully-connected layers and two 0.5 dropout layers was used, and all details were important: features from roi pooling were very high-dimensional (more than 13k), first layer reduced this to 1024, and second layer performed final classification. Adding more layers or removing intermediate bottleneck reduced quality.</li>\n<li>bigger backbones made a big difference, best model was the largest that could fit into 2080ti with a reasonable batch size: <code>resnext101_32x8d_wsl</code> from <a href=\"https://github.com/facebookresearch/WSL-Images\">https://github.com/facebookresearch/WSL-Images</a></li>\n<li>in order to train <code>resnext101_32x8d_wsl</code> on 2080ti, mixed precision training was required along with freezing first convolution and whole <code>layer1</code> (as I learned from Arthur Kuzin who did quite well in OpenImages, this is a trick used in mmdetection: <a href=\"https://github.com/open-mmlab/mmdetection/blob/6668bf0368b7ec6e88bc01aebdc281d2f79ef0cb/mmdet/models/backbones/resnet.py#L460\">https://github.com/open-mmlab/mmdetection/blob/6668bf0368b7ec6e88bc01aebdc281d2f79ef0cb/mmdet/models/backbones/resnet.py#L460</a>)</li>\n<li>another trick for reducing memory usage and making it train faster with cudnn.benchmark was limiting and bucketing number of targets in one batch.</li>\n<li>model was very sensitive to hyperparameters such as crop size and shape and batch size (and gradient accumulation wasn't enough to fix this).</li>\n<li>SGD with momentum performed significantly better than Adam, cosine schedule was used, weight decay was also quite important.</li>\n<li>quite large scale and color augmentations were used: hue/saturation/value, random brighness, contrast and gamma, all from Albumentations library.</li>\n<li>TTA (test-time-augmentation) of 4 different scales was used.</li>\n<li><code>resnext101_32x8d_wsl</code> took around 15 hours to train on one 2080ti.</li>\n</ul>\n\n<p>Best single model without pseudolabelling obtained public LB score of 0.935, although score varied quite a lot between folds, most folds were in 0.925 - 0.930 range. A blend of <code>resnet152</code> and <code>resnext101_32x8d_wsl</code> models across all folds scored 0.941 on the public LB.</p>\n\n<p>Overall, many improvement are possible here, from just using bigger models and freezing less layers, to more work on training schedule, augmentations, etc.</p>\n\n<p>See <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/main.py\">https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/main.py</a> for the training script, <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/models.py\">https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/models.py</a> for the models, and <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/dataset.py\">https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/dataset.py</a> for the dataset and augmentations.</p>\n\n<p>Here are validation F1 scores of several classification models (including psedulabeling, see below). <code>resnet50</code> is orange, <code>resnet152</code> is blue, <code>resnext101_32x8d_wsl</code> is violet, the same fine-tuned on pseduo-labels is red, and the same trained from scratch on pseduo-labels is green.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Ff06ac978504dcb33220f403a9b261ca8%2Fclassify-charts.jpg?generation=1572187271578401&amp;alt=media\" alt=\"\"></p>\n\n<h2>Pseudolabelling</h2>\n\n<p>Pseudolabelling is a technique where we take confident predictions of our model on test data, and add this to train dataset. Even though the model is already confident in such predictions, they are still useful and improve quality, because they allow the model to adapt better to different domain, as each book has it's own character and paper style, each author has different writing, etc.</p>\n\n<p>Here the simplest approach was chosen: most confident predictions were used for all test set, instead of splitting it by book. Top 80% most confident predictions from the blend were used, having accuracy &gt;99% according to validation. Next, two kinds of models were trained (all based on <code>resnext101_32x8d_wsl</code>):</p>\n\n<ul>\n<li>models from previous step fine-tuned for 5 epochs (compared to 50 epochs for training from scratch) with starting learning rate 10x smaller than initial learning rate.</li>\n<li>models trained from scratch with default settings.</li>\n</ul>\n\n<p>In both cases, models used both train and test data for training. Best single fine-tuned model scored 0.939 on the public LB. Best single from-scratch model scored 0.944 on the public LB (0.945 on private LB).</p>\n\n<p>See <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/pseudolabel.py\">https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/pseudolabel.py</a> for a script which filters confident prediction. This data is added in the regular classification train script.</p>\n\n<h2>Second level model</h2>\n\n<p>A simple blend worked already quite well, giving 0.943 public LB (without pseudolabelled from-scratch models). Adjusting coefficients of the models didn't improve the validation score, even though <code>resnext101_32x8d_wsl</code> models were noticeably better.</p>\n\n<p>Since all models were trained across all folds, it was possible to train a second level model, a blend of LightGBM and XGBoost. This model was inspired by Pavel Ostyakov's solution to Cdiscount’s Image Classification Challenge, which was a classification problem with 5k classes:\n<a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733</a></p>\n\n<p>Each of 4 model kinds from classification contributed classes and scores of top-3 predictions as features. Also max overlap with other bboxes was added. Then for each of all classes in top-3 predictions, and for a <code>seg_fp</code> class, we created one row with an extra feature <code>candidate</code>, which had a class as a value, and the target is binary: whether this candidate class was a true class which should be predicted. Then for each top-3 class, we added an extra binary feature which tells whether this class is a candidate class.</p>\n\n<p>Here is a simplified example with 1 model and top-2 predictions, all rows created for one character prediction (<code>seg_fp</code> was encoded as -1, <code>top0_s</code> means <code>top0_score</code>, <code>top0_is_c</code> means <code>top0_is_candidate</code>)::</p>\n\n<pre><code>top0_cls  top1_cls  top0_s  top1_s  candidate  top0_is_c  top1_is_c  y\n83        258       15.202  7.1246  83         True       False      True\n83        258       15.202  7.1246  258        False      True       False\n83        258       15.202  7.1246  -1         False      False      False\n</code></pre>\n\n<p>XGBoost and LighGBM models were trained across all folds, and then blended. It was better to first apply models to fold predictions on test and then blend them.</p>\n\n<p>Such blend gives 0.949 on public LB.</p>\n\n<p>I'm extremely bad at tuning such models, so there may be more improvements possible. Adjusting <code>seg_fp</code> ratio was tried and provided some boost on validation but didn't work on public LB.</p>\n\n<p>Second level features are defined in <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/level2_features.py\">https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/level2_features.py</a> and the models are built in <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/level2.py\">https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/level2.py</a></p>\n\n<h2>LB score summary </h2>\n\n<p>Scores on public LB (private LB scores are very well correlated):</p>\n\n<ul>\n<li><code>resnet50</code>, fold0: 0.916</li>\n<li><code>resnet152</code>, fold0: 0.926</li>\n<li><code>resnet152</code>, fold4: 0.932 (same model as above, best fold)</li>\n<li><code>resnext101_32x8d_wsl</code>, fold4: 0.934</li>\n<li><code>resnext101_32x8d_wsl</code>, fold4, fine-tuned with pseudo-labels: 0.939</li>\n<li><code>resnext101_32x8d_wsl</code>, fold4, re-trained with pseudo-labels: 0.944</li>\n<li>second-level model on top of all folds and models: 0.949</li>\n</ul>\n\n<h2>Discarded ideas</h2>\n\n<ul>\n<li>language model: a simple bi-LSTM language model was trained, but it achieved log loss of only ~4.5, while image-base model was at ~0.5, so it seemed that it would provide very little benefit. See <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/tree/master/kuzushiji/lm\">https://github.com/lopuhin/kaggle-kuzushiji-2019/tree/master/kuzushiji/lm</a></li>\n<li>kNN/metric learning: it's possible to use activations before the last layer as features, extract them from train and test, and then at inference time look closest (by cosine distance) example from train. This gave a minor boost over classification for single models, but inference time was quite high even with all optimizations, blending was less clear, so this was discarded. See <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/knn.py\">https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/knn.py</a></li>\n<li>adding an LSTM on top of the image model, where LSTM would work over symbols in the image using the same sequencing approach from the language model - I tried this only briefly but it worked worse than regular classification.</li>\n</ul>\n\n<h2>Running</h2>\n\n<p>See <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019#install\">https://github.com/lopuhin/kaggle-kuzushiji-2019#install</a> and <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019#run\">https://github.com/lopuhin/kaggle-kuzushiji-2019#run</a> and also the script which in theory contains all steps <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/run-all.sh\">https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/run-all.sh</a> but was not run in practice.</p>\n\n<h2>Hardware and libraries</h2>\n\n<p>Almost all models were trained on my home server with one 2080ti. <code>resnet152</code> classification models were trained on GCP with P100 GPUs as they required 16 GB of memory and I had some GCP credits. A few models towards the end were trained on vast.ai.</p>\n\n<p>All models are written with pytorch, detection models are based on torchvision. Apex is used for mixed precision training, and Albumentations for augmentations.</p>",
  "messages": [
    {
      "id": 649071,
      "postDate": "2019-10-14T22:55:12.150Z",
      "content": "<p>Thanks to all organizers and Kaggle for such an interesting dataset and competition, and congrats to everyone who finished 🎉 </p>\n\n<p>Below is an overview of my solution. Code is at <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019\">https://github.com/lopuhin/kaggle-kuzushiji-2019</a>, model weights for the best model are at <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/releases/tag/v1.0\">https://github.com/lopuhin/kaggle-kuzushiji-2019/releases/tag/v1.0</a></p>\n\n<p>General approach is as follows:</p>\n\n<ul>\n<li>Dataset is split into 5 folds by book.</li>\n<li>Class-agnostic bounding boxes are predicted for all characters using an object detection network (with <code>resnet152</code> backbone pretrained on ImageNet). Out-of-fold predictions are obtained for all 5 folds.</li>\n<li>A \"classification\" model is trained using OOF detection predictions. An extra class <code>seg_fp</code> (segmentation false-positive) is added for bounding boxes which have low overlap with ground truth boxes, so classification model can correct errors of segmentation model. Classification model is trained on all folds. Models with <code>resnet152</code> and <code>resnext101_32x8d_wsl</code> backbones are used, they are trained on large crops containing multiple symbols, using FPN and roi align with a classification head.</li>\n<li>Pseudolabelling is performed, in OCR terms this is similar to \"writer adaptation\", although here it is applied to the whole test for simplicity.</li>\n<li>A second level model is trained on classification predictions,  which creates the final submission.</li>\n</ul>\n\n<p>Why such approach was chosen? There are two other candidate approaches:</p>\n\n<ul>\n<li>End-to-end model which does detection and classification (e.g. Faster-RCNN). This may be possible with some effort, but here it seems that segmentation is quite easy, while classification is hard, and it's more convenient to tune a classification model alone without worrying about detection, also pipeline is easier and more flexible.</li>\n<li>A separate detection model, and then a classifier on single-character crops. This is probably the easiest approach to get a reasonable result, and makes it very easy to improve a classification model. Still I felt that using larger crops as inputs should provide better context for the model, so that it can see nearby symbols and would not suffer from not ideal crops. But it could be that classification on character crops can be better.</li>\n</ul>\n\n<p>Next come more details on each stage.</p>\n\n<h2>Segmentation</h2>\n\n<p>Segmentation into characters is done with a Faster-RCNN model with <code>resnet152</code> backbone trained with torchvision. Only one class is used, so it does not try to predict the character class. This model trains very fast and gives high quality boxes. Competition F1 metric (assuming\nperfect prediction for the classes) was around ~0.99 on validation.</p>\n\n<p>Some details:</p>\n\n<ul>\n<li>torchvision detection pipeline was adapted,</li>\n<li><code>resnet152</code> backbone worked a bit better than default <code>resnet50</code> (even though it was not pre-trained on COCO, doing this would offer another small boost),</li>\n<li>pipeline was modified to accept empty crops (crops without ground truth objects) to reduce amount of false positives,</li>\n<li>it was trained on 512x384 crops, with page height around 1500 px, and full pages were used for inference,</li>\n<li>augmentations used: scale, minor color augmentations (hue/saturation/value), Albumentations library was used.</li>\n</ul>\n\n<p>Overall many more improvements are possible here: using mmdetection, better models, pre-training on COCO, blending predictions from different folds for submission, TTA, separate model to discard out-of-page symbols, etc. Still it seemed that classification was more important.</p>\n\n<p>This is implemented in <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/tree/master/kuzushiji/segment\">https://github.com/lopuhin/kaggle-kuzushiji-2019/tree/master/kuzushiji/segment</a> (which is based on reference torchvision detection code), dataset is defined in <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/segment/dataset.py\">https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/segment/dataset.py</a></p>\n\n<p>Here are validation F1 scores (assuming perfect class prediction) for resnet50 pretrained on COCO (orange), resnet152 pretrained on ImageNet (green), and same but with a better training schedule (blue, used for the submission).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F3fdbb6b475cd1c5efa06572460c2b783%2Fsegment-charts.jpg?generation=1572187112654283&amp;alt=media\" alt=\"\"></p>\n\n<h2>Classification</h2>\n\n<p>Classification is performed by a model which gets as input a large crop from the image (512x768 from 2500-3000px high image) which contains multiple characters. It also receives as input boudning boxes predicted by segmentation model (these are out-of-fold predictions). This is similar to multi-class detection, but with frozen bounding boxes. ResNet base is used, <code>layer4</code> is discarded, features are extracted for each bounding box with <code>roi_align</code> from <code>layer2</code> and <code>layer3</code> and concatenated, and then passed into a classification head.</p>\n\n<p>Some details:</p>\n\n<ul>\n<li>surprisingly, details such as architecture, backbone and learning regime made a lot of difference, much more than usual.</li>\n<li>head with two fully-connected layers and two 0.5 dropout layers was used, and all details were important: features from roi pooling were very high-dimensional (more than 13k), first layer reduced this to 1024, and second layer performed final classification. Adding more layers or removing intermediate bottleneck reduced quality.</li>\n<li>bigger backbones made a big difference, best model was the largest that could fit into 2080ti with a reasonable batch size: <code>resnext101_32x8d_wsl</code> from <a href=\"https://github.com/facebookresearch/WSL-Images\">https://github.com/facebookresearch/WSL-Images</a></li>\n<li>in order to train <code>resnext101_32x8d_wsl</code> on 2080ti, mixed precision training was required along with freezing first convolution and whole <code>layer1</code> (as I learned from Arthur Kuzin who did quite well in OpenImages, this is a trick used in mmdetection: <a href=\"https://github.com/open-mmlab/mmdetection/blob/6668bf0368b7ec6e88bc01aebdc281d2f79ef0cb/mmdet/models/backbones/resnet.py#L460\">https://github.com/open-mmlab/mmdetection/blob/6668bf0368b7ec6e88bc01aebdc281d2f79ef0cb/mmdet/models/backbones/resnet.py#L460</a>)</li>\n<li>another trick for reducing memory usage and making it train faster with cudnn.benchmark was limiting and bucketing number of targets in one batch.</li>\n<li>model was very sensitive to hyperparameters such as crop size and shape and batch size (and gradient accumulation wasn't enough to fix this).</li>\n<li>SGD with momentum performed significantly better than Adam, cosine schedule was used, weight decay was also quite important.</li>\n<li>quite large scale and color augmentations were used: hue/saturation/value, random brighness, contrast and gamma, all from Albumentations library.</li>\n<li>TTA (test-time-augmentation) of 4 different scales was used.</li>\n<li><code>resnext101_32x8d_wsl</code> took around 15 hours to train on one 2080ti.</li>\n</ul>\n\n<p>Best single model without pseudolabelling obtained public LB score of 0.935, although score varied quite a lot between folds, most folds were in 0.925 - 0.930 range. A blend of <code>resnet152</code> and <code>resnext101_32x8d_wsl</code> models across all folds scored 0.941 on the public LB.</p>\n\n<p>Overall, many improvement are possible here, from just using bigger models and freezing less layers, to more work on training schedule, augmentations, etc.</p>\n\n<p>See <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/main.py\">https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/main.py</a> for the training script, <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/models.py\">https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/models.py</a> for the models, and <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/dataset.py\">https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/dataset.py</a> for the dataset and augmentations.</p>\n\n<p>Here are validation F1 scores of several classification models (including psedulabeling, see below). <code>resnet50</code> is orange, <code>resnet152</code> is blue, <code>resnext101_32x8d_wsl</code> is violet, the same fine-tuned on pseduo-labels is red, and the same trained from scratch on pseduo-labels is green.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Ff06ac978504dcb33220f403a9b261ca8%2Fclassify-charts.jpg?generation=1572187271578401&amp;alt=media\" alt=\"\"></p>\n\n<h2>Pseudolabelling</h2>\n\n<p>Pseudolabelling is a technique where we take confident predictions of our model on test data, and add this to train dataset. Even though the model is already confident in such predictions, they are still useful and improve quality, because they allow the model to adapt better to different domain, as each book has it's own character and paper style, each author has different writing, etc.</p>\n\n<p>Here the simplest approach was chosen: most confident predictions were used for all test set, instead of splitting it by book. Top 80% most confident predictions from the blend were used, having accuracy &gt;99% according to validation. Next, two kinds of models were trained (all based on <code>resnext101_32x8d_wsl</code>):</p>\n\n<ul>\n<li>models from previous step fine-tuned for 5 epochs (compared to 50 epochs for training from scratch) with starting learning rate 10x smaller than initial learning rate.</li>\n<li>models trained from scratch with default settings.</li>\n</ul>\n\n<p>In both cases, models used both train and test data for training. Best single fine-tuned model scored 0.939 on the public LB. Best single from-scratch model scored 0.944 on the public LB (0.945 on private LB).</p>\n\n<p>See <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/pseudolabel.py\">https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/pseudolabel.py</a> for a script which filters confident prediction. This data is added in the regular classification train script.</p>\n\n<h2>Second level model</h2>\n\n<p>A simple blend worked already quite well, giving 0.943 public LB (without pseudolabelled from-scratch models). Adjusting coefficients of the models didn't improve the validation score, even though <code>resnext101_32x8d_wsl</code> models were noticeably better.</p>\n\n<p>Since all models were trained across all folds, it was possible to train a second level model, a blend of LightGBM and XGBoost. This model was inspired by Pavel Ostyakov's solution to Cdiscount’s Image Classification Challenge, which was a classification problem with 5k classes:\n<a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733</a></p>\n\n<p>Each of 4 model kinds from classification contributed classes and scores of top-3 predictions as features. Also max overlap with other bboxes was added. Then for each of all classes in top-3 predictions, and for a <code>seg_fp</code> class, we created one row with an extra feature <code>candidate</code>, which had a class as a value, and the target is binary: whether this candidate class was a true class which should be predicted. Then for each top-3 class, we added an extra binary feature which tells whether this class is a candidate class.</p>\n\n<p>Here is a simplified example with 1 model and top-2 predictions, all rows created for one character prediction (<code>seg_fp</code> was encoded as -1, <code>top0_s</code> means <code>top0_score</code>, <code>top0_is_c</code> means <code>top0_is_candidate</code>)::</p>\n\n<pre><code>top0_cls  top1_cls  top0_s  top1_s  candidate  top0_is_c  top1_is_c  y\n83        258       15.202  7.1246  83         True       False      True\n83        258       15.202  7.1246  258        False      True       False\n83        258       15.202  7.1246  -1         False      False      False\n</code></pre>\n\n<p>XGBoost and LighGBM models were trained across all folds, and then blended. It was better to first apply models to fold predictions on test and then blend them.</p>\n\n<p>Such blend gives 0.949 on public LB.</p>\n\n<p>I'm extremely bad at tuning such models, so there may be more improvements possible. Adjusting <code>seg_fp</code> ratio was tried and provided some boost on validation but didn't work on public LB.</p>\n\n<p>Second level features are defined in <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/level2_features.py\">https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/level2_features.py</a> and the models are built in <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/level2.py\">https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/level2.py</a></p>\n\n<h2>LB score summary </h2>\n\n<p>Scores on public LB (private LB scores are very well correlated):</p>\n\n<ul>\n<li><code>resnet50</code>, fold0: 0.916</li>\n<li><code>resnet152</code>, fold0: 0.926</li>\n<li><code>resnet152</code>, fold4: 0.932 (same model as above, best fold)</li>\n<li><code>resnext101_32x8d_wsl</code>, fold4: 0.934</li>\n<li><code>resnext101_32x8d_wsl</code>, fold4, fine-tuned with pseudo-labels: 0.939</li>\n<li><code>resnext101_32x8d_wsl</code>, fold4, re-trained with pseudo-labels: 0.944</li>\n<li>second-level model on top of all folds and models: 0.949</li>\n</ul>\n\n<h2>Discarded ideas</h2>\n\n<ul>\n<li>language model: a simple bi-LSTM language model was trained, but it achieved log loss of only ~4.5, while image-base model was at ~0.5, so it seemed that it would provide very little benefit. See <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/tree/master/kuzushiji/lm\">https://github.com/lopuhin/kaggle-kuzushiji-2019/tree/master/kuzushiji/lm</a></li>\n<li>kNN/metric learning: it's possible to use activations before the last layer as features, extract them from train and test, and then at inference time look closest (by cosine distance) example from train. This gave a minor boost over classification for single models, but inference time was quite high even with all optimizations, blending was less clear, so this was discarded. See <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/knn.py\">https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/knn.py</a></li>\n<li>adding an LSTM on top of the image model, where LSTM would work over symbols in the image using the same sequencing approach from the language model - I tried this only briefly but it worked worse than regular classification.</li>\n</ul>\n\n<h2>Running</h2>\n\n<p>See <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019#install\">https://github.com/lopuhin/kaggle-kuzushiji-2019#install</a> and <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019#run\">https://github.com/lopuhin/kaggle-kuzushiji-2019#run</a> and also the script which in theory contains all steps <a href=\"https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/run-all.sh\">https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/run-all.sh</a> but was not run in practice.</p>\n\n<h2>Hardware and libraries</h2>\n\n<p>Almost all models were trained on my home server with one 2080ti. <code>resnet152</code> classification models were trained on GCP with P100 GPUs as they required 16 GB of memory and I had some GCP credits. A few models towards the end were trained on vast.ai.</p>\n\n<p>All models are written with pytorch, detection models are based on torchvision. Apex is used for mixed precision training, and Albumentations for augmentations.</p>",
      "rawMarkdown": "Thanks to all organizers and Kaggle for such an interesting dataset and competition, and congrats to everyone who finished 🎉 \n\nBelow is an overview of my solution. Code is at https://github.com/lopuhin/kaggle-kuzushiji-2019, model weights for the best model are at https://github.com/lopuhin/kaggle-kuzushiji-2019/releases/tag/v1.0\n\nGeneral approach is as follows:\n\n- Dataset is split into 5 folds by book.\n- Class-agnostic bounding boxes are predicted for all characters using an object detection network (with ``resnet152`` backbone pretrained on ImageNet). Out-of-fold predictions are obtained for all 5 folds.\n- A \"classification\" model is trained using OOF detection predictions. An extra class ``seg_fp`` (segmentation false-positive) is added for bounding boxes which have low overlap with ground truth boxes, so classification model can correct errors of segmentation model. Classification model is trained on all folds. Models with ``resnet152`` and ``resnext101_32x8d_wsl`` backbones are used, they are trained on large crops containing multiple symbols, using FPN and roi align with a classification head.\n- Pseudolabelling is performed, in OCR terms this is similar to \"writer adaptation\", although here it is applied to the whole test for simplicity.\n- A second level model is trained on classification predictions,  which creates the final submission.\n\nWhy such approach was chosen? There are two other candidate approaches:\n\n- End-to-end model which does detection and classification (e.g. Faster-RCNN). This may be possible with some effort, but here it seems that segmentation is quite easy, while classification is hard, and it's more convenient to tune a classification model alone without worrying about detection, also pipeline is easier and more flexible.\n- A separate detection model, and then a classifier on single-character crops. This is probably the easiest approach to get a reasonable result, and makes it very easy to improve a classification model. Still I felt that using larger crops as inputs should provide better context for the model, so that it can see nearby symbols and would not suffer from not ideal crops. But it could be that classification on character crops can be better.\n\nNext come more details on each stage.\n\nSegmentation\n------------\n\nSegmentation into characters is done with a Faster-RCNN model with ``resnet152`` backbone trained with torchvision. Only one class is used, so it does not try to predict the character class. This model trains very fast and gives high quality boxes. Competition F1 metric (assuming\nperfect prediction for the classes) was around ~0.99 on validation.\n\nSome details:\n\n* torchvision detection pipeline was adapted,\n* ``resnet152`` backbone worked a bit better than default ``resnet50`` (even though it was not pre-trained on COCO, doing this would offer another small boost),\n* pipeline was modified to accept empty crops (crops without ground truth objects) to reduce amount of false positives,\n* it was trained on 512x384 crops, with page height around 1500 px, and full pages were used for inference,\n* augmentations used: scale, minor color augmentations (hue/saturation/value), Albumentations library was used.\n\nOverall many more improvements are possible here: using mmdetection, better models, pre-training on COCO, blending predictions from different folds for submission, TTA, separate model to discard out-of-page symbols, etc. Still it seemed that classification was more important.\n\nThis is implemented in https://github.com/lopuhin/kaggle-kuzushiji-2019/tree/master/kuzushiji/segment (which is based on reference torchvision detection code), dataset is defined in https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/segment/dataset.py\n\nHere are validation F1 scores (assuming perfect class prediction) for resnet50 pretrained on COCO (orange), resnet152 pretrained on ImageNet (green), and same but with a better training schedule (blue, used for the submission).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F3fdbb6b475cd1c5efa06572460c2b783%2Fsegment-charts.jpg?generation=1572187112654283&amp;alt=media)\n\n\nClassification\n--------------\n\nClassification is performed by a model which gets as input a large crop from the image (512x768 from 2500-3000px high image) which contains multiple characters. It also receives as input boudning boxes predicted by segmentation model (these are out-of-fold predictions). This is similar to multi-class detection, but with frozen bounding boxes. ResNet base is used, ``layer4`` is discarded, features are extracted for each bounding box with ``roi_align`` from ``layer2`` and ``layer3`` and concatenated, and then passed into a classification head.\n\nSome details:\n\n* surprisingly, details such as architecture, backbone and learning regime made a lot of difference, much more than usual.\n* head with two fully-connected layers and two 0.5 dropout layers was used, and all details were important: features from roi pooling were very high-dimensional (more than 13k), first layer reduced this to 1024, and second layer performed final classification. Adding more layers or removing intermediate bottleneck reduced quality.\n* bigger backbones made a big difference, best model was the largest that could fit into 2080ti with a reasonable batch size: ``resnext101_32x8d_wsl`` from https://github.com/facebookresearch/WSL-Images\n* in order to train ``resnext101_32x8d_wsl`` on 2080ti, mixed precision training was required along with freezing first convolution and whole ``layer1`` (as I learned from Arthur Kuzin who did quite well in OpenImages, this is a trick used in mmdetection: https://github.com/open-mmlab/mmdetection/blob/6668bf0368b7ec6e88bc01aebdc281d2f79ef0cb/mmdet/models/backbones/resnet.py#L460)\n* another trick for reducing memory usage and making it train faster with cudnn.benchmark was limiting and bucketing number of targets in one batch.\n* model was very sensitive to hyperparameters such as crop size and shape and batch size (and gradient accumulation wasn't enough to fix this).\n* SGD with momentum performed significantly better than Adam, cosine schedule was used, weight decay was also quite important.\n* quite large scale and color augmentations were used: hue/saturation/value, random brighness, contrast and gamma, all from Albumentations library.\n* TTA (test-time-augmentation) of 4 different scales was used.\n* ``resnext101_32x8d_wsl`` took around 15 hours to train on one 2080ti.\n\nBest single model without pseudolabelling obtained public LB score of 0.935, although score varied quite a lot between folds, most folds were in 0.925 - 0.930 range. A blend of ``resnet152`` and ``resnext101_32x8d_wsl`` models across all folds scored 0.941 on the public LB.\n\nOverall, many improvement are possible here, from just using bigger models and freezing less layers, to more work on training schedule, augmentations, etc.\n\nSee https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/main.py for the training script, https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/models.py for the models, and https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/dataset.py for the dataset and augmentations.\n\nHere are validation F1 scores of several classification models (including psedulabeling, see below). ``resnet50`` is orange, ``resnet152`` is blue, ``resnext101_32x8d_wsl`` is violet, the same fine-tuned on pseduo-labels is red, and the same trained from scratch on pseduo-labels is green.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Ff06ac978504dcb33220f403a9b261ca8%2Fclassify-charts.jpg?generation=1572187271578401&amp;alt=media)\n\n\nPseudolabelling\n---------------\n\nPseudolabelling is a technique where we take confident predictions of our model on test data, and add this to train dataset. Even though the model is already confident in such predictions, they are still useful and improve quality, because they allow the model to adapt better to different domain, as each book has it's own character and paper style, each author has different writing, etc.\n\nHere the simplest approach was chosen: most confident predictions were used for all test set, instead of splitting it by book. Top 80% most confident predictions from the blend were used, having accuracy &gt;99% according to validation. Next, two kinds of models were trained (all based on ``resnext101_32x8d_wsl``):\n\n- models from previous step fine-tuned for 5 epochs (compared to 50 epochs for training from scratch) with starting learning rate 10x smaller than initial learning rate.\n- models trained from scratch with default settings.\n\nIn both cases, models used both train and test data for training. Best single fine-tuned model scored 0.939 on the public LB. Best single from-scratch model scored 0.944 on the public LB (0.945 on private LB).\n\nSee https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/pseudolabel.py for a script which filters confident prediction. This data is added in the regular classification train script.\n\nSecond level model\n------------------\n\nA simple blend worked already quite well, giving 0.943 public LB (without pseudolabelled from-scratch models). Adjusting coefficients of the models didn't improve the validation score, even though ``resnext101_32x8d_wsl`` models were noticeably better.\n\nSince all models were trained across all folds, it was possible to train a second level model, a blend of LightGBM and XGBoost. This model was inspired by Pavel Ostyakov's solution to Cdiscount’s Image Classification Challenge, which was a classification problem with 5k classes:\nhttps://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733\n\nEach of 4 model kinds from classification contributed classes and scores of top-3 predictions as features. Also max overlap with other bboxes was added. Then for each of all classes in top-3 predictions, and for a ``seg_fp`` class, we created one row with an extra feature ``candidate``, which had a class as a value, and the target is binary: whether this candidate class was a true class which should be predicted. Then for each top-3 class, we added an extra binary feature which tells whether this class is a candidate class.\n\nHere is a simplified example with 1 model and top-2 predictions, all rows created for one character prediction (``seg_fp`` was encoded as -1, ``top0_s`` means ``top0_score``, ``top0_is_c`` means ``top0_is_candidate``)::\n\n    top0_cls  top1_cls  top0_s  top1_s  candidate  top0_is_c  top1_is_c  y\n    83        258       15.202  7.1246  83         True       False      True\n    83        258       15.202  7.1246  258        False      True       False\n    83        258       15.202  7.1246  -1         False      False      False\n\nXGBoost and LighGBM models were trained across all folds, and then blended. It was better to first apply models to fold predictions on test and then blend them.\n\nSuch blend gives 0.949 on public LB.\n\nI'm extremely bad at tuning such models, so there may be more improvements possible. Adjusting ``seg_fp`` ratio was tried and provided some boost on validation but didn't work on public LB.\n\nSecond level features are defined in https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/level2_features.py and the models are built in https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/level2.py\n\nLB score summary \n------------------\n\nScores on public LB (private LB scores are very well correlated):\n\n- ``resnet50``, fold0: 0.916\n- ``resnet152``, fold0: 0.926\n- ``resnet152``, fold4: 0.932 (same model as above, best fold)\n- ``resnext101_32x8d_wsl``, fold4: 0.934\n- ``resnext101_32x8d_wsl``, fold4, fine-tuned with pseudo-labels: 0.939\n- ``resnext101_32x8d_wsl``, fold4, re-trained with pseudo-labels: 0.944\n- second-level model on top of all folds and models: 0.949\n\nDiscarded ideas\n---------------\n\n* language model: a simple bi-LSTM language model was trained, but it achieved log loss of only ~4.5, while image-base model was at ~0.5, so it seemed that it would provide very little benefit. See https://github.com/lopuhin/kaggle-kuzushiji-2019/tree/master/kuzushiji/lm\n* kNN/metric learning: it's possible to use activations before the last layer as features, extract them from train and test, and then at inference time look closest (by cosine distance) example from train. This gave a minor boost over classification for single models, but inference time was quite high even with all optimizations, blending was less clear, so this was discarded. See https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/knn.py\n* adding an LSTM on top of the image model, where LSTM would work over symbols in the image using the same sequencing approach from the language model - I tried this only briefly but it worked worse than regular classification.\n\nRunning\n----------\n\nSee https://github.com/lopuhin/kaggle-kuzushiji-2019#install and https://github.com/lopuhin/kaggle-kuzushiji-2019#run and also the script which in theory contains all steps https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/run-all.sh but was not run in practice.\n\nHardware and libraries\n----------------------\n\nAlmost all models were trained on my home server with one 2080ti. ``resnet152`` classification models were trained on GCP with P100 GPUs as they required 16 GB of memory and I had some GCP credits. A few models towards the end were trained on vast.ai.\n\nAll models are written with pytorch, detection models are based on torchvision. Apex is used for mixed precision training, and Albumentations for augmentations.",
      "votes": 75
    },
    {
      "id": 649111,
      "postDate": "2019-10-15T01:16:19.340Z",
      "content": "<p>Thank you so much for the solution!! And congratulations!</p>",
      "rawMarkdown": "Thank you so much for the solution!! And congratulations!",
      "votes": 5
    },
    {
      "id": 659378,
      "postDate": "2019-10-27T14:45:40.317Z",
      "content": "<p>Updated:</p>\n\n<ul>\n<li>added links to the source code for each section</li>\n<li>added some charts with validation F1 for different models</li>\n<li>added more single-model scores and a summary with LB scores at the end</li>\n<li>model weights</li>\n</ul>",
      "rawMarkdown": "Updated:\n\n- added links to the source code for each section\n- added some charts with validation F1 for different models\n- added more single-model scores and a summary with LB scores at the end\n- model weights",
      "votes": 1
    },
    {
      "id": 652660,
      "postDate": "2019-10-19T07:39:13.630Z",
      "content": "<p>Congratulations, and thank you for your sharing👍 </p>",
      "rawMarkdown": "Congratulations, and thank you for your sharing👍 ",
      "votes": 1
    },
    {
      "id": 651108,
      "postDate": "2019-10-17T03:44:40.253Z",
      "content": "<p>Thanks <a href=\"/lopuhin\">@lopuhin</a>  for sharing the great solution!  Especially, the approach of classification is very unique! I'm curious to know the effect of large crop which contains multiple characters. Have you tried the other crop size (smaller than 512x768)?</p>",
      "rawMarkdown": "Thanks @lopuhin  for sharing the great solution!  Especially, the approach of classification is very unique! I'm curious to know the effect of large crop which contains multiple characters. Have you tried the other crop size (smaller than 512x768)?",
      "votes": 1,
      "replies": [
        {
          "id": 651233,
          "postDate": "2019-10-17T07:58:15.110Z",
          "content": "<p>Thank you! Yes I tried smaller crop sizes and they performed a bit worse (e.g. I tried 512x512 and maybe smaller at the start), but note that changing crop size also changes the number of targets in one batch, which might require learning rate or batch size adjustment - this is one of the drawbacks of this approach, that crop size is not an independent hyperparameter, and I didn't do much tuning here.</p>",
          "rawMarkdown": "Thank you! Yes I tried smaller crop sizes and they performed a bit worse (e.g. I tried 512x512 and maybe smaller at the start), but note that changing crop size also changes the number of targets in one batch, which might require learning rate or batch size adjustment - this is one of the drawbacks of this approach, that crop size is not an independent hyperparameter, and I didn't do much tuning here.",
          "votes": 2
        },
        {
          "id": 651275,
          "postDate": "2019-10-17T09:17:52.023Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!",
          "votes": 1
        }
      ]
    },
    {
      "id": 654020,
      "postDate": "2019-10-21T10:15:15.783Z",
      "content": "<p>Congrats! 💪 🙌 </p>",
      "rawMarkdown": "Congrats! 💪 🙌 ",
      "votes": 2
    },
    {
      "id": 652408,
      "postDate": "2019-10-18T20:15:52.013Z",
      "content": "<p>Congrats on your scoring! Really enjoyed reading your solution, thank you for the work :)</p>",
      "rawMarkdown": "Congrats on your scoring! Really enjoyed reading your solution, thank you for the work :)",
      "votes": 2
    },
    {
      "id": 649445,
      "postDate": "2019-10-15T11:10:07.843Z",
      "content": "<p>Great solution! But I am a bit confused with your classification approach.\n1. The input is an image that may contain multiple characters, is this the output of the segmentation or a randomized cropping that takes bounding box information for each individual character present in the cropped image?\n2. Would it be different to take the bounding box of a character and input directly to a classification network to then extract features from <code>layer2</code> and <code>layer3</code>? </p>",
      "rawMarkdown": "Great solution! But I am a bit confused with your classification approach.\n1. The input is an image that may contain multiple characters, is this the output of the segmentation or a randomized cropping that takes bounding box information for each individual character present in the cropped image?\n2. Would it be different to take the bounding box of a character and input directly to a classification network to then extract features from `layer2` and `layer3`? \n",
      "votes": 2,
      "replies": [
        {
          "id": 649452,
          "postDate": "2019-10-15T11:19:42.697Z",
          "content": "<ol>\n<li>Yes, this is output of segmentation, which has a bounding box for each character. Alternatively it is possible to train on ground truth boxes, but I thought that using output from segmentation should be better as it is more similar to testing regime where bounding boxes are not perfect. No randomization is used here, although it could be beneficial, if I got your idea right.</li>\n<li>Yes, that would be different I think - if we take a crop around the character and pass into classification network, then the input would not contain any information about surrounding characters, while if we pass as input the whole image (or large crop) and then use roi pooling, then information about surrounding images would be used, because receptive field of convolutional network is already quite large.</li>\n</ol>",
          "rawMarkdown": "1. Yes, this is output of segmentation, which has a bounding box for each character. Alternatively it is possible to train on ground truth boxes, but I thought that using output from segmentation should be better as it is more similar to testing regime where bounding boxes are not perfect. No randomization is used here, although it could be beneficial, if I got your idea right.\n2. Yes, that would be different I think - if we take a crop around the character and pass into classification network, then the input would not contain any information about surrounding characters, while if we pass as input the whole image (or large crop) and then use roi pooling, then information about surrounding images would be used, because receptive field of convolutional network is already quite large.",
          "votes": 1
        },
        {
          "id": 649476,
          "postDate": "2019-10-15T12:11:17.787Z",
          "content": "<ol>\n<li>What I have been able to wrap my head around is that if a segmentation outputs individual bounding boxes for each detected character, how do you generate the cropped image that contains multiple characters?</li>\n</ol>",
          "rawMarkdown": "1. What I have been able to wrap my head around is that if a segmentation outputs individual bounding boxes for each detected character, how do you generate the cropped image that contains multiple characters?",
          "votes": 1
        },
        {
          "id": 649537,
          "postDate": "2019-10-15T14:00:32.010Z",
          "content": "<p>Sorry, I'm explaining this poorly - maybe this image would be helpful <a href=\"https://www.kaggle.com/c/kuzushiji-recognition/discussion/112204#647304\">https://www.kaggle.com/c/kuzushiji-recognition/discussion/112204#647304</a> - creating a crop containing multiple images is not a problem, what is more important, how the network is trained - it uses the same approach as detection networks use, but with fixed bounding boxes. So the image is passed through a backbone (resnet), and this backbone outputs a tensor which has lower spacial resolution but has more channels (2048). Then for each bounding box, we extract features from this tensor with roi align (a modification of roi pooling), and for each bbox these features are passed into a classification head.</p>",
          "rawMarkdown": "Sorry, I'm explaining this poorly - maybe this image would be helpful https://www.kaggle.com/c/kuzushiji-recognition/discussion/112204#647304 - creating a crop containing multiple images is not a problem, what is more important, how the network is trained - it uses the same approach as detection networks use, but with fixed bounding boxes. So the image is passed through a backbone (resnet), and this backbone outputs a tensor which has lower spacial resolution but has more channels (2048). Then for each bounding box, we extract features from this tensor with roi align (a modification of roi pooling), and for each bbox these features are passed into a classification head.",
          "votes": 2
        },
        {
          "id": 649618,
          "postDate": "2019-10-15T15:22:52.097Z",
          "content": "<p>Brilliant!</p>",
          "rawMarkdown": "Brilliant!",
          "votes": 1
        }
      ]
    },
    {
      "id": 649218,
      "postDate": "2019-10-15T04:55:33.787Z",
      "content": "<p>Congratulations\nGreat Write-Up\nThanks for Sharing Your Approach &amp; Insights, Code... <a href=\"/lopuhin\">@lopuhin</a> </p>",
      "rawMarkdown": "Congratulations\nGreat Write-Up\nThanks for Sharing Your Approach &amp; Insights, Code... @lopuhin ",
      "votes": 2
    },
    {
      "id": 649201,
      "postDate": "2019-10-15T04:09:35.603Z",
      "content": "<p>Thank you so much for sharing this. I'm a newbie in this area and so happy to hear great solutions of you guys. </p>\n\n<p>Mind asking you some basic concepts of the augmentation when a class contains less than 10 images? How did you determine and handle the number of images to augment? or you see it as an outlier and leave them as they be? </p>",
      "rawMarkdown": "Thank you so much for sharing this. I'm a newbie in this area and so happy to hear great solutions of you guys. \n\nMind asking you some basic concepts of the augmentation when a class contains less than 10 images? How did you determine and handle the number of images to augment? or you see it as an outlier and leave them as they be? ",
      "votes": 2,
      "replies": [
        {
          "id": 649333,
          "postDate": "2019-10-15T08:04:51.570Z",
          "content": "<p>Thank you! Great point regarding images with small number of samples - I didn't do anything special with them, this could be another potential area for improvement.\nRegarding augmentations, they are applied dynamically during training, and the sampling worked in this way: first I sampled a random page, and then I sampled a random location inside this page to make a crop. Augmentations at test time were applied to the whole page.</p>",
          "rawMarkdown": "Thank you! Great point regarding images with small number of samples - I didn't do anything special with them, this could be another potential area for improvement.\nRegarding augmentations, they are applied dynamically during training, and the sampling worked in this way: first I sampled a random page, and then I sampled a random location inside this page to make a crop. Augmentations at test time were applied to the whole page.",
          "votes": 1
        }
      ]
    },
    {
      "id": 649368,
      "postDate": "2019-10-15T08:52:04.470Z",
      "content": "<p><a href=\"/lopuhin\">@lopuhin</a> looks insane! I have to read it carefully and check the code.\nOur solution is very simple and we got over 0.9, but this is on a complete different level.</p>",
      "rawMarkdown": "@lopuhin looks insane! I have to read it carefully and check the code.\nOur solution is very simple and we got over 0.9, but this is on a complete different level.",
      "votes": 1,
      "replies": [
        {
          "id": 649371,
          "postDate": "2019-10-15T08:59:24.910Z",
          "content": "<p>Thanks! Looks forward to learning about your solution more, and for everyone else who would share more details.\nOne thing I was really unsure was (and still is), whether is would be better to just use single character crops for classification, maybe with some extra margin for context.</p>",
          "rawMarkdown": "Thanks! Looks forward to learning about your solution more, and for everyone else who would share more details.\nOne thing I was really unsure was (and still is), whether is would be better to just use single character crops for classification, maybe with some extra margin for context.",
          "votes": 2
        }
      ]
    },
    {
      "id": 675396,
      "postDate": "2019-11-18T03:05:42.730Z",
      "content": "<p>Hello, </p>\n\n<p>Organizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  </p>\n\n<p>Best, </p>\n\n<p>Alex.  </p>",
      "rawMarkdown": "Hello, \n\nOrganizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  \n\nBest, \n\nAlex.  ",
      "replies": [
        {
          "id": 675509,
          "postDate": "2019-11-18T06:30:24.787Z",
          "content": "<p>Hello <a href=\"/thenuttynetter\">@thenuttynetter</a> , \nHere are results for the best model, all times without TTA on 2080 ti, don't include image loading or resizing (but it would be quite small). All images are resized to max height 2528 px.</p>\n\n<ol>\n<li>Detection, using <code>resnet152</code>: ~230-40 ms per image (not using fp16).</li>\n<li>Classification, using <code>resnext101_32x8d_wsl</code>: ~210-220 ms per image (using fp16 although without native extensions).</li>\n</ol>\n\n<p>Overall time would be around 440-460 ms per image, without pre/post processing, and without TTA.</p>",
          "rawMarkdown": "\nHello @thenuttynetter , \nHere are results for the best model, all times without TTA on 2080 ti, don't include image loading or resizing (but it would be quite small). All images are resized to max height 2528 px.\n\n1. Detection, using `resnet152`: ~230-40 ms per image (not using fp16).\n2. Classification, using `resnext101_32x8d_wsl`: ~210-220 ms per image (using fp16 although without native extensions).\n\nOverall time would be around 440-460 ms per image, without pre/post processing, and without TTA.",
          "votes": 1
        }
      ]
    },
    {
      "id": 657348,
      "postDate": "2019-10-25T03:52:12.477Z",
      "content": "<p>Congrats!\nAmazing solution</p>",
      "rawMarkdown": "Congrats!\nAmazing solution"
    },
    {
      "id": 655985,
      "postDate": "2019-10-23T18:52:33.200Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 650690,
      "postDate": "2019-10-16T15:48:30.420Z",
      "content": "<p>Great solution, thank you!</p>",
      "rawMarkdown": "Great solution, thank you!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 649111,
      "author_name": "tkasasagi",
      "author_url": "",
      "post_date": "2019-10-15T01:16:19.340000",
      "content": "<p>Thank you so much for the solution!! And congratulations!</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 659378,
      "author_name": "Konstantin Lopukhin",
      "author_url": "",
      "post_date": "2019-10-27T14:45:40.317000",
      "content": "<p>Updated:</p>\n\n<ul>\n<li>added links to the source code for each section</li>\n<li>added some charts with validation F1 for different models</li>\n<li>added more single-model scores and a summary with LB scores at the end</li>\n<li>model weights</li>\n</ul>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 652660,
      "author_name": "Sharon",
      "author_url": "",
      "post_date": "2019-10-19T07:39:13.630000",
      "content": "<p>Congratulations, and thank you for your sharing👍 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 651108,
      "author_name": "K_mat",
      "author_url": "",
      "post_date": "2019-10-17T03:44:40.253000",
      "content": "<p>Thanks <a href=\"/lopuhin\">@lopuhin</a>  for sharing the great solution!  Especially, the approach of classification is very unique! I'm curious to know the effect of large crop which contains multiple characters. Have you tried the other crop size (smaller than 512x768)?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 651233,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-10-17T07:58:15.110000",
          "content": "<p>Thank you! Yes I tried smaller crop sizes and they performed a bit worse (e.g. I tried 512x512 and maybe smaller at the start), but note that changing crop size also changes the number of targets in one batch, which might require learning rate or batch size adjustment - this is one of the drawbacks of this approach, that crop size is not an independent hyperparameter, and I didn't do much tuning here.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 651275,
          "author_name": "K_mat",
          "author_url": "",
          "post_date": "2019-10-17T09:17:52.023000",
          "content": "<p>Thank you!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 654020,
      "author_name": "Iván de Prado",
      "author_url": "",
      "post_date": "2019-10-21T10:15:15.783000",
      "content": "<p>Congrats! 💪 🙌 </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 652408,
      "author_name": "Dmytro Danevskyi",
      "author_url": "",
      "post_date": "2019-10-18T20:15:52.013000",
      "content": "<p>Congrats on your scoring! Really enjoyed reading your solution, thank you for the work :)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 649445,
      "author_name": "LintangSutawika",
      "author_url": "",
      "post_date": "2019-10-15T11:10:07.843000",
      "content": "<p>Great solution! But I am a bit confused with your classification approach.\n1. The input is an image that may contain multiple characters, is this the output of the segmentation or a randomized cropping that takes bounding box information for each individual character present in the cropped image?\n2. Would it be different to take the bounding box of a character and input directly to a classification network to then extract features from <code>layer2</code> and <code>layer3</code>? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 649452,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-10-15T11:19:42.697000",
          "content": "<ol>\n<li>Yes, this is output of segmentation, which has a bounding box for each character. Alternatively it is possible to train on ground truth boxes, but I thought that using output from segmentation should be better as it is more similar to testing regime where bounding boxes are not perfect. No randomization is used here, although it could be beneficial, if I got your idea right.</li>\n<li>Yes, that would be different I think - if we take a crop around the character and pass into classification network, then the input would not contain any information about surrounding characters, while if we pass as input the whole image (or large crop) and then use roi pooling, then information about surrounding images would be used, because receptive field of convolutional network is already quite large.</li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 649476,
          "author_name": "LintangSutawika",
          "author_url": "",
          "post_date": "2019-10-15T12:11:17.787000",
          "content": "<ol>\n<li>What I have been able to wrap my head around is that if a segmentation outputs individual bounding boxes for each detected character, how do you generate the cropped image that contains multiple characters?</li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 649537,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-10-15T14:00:32.010000",
          "content": "<p>Sorry, I'm explaining this poorly - maybe this image would be helpful <a href=\"https://www.kaggle.com/c/kuzushiji-recognition/discussion/112204#647304\">https://www.kaggle.com/c/kuzushiji-recognition/discussion/112204#647304</a> - creating a crop containing multiple images is not a problem, what is more important, how the network is trained - it uses the same approach as detection networks use, but with fixed bounding boxes. So the image is passed through a backbone (resnet), and this backbone outputs a tensor which has lower spacial resolution but has more channels (2048). Then for each bounding box, we extract features from this tensor with roi align (a modification of roi pooling), and for each bbox these features are passed into a classification head.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 649618,
          "author_name": "LintangSutawika",
          "author_url": "",
          "post_date": "2019-10-15T15:22:52.097000",
          "content": "<p>Brilliant!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 649218,
      "author_name": "Ailurophile",
      "author_url": "",
      "post_date": "2019-10-15T04:55:33.787000",
      "content": "<p>Congratulations\nGreat Write-Up\nThanks for Sharing Your Approach &amp; Insights, Code... <a href=\"/lopuhin\">@lopuhin</a> </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 649201,
      "author_name": "Tanyapohn",
      "author_url": "",
      "post_date": "2019-10-15T04:09:35.603000",
      "content": "<p>Thank you so much for sharing this. I'm a newbie in this area and so happy to hear great solutions of you guys. </p>\n\n<p>Mind asking you some basic concepts of the augmentation when a class contains less than 10 images? How did you determine and handle the number of images to augment? or you see it as an outlier and leave them as they be? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 649333,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-10-15T08:04:51.570000",
          "content": "<p>Thank you! Great point regarding images with small number of samples - I didn't do anything special with them, this could be another potential area for improvement.\nRegarding augmentations, they are applied dynamically during training, and the sampling worked in this way: first I sampled a random page, and then I sampled a random location inside this page to make a crop. Augmentations at test time were applied to the whole page.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 649368,
      "author_name": "Nanashi",
      "author_url": "",
      "post_date": "2019-10-15T08:52:04.470000",
      "content": "<p><a href=\"/lopuhin\">@lopuhin</a> looks insane! I have to read it carefully and check the code.\nOur solution is very simple and we got over 0.9, but this is on a complete different level.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 649371,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-10-15T08:59:24.910000",
          "content": "<p>Thanks! Looks forward to learning about your solution more, and for everyone else who would share more details.\nOne thing I was really unsure was (and still is), whether is would be better to just use single character crops for classification, maybe with some extra margin for context.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 675396,
      "author_name": "TheNuttyNetter",
      "author_url": "",
      "post_date": "2019-11-18T03:05:42.730000",
      "content": "<p>Hello, </p>\n\n<p>Organizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  </p>\n\n<p>Best, </p>\n\n<p>Alex.  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 675509,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-11-18T06:30:24.787000",
          "content": "<p>Hello <a href=\"/thenuttynetter\">@thenuttynetter</a> , \nHere are results for the best model, all times without TTA on 2080 ti, don't include image loading or resizing (but it would be quite small). All images are resized to max height 2528 px.</p>\n\n<ol>\n<li>Detection, using <code>resnet152</code>: ~230-40 ms per image (not using fp16).</li>\n<li>Classification, using <code>resnext101_32x8d_wsl</code>: ~210-220 ms per image (using fp16 although without native extensions).</li>\n</ol>\n\n<p>Overall time would be around 440-460 ms per image, without pre/post processing, and without TTA.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 657348,
      "author_name": "wakas",
      "author_url": "",
      "post_date": "2019-10-25T03:52:12.477000",
      "content": "<p>Congrats!\nAmazing solution</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 655985,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-10-23T18:52:33.200000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 650690,
      "author_name": "Ma Yuxuan",
      "author_url": "",
      "post_date": "2019-10-16T15:48:30.420000",
      "content": "<p>Great solution, thank you!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "649071": "Thanks to all organizers and Kaggle for such an interesting dataset and competition, and congrats to everyone who finished 🎉 \n\nBelow is an overview of my solution. Code is at https://github.com/lopuhin/kaggle-kuzushiji-2019, model weights for the best model are at https://github.com/lopuhin/kaggle-kuzushiji-2019/releases/tag/v1.0\n\nGeneral approach is as follows:\n\n- Dataset is split into 5 folds by book.\n- Class-agnostic bounding boxes are predicted for all characters using an object detection network (with ``resnet152`` backbone pretrained on ImageNet). Out-of-fold predictions are obtained for all 5 folds.\n- A \"classification\" model is trained using OOF detection predictions. An extra class ``seg_fp`` (segmentation false-positive) is added for bounding boxes which have low overlap with ground truth boxes, so classification model can correct errors of segmentation model. Classification model is trained on all folds. Models with ``resnet152`` and ``resnext101_32x8d_wsl`` backbones are used, they are trained on large crops containing multiple symbols, using FPN and roi align with a classification head.\n- Pseudolabelling is performed, in OCR terms this is similar to \"writer adaptation\", although here it is applied to the whole test for simplicity.\n- A second level model is trained on classification predictions,  which creates the final submission.\n\nWhy such approach was chosen? There are two other candidate approaches:\n\n- End-to-end model which does detection and classification (e.g. Faster-RCNN). This may be possible with some effort, but here it seems that segmentation is quite easy, while classification is hard, and it's more convenient to tune a classification model alone without worrying about detection, also pipeline is easier and more flexible.\n- A separate detection model, and then a classifier on single-character crops. This is probably the easiest approach to get a reasonable result, and makes it very easy to improve a classification model. Still I felt that using larger crops as inputs should provide better context for the model, so that it can see nearby symbols and would not suffer from not ideal crops. But it could be that classification on character crops can be better.\n\nNext come more details on each stage.\n\nSegmentation\n------------\n\nSegmentation into characters is done with a Faster-RCNN model with ``resnet152`` backbone trained with torchvision. Only one class is used, so it does not try to predict the character class. This model trains very fast and gives high quality boxes. Competition F1 metric (assuming\nperfect prediction for the classes) was around ~0.99 on validation.\n\nSome details:\n\n* torchvision detection pipeline was adapted,\n* ``resnet152`` backbone worked a bit better than default ``resnet50`` (even though it was not pre-trained on COCO, doing this would offer another small boost),\n* pipeline was modified to accept empty crops (crops without ground truth objects) to reduce amount of false positives,\n* it was trained on 512x384 crops, with page height around 1500 px, and full pages were used for inference,\n* augmentations used: scale, minor color augmentations (hue/saturation/value), Albumentations library was used.\n\nOverall many more improvements are possible here: using mmdetection, better models, pre-training on COCO, blending predictions from different folds for submission, TTA, separate model to discard out-of-page symbols, etc. Still it seemed that classification was more important.\n\nThis is implemented in https://github.com/lopuhin/kaggle-kuzushiji-2019/tree/master/kuzushiji/segment (which is based on reference torchvision detection code), dataset is defined in https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/segment/dataset.py\n\nHere are validation F1 scores (assuming perfect class prediction) for resnet50 pretrained on COCO (orange), resnet152 pretrained on ImageNet (green), and same but with a better training schedule (blue, used for the submission).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2F3fdbb6b475cd1c5efa06572460c2b783%2Fsegment-charts.jpg?generation=1572187112654283&amp;alt=media)\n\n\nClassification\n--------------\n\nClassification is performed by a model which gets as input a large crop from the image (512x768 from 2500-3000px high image) which contains multiple characters. It also receives as input boudning boxes predicted by segmentation model (these are out-of-fold predictions). This is similar to multi-class detection, but with frozen bounding boxes. ResNet base is used, ``layer4`` is discarded, features are extracted for each bounding box with ``roi_align`` from ``layer2`` and ``layer3`` and concatenated, and then passed into a classification head.\n\nSome details:\n\n* surprisingly, details such as architecture, backbone and learning regime made a lot of difference, much more than usual.\n* head with two fully-connected layers and two 0.5 dropout layers was used, and all details were important: features from roi pooling were very high-dimensional (more than 13k), first layer reduced this to 1024, and second layer performed final classification. Adding more layers or removing intermediate bottleneck reduced quality.\n* bigger backbones made a big difference, best model was the largest that could fit into 2080ti with a reasonable batch size: ``resnext101_32x8d_wsl`` from https://github.com/facebookresearch/WSL-Images\n* in order to train ``resnext101_32x8d_wsl`` on 2080ti, mixed precision training was required along with freezing first convolution and whole ``layer1`` (as I learned from Arthur Kuzin who did quite well in OpenImages, this is a trick used in mmdetection: https://github.com/open-mmlab/mmdetection/blob/6668bf0368b7ec6e88bc01aebdc281d2f79ef0cb/mmdet/models/backbones/resnet.py#L460)\n* another trick for reducing memory usage and making it train faster with cudnn.benchmark was limiting and bucketing number of targets in one batch.\n* model was very sensitive to hyperparameters such as crop size and shape and batch size (and gradient accumulation wasn't enough to fix this).\n* SGD with momentum performed significantly better than Adam, cosine schedule was used, weight decay was also quite important.\n* quite large scale and color augmentations were used: hue/saturation/value, random brighness, contrast and gamma, all from Albumentations library.\n* TTA (test-time-augmentation) of 4 different scales was used.\n* ``resnext101_32x8d_wsl`` took around 15 hours to train on one 2080ti.\n\nBest single model without pseudolabelling obtained public LB score of 0.935, although score varied quite a lot between folds, most folds were in 0.925 - 0.930 range. A blend of ``resnet152`` and ``resnext101_32x8d_wsl`` models across all folds scored 0.941 on the public LB.\n\nOverall, many improvement are possible here, from just using bigger models and freezing less layers, to more work on training schedule, augmentations, etc.\n\nSee https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/main.py for the training script, https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/models.py for the models, and https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/dataset.py for the dataset and augmentations.\n\nHere are validation F1 scores of several classification models (including psedulabeling, see below). ``resnet50`` is orange, ``resnet152`` is blue, ``resnext101_32x8d_wsl`` is violet, the same fine-tuned on pseduo-labels is red, and the same trained from scratch on pseduo-labels is green.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F19390%2Ff06ac978504dcb33220f403a9b261ca8%2Fclassify-charts.jpg?generation=1572187271578401&amp;alt=media)\n\n\nPseudolabelling\n---------------\n\nPseudolabelling is a technique where we take confident predictions of our model on test data, and add this to train dataset. Even though the model is already confident in such predictions, they are still useful and improve quality, because they allow the model to adapt better to different domain, as each book has it's own character and paper style, each author has different writing, etc.\n\nHere the simplest approach was chosen: most confident predictions were used for all test set, instead of splitting it by book. Top 80% most confident predictions from the blend were used, having accuracy &gt;99% according to validation. Next, two kinds of models were trained (all based on ``resnext101_32x8d_wsl``):\n\n- models from previous step fine-tuned for 5 epochs (compared to 50 epochs for training from scratch) with starting learning rate 10x smaller than initial learning rate.\n- models trained from scratch with default settings.\n\nIn both cases, models used both train and test data for training. Best single fine-tuned model scored 0.939 on the public LB. Best single from-scratch model scored 0.944 on the public LB (0.945 on private LB).\n\nSee https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/pseudolabel.py for a script which filters confident prediction. This data is added in the regular classification train script.\n\nSecond level model\n------------------\n\nA simple blend worked already quite well, giving 0.943 public LB (without pseudolabelled from-scratch models). Adjusting coefficients of the models didn't improve the validation score, even though ``resnext101_32x8d_wsl`` models were noticeably better.\n\nSince all models were trained across all folds, it was possible to train a second level model, a blend of LightGBM and XGBoost. This model was inspired by Pavel Ostyakov's solution to Cdiscount’s Image Classification Challenge, which was a classification problem with 5k classes:\nhttps://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733\n\nEach of 4 model kinds from classification contributed classes and scores of top-3 predictions as features. Also max overlap with other bboxes was added. Then for each of all classes in top-3 predictions, and for a ``seg_fp`` class, we created one row with an extra feature ``candidate``, which had a class as a value, and the target is binary: whether this candidate class was a true class which should be predicted. Then for each top-3 class, we added an extra binary feature which tells whether this class is a candidate class.\n\nHere is a simplified example with 1 model and top-2 predictions, all rows created for one character prediction (``seg_fp`` was encoded as -1, ``top0_s`` means ``top0_score``, ``top0_is_c`` means ``top0_is_candidate``)::\n\n    top0_cls  top1_cls  top0_s  top1_s  candidate  top0_is_c  top1_is_c  y\n    83        258       15.202  7.1246  83         True       False      True\n    83        258       15.202  7.1246  258        False      True       False\n    83        258       15.202  7.1246  -1         False      False      False\n\nXGBoost and LighGBM models were trained across all folds, and then blended. It was better to first apply models to fold predictions on test and then blend them.\n\nSuch blend gives 0.949 on public LB.\n\nI'm extremely bad at tuning such models, so there may be more improvements possible. Adjusting ``seg_fp`` ratio was tried and provided some boost on validation but didn't work on public LB.\n\nSecond level features are defined in https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/level2_features.py and the models are built in https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/level2.py\n\nLB score summary \n------------------\n\nScores on public LB (private LB scores are very well correlated):\n\n- ``resnet50``, fold0: 0.916\n- ``resnet152``, fold0: 0.926\n- ``resnet152``, fold4: 0.932 (same model as above, best fold)\n- ``resnext101_32x8d_wsl``, fold4: 0.934\n- ``resnext101_32x8d_wsl``, fold4, fine-tuned with pseudo-labels: 0.939\n- ``resnext101_32x8d_wsl``, fold4, re-trained with pseudo-labels: 0.944\n- second-level model on top of all folds and models: 0.949\n\nDiscarded ideas\n---------------\n\n* language model: a simple bi-LSTM language model was trained, but it achieved log loss of only ~4.5, while image-base model was at ~0.5, so it seemed that it would provide very little benefit. See https://github.com/lopuhin/kaggle-kuzushiji-2019/tree/master/kuzushiji/lm\n* kNN/metric learning: it's possible to use activations before the last layer as features, extract them from train and test, and then at inference time look closest (by cosine distance) example from train. This gave a minor boost over classification for single models, but inference time was quite high even with all optimizations, blending was less clear, so this was discarded. See https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/kuzushiji/classify/knn.py\n* adding an LSTM on top of the image model, where LSTM would work over symbols in the image using the same sequencing approach from the language model - I tried this only briefly but it worked worse than regular classification.\n\nRunning\n----------\n\nSee https://github.com/lopuhin/kaggle-kuzushiji-2019#install and https://github.com/lopuhin/kaggle-kuzushiji-2019#run and also the script which in theory contains all steps https://github.com/lopuhin/kaggle-kuzushiji-2019/blob/master/run-all.sh but was not run in practice.\n\nHardware and libraries\n----------------------\n\nAlmost all models were trained on my home server with one 2080ti. ``resnet152`` classification models were trained on GCP with P100 GPUs as they required 16 GB of memory and I had some GCP credits. A few models towards the end were trained on vast.ai.\n\nAll models are written with pytorch, detection models are based on torchvision. Apex is used for mixed precision training, and Albumentations for augmentations.",
    "649111": "Thank you so much for the solution!! And congratulations!",
    "659378": "Updated:\n\n- added links to the source code for each section\n- added some charts with validation F1 for different models\n- added more single-model scores and a summary with LB scores at the end\n- model weights",
    "652660": "Congratulations, and thank you for your sharing👍 ",
    "651108": "Thanks @lopuhin  for sharing the great solution!  Especially, the approach of classification is very unique! I'm curious to know the effect of large crop which contains multiple characters. Have you tried the other crop size (smaller than 512x768)?",
    "654020": "Congrats! 💪 🙌 ",
    "652408": "Congrats on your scoring! Really enjoyed reading your solution, thank you for the work :)",
    "649445": "Great solution! But I am a bit confused with your classification approach.\n1. The input is an image that may contain multiple characters, is this the output of the segmentation or a randomized cropping that takes bounding box information for each individual character present in the cropped image?\n2. Would it be different to take the bounding box of a character and input directly to a classification network to then extract features from `layer2` and `layer3`? \n",
    "649218": "Congratulations\nGreat Write-Up\nThanks for Sharing Your Approach &amp; Insights, Code... @lopuhin ",
    "649201": "Thank you so much for sharing this. I'm a newbie in this area and so happy to hear great solutions of you guys. \n\nMind asking you some basic concepts of the augmentation when a class contains less than 10 images? How did you determine and handle the number of images to augment? or you see it as an outlier and leave them as they be? ",
    "649368": "@lopuhin looks insane! I have to read it carefully and check the code.\nOur solution is very simple and we got over 0.9, but this is on a complete different level.",
    "675396": "Hello, \n\nOrganizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  \n\nBest, \n\nAlex.  ",
    "657348": "Congrats!\nAmazing solution",
    "655985": "",
    "650690": "Great solution, thank you!"
  }
}