{
  "id": 113419,
  "title": "8th place solution: Two stage & kuzushiji data augmentation",
  "url": "/competitions/kuzushiji-recognition/writeups/t-hanya-8th-place-solution-two-stage-kuzushiji-dat",
  "author_name": "",
  "post_date": "2019-10-19T10:57:20.040Z",
  "votes": 14,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thanks to organizers and Kaggle for such an interesting competition and congrats to everyone who enjoyed this challenge!</p>\n\n<p>My approach is two stage pipeline. I used CenterNet [1] for character detection, and MobileNetV3 [2] for classification. The overview of my approach is shown in the figure below. All models were trained with single GTX 970 GPU installed on my home server, so my solution is relative resource efficient.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2Fb352a774b803ab9abedbd8d1b4333437%2F2019-10-19%2013.51.21.png?generation=1571478495979958&amp;alt=media\" alt=\"\"></p>\n\n<h2>Detection</h2>\n\n<p>I chose CenterNet because of its anchor-free simple design. Kuzushiji characters have wide aspect ratio range and relative small to page size, so it seemed to be difficult to find good anchor box setting. </p>\n\n<ul>\n<li>Architecture: ResNet18 + U-Net</li>\n<li>Data: Full training set to train single model</li>\n<li>TTA: Scale adjustment -&gt; Multi-scale + Bounding box voting (bbox voting is proposed in [3])</li>\n</ul>\n\n<h2>Classification</h2>\n\n<p>I combined various data augmentation to improve accuracy.  (black grid lines are only for visualization)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2Fcacf64aafe6e31d4e81bb381fd3e9d65%2F2019-10-19%2016.44.51.png?generation=1571478553323751&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2F47aa0cd704c01ebf1bf7d5466f10d669%2F2019-10-19%2016.44.59.png?generation=1571478576515188&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2F0718a01671e0ab60beb6681908f24b42%2F2019-10-19%2016.47.18.png?generation=1571478604246238&amp;alt=media\" alt=\"\"></p>\n\n<p>Other normal augmentation operations such as color, brightness, contrast, rotation and noise are also applied. Actual input images for training are as follows. These operations are implemented with <code>albumentations</code> library.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2F3194d2fb645ce38501df86ef26596315%2F2019-10-19%2016.53.23.png?generation=1571478627485652&amp;alt=media\" alt=\"\"></p>\n\n<ul>\n<li>Architecture: MobileNetV3 Large</li>\n<li>Loss: Standard softmax cross entropy loss</li>\n<li>Training: Use full training set to train single model</li>\n<li>Fine-tuning: Use full training set + pseudo label (filtered by score &gt; 0.9) with reduced data augmentation.</li>\n<li>No TTA (I tried multi-scale inference, but its contribution was very small.)</li>\n</ul>\n\n<h2>Local validation</h2>\n\n<p>As discussed, group splitting by book title is important. The appearance of characters varies largely depending on the book. I used group split by book title for local validation to select hyper parameters, data augmentation types and threshold values for testing. I set the score threshold for detector to be 0.3 (&lt;- default = 0.5) to reduce false negatives predictions, and score threshold for classification to be 0.5 to reduce false positive predictions.</p>\n\n<h2>Discarded ideas</h2>\n\n<ul>\n<li>[Classifier] Oversampling of minor class and class balanced loss [4]. I tried these techniques to handle the class imbalance problem, but its contribution was very small.</li>\n<li>[Classifier] Sequence based training to use context (Reading order prediction -&gt; CNN + Bi-LSTM), but did not work.</li>\n<li>[Detector/Classifier] Collect large amount of unlabeled images from CODH web site, and use it for representation learning or semi-supervised learning. I gave up this idea because it requires huge computational resources and disk spaces.</li>\n</ul>\n\n<h2>Code</h2>\n\n<p>All models were written with Chainer. Please check my github repo for more details.</p>\n\n<ul>\n<li><a href=\"https://github.com/t-hanya/kuzushiji-recognition\">https://github.com/t-hanya/kuzushiji-recognition</a></li>\n</ul>\n\n<h2>References</h2>\n\n<ul>\n<li>[1] Objects as Points: <a href=\"https://arxiv.org/abs/1904.07850\">https://arxiv.org/abs/1904.07850</a></li>\n<li>[2] Searching for MobileNetV3: <a href=\"https://arxiv.org/abs/1905.02244\">https://arxiv.org/abs/1905.02244</a></li>\n<li>[3] Object detection via a multi-region &amp; semantic segmentation-aware CNN model: <a href=\"https://arxiv.org/abs/1505.01749\">https://arxiv.org/abs/1505.01749</a></li>\n<li>[4] Class-Balanced Loss Based on Effective Number of Samples: <a href=\"https://arxiv.org/abs/1901.05555\">https://arxiv.org/abs/1901.05555</a></li>\n</ul>",
  "messages": [
    {
      "id": "652727",
      "postDate": "10/19/2019 10:25:03",
      "content": "<p>Thanks to organizers and Kaggle for such an interesting competition and congrats to everyone who enjoyed this challenge!</p>\n\n<p>My approach is two stage pipeline. I used CenterNet [1] for character detection, and MobileNetV3 [2] for classification. The overview of my approach is shown in the figure below. All models were trained with single GTX 970 GPU installed on my home server, so my solution is relative resource efficient.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2Fb352a774b803ab9abedbd8d1b4333437%2F2019-10-19%2013.51.21.png?generation=1571478495979958&amp;alt=media\" alt=\"\"></p>\n\n<h2>Detection</h2>\n\n<p>I chose CenterNet because of its anchor-free simple design. Kuzushiji characters have wide aspect ratio range and relative small to page size, so it seemed to be difficult to find good anchor box setting. </p>\n\n<ul>\n<li>Architecture: ResNet18 + U-Net</li>\n<li>Data: Full training set to train single model</li>\n<li>TTA: Scale adjustment -&gt; Multi-scale + Bounding box voting (bbox voting is proposed in [3])</li>\n</ul>\n\n<h2>Classification</h2>\n\n<p>I combined various data augmentation to improve accuracy.  (black grid lines are only for visualization)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2Fcacf64aafe6e31d4e81bb381fd3e9d65%2F2019-10-19%2016.44.51.png?generation=1571478553323751&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2F47aa0cd704c01ebf1bf7d5466f10d669%2F2019-10-19%2016.44.59.png?generation=1571478576515188&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2F0718a01671e0ab60beb6681908f24b42%2F2019-10-19%2016.47.18.png?generation=1571478604246238&amp;alt=media\" alt=\"\"></p>\n\n<p>Other normal augmentation operations such as color, brightness, contrast, rotation and noise are also applied. Actual input images for training are as follows. These operations are implemented with <code>albumentations</code> library.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2F3194d2fb645ce38501df86ef26596315%2F2019-10-19%2016.53.23.png?generation=1571478627485652&amp;alt=media\" alt=\"\"></p>\n\n<ul>\n<li>Architecture: MobileNetV3 Large</li>\n<li>Loss: Standard softmax cross entropy loss</li>\n<li>Training: Use full training set to train single model</li>\n<li>Fine-tuning: Use full training set + pseudo label (filtered by score &gt; 0.9) with reduced data augmentation.</li>\n<li>No TTA (I tried multi-scale inference, but its contribution was very small.)</li>\n</ul>\n\n<h2>Local validation</h2>\n\n<p>As discussed, group splitting by book title is important. The appearance of characters varies largely depending on the book. I used group split by book title for local validation to select hyper parameters, data augmentation types and threshold values for testing. I set the score threshold for detector to be 0.3 (&lt;- default = 0.5) to reduce false negatives predictions, and score threshold for classification to be 0.5 to reduce false positive predictions.</p>\n\n<h2>Discarded ideas</h2>\n\n<ul>\n<li>[Classifier] Oversampling of minor class and class balanced loss [4]. I tried these techniques to handle the class imbalance problem, but its contribution was very small.</li>\n<li>[Classifier] Sequence based training to use context (Reading order prediction -&gt; CNN + Bi-LSTM), but did not work.</li>\n<li>[Detector/Classifier] Collect large amount of unlabeled images from CODH web site, and use it for representation learning or semi-supervised learning. I gave up this idea because it requires huge computational resources and disk spaces.</li>\n</ul>\n\n<h2>Code</h2>\n\n<p>All models were written with Chainer. Please check my github repo for more details.</p>\n\n<ul>\n<li><a href=\"https://github.com/t-hanya/kuzushiji-recognition\">https://github.com/t-hanya/kuzushiji-recognition</a></li>\n</ul>\n\n<h2>References</h2>\n\n<ul>\n<li>[1] Objects as Points: <a href=\"https://arxiv.org/abs/1904.07850\">https://arxiv.org/abs/1904.07850</a></li>\n<li>[2] Searching for MobileNetV3: <a href=\"https://arxiv.org/abs/1905.02244\">https://arxiv.org/abs/1905.02244</a></li>\n<li>[3] Object detection via a multi-region &amp; semantic segmentation-aware CNN model: <a href=\"https://arxiv.org/abs/1505.01749\">https://arxiv.org/abs/1505.01749</a></li>\n<li>[4] Class-Balanced Loss Based on Effective Number of Samples: <a href=\"https://arxiv.org/abs/1901.05555\">https://arxiv.org/abs/1901.05555</a></li>\n</ul>",
      "rawMarkdown": "Thanks to organizers and Kaggle for such an interesting competition and congrats to everyone who enjoyed this challenge!\n\nMy approach is two stage pipeline. I used CenterNet [1] for character detection, and MobileNetV3 [2] for classification. The overview of my approach is shown in the figure below. All models were trained with single GTX 970 GPU installed on my home server, so my solution is relative resource efficient.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2Fb352a774b803ab9abedbd8d1b4333437%2F2019-10-19%2013.51.21.png?generation=1571478495979958&amp;alt=media)\n\n## Detection\n\nI chose CenterNet because of its anchor-free simple design. Kuzushiji characters have wide aspect ratio range and relative small to page size, so it seemed to be difficult to find good anchor box setting. \n\n* Architecture: ResNet18 + U-Net\n* Data: Full training set to train single model\n* TTA: Scale adjustment -&gt; Multi-scale + Bounding box voting (bbox voting is proposed in [3])\n\n## Classification\n\nI combined various data augmentation to improve accuracy.  (black grid lines are only for visualization)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2Fcacf64aafe6e31d4e81bb381fd3e9d65%2F2019-10-19%2016.44.51.png?generation=1571478553323751&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2F47aa0cd704c01ebf1bf7d5466f10d669%2F2019-10-19%2016.44.59.png?generation=1571478576515188&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2F0718a01671e0ab60beb6681908f24b42%2F2019-10-19%2016.47.18.png?generation=1571478604246238&amp;alt=media)\n\nOther normal augmentation operations such as color, brightness, contrast, rotation and noise are also applied. Actual input images for training are as follows. These operations are implemented with `albumentations` library.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2F3194d2fb645ce38501df86ef26596315%2F2019-10-19%2016.53.23.png?generation=1571478627485652&amp;alt=media)\n\n* Architecture: MobileNetV3 Large\n* Loss: Standard softmax cross entropy loss\n* Training: Use full training set to train single model\n* Fine-tuning: Use full training set + pseudo label (filtered by score &gt; 0.9) with reduced data augmentation.\n* No TTA (I tried multi-scale inference, but its contribution was very small.)\n\n## Local validation\n\nAs discussed, group splitting by book title is important. The appearance of characters varies largely depending on the book. I used group split by book title for local validation to select hyper parameters, data augmentation types and threshold values for testing. I set the score threshold for detector to be 0.3 (&lt;- default = 0.5) to reduce false negatives predictions, and score threshold for classification to be 0.5 to reduce false positive predictions.\n\n## Discarded ideas\n\n* [Classifier] Oversampling of minor class and class balanced loss [4]. I tried these techniques to handle the class imbalance problem, but its contribution was very small.\n* [Classifier] Sequence based training to use context (Reading order prediction -&gt; CNN + Bi-LSTM), but did not work.\n* [Detector/Classifier] Collect large amount of unlabeled images from CODH web site, and use it for representation learning or semi-supervised learning. I gave up this idea because it requires huge computational resources and disk spaces.\n\n## Code\n\nAll models were written with Chainer. Please check my github repo for more details.\n\n- https://github.com/t-hanya/kuzushiji-recognition\n\n## References\n\n- [1] Objects as Points: https://arxiv.org/abs/1904.07850\n- [2] Searching for MobileNetV3: https://arxiv.org/abs/1905.02244\n- [3] Object detection via a multi-region &amp; semantic segmentation-aware CNN model: https://arxiv.org/abs/1505.01749\n- [4] Class-Balanced Loss Based on Effective Number of Samples: https://arxiv.org/abs/1901.05555",
      "votes": null
    },
    {
      "id": "652777",
      "postDate": "10/19/2019 12:00:17",
      "content": "<p>Congratulations\nGreat Work, Great Write-Up\nThanks for Sharing your Approach &amp; Insights <a href=\"/toshinori\">@toshinori</a> </p>",
      "rawMarkdown": "Congratulations\nGreat Work, Great Write-Up\nThanks for Sharing your Approach &amp; Insights @toshinori",
      "votes": null
    },
    {
      "id": "665120",
      "postDate": "11/04/2019 16:42:31",
      "content": "<p>Great work!!</p>",
      "rawMarkdown": "Great work!!",
      "votes": null
    },
    {
      "id": "675395",
      "postDate": "11/18/2019 03:05:10",
      "content": "<p>Hello, </p>\n\n<p>Organizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  </p>\n\n<p>Best, </p>\n\n<p>Alex.  </p>",
      "rawMarkdown": "Hello, \n\nOrganizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  \n\nBest, \n\nAlex.",
      "votes": null
    },
    {
      "id": "679626",
      "postDate": "11/23/2019 03:58:15",
      "content": "<p><a href=\"/thenuttynetter\">@thenuttynetter</a> thanks for your question.</p>\n\n<p>I measured inference time on test image set under exactly the same setting as my best submission (using same model parameter, including TTA). This value includes pre-processing such as resizing and cropping, but does not include image loading time. I used single GTX970 GPU.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2F3a9af0cf6a96e6b2813bd19a8a3cbb92%2Finference%20time.png?generation=1574480971325485&amp;alt=media\" alt=\"\"></p>\n\n<h2>Example:</h2>\n\n<p>1.26 sec (292 predicted characters)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2Fc7c971a9a515c3713b3de8529b8ca32f%2Fsample-prediction.png?generation=1574481017493840&amp;alt=media\" width=\"400\"></p>",
      "rawMarkdown": "thenuttynetter thanks for your question.\n\nI measured inference time on test image set under exactly the same setting as my best submission (using same model parameter, including TTA). This value includes pre-processing such as resizing and cropping, but does not include image loading time. I used single GTX970 GPU.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2F3a9af0cf6a96e6b2813bd19a8a3cbb92%2Finference%20time.png?generation=1574480971325485&amp;alt=media)\n\n## Example:\n\n1.26 sec (292 predicted characters)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2Fc7c971a9a515c3713b3de8529b8ca32f%2Fsample-prediction.png?generation=1574481017493840&amp;alt=media\" width=\"400\">",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 652777,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "10/19/2019 12:00:17",
      "content": "<p>Congratulations\nGreat Work, Great Write-Up\nThanks for Sharing your Approach &amp; Insights <a href=\"/toshinori\">@toshinori</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 665120,
      "author_name": "pradeepmuniasamy",
      "author_url": "",
      "post_date": "11/04/2019 16:42:31",
      "content": "<p>Great work!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 675395,
      "author_name": "thenuttynetter",
      "author_url": "",
      "post_date": "11/18/2019 03:05:10",
      "content": "<p>Hello, </p>\n\n<p>Organizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  </p>\n\n<p>Best, </p>\n\n<p>Alex.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 679626,
          "author_name": "toshinori",
          "author_url": "",
          "post_date": "11/23/2019 03:58:15",
          "content": "<p><a href=\"/thenuttynetter\">@thenuttynetter</a> thanks for your question.</p>\n\n<p>I measured inference time on test image set under exactly the same setting as my best submission (using same model parameter, including TTA). This value includes pre-processing such as resizing and cropping, but does not include image loading time. I used single GTX970 GPU.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2F3a9af0cf6a96e6b2813bd19a8a3cbb92%2Finference%20time.png?generation=1574480971325485&amp;alt=media\" alt=\"\"></p>\n\n<h2>Example:</h2>\n\n<p>1.26 sec (292 predicted characters)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2Fc7c971a9a515c3713b3de8529b8ca32f%2Fsample-prediction.png?generation=1574481017493840&amp;alt=media\" width=\"400\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "652727": "Thanks to organizers and Kaggle for such an interesting competition and congrats to everyone who enjoyed this challenge!\n\nMy approach is two stage pipeline. I used CenterNet [1] for character detection, and MobileNetV3 [2] for classification. The overview of my approach is shown in the figure below. All models were trained with single GTX 970 GPU installed on my home server, so my solution is relative resource efficient.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2Fb352a774b803ab9abedbd8d1b4333437%2F2019-10-19%2013.51.21.png?generation=1571478495979958&amp;alt=media)\n\n## Detection\n\nI chose CenterNet because of its anchor-free simple design. Kuzushiji characters have wide aspect ratio range and relative small to page size, so it seemed to be difficult to find good anchor box setting. \n\n* Architecture: ResNet18 + U-Net\n* Data: Full training set to train single model\n* TTA: Scale adjustment -&gt; Multi-scale + Bounding box voting (bbox voting is proposed in [3])\n\n## Classification\n\nI combined various data augmentation to improve accuracy.  (black grid lines are only for visualization)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2Fcacf64aafe6e31d4e81bb381fd3e9d65%2F2019-10-19%2016.44.51.png?generation=1571478553323751&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2F47aa0cd704c01ebf1bf7d5466f10d669%2F2019-10-19%2016.44.59.png?generation=1571478576515188&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2F0718a01671e0ab60beb6681908f24b42%2F2019-10-19%2016.47.18.png?generation=1571478604246238&amp;alt=media)\n\nOther normal augmentation operations such as color, brightness, contrast, rotation and noise are also applied. Actual input images for training are as follows. These operations are implemented with `albumentations` library.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2F3194d2fb645ce38501df86ef26596315%2F2019-10-19%2016.53.23.png?generation=1571478627485652&amp;alt=media)\n\n* Architecture: MobileNetV3 Large\n* Loss: Standard softmax cross entropy loss\n* Training: Use full training set to train single model\n* Fine-tuning: Use full training set + pseudo label (filtered by score &gt; 0.9) with reduced data augmentation.\n* No TTA (I tried multi-scale inference, but its contribution was very small.)\n\n## Local validation\n\nAs discussed, group splitting by book title is important. The appearance of characters varies largely depending on the book. I used group split by book title for local validation to select hyper parameters, data augmentation types and threshold values for testing. I set the score threshold for detector to be 0.3 (&lt;- default = 0.5) to reduce false negatives predictions, and score threshold for classification to be 0.5 to reduce false positive predictions.\n\n## Discarded ideas\n\n* [Classifier] Oversampling of minor class and class balanced loss [4]. I tried these techniques to handle the class imbalance problem, but its contribution was very small.\n* [Classifier] Sequence based training to use context (Reading order prediction -&gt; CNN + Bi-LSTM), but did not work.\n* [Detector/Classifier] Collect large amount of unlabeled images from CODH web site, and use it for representation learning or semi-supervised learning. I gave up this idea because it requires huge computational resources and disk spaces.\n\n## Code\n\nAll models were written with Chainer. Please check my github repo for more details.\n\n- https://github.com/t-hanya/kuzushiji-recognition\n\n## References\n\n- [1] Objects as Points: https://arxiv.org/abs/1904.07850\n- [2] Searching for MobileNetV3: https://arxiv.org/abs/1905.02244\n- [3] Object detection via a multi-region &amp; semantic segmentation-aware CNN model: https://arxiv.org/abs/1505.01749\n- [4] Class-Balanced Loss Based on Effective Number of Samples: https://arxiv.org/abs/1901.05555",
    "652777": "Congratulations\nGreat Work, Great Write-Up\nThanks for Sharing your Approach &amp; Insights @toshinori",
    "665120": "Great work!!",
    "675395": "Hello, \n\nOrganizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  \n\nBest, \n\nAlex.",
    "679626": "thenuttynetter thanks for your question.\n\nI measured inference time on test image set under exactly the same setting as my best submission (using same model parameter, including TTA). This value includes pre-processing such as resizing and cropping, but does not include image loading time. I used single GTX970 GPU.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2F3a9af0cf6a96e6b2813bd19a8a3cbb92%2Finference%20time.png?generation=1574480971325485&amp;alt=media)\n\n## Example:\n\n1.26 sec (292 predicted characters)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F394597%2Fc7c971a9a515c3713b3de8529b8ca32f%2Fsample-prediction.png?generation=1574481017493840&amp;alt=media\" width=\"400\">"
  },
  "source": "meta"
}