{
  "id": 112771,
  "title": "5th place solution overview: One stage CenterNet",
  "url": "/competitions/kuzushiji-recognition/writeups/see-5th-place-solution-overview-one-stage-centerne",
  "author_name": "",
  "post_date": "2019-10-15T17:34:48.457Z",
  "votes": 39,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Thanks to the Center for Open Data in the Humanities and Kaggle for this great challenge. Congratulations to the winners and everyone who had fun in this challenge. Here are a few details of my approach:</p>\n\n<p>I already created a port of <a href=\"https://arxiv.org/abs/1904.07850\">CenterNet</a> for Keras: <a href=\"https://github.com/see--/keras-centernet\">https://github.com/see--/keras-centernet</a> but I never really trained a model with this repo. So I decided to get the training running as well. I reimplemented CenterNet from scratch with the following modifications:</p>\n\n<ul>\n<li><p>Replace the Hourglass network with ResNet50 and ResNet101 + <a href=\"http://presentations.cocodataset.org/COCO17-Stuff-FAIR.pdf\">FPN</a> decoder. Hourglass is too demanding in terms of hardware. You could probably still get better results.</p></li>\n<li><p>Use a simpler loss function. I replaced the penalty-reduced pixel-wise logistic regression with focal loss by binary cross-entropy. This worked well for me and removed a few hyperparameters.</p></li>\n<li><p>I also removed the heads that regressed the height &amp; width and the x &amp; y offsets. You can regress them easily with L2 loss but the metric only wants centers. We also don't need to be more accurate than +-2 pixels (which is what the offsets are used for). This gave a little speed-up.</p></li>\n<li><p>CenterNet uses a gaussian around the the centers. I am not sure this is needed. I still use the following 3x3 kernel: \n<code>\npos_kernel = np.float32([[0.5, 0.75, 0.5], [0.75, 1.0, 0.75], [0.5, 0.75, 0.5]])\n</code>\nBut using only the exact center gave very similar results (i.e. <code>\npos_kernel = np.float32([[0.0, 0.0, 0.0],\n                         [0.0, 1.0, 0.0],\n                         [0.0, 0.0, 0.0]])\n</code>). One just needs to reduce the threshold.</p></li>\n</ul>\n\n<p>A single ResNet101 that I trained for 125 epochs could reach: <code>0.938</code>. I used the improved version from the Bag of Tricks for Image Classification paper: <a href=\"https://arxiv.org/abs/1812.01187\">https://arxiv.org/abs/1812.01187</a>. Thanks to <code>timm</code> this is really simple: <code>timm.create_model('gluon_resnet101_v1d', pretrained=pretrained)</code>.</p>\n\n<p>Albumentations was also super useful. Getting the geometric transforms for bounding boxes correct is <a href=\"https://github.com/poodarchu/learn_aug_for_object_detection.numpy\">not trivial</a>. I used the following augmentations (without validation):\n<code>\n    self.aug = Compose([\n      ShiftScaleRotate(p=0.9, rotate_limit=10,\n          scale_limit=0.2, border_mode=cv2.BORDER_CONSTANT),\n      RandomCrop(512, 512, p=1.0),\n      ToGray(),\n      CLAHE(),\n      GaussNoise(),\n      GaussianBlur(),\n      RandomBrightnessContrast(),\n      RandomGamma(),\n      RGBShift(),\n      HueSaturationValue(),\n    ], bbox_params=BboxParams(format='coco', min_visibility=0.75))\n</code></p>\n\n<p>A high <code>min_visibility</code> is set to not predict partial letters (that is how the data was labeled).</p>\n\n<p>Finally, scale TTA (with scales <code>[0.75, 1.0, 1.25]</code>) helped a bit <code>0.938</code> -&gt; <code>0.939</code>.</p>\n\n<p><strong>What did not work</strong></p>\n\n<ul>\n<li><p>Multi-stage approaches with separate detection and classification. I only got ~<code>0.85</code>. One stage worked always better. I think this makes sense as the model can learn some kind of language model from the images.</p></li>\n<li><p>Increasing the resolution even further. I am running predictions at 1536x1536 resolution but I didn't see any improvement  from higher resolutions.</p></li>\n<li><p>Creating more letters. This was a popular trick in the whale competition. We double the number of classes by adding flipped versions. However, the validation score was about the same so I removed it as it added code &amp; complexity.</p></li>\n</ul>\n\n<p><strong>Hardware</strong></p>\n\n<p>The final model can be trained in ~8 hours on a single V100 VM. I mostly used vast.ai and GCP.</p>\n\n<p>Thanks for reading. You can find the code and trained weights on GitHub:</p>\n\n<ul>\n<li><a href=\"https://github.com/see--/kuzushiji-recognition\">https://github.com/see--/kuzushiji-recognition</a></li>\n<li><a href=\"https://github.com/see--/kuzushiji-recognition/releases/download/v0.1/f00-ep-0125-val_hm_acc-0.9944-val_classes_acc-0.4986.pth\">https://github.com/see--/kuzushiji-recognition/releases/download/v0.1/f00-ep-0125-val_hm_acc-0.9944-val_classes_acc-0.4986.pth</a></li>\n</ul>",
  "messages": [
    {
      "id": "649386",
      "postDate": "10/15/2019 09:23:33",
      "content": "<p>Thanks to the Center for Open Data in the Humanities and Kaggle for this great challenge. Congratulations to the winners and everyone who had fun in this challenge. Here are a few details of my approach:</p>\n\n<p>I already created a port of <a href=\"https://arxiv.org/abs/1904.07850\">CenterNet</a> for Keras: <a href=\"https://github.com/see--/keras-centernet\">https://github.com/see--/keras-centernet</a> but I never really trained a model with this repo. So I decided to get the training running as well. I reimplemented CenterNet from scratch with the following modifications:</p>\n\n<ul>\n<li><p>Replace the Hourglass network with ResNet50 and ResNet101 + <a href=\"http://presentations.cocodataset.org/COCO17-Stuff-FAIR.pdf\">FPN</a> decoder. Hourglass is too demanding in terms of hardware. You could probably still get better results.</p></li>\n<li><p>Use a simpler loss function. I replaced the penalty-reduced pixel-wise logistic regression with focal loss by binary cross-entropy. This worked well for me and removed a few hyperparameters.</p></li>\n<li><p>I also removed the heads that regressed the height &amp; width and the x &amp; y offsets. You can regress them easily with L2 loss but the metric only wants centers. We also don't need to be more accurate than +-2 pixels (which is what the offsets are used for). This gave a little speed-up.</p></li>\n<li><p>CenterNet uses a gaussian around the the centers. I am not sure this is needed. I still use the following 3x3 kernel: \n<code>\npos_kernel = np.float32([[0.5, 0.75, 0.5], [0.75, 1.0, 0.75], [0.5, 0.75, 0.5]])\n</code>\nBut using only the exact center gave very similar results (i.e. <code>\npos_kernel = np.float32([[0.0, 0.0, 0.0],\n                         [0.0, 1.0, 0.0],\n                         [0.0, 0.0, 0.0]])\n</code>). One just needs to reduce the threshold.</p></li>\n</ul>\n\n<p>A single ResNet101 that I trained for 125 epochs could reach: <code>0.938</code>. I used the improved version from the Bag of Tricks for Image Classification paper: <a href=\"https://arxiv.org/abs/1812.01187\">https://arxiv.org/abs/1812.01187</a>. Thanks to <code>timm</code> this is really simple: <code>timm.create_model('gluon_resnet101_v1d', pretrained=pretrained)</code>.</p>\n\n<p>Albumentations was also super useful. Getting the geometric transforms for bounding boxes correct is <a href=\"https://github.com/poodarchu/learn_aug_for_object_detection.numpy\">not trivial</a>. I used the following augmentations (without validation):\n<code>\n    self.aug = Compose([\n      ShiftScaleRotate(p=0.9, rotate_limit=10,\n          scale_limit=0.2, border_mode=cv2.BORDER_CONSTANT),\n      RandomCrop(512, 512, p=1.0),\n      ToGray(),\n      CLAHE(),\n      GaussNoise(),\n      GaussianBlur(),\n      RandomBrightnessContrast(),\n      RandomGamma(),\n      RGBShift(),\n      HueSaturationValue(),\n    ], bbox_params=BboxParams(format='coco', min_visibility=0.75))\n</code></p>\n\n<p>A high <code>min_visibility</code> is set to not predict partial letters (that is how the data was labeled).</p>\n\n<p>Finally, scale TTA (with scales <code>[0.75, 1.0, 1.25]</code>) helped a bit <code>0.938</code> -&gt; <code>0.939</code>.</p>\n\n<p><strong>What did not work</strong></p>\n\n<ul>\n<li><p>Multi-stage approaches with separate detection and classification. I only got ~<code>0.85</code>. One stage worked always better. I think this makes sense as the model can learn some kind of language model from the images.</p></li>\n<li><p>Increasing the resolution even further. I am running predictions at 1536x1536 resolution but I didn't see any improvement  from higher resolutions.</p></li>\n<li><p>Creating more letters. This was a popular trick in the whale competition. We double the number of classes by adding flipped versions. However, the validation score was about the same so I removed it as it added code &amp; complexity.</p></li>\n</ul>\n\n<p><strong>Hardware</strong></p>\n\n<p>The final model can be trained in ~8 hours on a single V100 VM. I mostly used vast.ai and GCP.</p>\n\n<p>Thanks for reading. You can find the code and trained weights on GitHub:</p>\n\n<ul>\n<li><a href=\"https://github.com/see--/kuzushiji-recognition\">https://github.com/see--/kuzushiji-recognition</a></li>\n<li><a href=\"https://github.com/see--/kuzushiji-recognition/releases/download/v0.1/f00-ep-0125-val_hm_acc-0.9944-val_classes_acc-0.4986.pth\">https://github.com/see--/kuzushiji-recognition/releases/download/v0.1/f00-ep-0125-val_hm_acc-0.9944-val_classes_acc-0.4986.pth</a></li>\n</ul>",
      "rawMarkdown": "Thanks to the Center for Open Data in the Humanities and Kaggle for this great challenge. Congratulations to the winners and everyone who had fun in this challenge. Here are a few details of my approach:\n\nI already created a port of [CenterNet](https://arxiv.org/abs/1904.07850) for Keras: https://github.com/see--/keras-centernet but I never really trained a model with this repo. So I decided to get the training running as well. I reimplemented CenterNet from scratch with the following modifications:\n\n* Replace the Hourglass network with ResNet50 and ResNet101 + [FPN](http://presentations.cocodataset.org/COCO17-Stuff-FAIR.pdf) decoder. Hourglass is too demanding in terms of hardware. You could probably still get better results.\n\n* Use a simpler loss function. I replaced the penalty-reduced pixel-wise logistic regression with focal loss by binary cross-entropy. This worked well for me and removed a few hyperparameters.\n\n* I also removed the heads that regressed the height &amp; width and the x &amp; y offsets. You can regress them easily with L2 loss but the metric only wants centers. We also don't need to be more accurate than +-2 pixels (which is what the offsets are used for). This gave a little speed-up.\n\n* CenterNet uses a gaussian around the the centers. I am not sure this is needed. I still use the following 3x3 kernel: \n```\npos_kernel = np.float32([[0.5, 0.75, 0.5], [0.75, 1.0, 0.75], [0.5, 0.75, 0.5]])\n```\nBut using only the exact center gave very similar results (i.e. ```\npos_kernel = np.float32([[0.0, 0.0, 0.0],\n                             [0.0, 1.0, 0.0],\n                             [0.0, 0.0, 0.0]])\n```). One just needs to reduce the threshold.\n\nA single ResNet101 that I trained for 125 epochs could reach: `0.938`. I used the improved version from the Bag of Tricks for Image Classification paper: https://arxiv.org/abs/1812.01187. Thanks to `timm` this is really simple: `timm.create_model('gluon_resnet101_v1d', pretrained=pretrained)`.\n\nAlbumentations was also super useful. Getting the geometric transforms for bounding boxes correct is [not trivial](https://github.com/poodarchu/learn_aug_for_object_detection.numpy). I used the following augmentations (without validation):\n```\n    self.aug = Compose([\n      ShiftScaleRotate(p=0.9, rotate_limit=10,\n          scale_limit=0.2, border_mode=cv2.BORDER_CONSTANT),\n      RandomCrop(512, 512, p=1.0),\n      ToGray(),\n      CLAHE(),\n      GaussNoise(),\n      GaussianBlur(),\n      RandomBrightnessContrast(),\n      RandomGamma(),\n      RGBShift(),\n      HueSaturationValue(),\n    ], bbox_params=BboxParams(format='coco', min_visibility=0.75))\n```\n\nA high `min_visibility` is set to not predict partial letters (that is how the data was labeled).\n\nFinally, scale TTA (with scales `[0.75, 1.0, 1.25]`) helped a bit `0.938` -&gt; `0.939`.\n\n**What did not work**\n\n* Multi-stage approaches with separate detection and classification. I only got ~`0.85`. One stage worked always better. I think this makes sense as the model can learn some kind of language model from the images.\n\n* Increasing the resolution even further. I am running predictions at 1536x1536 resolution but I didn't see any improvement  from higher resolutions.\n\n* Creating more letters. This was a popular trick in the whale competition. We double the number of classes by adding flipped versions. However, the validation score was about the same so I removed it as it added code &amp; complexity.\n\n**Hardware**\n\nThe final model can be trained in ~8 hours on a single V100 VM. I mostly used vast.ai and GCP.\n\nThanks for reading. You can find the code and trained weights on GitHub:\n\n* https://github.com/see--/kuzushiji-recognition\n* https://github.com/see--/kuzushiji-recognition/releases/download/v0.1/f00-ep-0125-val_hm_acc-0.9944-val_classes_acc-0.4986.pth",
      "votes": null
    },
    {
      "id": "649431",
      "postDate": "10/15/2019 10:41:22",
      "content": "<p>Thanks for sharing, this is a really elegant solution and a very strong single model score, congrats! I'm curious what validation split did you use?</p>",
      "rawMarkdown": "Thanks for sharing, this is a really elegant solution and a very strong single model score, congrats! I'm curious what validation split did you use?",
      "votes": null
    },
    {
      "id": "649438",
      "postDate": "10/15/2019 10:49:39",
      "content": "<p>Thanks! I used 20% (randomly selected) as validation data. The final model is trained with all data. This gave a big improvement: From <code>0.907</code> -&gt; <code>0.938</code>. Probably because some letters never occurred during training with validation data.</p>",
      "rawMarkdown": "Thanks! I used 20% (randomly selected) as validation data. The final model is trained with all data. This gave a big improvement: From `0.907` -&gt; `0.938`. Probably because some letters never occurred during training with validation data.",
      "votes": null
    },
    {
      "id": "649450",
      "postDate": "10/15/2019 11:17:41",
      "content": "<p>Thank you so much!!!</p>",
      "rawMarkdown": "Thank you so much!!!",
      "votes": null
    },
    {
      "id": "649596",
      "postDate": "10/15/2019 14:58:26",
      "content": "<p>Congratulations\nGreat Write-Up\nThanks for Sharing your Approach &amp; Insights.. <a href=\"/seesee\">@seesee</a> </p>",
      "rawMarkdown": "Congratulations\nGreat Write-Up\nThanks for Sharing your Approach &amp; Insights.. @seesee",
      "votes": null
    },
    {
      "id": "675400",
      "postDate": "11/18/2019 03:09:11",
      "content": "<p>Hello, </p>\n\n<p>Organizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  </p>\n\n<p>Best, </p>\n\n<p>Alex.  </p>",
      "rawMarkdown": "Hello, \n\nOrganizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  \n\nBest, \n\nAlex.",
      "votes": null
    },
    {
      "id": "675692",
      "postDate": "11/18/2019 12:35:06",
      "content": "<p><a href=\"/seesee\">@seesee</a> Thanks for sharing ....</p>",
      "rawMarkdown": "seesee Thanks for sharing ....",
      "votes": null
    },
    {
      "id": "688281",
      "postDate": "12/05/2019 12:25:52",
      "content": "<p>'Hourglass is too demanding in terms of hardware', if I just use free kernel and pre-trained weights to train just part of network, can it work? <a href=\"/seesee\">@seesee</a> </p>",
      "rawMarkdown": "'Hourglass is too demanding in terms of hardware', if I just use free kernel and pre-trained weights to train just part of network, can it work? @seesee",
      "votes": null
    },
    {
      "id": "706233",
      "postDate": "12/30/2019 05:20:09",
      "content": "<p>Hi see--:\nI'm using your keras centernet, it's wonderful, may I ask if this is the weight you train or get from original repo? <a href=\"/seesee\">@seesee</a> </p>\n\n<p>CTDET_COCO_WEIGHTS_PATH =('<a href=\"https://github.com/see-/kerascenternet/\">https://github.com/see-/kerascenternet/</a>'releases/download/0.1.0/ctdet_coco_hg.hdf5')</p>",
      "rawMarkdown": "Hi see--:\nI'm using your keras centernet, it's wonderful, may I ask if this is the weight you train or get from original repo? @seesee \n\nCTDET_COCO_WEIGHTS_PATH =('https://github.com/see-/kerascenternet/'releases/download/0.1.0/ctdet_coco_hg.hdf5')",
      "votes": null
    },
    {
      "id": "706646",
      "postDate": "12/30/2019 16:49:10",
      "content": "<p>Hi <a href=\"/diegojohnson\">@diegojohnson</a>,</p>\n\n<p>I glad that you find it useful ;)</p>\n\n<blockquote>\n  <p>'Hourglass is too demanding in terms of hardware', if I just use free kernel and pre-trained weights to train just part of network, can it work</p>\n</blockquote>\n\n<p>From my experience, training the whole model is always better than freezing parts.</p>\n\n<blockquote>\n  <p>I'm using your keras centernet, it's wonderful, may I ask if this is the weight you train or get from original repo?</p>\n</blockquote>\n\n<p>The weights are from the original repo. I hope I find some time to update the code to tf 2.0 and add the training part. It's much more efficient and simpler to use with keras.</p>",
      "rawMarkdown": "Hi @diegojohnson,\n\nI glad that you find it useful ;)\n\n&gt;  'Hourglass is too demanding in terms of hardware', if I just use free kernel and pre-trained weights to train just part of network, can it work\n\nFrom my experience, training the whole model is always better than freezing parts.\n\n&gt; I'm using your keras centernet, it's wonderful, may I ask if this is the weight you train or get from original repo?\n\nThe weights are from the original repo. I hope I find some time to update the code to tf 2.0 and add the training part. It's much more efficient and simpler to use with keras.",
      "votes": null
    },
    {
      "id": "706648",
      "postDate": "12/30/2019 16:52:31",
      "content": "<p>A rough guess is faster than 10 FPS on a 1080 TI without TTA.</p>",
      "rawMarkdown": "A rough guess is faster than 10 FPS on a 1080 TI without TTA.",
      "votes": null
    },
    {
      "id": "706657",
      "postDate": "12/30/2019 17:08:22",
      "content": "<p>thanks very much.😃 </p>",
      "rawMarkdown": "thanks very much.😃",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 649431,
      "author_name": "lopuhin",
      "author_url": "",
      "post_date": "10/15/2019 10:41:22",
      "content": "<p>Thanks for sharing, this is a really elegant solution and a very strong single model score, congrats! I'm curious what validation split did you use?</p>",
      "votes": null,
      "replies": [
        {
          "id": 649438,
          "author_name": "seesee",
          "author_url": "",
          "post_date": "10/15/2019 10:49:39",
          "content": "<p>Thanks! I used 20% (randomly selected) as validation data. The final model is trained with all data. This gave a big improvement: From <code>0.907</code> -&gt; <code>0.938</code>. Probably because some letters never occurred during training with validation data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 649450,
      "author_name": "tkasasagi",
      "author_url": "",
      "post_date": "10/15/2019 11:17:41",
      "content": "<p>Thank you so much!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 649596,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "10/15/2019 14:58:26",
      "content": "<p>Congratulations\nGreat Write-Up\nThanks for Sharing your Approach &amp; Insights.. <a href=\"/seesee\">@seesee</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 675400,
      "author_name": "thenuttynetter",
      "author_url": "",
      "post_date": "11/18/2019 03:09:11",
      "content": "<p>Hello, </p>\n\n<p>Organizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  </p>\n\n<p>Best, </p>\n\n<p>Alex.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 706648,
          "author_name": "seesee",
          "author_url": "",
          "post_date": "12/30/2019 16:52:31",
          "content": "<p>A rough guess is faster than 10 FPS on a 1080 TI without TTA.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 675692,
      "author_name": "saurav9786",
      "author_url": "",
      "post_date": "11/18/2019 12:35:06",
      "content": "<p><a href=\"/seesee\">@seesee</a> Thanks for sharing ....</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 688281,
      "author_name": "diegojohnson",
      "author_url": "",
      "post_date": "12/05/2019 12:25:52",
      "content": "<p>'Hourglass is too demanding in terms of hardware', if I just use free kernel and pre-trained weights to train just part of network, can it work? <a href=\"/seesee\">@seesee</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 706233,
      "author_name": "diegojohnson",
      "author_url": "",
      "post_date": "12/30/2019 05:20:09",
      "content": "<p>Hi see--:\nI'm using your keras centernet, it's wonderful, may I ask if this is the weight you train or get from original repo? <a href=\"/seesee\">@seesee</a> </p>\n\n<p>CTDET_COCO_WEIGHTS_PATH =('<a href=\"https://github.com/see-/kerascenternet/\">https://github.com/see-/kerascenternet/</a>'releases/download/0.1.0/ctdet_coco_hg.hdf5')</p>",
      "votes": null,
      "replies": [
        {
          "id": 706646,
          "author_name": "seesee",
          "author_url": "",
          "post_date": "12/30/2019 16:49:10",
          "content": "<p>Hi <a href=\"/diegojohnson\">@diegojohnson</a>,</p>\n\n<p>I glad that you find it useful ;)</p>\n\n<blockquote>\n  <p>'Hourglass is too demanding in terms of hardware', if I just use free kernel and pre-trained weights to train just part of network, can it work</p>\n</blockquote>\n\n<p>From my experience, training the whole model is always better than freezing parts.</p>\n\n<blockquote>\n  <p>I'm using your keras centernet, it's wonderful, may I ask if this is the weight you train or get from original repo?</p>\n</blockquote>\n\n<p>The weights are from the original repo. I hope I find some time to update the code to tf 2.0 and add the training part. It's much more efficient and simpler to use with keras.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 706657,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "12/30/2019 17:08:22",
          "content": "<p>thanks very much.😃 </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "649386": "Thanks to the Center for Open Data in the Humanities and Kaggle for this great challenge. Congratulations to the winners and everyone who had fun in this challenge. Here are a few details of my approach:\n\nI already created a port of [CenterNet](https://arxiv.org/abs/1904.07850) for Keras: https://github.com/see--/keras-centernet but I never really trained a model with this repo. So I decided to get the training running as well. I reimplemented CenterNet from scratch with the following modifications:\n\n* Replace the Hourglass network with ResNet50 and ResNet101 + [FPN](http://presentations.cocodataset.org/COCO17-Stuff-FAIR.pdf) decoder. Hourglass is too demanding in terms of hardware. You could probably still get better results.\n\n* Use a simpler loss function. I replaced the penalty-reduced pixel-wise logistic regression with focal loss by binary cross-entropy. This worked well for me and removed a few hyperparameters.\n\n* I also removed the heads that regressed the height &amp; width and the x &amp; y offsets. You can regress them easily with L2 loss but the metric only wants centers. We also don't need to be more accurate than +-2 pixels (which is what the offsets are used for). This gave a little speed-up.\n\n* CenterNet uses a gaussian around the the centers. I am not sure this is needed. I still use the following 3x3 kernel: \n```\npos_kernel = np.float32([[0.5, 0.75, 0.5], [0.75, 1.0, 0.75], [0.5, 0.75, 0.5]])\n```\nBut using only the exact center gave very similar results (i.e. ```\npos_kernel = np.float32([[0.0, 0.0, 0.0],\n                             [0.0, 1.0, 0.0],\n                             [0.0, 0.0, 0.0]])\n```). One just needs to reduce the threshold.\n\nA single ResNet101 that I trained for 125 epochs could reach: `0.938`. I used the improved version from the Bag of Tricks for Image Classification paper: https://arxiv.org/abs/1812.01187. Thanks to `timm` this is really simple: `timm.create_model('gluon_resnet101_v1d', pretrained=pretrained)`.\n\nAlbumentations was also super useful. Getting the geometric transforms for bounding boxes correct is [not trivial](https://github.com/poodarchu/learn_aug_for_object_detection.numpy). I used the following augmentations (without validation):\n```\n    self.aug = Compose([\n      ShiftScaleRotate(p=0.9, rotate_limit=10,\n          scale_limit=0.2, border_mode=cv2.BORDER_CONSTANT),\n      RandomCrop(512, 512, p=1.0),\n      ToGray(),\n      CLAHE(),\n      GaussNoise(),\n      GaussianBlur(),\n      RandomBrightnessContrast(),\n      RandomGamma(),\n      RGBShift(),\n      HueSaturationValue(),\n    ], bbox_params=BboxParams(format='coco', min_visibility=0.75))\n```\n\nA high `min_visibility` is set to not predict partial letters (that is how the data was labeled).\n\nFinally, scale TTA (with scales `[0.75, 1.0, 1.25]`) helped a bit `0.938` -&gt; `0.939`.\n\n**What did not work**\n\n* Multi-stage approaches with separate detection and classification. I only got ~`0.85`. One stage worked always better. I think this makes sense as the model can learn some kind of language model from the images.\n\n* Increasing the resolution even further. I am running predictions at 1536x1536 resolution but I didn't see any improvement  from higher resolutions.\n\n* Creating more letters. This was a popular trick in the whale competition. We double the number of classes by adding flipped versions. However, the validation score was about the same so I removed it as it added code &amp; complexity.\n\n**Hardware**\n\nThe final model can be trained in ~8 hours on a single V100 VM. I mostly used vast.ai and GCP.\n\nThanks for reading. You can find the code and trained weights on GitHub:\n\n* https://github.com/see--/kuzushiji-recognition\n* https://github.com/see--/kuzushiji-recognition/releases/download/v0.1/f00-ep-0125-val_hm_acc-0.9944-val_classes_acc-0.4986.pth",
    "649431": "Thanks for sharing, this is a really elegant solution and a very strong single model score, congrats! I'm curious what validation split did you use?",
    "649438": "Thanks! I used 20% (randomly selected) as validation data. The final model is trained with all data. This gave a big improvement: From `0.907` -&gt; `0.938`. Probably because some letters never occurred during training with validation data.",
    "649450": "Thank you so much!!!",
    "649596": "Congratulations\nGreat Write-Up\nThanks for Sharing your Approach &amp; Insights.. @seesee",
    "675400": "Hello, \n\nOrganizer here.  If you wouldn't mind, can you say how much time your method takes for inference per-page-image (even a ballpark or rough estimate is fine)?  \n\nBest, \n\nAlex.",
    "675692": "seesee Thanks for sharing ....",
    "688281": "'Hourglass is too demanding in terms of hardware', if I just use free kernel and pre-trained weights to train just part of network, can it work? @seesee",
    "706233": "Hi see--:\nI'm using your keras centernet, it's wonderful, may I ask if this is the weight you train or get from original repo? @seesee \n\nCTDET_COCO_WEIGHTS_PATH =('https://github.com/see-/kerascenternet/'releases/download/0.1.0/ctdet_coco_hg.hdf5')",
    "706646": "Hi @diegojohnson,\n\nI glad that you find it useful ;)\n\n&gt;  'Hourglass is too demanding in terms of hardware', if I just use free kernel and pre-trained weights to train just part of network, can it work\n\nFrom my experience, training the whole model is always better than freezing parts.\n\n&gt; I'm using your keras centernet, it's wonderful, may I ask if this is the weight you train or get from original repo?\n\nThe weights are from the original repo. I hope I find some time to update the code to tf 2.0 and add the training part. It's much more efficient and simpler to use with keras.",
    "706648": "A rough guess is faster than 10 FPS on a 1080 TI without TTA.",
    "706657": "thanks very much.😃"
  },
  "source": "meta"
}