{
  "id": 243766,
  "title": "10th solution private 0.61 : TNT/VIT/CAIT ensemble + tricks",
  "url": "/competitions/bms-molecular-translation/writeups/inchi-1s-c2h6-c1-2-h1-2h3-10th-solution-private-0-",
  "author_name": "",
  "post_date": "2021-06-10T05:06:28.217Z",
  "votes": 60,
  "comment_count": 14,
  "views": 0,
  "content": "<p>to be updated ….</p>\n<p>there are two part:</p>\n<ul>\n<li>part one: TNT/VIT/CAIT ensemble model (public/private LB 0.71/0.71)<br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/243766\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/243766</a> (this post)</li>\n<li>part two: OCR + MolBuilder + object detection CNN (public/private LB 0.62/0.61)<br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/243809\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/243809</a></li>\n</ul>\n<hr>\n<p>TNT/VIT/CAIT ensemble model (public/private LB 0.71/0.71)</p>\n<ul>\n<li><p>we train 5 models : 3x patched based VIT (patch size 16, image scale = 0.8,0.9,0.1), 1x TNT with image size 384 and 1x CAIT with image size 240. For details, please refer to <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231190\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/231190</a><br>\n(local CV 0.78) </p></li>\n<li><p>we do 5 iterations of image perturbation. if an image has gives an invalid inchi (verification by rdkit), we perturb the image and try inference again.<br>\n( local CV 0.76) </p></li>\n<li><p>finally, we apply rdkit normalisation.<br>\n(public/private LB 0.73/0.74, local CV 0.72) </p></li>\n</ul>\n<p>The ensemble is our top score submission. We also have submissions for each of the individual models. We then sort all submission cvs by LB score. To create the final submission:</p>\n<ul>\n<li>1. We start off will predictions from the top score submission csv. </li>\n<li>2. For a current prediction has invalid inchi, it will be replaced from the corresponding one from the next higher score submission csv if ( and only if ) that corresponding one is valid</li>\n<li>3. The process is repeated for all csvs. <br>\nFinal LB score is (public/private LB 0.71/0.71, local CV 0.70)</li>\n</ul>\n<p><img src=\"https://i.ibb.co/TYt5h7m/Selection-193.png\" alt=\"\"></p>\n<p>we use random shift and scale to perturb image</p>\n<hr>\n<p>here is a tip on ensemble strategy:<br>\n<img src=\"https://i.ibb.co/F3Yns33/Selection-202.png\" alt=\"\"></p>\n<p>we note that lb score is correlated to the number of invalid inchi (verified from rdkit). If you have fewer invalid inchi, you are likely to better LB score.</p>\n<p>instead of making many submissions, you can use the number of invalid inchi to estimate your lb score. To know the limit of LB score of your current method, you can extrapolate your graph to the lowest number of possible inchi. If that doesn't work well, you can change your ensemble strategy so that you have a steeper graph. (i.e. better lb score per invalid inchi)</p>\n<hr>\n<h2>Acknowledgment</h2>\n<p>I am grateful to Z by HP &amp; NVIDIA for sponsoring me a Z8 Workstation with dual RTX8000 GPU for this competition. In the final days, I need to run TNT/VIT/CAIT ensemble (5 transformers) on the 1.6 million test images. Some of the test images are large and have long sequences. Besides high computative power, we also need high GPU memory.</p>\n<p>One RTX8000 GPU has 48GB onboard. It takes about 7 hours to process all images using the workstation. This is very fast and definitely help our team to secure a gold medal of 10th placing in the final ranking!</p>",
  "messages": [
    {
      "id": "1334981",
      "postDate": "06/04/2021 00:11:46",
      "content": "<p>to be updated ….</p>\n<p>there are two part:</p>\n<ul>\n<li>part one: TNT/VIT/CAIT ensemble model (public/private LB 0.71/0.71)<br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/243766\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/243766</a> (this post)</li>\n<li>part two: OCR + MolBuilder + object detection CNN (public/private LB 0.62/0.61)<br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/243809\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/243809</a></li>\n</ul>\n<hr>\n<p>TNT/VIT/CAIT ensemble model (public/private LB 0.71/0.71)</p>\n<ul>\n<li><p>we train 5 models : 3x patched based VIT (patch size 16, image scale = 0.8,0.9,0.1), 1x TNT with image size 384 and 1x CAIT with image size 240. For details, please refer to <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231190\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/231190</a><br>\n(local CV 0.78) </p></li>\n<li><p>we do 5 iterations of image perturbation. if an image has gives an invalid inchi (verification by rdkit), we perturb the image and try inference again.<br>\n( local CV 0.76) </p></li>\n<li><p>finally, we apply rdkit normalisation.<br>\n(public/private LB 0.73/0.74, local CV 0.72) </p></li>\n</ul>\n<p>The ensemble is our top score submission. We also have submissions for each of the individual models. We then sort all submission cvs by LB score. To create the final submission:</p>\n<ul>\n<li>1. We start off will predictions from the top score submission csv. </li>\n<li>2. For a current prediction has invalid inchi, it will be replaced from the corresponding one from the next higher score submission csv if ( and only if ) that corresponding one is valid</li>\n<li>3. The process is repeated for all csvs. <br>\nFinal LB score is (public/private LB 0.71/0.71, local CV 0.70)</li>\n</ul>\n<p><img src=\"https://i.ibb.co/TYt5h7m/Selection-193.png\" alt=\"\"></p>\n<p>we use random shift and scale to perturb image</p>\n<hr>\n<p>here is a tip on ensemble strategy:<br>\n<img src=\"https://i.ibb.co/F3Yns33/Selection-202.png\" alt=\"\"></p>\n<p>we note that lb score is correlated to the number of invalid inchi (verified from rdkit). If you have fewer invalid inchi, you are likely to better LB score.</p>\n<p>instead of making many submissions, you can use the number of invalid inchi to estimate your lb score. To know the limit of LB score of your current method, you can extrapolate your graph to the lowest number of possible inchi. If that doesn't work well, you can change your ensemble strategy so that you have a steeper graph. (i.e. better lb score per invalid inchi)</p>\n<hr>\n<h2>Acknowledgment</h2>\n<p>I am grateful to Z by HP &amp; NVIDIA for sponsoring me a Z8 Workstation with dual RTX8000 GPU for this competition. In the final days, I need to run TNT/VIT/CAIT ensemble (5 transformers) on the 1.6 million test images. Some of the test images are large and have long sequences. Besides high computative power, we also need high GPU memory.</p>\n<p>One RTX8000 GPU has 48GB onboard. It takes about 7 hours to process all images using the workstation. This is very fast and definitely help our team to secure a gold medal of 10th placing in the final ranking!</p>",
      "rawMarkdown": "to be updated ....\n\nthere are two part:\n- part one: TNT/VIT/CAIT ensemble model (public/private LB 0.71/0.71)\nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/243766 (this post)\n- part two: OCR + MolBuilder + object detection CNN (public/private LB 0.62/0.61)\nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/243809\n\n\n---\n\nTNT/VIT/CAIT ensemble model (public/private LB 0.71/0.71)\n\n- we train 5 models : 3x patched based VIT (patch size 16, image scale = 0.8,0.9,0.1), 1x TNT with image size 384 and 1x CAIT with image size 240. For details, please refer to https://www.kaggle.com/c/bms-molecular-translation/discussion/231190\n(local CV 0.78) \n\n- we do 5 iterations of image perturbation. if an image has gives an invalid inchi (verification by rdkit), we perturb the image and try inference again.\n( local CV 0.76) \n\n- finally, we apply rdkit normalisation.\n(public/private LB 0.73/0.74, local CV 0.72) \n\nThe ensemble is our top score submission. We also have submissions for each of the individual models. We then sort all submission cvs by LB score. To create the final submission:\n- 1. We start off will predictions from the top score submission csv. \n- 2. For a current prediction has invalid inchi, it will be replaced from the corresponding one from the next higher score submission csv if ( and only if ) that corresponding one is valid\n- 3. The process is repeated for all csvs. \nFinal LB score is (public/private LB 0.71/0.71, local CV 0.70)\n\n![](https://i.ibb.co/TYt5h7m/Selection-193.png)\n\nwe use random shift and scale to perturb image\n\n---\n\nhere is a tip on ensemble strategy:\n![](https://i.ibb.co/F3Yns33/Selection-202.png)\n\nwe note that lb score is correlated to the number of invalid inchi (verified from rdkit). If you have fewer invalid inchi, you are likely to better LB score.\n\ninstead of making many submissions, you can use the number of invalid inchi to estimate your lb score. To know the limit of LB score of your current method, you can extrapolate your graph to the lowest number of possible inchi. If that doesn't work well, you can change your ensemble strategy so that you have a steeper graph. (i.e. better lb score per invalid inchi)\n\n---\n\nAcknowledgment\n---\n\nI am grateful to Z by HP & NVIDIA for sponsoring me a Z8 Workstation with dual RTX8000 GPU for this competition. In the final days, I need to run TNT/VIT/CAIT ensemble (5 transformers) on the 1.6 million test images. Some of the test images are large and have long sequences. Besides high computative power, we also need high GPU memory.\n\nOne RTX8000 GPU has 48GB onboard. It takes about 7 hours to process all images using the workstation. This is very fast and definitely help our team to secure a gold medal of 10th placing in the final ranking!",
      "votes": null
    },
    {
      "id": "1335009",
      "postDate": "06/04/2021 00:38:00",
      "content": "<p>the kaggle dataset is rendered by indigo toolkit (use old version 1.4.0-beta)<br>\n<a href=\"https://lifescience.opensource.epam.com/indigo/\" target=\"_blank\">https://lifescience.opensource.epam.com/indigo/</a></p>\n<p><img src=\"https://i.ibb.co/ZzSv1mV/Selection-194.png\" alt=\"\"></p>",
      "rawMarkdown": "the kaggle dataset is rendered by indigo toolkit (use old version 1.4.0-beta)\nhttps://lifescience.opensource.epam.com/indigo/\n\n\n![](https://i.ibb.co/ZzSv1mV/Selection-194.png)",
      "votes": null
    },
    {
      "id": "1335014",
      "postDate": "06/04/2021 00:44:50",
      "content": "<p>Wow really missed that. We were training with 1m extra images created by rdkit. It didn't hurt but did't help much.</p>",
      "rawMarkdown": "Wow really missed that. We were training with 1m extra images created by rdkit. It didn't hurt but did't help much.",
      "votes": null
    },
    {
      "id": "1335018",
      "postDate": "06/04/2021 00:55:21",
      "content": "<p><img src=\"https://i.ibb.co/4Jqh27S/Selection-195.png\" alt=\"\"></p>",
      "rawMarkdown": "![](https://i.ibb.co/4Jqh27S/Selection-195.png)",
      "votes": null
    },
    {
      "id": "1335019",
      "postDate": "06/04/2021 00:55:59",
      "content": "<p>This means your model outperforms ours. A lot to learn from different teams! Big congrats to GM!</p>",
      "rawMarkdown": "This means your model outperforms ours. A lot to learn from different teams! Big congrats to GM!",
      "votes": null
    },
    {
      "id": "1335021",
      "postDate": "06/04/2021 00:56:41",
      "content": "<p>i trained with extra 2 million indigo images. i didn't exactly analyse/compare the results (due to lack of time). But it seems that it didn't help much in local cross validation.</p>\n<p>I may think that learning super resolution is a good solution, but i didn't have time to implement. </p>\n<p>If we validate only using clean indigo images, local CV score is about 0.65 (instead of 1.20 with kaggle noisy images)</p>\n<hr>\n<p>with the original indigo image, we are able to analyze/model the domain difference of test and train data. In fact, our local CV and public are almost identical. This makes it easier to design our experiments.</p>",
      "rawMarkdown": "i trained with extra 2 million indigo images. i didn't exactly analyse/compare the results (due to lack of time). But it seems that it didn't help much in local cross validation.\n\nI may think that learning super resolution is a good solution, but i didn't have time to implement. \n\nIf we validate only using clean indigo images, local CV score is about 0.65 (instead of 1.20 with kaggle noisy images)\n\n---\n\nwith the original indigo image, we are able to analyze/model the domain difference of test and train data. In fact, our local CV and public are almost identical. This makes it easier to design our experiments.",
      "votes": null
    },
    {
      "id": "1335036",
      "postDate": "06/04/2021 01:20:39",
      "content": "<p>example of results:<br>\n<img src=\"https://i.ibb.co/ZzL5HLk/Selection-196.png\" alt=\"\"></p>",
      "rawMarkdown": "example of results:\n![](https://i.ibb.co/ZzL5HLk/Selection-196.png)",
      "votes": null
    },
    {
      "id": "1335044",
      "postDate": "06/04/2021 01:33:55",
      "content": "<p>Big congrats and thank you for sharing your code!</p>",
      "rawMarkdown": "Big congrats and thank you for sharing your code!",
      "votes": null
    },
    {
      "id": "1335342",
      "postDate": "06/04/2021 06:54:25",
      "content": "<p>Big Congratulations <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , I am sure a lot of people have got their medals owing to your pipeline . Using object detection is another good idea , great work</p>",
      "rawMarkdown": "Big Congratulations @hengck23 , I am sure a lot of people have got their medals owing to your pipeline . Using object detection is another good idea , great work",
      "votes": null
    },
    {
      "id": "1335432",
      "postDate": "06/04/2021 07:55:12",
      "content": "<p>Dear frog bro</p>\n<p>why I cannot get a score as same as yours <br>\nI also training a VIT from scratch on TPU , I think the score should be more near , <br>\nbut after 10 epoch  I only get 88% Acc,</p>\n<pre><code># transformer decoder params\nnum_layer=4\nd_model=768\ndff= 512\nnum_heads=8\ntarget_vocab_size=197\ndropout_rate=0.1 \n\n# vision encoder params\nimage_size = (IMG_SHAPE[0], IMG_SHAPE[1])  # We'll resize input images to this size\npatch_size = 32  # Size of the patches to be extract from the input images\nnum_patches = (image_size[0] // patch_size) * (image_size[1] // patch_size) \nprojection_dim = d_model\nvit_heads = 12\ntransformer_units = [\n    projection_dim * 2,\n    projection_dim,\n]  # Size of the transformer layers\ntransformer_layers = 12\n</code></pre>\n<p>I use gradient clip to keep stable</p>\n<pre><code>gradients, _ = tf.clip_by_global_norm(gradients, 10.0)\n</code></pre>\n<p>Here is my notebook:  maybe you can take a look at the model part<br>\n<a href=\"https://www.kaggle.com/drzhuzhe/bms-vision-transformer?scriptVersionId=64760387\" target=\"_blank\">https://www.kaggle.com/drzhuzhe/bms-vision-transformer?scriptVersionId=64760387</a></p>\n<p>training is using spike learning rate</p>\n<p>It's the first time i implement Transformer via bare hand </p>",
      "rawMarkdown": "Dear frog bro\n\n\nwhy I cannot get a score as same as yours \nI also training a VIT from scratch on TPU , I think the score should be more near , \nbut after 10 epoch  I only get 88% Acc,\n\n```\n# transformer decoder params\nnum_layer=4\nd_model=768\ndff= 512\nnum_heads=8\ntarget_vocab_size=197\ndropout_rate=0.1 \n\n# vision encoder params\nimage_size = (IMG_SHAPE[0], IMG_SHAPE[1])  # We'll resize input images to this size\npatch_size = 32  # Size of the patches to be extract from the input images\nnum_patches = (image_size[0] // patch_size) * (image_size[1] // patch_size) \nprojection_dim = d_model\nvit_heads = 12\ntransformer_units = [\n    projection_dim * 2,\n    projection_dim,\n]  # Size of the transformer layers\ntransformer_layers = 12\n```\n\n I use gradient clip to keep stable\n\n```\ngradients, _ = tf.clip_by_global_norm(gradients, 10.0)\n```\n\nHere is my notebook:  maybe you can take a look at the model part\nhttps://www.kaggle.com/drzhuzhe/bms-vision-transformer?scriptVersionId=64760387\n\ntraining is using spike learning rate\n\nIt's the first time i implement Transformer via bare hand",
      "votes": null
    },
    {
      "id": "1335742",
      "postDate": "06/04/2021 12:16:04",
      "content": "<p>I've been looking for this for a long time and tried different renderers, but I couldn't find the one used in the competition, so I gave up. :D</p>",
      "rawMarkdown": "I've been looking for this for a long time and tried different renderers, but I couldn't find the one used in the competition, so I gave up. :D",
      "votes": null
    },
    {
      "id": "1336028",
      "postDate": "06/04/2021 15:35:41",
      "content": "<p>Congrats for the gold medal ended, and thanks for sharing, always! ;)</p>",
      "rawMarkdown": "Congrats for the gold medal ended, and thanks for sharing, always! ;)",
      "votes": null
    },
    {
      "id": "1336111",
      "postDate": "06/04/2021 16:20:03",
      "content": "<p>Ha-ha, I wonder how many other teams did discover this as well :)</p>",
      "rawMarkdown": "Ha-ha, I wonder how many other teams did discover this as well :)",
      "votes": null
    },
    {
      "id": "1336227",
      "postDate": "06/04/2021 18:21:45",
      "content": "<p>Always a pleasure reading your posts. Congratulations on your place!</p>",
      "rawMarkdown": "Always a pleasure reading your posts. Congratulations on your place!",
      "votes": null
    },
    {
      "id": "1343154",
      "postDate": "06/10/2021 04:43:21",
      "content": "<p>we also make some modifications to our patch-based transformer.<br>\nwe note that resize artifacts are only present in the resized images.<br>\nwe want the model to handle resized and original images differently.</p>\n<p><img src=\"https://i.ibb.co/nC5tC3N/Selection-201.png\" alt=\"\"></p>",
      "rawMarkdown": "we also make some modifications to our patch-based transformer.\nwe note that resize artifacts are only present in the resized images.\nwe want the model to handle resized and original images differently.\n\n![](https://i.ibb.co/nC5tC3N/Selection-201.png)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1335009,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/04/2021 00:38:00",
      "content": "<p>the kaggle dataset is rendered by indigo toolkit (use old version 1.4.0-beta)<br>\n<a href=\"https://lifescience.opensource.epam.com/indigo/\" target=\"_blank\">https://lifescience.opensource.epam.com/indigo/</a></p>\n<p><img src=\"https://i.ibb.co/ZzSv1mV/Selection-194.png\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1335014,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "06/04/2021 00:44:50",
          "content": "<p>Wow really missed that. We were training with 1m extra images created by rdkit. It didn't hurt but did't help much.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1335019,
          "author_name": "houndcl",
          "author_url": "",
          "post_date": "06/04/2021 00:55:59",
          "content": "<p>This means your model outperforms ours. A lot to learn from different teams! Big congrats to GM!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1335021,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "06/04/2021 00:56:41",
          "content": "<p>i trained with extra 2 million indigo images. i didn't exactly analyse/compare the results (due to lack of time). But it seems that it didn't help much in local cross validation.</p>\n<p>I may think that learning super resolution is a good solution, but i didn't have time to implement. </p>\n<p>If we validate only using clean indigo images, local CV score is about 0.65 (instead of 1.20 with kaggle noisy images)</p>\n<hr>\n<p>with the original indigo image, we are able to analyze/model the domain difference of test and train data. In fact, our local CV and public are almost identical. This makes it easier to design our experiments.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1335742,
          "author_name": "sorkun",
          "author_url": "",
          "post_date": "06/04/2021 12:16:04",
          "content": "<p>I've been looking for this for a long time and tried different renderers, but I couldn't find the one used in the competition, so I gave up. :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1336111,
          "author_name": "stassl",
          "author_url": "",
          "post_date": "06/04/2021 16:20:03",
          "content": "<p>Ha-ha, I wonder how many other teams did discover this as well :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335018,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/04/2021 00:55:21",
      "content": "<p><img src=\"https://i.ibb.co/4Jqh27S/Selection-195.png\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1335036,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "06/04/2021 01:20:39",
          "content": "<p>example of results:<br>\n<img src=\"https://i.ibb.co/ZzL5HLk/Selection-196.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335044,
      "author_name": "hirotetsu",
      "author_url": "",
      "post_date": "06/04/2021 01:33:55",
      "content": "<p>Big congrats and thank you for sharing your code!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1335342,
      "author_name": "tanulsingh077",
      "author_url": "",
      "post_date": "06/04/2021 06:54:25",
      "content": "<p>Big Congratulations <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , I am sure a lot of people have got their medals owing to your pipeline . Using object detection is another good idea , great work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1335432,
      "author_name": "drzhuzhe",
      "author_url": "",
      "post_date": "06/04/2021 07:55:12",
      "content": "<p>Dear frog bro</p>\n<p>why I cannot get a score as same as yours <br>\nI also training a VIT from scratch on TPU , I think the score should be more near , <br>\nbut after 10 epoch  I only get 88% Acc,</p>\n<pre><code># transformer decoder params\nnum_layer=4\nd_model=768\ndff= 512\nnum_heads=8\ntarget_vocab_size=197\ndropout_rate=0.1 \n\n# vision encoder params\nimage_size = (IMG_SHAPE[0], IMG_SHAPE[1])  # We'll resize input images to this size\npatch_size = 32  # Size of the patches to be extract from the input images\nnum_patches = (image_size[0] // patch_size) * (image_size[1] // patch_size) \nprojection_dim = d_model\nvit_heads = 12\ntransformer_units = [\n    projection_dim * 2,\n    projection_dim,\n]  # Size of the transformer layers\ntransformer_layers = 12\n</code></pre>\n<p>I use gradient clip to keep stable</p>\n<pre><code>gradients, _ = tf.clip_by_global_norm(gradients, 10.0)\n</code></pre>\n<p>Here is my notebook:  maybe you can take a look at the model part<br>\n<a href=\"https://www.kaggle.com/drzhuzhe/bms-vision-transformer?scriptVersionId=64760387\" target=\"_blank\">https://www.kaggle.com/drzhuzhe/bms-vision-transformer?scriptVersionId=64760387</a></p>\n<p>training is using spike learning rate</p>\n<p>It's the first time i implement Transformer via bare hand </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1336028,
      "author_name": "haqishen",
      "author_url": "",
      "post_date": "06/04/2021 15:35:41",
      "content": "<p>Congrats for the gold medal ended, and thanks for sharing, always! ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1336227,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "06/04/2021 18:21:45",
      "content": "<p>Always a pleasure reading your posts. Congratulations on your place!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1343154,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/10/2021 04:43:21",
      "content": "<p>we also make some modifications to our patch-based transformer.<br>\nwe note that resize artifacts are only present in the resized images.<br>\nwe want the model to handle resized and original images differently.</p>\n<p><img src=\"https://i.ibb.co/nC5tC3N/Selection-201.png\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1334981": "to be updated ....\n\nthere are two part:\n- part one: TNT/VIT/CAIT ensemble model (public/private LB 0.71/0.71)\nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/243766 (this post)\n- part two: OCR + MolBuilder + object detection CNN (public/private LB 0.62/0.61)\nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/243809\n\n\n---\n\nTNT/VIT/CAIT ensemble model (public/private LB 0.71/0.71)\n\n- we train 5 models : 3x patched based VIT (patch size 16, image scale = 0.8,0.9,0.1), 1x TNT with image size 384 and 1x CAIT with image size 240. For details, please refer to https://www.kaggle.com/c/bms-molecular-translation/discussion/231190\n(local CV 0.78) \n\n- we do 5 iterations of image perturbation. if an image has gives an invalid inchi (verification by rdkit), we perturb the image and try inference again.\n( local CV 0.76) \n\n- finally, we apply rdkit normalisation.\n(public/private LB 0.73/0.74, local CV 0.72) \n\nThe ensemble is our top score submission. We also have submissions for each of the individual models. We then sort all submission cvs by LB score. To create the final submission:\n- 1. We start off will predictions from the top score submission csv. \n- 2. For a current prediction has invalid inchi, it will be replaced from the corresponding one from the next higher score submission csv if ( and only if ) that corresponding one is valid\n- 3. The process is repeated for all csvs. \nFinal LB score is (public/private LB 0.71/0.71, local CV 0.70)\n\n![](https://i.ibb.co/TYt5h7m/Selection-193.png)\n\nwe use random shift and scale to perturb image\n\n---\n\nhere is a tip on ensemble strategy:\n![](https://i.ibb.co/F3Yns33/Selection-202.png)\n\nwe note that lb score is correlated to the number of invalid inchi (verified from rdkit). If you have fewer invalid inchi, you are likely to better LB score.\n\ninstead of making many submissions, you can use the number of invalid inchi to estimate your lb score. To know the limit of LB score of your current method, you can extrapolate your graph to the lowest number of possible inchi. If that doesn't work well, you can change your ensemble strategy so that you have a steeper graph. (i.e. better lb score per invalid inchi)\n\n---\n\nAcknowledgment\n---\n\nI am grateful to Z by HP & NVIDIA for sponsoring me a Z8 Workstation with dual RTX8000 GPU for this competition. In the final days, I need to run TNT/VIT/CAIT ensemble (5 transformers) on the 1.6 million test images. Some of the test images are large and have long sequences. Besides high computative power, we also need high GPU memory.\n\nOne RTX8000 GPU has 48GB onboard. It takes about 7 hours to process all images using the workstation. This is very fast and definitely help our team to secure a gold medal of 10th placing in the final ranking!",
    "1335009": "the kaggle dataset is rendered by indigo toolkit (use old version 1.4.0-beta)\nhttps://lifescience.opensource.epam.com/indigo/\n\n\n![](https://i.ibb.co/ZzSv1mV/Selection-194.png)",
    "1335014": "Wow really missed that. We were training with 1m extra images created by rdkit. It didn't hurt but did't help much.",
    "1335018": "![](https://i.ibb.co/4Jqh27S/Selection-195.png)",
    "1335019": "This means your model outperforms ours. A lot to learn from different teams! Big congrats to GM!",
    "1335021": "i trained with extra 2 million indigo images. i didn't exactly analyse/compare the results (due to lack of time). But it seems that it didn't help much in local cross validation.\n\nI may think that learning super resolution is a good solution, but i didn't have time to implement. \n\nIf we validate only using clean indigo images, local CV score is about 0.65 (instead of 1.20 with kaggle noisy images)\n\n---\n\nwith the original indigo image, we are able to analyze/model the domain difference of test and train data. In fact, our local CV and public are almost identical. This makes it easier to design our experiments.",
    "1335036": "example of results:\n![](https://i.ibb.co/ZzL5HLk/Selection-196.png)",
    "1335044": "Big congrats and thank you for sharing your code!",
    "1335342": "Big Congratulations @hengck23 , I am sure a lot of people have got their medals owing to your pipeline . Using object detection is another good idea , great work",
    "1335432": "Dear frog bro\n\n\nwhy I cannot get a score as same as yours \nI also training a VIT from scratch on TPU , I think the score should be more near , \nbut after 10 epoch  I only get 88% Acc,\n\n```\n# transformer decoder params\nnum_layer=4\nd_model=768\ndff= 512\nnum_heads=8\ntarget_vocab_size=197\ndropout_rate=0.1 \n\n# vision encoder params\nimage_size = (IMG_SHAPE[0], IMG_SHAPE[1])  # We'll resize input images to this size\npatch_size = 32  # Size of the patches to be extract from the input images\nnum_patches = (image_size[0] // patch_size) * (image_size[1] // patch_size) \nprojection_dim = d_model\nvit_heads = 12\ntransformer_units = [\n    projection_dim * 2,\n    projection_dim,\n]  # Size of the transformer layers\ntransformer_layers = 12\n```\n\n I use gradient clip to keep stable\n\n```\ngradients, _ = tf.clip_by_global_norm(gradients, 10.0)\n```\n\nHere is my notebook:  maybe you can take a look at the model part\nhttps://www.kaggle.com/drzhuzhe/bms-vision-transformer?scriptVersionId=64760387\n\ntraining is using spike learning rate\n\nIt's the first time i implement Transformer via bare hand",
    "1335742": "I've been looking for this for a long time and tried different renderers, but I couldn't find the one used in the competition, so I gave up. :D",
    "1336028": "Congrats for the gold medal ended, and thanks for sharing, always! ;)",
    "1336111": "Ha-ha, I wonder how many other teams did discover this as well :)",
    "1336227": "Always a pleasure reading your posts. Congratulations on your place!",
    "1343154": "we also make some modifications to our patch-based transformer.\nwe note that resize artifacts are only present in the resized images.\nwe want the model to handle resized and original images differently.\n\n![](https://i.ibb.co/nC5tC3N/Selection-201.png)"
  },
  "source": "meta"
}