{
  "id": 243943,
  "title": "5th Place Solution",
  "url": "/competitions/bms-molecular-translation/writeups/all-data-are-ext-5th-place-solution",
  "author_name": "",
  "post_date": "2021-06-04T15:26:03.257Z",
  "votes": 53,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>This is absolutely a tough one, so congratulations to all those who persevered in this competition until the end.</p>\n<p>Thanks to the organizers and congrats to all the winners and my wonderful teammates <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> <a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a> and we must thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> we couldn't have achieved this without your sharing. Thank you so much!</p>\n<h1>Summary</h1>\n<ul>\n<li>We didn't use Transformer Encoder, our architecture is just CNN (384x384 &amp; 512x512) + Transformer Decoder (12~16 layers)</li>\n<li>We trained 5 backbones. EffNet B3/B5/B7, ResNet200D, eca-nfnet-l0</li>\n<li>Noise Injection for regularization.</li>\n<li>Training with some ext data which generated by rdkit. (But not very helpful)</li>\n<li>Ensemble during decode phase.</li>\n<li>Handling invalid predictions (most crucial trick for going under 0.6).</li>\n</ul>\n<h1>Noise Injection</h1>\n<p>This is a very old regularization technique, if I remember correctly someone mentioned it in a paper before 2000. In the field of Image Caption, the accuracy of prediction for the next character is quite high while ensuring that the previous sequence is exactly correct. But once a character is predicted incorrectly, then the next prediction failure rate gradually increases until it runs completely off.</p>\n<p>So I came up with noise injection, which is done by randomly replacing GT characters with other characters during training. The replaced characters are ignored in the calculation of the loss, but this does not mean that this technique has no effect on the prediction. Although the character itself is not computed as part of the loss, the model is then forced to correctly predict the next character on the basis that the previous one was wrong.</p>\n<p>In addition to this, the technique itself has a regularization effect. This allows us to train 12-16 layers of transformer decoder.</p>\n<h1>Official Ext Data</h1>\n<p>The host released 10 million InChI without images as ext data during the competition.</p>\n<p>We found that rdkit can use InChI to generate images, so we generated some of the images (about 1~2 million) as ext data and added them to the training. But later we found that this part of data did not bring much improvement. However, since it did not degrade the performance either, we kept it.</p>\n<h1>Handling Invalid Predictions</h1>\n<p>After training all models we ensemble them in the decode phase.</p>\n<p>The LB of the original prediction is 0.62, then we use rdkit to norm the prediction and get LB 0.60</p>\n<p>At this point we find that there are about 14000 rows in the predicted InChI that are not valid. To replace these invalid predictions, we use 3 methods to find their valid predictions:</p>\n<ul>\n<li>Prediction generated by single model (fixed 5k~ rows)</li>\n<li>Search for top 12 predictions by ensemble (fixed 2k5~ rows)</li>\n<li>For the numerical part of the chemical formula, we use top2 predicted values to decode (fixed 500~ rows)<ul>\n<li>For example, if the original prediction is C12H3, and the top2 predictions for the second and fourth characters are 11 and 4, then we replace them with C11H3, and C12H4 and then decode the next part.</li></ul></li>\n</ul>\n<p>We discovered that replacing invalid predictions with valid predictions could significantly improve LB the night before the competition deadline, due to lack of time, we only had time to dealing with 8000 rows out of the 14000 invalid predictions in the original prediction, which improved our LB score from 0.60 to 0.57</p>\n<h1>Acknowledge</h1>\n<p>Special Thanks to Z by HP &amp; NVIDIA for sponsoring me a Z8G4 Workstation with dual RTX6000 GPU and a ZBook with RTX5000 GPU.<br>\nThis is without a doubt a competition with the biggest dataset in the last half years in kaggle.<br>\nSo I tried pytorch's DDP parallel training on my dual RTX6000 GPUs and the experience was great, basically it can be twice as fast as single GPU training.</p>",
  "messages": [
    {
      "id": "1335989",
      "postDate": "06/04/2021 15:12:38",
      "content": "<p>Hi all,</p>\n<p>This is absolutely a tough one, so congratulations to all those who persevered in this competition until the end.</p>\n<p>Thanks to the organizers and congrats to all the winners and my wonderful teammates <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> <a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a> and we must thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> we couldn't have achieved this without your sharing. Thank you so much!</p>\n<h1>Summary</h1>\n<ul>\n<li>We didn't use Transformer Encoder, our architecture is just CNN (384x384 &amp; 512x512) + Transformer Decoder (12~16 layers)</li>\n<li>We trained 5 backbones. EffNet B3/B5/B7, ResNet200D, eca-nfnet-l0</li>\n<li>Noise Injection for regularization.</li>\n<li>Training with some ext data which generated by rdkit. (But not very helpful)</li>\n<li>Ensemble during decode phase.</li>\n<li>Handling invalid predictions (most crucial trick for going under 0.6).</li>\n</ul>\n<h1>Noise Injection</h1>\n<p>This is a very old regularization technique, if I remember correctly someone mentioned it in a paper before 2000. In the field of Image Caption, the accuracy of prediction for the next character is quite high while ensuring that the previous sequence is exactly correct. But once a character is predicted incorrectly, then the next prediction failure rate gradually increases until it runs completely off.</p>\n<p>So I came up with noise injection, which is done by randomly replacing GT characters with other characters during training. The replaced characters are ignored in the calculation of the loss, but this does not mean that this technique has no effect on the prediction. Although the character itself is not computed as part of the loss, the model is then forced to correctly predict the next character on the basis that the previous one was wrong.</p>\n<p>In addition to this, the technique itself has a regularization effect. This allows us to train 12-16 layers of transformer decoder.</p>\n<h1>Official Ext Data</h1>\n<p>The host released 10 million InChI without images as ext data during the competition.</p>\n<p>We found that rdkit can use InChI to generate images, so we generated some of the images (about 1~2 million) as ext data and added them to the training. But later we found that this part of data did not bring much improvement. However, since it did not degrade the performance either, we kept it.</p>\n<h1>Handling Invalid Predictions</h1>\n<p>After training all models we ensemble them in the decode phase.</p>\n<p>The LB of the original prediction is 0.62, then we use rdkit to norm the prediction and get LB 0.60</p>\n<p>At this point we find that there are about 14000 rows in the predicted InChI that are not valid. To replace these invalid predictions, we use 3 methods to find their valid predictions:</p>\n<ul>\n<li>Prediction generated by single model (fixed 5k~ rows)</li>\n<li>Search for top 12 predictions by ensemble (fixed 2k5~ rows)</li>\n<li>For the numerical part of the chemical formula, we use top2 predicted values to decode (fixed 500~ rows)<ul>\n<li>For example, if the original prediction is C12H3, and the top2 predictions for the second and fourth characters are 11 and 4, then we replace them with C11H3, and C12H4 and then decode the next part.</li></ul></li>\n</ul>\n<p>We discovered that replacing invalid predictions with valid predictions could significantly improve LB the night before the competition deadline, due to lack of time, we only had time to dealing with 8000 rows out of the 14000 invalid predictions in the original prediction, which improved our LB score from 0.60 to 0.57</p>\n<h1>Acknowledge</h1>\n<p>Special Thanks to Z by HP &amp; NVIDIA for sponsoring me a Z8G4 Workstation with dual RTX6000 GPU and a ZBook with RTX5000 GPU.<br>\nThis is without a doubt a competition with the biggest dataset in the last half years in kaggle.<br>\nSo I tried pytorch's DDP parallel training on my dual RTX6000 GPUs and the experience was great, basically it can be twice as fast as single GPU training.</p>",
      "rawMarkdown": "Hi all,\n\nThis is absolutely a tough one, so congratulations to all those who persevered in this competition until the end.\n\nThanks to the organizers and congrats to all the winners and my wonderful teammates @boliu0 @garybios and we must thanks to @hengck23 and @yasufuminakama we couldn't have achieved this without your sharing. Thank you so much!\n\n# Summary\n\n* We didn't use Transformer Encoder, our architecture is just CNN (384x384 & 512x512) + Transformer Decoder (12~16 layers)\n* We trained 5 backbones. EffNet B3/B5/B7, ResNet200D, eca-nfnet-l0\n* Noise Injection for regularization.\n* Training with some ext data which generated by rdkit. (But not very helpful)\n* Ensemble during decode phase.\n* Handling invalid predictions (most crucial trick for going under 0.6).\n\n# Noise Injection\n\nThis is a very old regularization technique, if I remember correctly someone mentioned it in a paper before 2000. In the field of Image Caption, the accuracy of prediction for the next character is quite high while ensuring that the previous sequence is exactly correct. But once a character is predicted incorrectly, then the next prediction failure rate gradually increases until it runs completely off.\n\nSo I came up with noise injection, which is done by randomly replacing GT characters with other characters during training. The replaced characters are ignored in the calculation of the loss, but this does not mean that this technique has no effect on the prediction. Although the character itself is not computed as part of the loss, the model is then forced to correctly predict the next character on the basis that the previous one was wrong.\n\nIn addition to this, the technique itself has a regularization effect. This allows us to train 12-16 layers of transformer decoder.\n\n# Official Ext Data\n\nThe host released 10 million InChI without images as ext data during the competition.\n\nWe found that rdkit can use InChI to generate images, so we generated some of the images (about 1~2 million) as ext data and added them to the training. But later we found that this part of data did not bring much improvement. However, since it did not degrade the performance either, we kept it.\n\n# Handling Invalid Predictions\n\nAfter training all models we ensemble them in the decode phase.\n\nThe LB of the original prediction is 0.62, then we use rdkit to norm the prediction and get LB 0.60\n\nAt this point we find that there are about 14000 rows in the predicted InChI that are not valid. To replace these invalid predictions, we use 3 methods to find their valid predictions:\n\n* Prediction generated by single model (fixed 5k~ rows)\n* Search for top 12 predictions by ensemble (fixed 2k5~ rows)\n* For the numerical part of the chemical formula, we use top2 predicted values to decode (fixed 500~ rows)\n    * For example, if the original prediction is C12H3, and the top2 predictions for the second and fourth characters are 11 and 4, then we replace them with C11H3, and C12H4 and then decode the next part.\n\nWe discovered that replacing invalid predictions with valid predictions could significantly improve LB the night before the competition deadline, due to lack of time, we only had time to dealing with 8000 rows out of the 14000 invalid predictions in the original prediction, which improved our LB score from 0.60 to 0.57\n\n# Acknowledge\n\nSpecial Thanks to Z by HP & NVIDIA for sponsoring me a Z8G4 Workstation with dual RTX6000 GPU and a ZBook with RTX5000 GPU.\nThis is without a doubt a competition with the biggest dataset in the last half years in kaggle.\nSo I tried pytorch's DDP parallel training on my dual RTX6000 GPUs and the experience was great, basically it can be twice as fast as single GPU training.",
      "votes": null
    },
    {
      "id": "1335999",
      "postDate": "06/04/2021 15:18:02",
      "content": "<p>Clever noise injection! Out of interest, what was the impact of this on your scores at the time?</p>",
      "rawMarkdown": "Clever noise injection! Out of interest, what was the impact of this on your scores at the time?",
      "votes": null
    },
    {
      "id": "1336018",
      "postDate": "06/04/2021 15:31:30",
      "content": "<p>It decreased our score by 10% or more when we were doing experiments by small backbones and small images.</p>",
      "rawMarkdown": "It decreased our score by 10% or more when we were doing experiments by small backbones and small images.",
      "votes": null
    },
    {
      "id": "1336021",
      "postDate": "06/04/2021 15:33:59",
      "content": "<p>Nice. Congratulations!</p>",
      "rawMarkdown": "Nice. Congratulations!",
      "votes": null
    },
    {
      "id": "1336026",
      "postDate": "06/04/2021 15:35:29",
      "content": "<p>Congrats on another great finish <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> and team!! You guys always prove we don't need enough submissions to win a competition. </p>\n<blockquote>\n  <p>Search for top 12 predictions by ensemble (fixed 2k5~ rows)</p>\n</blockquote>\n<p>Could you explain a bit more about this please, I didn't get it clearly.</p>",
      "rawMarkdown": "Congrats on another great finish @haqishen and team!! You guys always prove we don't need enough submissions to win a competition. \n> Search for top 12 predictions by ensemble (fixed 2k5~ rows)\n\nCould you explain a bit more about this please, I didn't get it clearly.",
      "votes": null
    },
    {
      "id": "1336107",
      "postDate": "06/04/2021 16:17:13",
      "content": "<p>We didn't have time to implement a correct beam search, so here our top 12 is only based on the original prediction confidence for all characters. Then pick up 12 characters with lowerest confidence, replace them with second highest confidence character then do the rest decoding, one by one. As a result we got 12 predicted sequence for each invalid sample.</p>",
      "rawMarkdown": "We didn't have time to implement a correct beam search, so here our top 12 is only based on the original prediction confidence for all characters. Then pick up 12 characters with lowerest confidence, replace them with second highest confidence character then do the rest decoding, one by one. As a result we got 12 predicted sequence for each invalid sample.",
      "votes": null
    },
    {
      "id": "1336155",
      "postDate": "06/04/2021 17:05:46",
      "content": "<p>Got it, Smart and time-efficient strategy I also thought of doing the same, unfortunately, it was too late for us to implement that :)  Congrats again on a great finish.  </p>",
      "rawMarkdown": "Got it, Smart and time-efficient strategy I also thought of doing the same, unfortunately, it was too late for us to implement that :)  Congrats again on a great finish.",
      "votes": null
    },
    {
      "id": "1336386",
      "postDate": "06/04/2021 21:00:19",
      "content": "<p>Wonderful summary. Great job!</p>",
      "rawMarkdown": "Wonderful summary. Great job!",
      "votes": null
    },
    {
      "id": "1336464",
      "postDate": "06/05/2021 00:13:41",
      "content": "<p>thx for sharing. </p>",
      "rawMarkdown": "thx for sharing.",
      "votes": null
    },
    {
      "id": "1336551",
      "postDate": "06/05/2021 03:16:33",
      "content": "<p>Congratulations! The noise injection idea is marvelous</p>",
      "rawMarkdown": "Congratulations! The noise injection idea is marvelous",
      "votes": null
    },
    {
      "id": "1336934",
      "postDate": "06/05/2021 09:50:17",
      "content": "<p>A wonderful write up! I like the noise injection part, it is pretty creative.</p>",
      "rawMarkdown": "A wonderful write up! I like the noise injection part, it is pretty creative.",
      "votes": null
    },
    {
      "id": "1338133",
      "postDate": "06/06/2021 07:56:11",
      "content": "<p>Noise injection is really smart!</p>",
      "rawMarkdown": "Noise injection is really smart!",
      "votes": null
    },
    {
      "id": "1339118",
      "postDate": "06/07/2021 03:54:35",
      "content": "<p>Congrats!<br>\nOur models also adopted Noise Injection.<br>\nWe randomly replaced 10-15%, but what percentage did you adopt?</p>",
      "rawMarkdown": "Congrats!\nOur models also adopted Noise Injection.\nWe randomly replaced 10-15%, but what percentage did you adopt?",
      "votes": null
    },
    {
      "id": "1339388",
      "postDate": "06/07/2021 08:07:35",
      "content": "<p>Great idea about the noise injection. We actually tried to solve the same problem by removing teacher forcing during the training. However, with this approach both the advantage and the downside that it would use all previous steps in the gradient calculation, which should help the model, but requires more memory.</p>",
      "rawMarkdown": "Great idea about the noise injection. We actually tried to solve the same problem by removing teacher forcing during the training. However, with this approach both the advantage and the downside that it would use all previous steps in the gradient calculation, which should help the model, but requires more memory.",
      "votes": null
    },
    {
      "id": "1340508",
      "postDate": "06/08/2021 03:51:17",
      "content": "<p>Congrats to you, too!<br>\nWe use 20~25% in our training.</p>",
      "rawMarkdown": "Congrats to you, too!\nWe use 20~25% in our training.",
      "votes": null
    },
    {
      "id": "1341769",
      "postDate": "06/09/2021 01:32:09",
      "content": "<p>great!<br>\nthanks!</p>",
      "rawMarkdown": "great!\nthanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1335999,
      "author_name": "fergusoci",
      "author_url": "",
      "post_date": "06/04/2021 15:18:02",
      "content": "<p>Clever noise injection! Out of interest, what was the impact of this on your scores at the time?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1336018,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "06/04/2021 15:31:30",
          "content": "<p>It decreased our score by 10% or more when we were doing experiments by small backbones and small images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1336021,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "06/04/2021 15:33:59",
          "content": "<p>Nice. Congratulations!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1336026,
      "author_name": "nischaydnk",
      "author_url": "",
      "post_date": "06/04/2021 15:35:29",
      "content": "<p>Congrats on another great finish <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> and team!! You guys always prove we don't need enough submissions to win a competition. </p>\n<blockquote>\n  <p>Search for top 12 predictions by ensemble (fixed 2k5~ rows)</p>\n</blockquote>\n<p>Could you explain a bit more about this please, I didn't get it clearly.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1336107,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "06/04/2021 16:17:13",
          "content": "<p>We didn't have time to implement a correct beam search, so here our top 12 is only based on the original prediction confidence for all characters. Then pick up 12 characters with lowerest confidence, replace them with second highest confidence character then do the rest decoding, one by one. As a result we got 12 predicted sequence for each invalid sample.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1336155,
          "author_name": "nischaydnk",
          "author_url": "",
          "post_date": "06/04/2021 17:05:46",
          "content": "<p>Got it, Smart and time-efficient strategy I also thought of doing the same, unfortunately, it was too late for us to implement that :)  Congrats again on a great finish.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1336386,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "06/04/2021 21:00:19",
      "content": "<p>Wonderful summary. Great job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1336464,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "06/05/2021 00:13:41",
      "content": "<p>thx for sharing. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1336551,
      "author_name": "wowfattie",
      "author_url": "",
      "post_date": "06/05/2021 03:16:33",
      "content": "<p>Congratulations! The noise injection idea is marvelous</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1336934,
      "author_name": "alexlwh",
      "author_url": "",
      "post_date": "06/05/2021 09:50:17",
      "content": "<p>A wonderful write up! I like the noise injection part, it is pretty creative.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1338133,
      "author_name": "vigneshbaskaran",
      "author_url": "",
      "post_date": "06/06/2021 07:56:11",
      "content": "<p>Noise injection is really smart!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1339118,
      "author_name": "atsunorifujita",
      "author_url": "",
      "post_date": "06/07/2021 03:54:35",
      "content": "<p>Congrats!<br>\nOur models also adopted Noise Injection.<br>\nWe randomly replaced 10-15%, but what percentage did you adopt?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1340508,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "06/08/2021 03:51:17",
          "content": "<p>Congrats to you, too!<br>\nWe use 20~25% in our training.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1341769,
          "author_name": "atsunorifujita",
          "author_url": "",
          "post_date": "06/09/2021 01:32:09",
          "content": "<p>great!<br>\nthanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1339388,
      "author_name": "igorshulgan",
      "author_url": "",
      "post_date": "06/07/2021 08:07:35",
      "content": "<p>Great idea about the noise injection. We actually tried to solve the same problem by removing teacher forcing during the training. However, with this approach both the advantage and the downside that it would use all previous steps in the gradient calculation, which should help the model, but requires more memory.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1335989": "Hi all,\n\nThis is absolutely a tough one, so congratulations to all those who persevered in this competition until the end.\n\nThanks to the organizers and congrats to all the winners and my wonderful teammates @boliu0 @garybios and we must thanks to @hengck23 and @yasufuminakama we couldn't have achieved this without your sharing. Thank you so much!\n\n# Summary\n\n* We didn't use Transformer Encoder, our architecture is just CNN (384x384 & 512x512) + Transformer Decoder (12~16 layers)\n* We trained 5 backbones. EffNet B3/B5/B7, ResNet200D, eca-nfnet-l0\n* Noise Injection for regularization.\n* Training with some ext data which generated by rdkit. (But not very helpful)\n* Ensemble during decode phase.\n* Handling invalid predictions (most crucial trick for going under 0.6).\n\n# Noise Injection\n\nThis is a very old regularization technique, if I remember correctly someone mentioned it in a paper before 2000. In the field of Image Caption, the accuracy of prediction for the next character is quite high while ensuring that the previous sequence is exactly correct. But once a character is predicted incorrectly, then the next prediction failure rate gradually increases until it runs completely off.\n\nSo I came up with noise injection, which is done by randomly replacing GT characters with other characters during training. The replaced characters are ignored in the calculation of the loss, but this does not mean that this technique has no effect on the prediction. Although the character itself is not computed as part of the loss, the model is then forced to correctly predict the next character on the basis that the previous one was wrong.\n\nIn addition to this, the technique itself has a regularization effect. This allows us to train 12-16 layers of transformer decoder.\n\n# Official Ext Data\n\nThe host released 10 million InChI without images as ext data during the competition.\n\nWe found that rdkit can use InChI to generate images, so we generated some of the images (about 1~2 million) as ext data and added them to the training. But later we found that this part of data did not bring much improvement. However, since it did not degrade the performance either, we kept it.\n\n# Handling Invalid Predictions\n\nAfter training all models we ensemble them in the decode phase.\n\nThe LB of the original prediction is 0.62, then we use rdkit to norm the prediction and get LB 0.60\n\nAt this point we find that there are about 14000 rows in the predicted InChI that are not valid. To replace these invalid predictions, we use 3 methods to find their valid predictions:\n\n* Prediction generated by single model (fixed 5k~ rows)\n* Search for top 12 predictions by ensemble (fixed 2k5~ rows)\n* For the numerical part of the chemical formula, we use top2 predicted values to decode (fixed 500~ rows)\n    * For example, if the original prediction is C12H3, and the top2 predictions for the second and fourth characters are 11 and 4, then we replace them with C11H3, and C12H4 and then decode the next part.\n\nWe discovered that replacing invalid predictions with valid predictions could significantly improve LB the night before the competition deadline, due to lack of time, we only had time to dealing with 8000 rows out of the 14000 invalid predictions in the original prediction, which improved our LB score from 0.60 to 0.57\n\n# Acknowledge\n\nSpecial Thanks to Z by HP & NVIDIA for sponsoring me a Z8G4 Workstation with dual RTX6000 GPU and a ZBook with RTX5000 GPU.\nThis is without a doubt a competition with the biggest dataset in the last half years in kaggle.\nSo I tried pytorch's DDP parallel training on my dual RTX6000 GPUs and the experience was great, basically it can be twice as fast as single GPU training.",
    "1335999": "Clever noise injection! Out of interest, what was the impact of this on your scores at the time?",
    "1336018": "It decreased our score by 10% or more when we were doing experiments by small backbones and small images.",
    "1336021": "Nice. Congratulations!",
    "1336026": "Congrats on another great finish @haqishen and team!! You guys always prove we don't need enough submissions to win a competition. \n> Search for top 12 predictions by ensemble (fixed 2k5~ rows)\n\nCould you explain a bit more about this please, I didn't get it clearly.",
    "1336107": "We didn't have time to implement a correct beam search, so here our top 12 is only based on the original prediction confidence for all characters. Then pick up 12 characters with lowerest confidence, replace them with second highest confidence character then do the rest decoding, one by one. As a result we got 12 predicted sequence for each invalid sample.",
    "1336155": "Got it, Smart and time-efficient strategy I also thought of doing the same, unfortunately, it was too late for us to implement that :)  Congrats again on a great finish.",
    "1336386": "Wonderful summary. Great job!",
    "1336464": "thx for sharing.",
    "1336551": "Congratulations! The noise injection idea is marvelous",
    "1336934": "A wonderful write up! I like the noise injection part, it is pretty creative.",
    "1338133": "Noise injection is really smart!",
    "1339118": "Congrats!\nOur models also adopted Noise Injection.\nWe randomly replaced 10-15%, but what percentage did you adopt?",
    "1339388": "Great idea about the noise injection. We actually tried to solve the same problem by removing teacher forcing during the training. However, with this approach both the advantage and the downside that it would use all previous steps in the gradient calculation, which should help the model, but requires more memory.",
    "1340508": "Congrats to you, too!\nWe use 20~25% in our training.",
    "1341769": "great!\nthanks!"
  },
  "source": "meta"
}