{
  "id": 243824,
  "title": "12th place...",
  "url": "/competitions/bms-molecular-translation/discussion/243824",
  "author_name": "fergusoci",
  "post_date": "2021-06-04T06:32:33.140000",
  "votes": 86,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Uuuuff. Kaggle is tough! The last two months have felt like a really hard slog. I was trying to go for a solo gold but fell just short - congratulations to all the winners, and everyone in the medal zone, it was a pleasure to compete with you!</p>\n<p>Some recognition before I go into solution details: <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> - delighted you made it to the gold zone. My strongest models were based on the solutions which you so generously opensourced early on. I learnt a huge amount from team mates in the last competition that I entered (@christofhenkel, <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>, <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>, <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a>) - I don't think I would have done as well had it not been for that. Cheers.</p>\n<p><strong>Solution strategy:</strong></p>\n<p>Given the size of both the dataset and the highest performing models, it became clear early on that I would need to make a strategy call: train multiple smaller models and ensemble them, or train a single large model and try to push the limit of what that could do. </p>\n<p>I felt that this competition was similar to <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles\" target=\"_blank\">Lyft</a> in terms of data size. My main learning from that competition was that patience can be a virtue: there is value in simply training a model for far longer than you think you should. Improvements are slow, but they are steady. </p>\n<p>To this end, I focussed on a single VIT transformer (based on <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, 384 x 384 inputs). My best single checkpoint ultimately made it to 0.67 LB. Each epoch took c. 24 hours (!!!). Ensembled they reached 0.63.</p>\n<p>When experimenting early on I trained a LSTM+attention model based on <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> 's code with a resnest200 backbone. This reached 1.11 LB so I included it in my final ensemble for a small amount of diversification. </p>\n<p><strong>Data:</strong></p>\n<p>Both models were fit using all of the data provided, plus an additional 10% samples per epoch of randomly generated images from <code>extra_approved_InChIs.csv</code>. I based the image generation on <a href=\"https://www.kaggle.com/stainsby\" target=\"_blank\">@stainsby</a> 's <a href=\"https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v3\" target=\"_blank\">notebook</a>. Augmentation comprised random rotation (plus coarse dropout in the case of the LSTM model). The final ensemble set also included models that had been fit on pseudo-labelled test data.</p>\n<p><strong>Ensembling strategy:</strong></p>\n<p>Throughout the competition I had been using <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a>'s InChI normalisation as the basis for ensembling: each prediction was categorized as 1) normalized, where normalized result = prediction; 2) normalized, with modifications; 3) normalization fails. </p>\n<p>For each image in the test set I isolated the predictions with the lowest normalization category. Within this group I used simple voting, with ties decided by model preference (i.e. model were ordered highest preference to lowest preference based on their LB scores).</p>\n<p>I ensembled 7 checkpoints for the VIT model + 2 for the Resnest200/LSTM model. Each sample was computed twice per checkpoint: once in its original form, once flipped L-&gt;R.</p>\n<p>Congrats, all.</p>",
  "messages": [
    {
      "id": 1335307,
      "postDate": "2021-06-04T06:32:33.140Z",
      "content": "<p>Uuuuff. Kaggle is tough! The last two months have felt like a really hard slog. I was trying to go for a solo gold but fell just short - congratulations to all the winners, and everyone in the medal zone, it was a pleasure to compete with you!</p>\n<p>Some recognition before I go into solution details: <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> - delighted you made it to the gold zone. My strongest models were based on the solutions which you so generously opensourced early on. I learnt a huge amount from team mates in the last competition that I entered (@christofhenkel, <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>, <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>, <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a>) - I don't think I would have done as well had it not been for that. Cheers.</p>\n<p><strong>Solution strategy:</strong></p>\n<p>Given the size of both the dataset and the highest performing models, it became clear early on that I would need to make a strategy call: train multiple smaller models and ensemble them, or train a single large model and try to push the limit of what that could do. </p>\n<p>I felt that this competition was similar to <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles\" target=\"_blank\">Lyft</a> in terms of data size. My main learning from that competition was that patience can be a virtue: there is value in simply training a model for far longer than you think you should. Improvements are slow, but they are steady. </p>\n<p>To this end, I focussed on a single VIT transformer (based on <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, 384 x 384 inputs). My best single checkpoint ultimately made it to 0.67 LB. Each epoch took c. 24 hours (!!!). Ensembled they reached 0.63.</p>\n<p>When experimenting early on I trained a LSTM+attention model based on <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> 's code with a resnest200 backbone. This reached 1.11 LB so I included it in my final ensemble for a small amount of diversification. </p>\n<p><strong>Data:</strong></p>\n<p>Both models were fit using all of the data provided, plus an additional 10% samples per epoch of randomly generated images from <code>extra_approved_InChIs.csv</code>. I based the image generation on <a href=\"https://www.kaggle.com/stainsby\" target=\"_blank\">@stainsby</a> 's <a href=\"https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v3\" target=\"_blank\">notebook</a>. Augmentation comprised random rotation (plus coarse dropout in the case of the LSTM model). The final ensemble set also included models that had been fit on pseudo-labelled test data.</p>\n<p><strong>Ensembling strategy:</strong></p>\n<p>Throughout the competition I had been using <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a>'s InChI normalisation as the basis for ensembling: each prediction was categorized as 1) normalized, where normalized result = prediction; 2) normalized, with modifications; 3) normalization fails. </p>\n<p>For each image in the test set I isolated the predictions with the lowest normalization category. Within this group I used simple voting, with ties decided by model preference (i.e. model were ordered highest preference to lowest preference based on their LB scores).</p>\n<p>I ensembled 7 checkpoints for the VIT model + 2 for the Resnest200/LSTM model. Each sample was computed twice per checkpoint: once in its original form, once flipped L-&gt;R.</p>\n<p>Congrats, all.</p>",
      "rawMarkdown": "Uuuuff. Kaggle is tough! The last two months have felt like a really hard slog. I was trying to go for a solo gold but fell just short - congratulations to all the winners, and everyone in the medal zone, it was a pleasure to compete with you!\n\nSome recognition before I go into solution details: @hengck23 - delighted you made it to the gold zone. My strongest models were based on the solutions which you so generously opensourced early on. I learnt a huge amount from team mates in the last competition that I entered (@christofhenkel, @ilu000, @philippsinger, @robikscube) - I don't think I would have done as well had it not been for that. Cheers.\n\n**Solution strategy:**\n\nGiven the size of both the dataset and the highest performing models, it became clear early on that I would need to make a strategy call: train multiple smaller models and ensemble them, or train a single large model and try to push the limit of what that could do. \n\nI felt that this competition was similar to [Lyft](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles) in terms of data size. My main learning from that competition was that patience can be a virtue: there is value in simply training a model for far longer than you think you should. Improvements are slow, but they are steady. \n\nTo this end, I focussed on a single VIT transformer (based on @hengck23, 384 x 384 inputs). My best single checkpoint ultimately made it to 0.67 LB. Each epoch took c. 24 hours (!!!). Ensembled they reached 0.63.\n\nWhen experimenting early on I trained a LSTM+attention model based on @yasufuminakama 's code with a resnest200 backbone. This reached 1.11 LB so I included it in my final ensemble for a small amount of diversification. \n\n**Data:**\n\nBoth models were fit using all of the data provided, plus an additional 10% samples per epoch of randomly generated images from `extra_approved_InChIs.csv`. I based the image generation on @stainsby 's [notebook](https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v3). Augmentation comprised random rotation (plus coarse dropout in the case of the LSTM model). The final ensemble set also included models that had been fit on pseudo-labelled test data.\n\n**Ensembling strategy:**\n\nThroughout the competition I had been using @nofreewill's InChI normalisation as the basis for ensembling: each prediction was categorized as 1) normalized, where normalized result = prediction; 2) normalized, with modifications; 3) normalization fails. \n\nFor each image in the test set I isolated the predictions with the lowest normalization category. Within this group I used simple voting, with ties decided by model preference (i.e. model were ordered highest preference to lowest preference based on their LB scores).\n\nI ensembled 7 checkpoints for the VIT model + 2 for the Resnest200/LSTM model. Each sample was computed twice per checkpoint: once in its original form, once flipped L->R.\n\nCongrats, all.",
      "votes": 86
    },
    {
      "id": 1335928,
      "postDate": "2021-06-04T14:16:54.170Z",
      "content": "<p>I am sure you are disappointed Ciara but you can be proud of an amazing performance in a very competitive competition as a solo competitor fighting against talent and hardware! Great job!</p>",
      "rawMarkdown": "I am sure you are disappointed Ciara but you can be proud of an amazing performance in a very competitive competition as a solo competitor fighting against talent and hardware! Great job!",
      "votes": 11,
      "replies": [
        {
          "id": 1335944,
          "postDate": "2021-06-04T14:28:34.873Z",
          "content": "<p>Cheers. I'm going to go sleep for a week. Might have the capacity for a brighter outlook on it after that!</p>",
          "rawMarkdown": "Cheers. I'm going to go sleep for a week. Might have the capacity for a brighter outlook on it after that!",
          "votes": 12
        }
      ]
    },
    {
      "id": 1335663,
      "postDate": "2021-06-04T11:00:09.060Z",
      "content": "<p>Amazing work! Not the gold medal but you gained the respect of me and many kagglers by having a such good solo score. Also, all other solo competitors below 1.00 did a great job. Congratulations all!</p>",
      "rawMarkdown": "Amazing work! Not the gold medal but you gained the respect of me and many kagglers by having a such good solo score. Also, all other solo competitors below 1.00 did a great job. Congratulations all!",
      "votes": 3
    },
    {
      "id": 1335404,
      "postDate": "2021-06-04T07:38:14.183Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> for your outstanding result, especially considering the high competition in this challenge (17 GMs in the top 11 teams). I was rooting for you throughout the full competition! I'm sure you will attain the solo-gold medal very soon!</p>",
      "rawMarkdown": "Congratulations @fergusoci for your outstanding result, especially considering the high competition in this challenge (17 GMs in the top 11 teams). I was rooting for you throughout the full competition! I'm sure you will attain the solo-gold medal very soon!",
      "votes": 3
    },
    {
      "id": 1336319,
      "postDate": "2021-06-04T19:30:25.930Z",
      "content": "<blockquote>\n  <p>Uuuuff. Kaggle is tough! The last two months have felt like a really hard slog. I was trying to go for a solo gold but fell just short </p>\n</blockquote>\n<p>I've been there exactly one year ago in the tweet sentiment competition.  One rank short of gold felt like the worst possible outcome at the time…</p>\n<p>Reality is that even if not gold, this is a huge achievement.  As you said Kaggle competitions are tough, especially for solo competitors.  You can be proud of your result.</p>\n<p>I hope we will see you enter new competitions soon.</p>",
      "rawMarkdown": "> Uuuuff. Kaggle is tough! The last two months have felt like a really hard slog. I was trying to go for a solo gold but fell just short \n\nI've been there exactly one year ago in the tweet sentiment competition.  One rank short of gold felt like the worst possible outcome at the time...\n\nReality is that even if not gold, this is a huge achievement.  As you said Kaggle competitions are tough, especially for solo competitors.  You can be proud of your result.\n\nI hope we will see you enter new competitions soon.",
      "votes": 4,
      "replies": [
        {
          "id": 1336338,
          "postDate": "2021-06-04T19:44:09.107Z",
          "content": "<p>Well, fair play for continuing to make solo attempts after that. I'm not sure whether I have the capacity for another one! Life-ex-kaggle circumstances makes it pretty difficult…</p>",
          "rawMarkdown": "Well, fair play for continuing to make solo attempts after that. I'm not sure whether I have the capacity for another one! Life-ex-kaggle circumstances makes it pretty difficult...",
          "votes": 4
        },
        {
          "id": 1336937,
          "postDate": "2021-06-05T09:52:39.750Z",
          "content": "<p>I'm sorry… ! :(</p>",
          "rawMarkdown": "I'm sorry... ! :("
        }
      ]
    },
    {
      "id": 1335732,
      "postDate": "2021-06-04T12:03:50.877Z",
      "content": "<p>Congrats for your strong finish. Single model with 0.63 is amazing!<br>\nAnd don't give up all hope, there is still a little chance, as you know🤗</p>",
      "rawMarkdown": "Congrats for your strong finish. Single model with 0.63 is amazing!\nAnd don't give up all hope, there is still a little chance, as you know🤗",
      "votes": 4,
      "replies": [
        {
          "id": 1335769,
          "postDate": "2021-06-04T12:29:42.400Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a>. Congrats to you guys, great result. Fingers crossed we get a fair result, whatever that is.</p>",
          "rawMarkdown": "Thanks @bamps53. Congrats to you guys, great result. Fingers crossed we get a fair result, whatever that is.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1343650,
      "postDate": "2021-06-10T10:49:48.803Z",
      "content": "<p><a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> , congrats on such an outstanding result. If it’s possible, then could you please tell more detailed information about lstm model with 1.11 LB? How long did you train it and what parameters did you use?</p>",
      "rawMarkdown": "@fergusoci , congrats on such an outstanding result. If it’s possible, then could you please tell more detailed information about lstm model with 1.11 LB? How long did you train it and what parameters did you use?",
      "votes": 1,
      "replies": [
        {
          "id": 1343753,
          "postDate": "2021-06-10T12:12:50.393Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/antonymipt\" target=\"_blank\">@antonymipt</a>. Details for that one were as follows:</p>\n<ul>\n<li>Encoder backbone: resnest200e (timm)</li>\n<li>Optimizer: Adam</li>\n<li>Scheduler: cosine annealing</li>\n<li>Input size: 320 x 320</li>\n<li>Batch size: 32</li>\n<li>LR: 15e-5</li>\n<li>Attention dim: 512</li>\n<li>Embedding dim: 512</li>\n<li>Decoder dim: 2048</li>\n<li>mixed precision</li>\n<li># epochs: 20, each epoch contained all of the training data + additional randomly generated samples from <code>extra_approved_InChIs.csv</code>. Augmentation: Coarse dropout, random rotations</li>\n<li>5 additional epochs on pseudo labelled test data with a lower LR (5e-5)</li>\n</ul>\n<p>The choice of backbone was the biggest differentiator here. I experimented with many (efficientnets, efficientnet v2, different sizes of resnext, resnest, also different attention/embedding/decoder dims). Resnest200 took a long time to train, but the additional capacity meant it could achieve higher accuracy in the end.</p>",
          "rawMarkdown": "Hi @antonymipt. Details for that one were as follows:\n\n- Encoder backbone: resnest200e (timm)\n- Optimizer: Adam\n- Scheduler: cosine annealing\n- Input size: 320 x 320\n- Batch size: 32\n- LR: 15e-5\n- Attention dim: 512\n- Embedding dim: 512\n- Decoder dim: 2048\n- mixed precision\n- # epochs: 20, each epoch contained all of the training data + additional randomly generated samples from `extra_approved_InChIs.csv`. Augmentation: Coarse dropout, random rotations\n- 5 additional epochs on pseudo labelled test data with a lower LR (5e-5)\n\nThe choice of backbone was the biggest differentiator here. I experimented with many (efficientnets, efficientnet v2, different sizes of resnext, resnest, also different attention/embedding/decoder dims). Resnest200 took a long time to train, but the additional capacity meant it could achieve higher accuracy in the end.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1336385,
      "postDate": "2021-06-04T20:56:30.663Z",
      "content": "<p><a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> You should be incredibly proud of your performance in this competition.  I (like others) was keeping my eye on this competition rooting for you to get solo gold, it's a bummer that it didn't happen this time. I'm confident that with your skills you will get it soon enough!</p>",
      "rawMarkdown": "@fergusoci You should be incredibly proud of your performance in this competition.  I (like others) was keeping my eye on this competition rooting for you to get solo gold, it's a bummer that it didn't happen this time. I'm confident that with your skills you will get it soon enough!",
      "votes": 1,
      "replies": [
        {
          "id": 1336400,
          "postDate": "2021-06-04T21:32:06.060Z",
          "content": "<p>Cheers, Rob</p>",
          "rawMarkdown": "Cheers, Rob",
          "votes": 2
        }
      ]
    },
    {
      "id": 1358556,
      "postDate": "2021-06-20T15:11:18.873Z",
      "content": "<p>I'm very impressed with your amazing solo result, especially given that each epoch took 24 hrs. Kaggle can be a lot about luck, with luck favoring better data scientists, but also those with fast hardware for more rolls of the dice. So… you must be an amazing data scientist!</p>",
      "rawMarkdown": "I'm very impressed with your amazing solo result, especially given that each epoch took 24 hrs. Kaggle can be a lot about luck, with luck favoring better data scientists, but also those with fast hardware for more rolls of the dice. So... you must be an amazing data scientist!",
      "votes": 2
    },
    {
      "id": 1340791,
      "postDate": "2021-06-08T09:09:14.690Z",
      "content": "<p>hopefully, you will get what you deserved in near future.</p>",
      "rawMarkdown": "hopefully, you will get what you deserved in near future.",
      "votes": 2
    },
    {
      "id": 1336502,
      "postDate": "2021-06-05T01:53:11.107Z",
      "content": "<p>Thx for sharing.  Also thanks  <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> here for his sharing, though I had no hardware available to reproduce. </p>\n<p>I tried a modified version of VIT, It ran timeout on Kaggle, giving up  However learned the methodology.</p>",
      "rawMarkdown": "Thx for sharing.  Also thanks  @hengck23 here for his sharing, though I had no hardware available to reproduce. \n\nI tried a modified version of VIT, It ran timeout on Kaggle, giving up  However learned the methodology.",
      "votes": 2
    },
    {
      "id": 1336439,
      "postDate": "2021-06-04T22:48:07.227Z",
      "content": "<p>Probably one of the most difficult competitions to achieve a solo gold in, and you got so close! You should be proud of your efforts - and will surely achieve that solo gold in the near future. </p>",
      "rawMarkdown": "Probably one of the most difficult competitions to achieve a solo gold in, and you got so close! You should be proud of your efforts - and will surely achieve that solo gold in the near future. ",
      "votes": 2
    },
    {
      "id": 1336402,
      "postDate": "2021-06-04T21:32:35.027Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> , I participated in the competition early but gave up after three submissions since training one model was taking about 10 days and I didnt have enough computational resources in hand.<br>\nBut after that , I spectated the leaderboard and I was mainly watching you go up and hoping that you win it.<br>\nYour final placement is an inspiration for everyone ( What you have achieved is so hard, It was like hunger games ).<br>\nCongratulations, You really stole the spotlight in this competition .</p>",
      "rawMarkdown": "Hello @fergusoci , I participated in the competition early but gave up after three submissions since training one model was taking about 10 days and I didnt have enough computational resources in hand.\nBut after that , I spectated the leaderboard and I was mainly watching you go up and hoping that you win it.\nYour final placement is an inspiration for everyone ( What you have achieved is so hard, It was like hunger games ).\nCongratulations, You really stole the spotlight in this competition .\n",
      "votes": 2
    },
    {
      "id": 1335875,
      "postDate": "2021-06-04T13:43:15.833Z",
      "content": "<p>It was a tough competition, other top teams didn't achieve the score you did (with less than 24h per epoch). Congrats!</p>",
      "rawMarkdown": "It was a tough competition, other top teams didn't achieve the score you did (with less than 24h per epoch). Congrats!",
      "votes": 2
    },
    {
      "id": 1335833,
      "postDate": "2021-06-04T13:12:52.163Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a>  and thanks a lot for sharing you strategy, great job </p>",
      "rawMarkdown": "Congrats @fergusoci  and thanks a lot for sharing you strategy, great job ",
      "votes": 2
    },
    {
      "id": 1335649,
      "postDate": "2021-06-04T10:45:34.613Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> you fought really hard , to finish 12th place solo in such a tough competition is an amazing feat.. <br>\nThanks for the writeup . Can you tell us the training strategy for your best ViT model (0.67 Lb is really great)</p>",
      "rawMarkdown": "Congrats @fergusoci you fought really hard , to finish 12th place solo in such a tough competition is an amazing feat.. \nThanks for the writeup . Can you tell us the training strategy for your best ViT model (0.67 Lb is really great)",
      "votes": 2,
      "replies": [
        {
          "id": 1335733,
          "postDate": "2021-06-04T12:03:54.760Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a>. Details for ViT training as follows:</p>\n<p><strong>Model details:</strong></p>\n<p>encoder backbone = timm <code>vit_deit_base_patch16_384</code> pretrained<br>\ntransformer decoder = <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> TransformerDecode class. 8 attention heads; 6 layers; vocab size 300<br>\ninput size = (384, 384)<br>\noptimizer = Adam<br>\nlr = 15e-5<br>\nbatch size = 24<br>\nlr scheduler = cosine annealing</p>\n<p><strong>Training:</strong></p>\n<p>Mixed precision for 21 epochs (originally attempting 40 epochs). <br>\nAt this point I ran into NaN losses.<br>\nSwitched to full precision for 10 epochs, lowered batch size to 18.<br>\nCalculated pseudo labels for the test set based on ensemble of predictions from this + other models available at this point.<br>\nTrained for a further 5 epochs on the test set with a low lr.</p>\n<p>Things got interesting at this point. I still had time left, but not enough to train other models from scratch. The best option was to continue to train the ViT model. I decided to continue with the non-pseudo labelled version (the one that had been trained for 21 + 10 epochs). It went as follows:</p>\n<p>Train for 2 epochs. Power outage for a whole day. Attempt to train for another 8 epochs. Machine crash after epoch 2 (I assume load-related… At this point the whole thing looked like it was about to lift off). Train for another 3 epochs before time ran out.</p>",
          "rawMarkdown": "Thanks @tanulsingh077. Details for ViT training as follows:\n\n**Model details:**\n\nencoder backbone = timm `vit_deit_base_patch16_384` pretrained\ntransformer decoder = @hengck23 TransformerDecode class. 8 attention heads; 6 layers; vocab size 300\ninput size = (384, 384)\noptimizer = Adam\nlr = 15e-5\nbatch size = 24\nlr scheduler = cosine annealing\n\n**Training:**\n\nMixed precision for 21 epochs (originally attempting 40 epochs). \nAt this point I ran into NaN losses.\nSwitched to full precision for 10 epochs, lowered batch size to 18.\nCalculated pseudo labels for the test set based on ensemble of predictions from this + other models available at this point.\nTrained for a further 5 epochs on the test set with a low lr.\n\nThings got interesting at this point. I still had time left, but not enough to train other models from scratch. The best option was to continue to train the ViT model. I decided to continue with the non-pseudo labelled version (the one that had been trained for 21 + 10 epochs). It went as follows:\n\nTrain for 2 epochs. Power outage for a whole day. Attempt to train for another 8 epochs. Machine crash after epoch 2 (I assume load-related... At this point the whole thing looked like it was about to lift off). Train for another 3 epochs before time ran out.",
          "votes": 5
        },
        {
          "id": 1335766,
          "postDate": "2021-06-04T12:26:41.150Z",
          "content": "<p>Great Job , We also tried with Pseudo labels and increased Img Size , but were not able to squeeze out much improvement  , perhaps we should not have given up on our model and let it train 😉</p>\n<p>Its good to know that we were on the right track</p>",
          "rawMarkdown": "Great Job , We also tried with Pseudo labels and increased Img Size , but were not able to squeeze out much improvement  , perhaps we should not have given up on our model and let it train 😉\n\nIts good to know that we were on the right track"
        }
      ]
    },
    {
      "id": 1335418,
      "postDate": "2021-06-04T07:47:50.227Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> . I had my fingers crossed for you to get a solo gold in this competitive challenge with a lot of large teams in the gold zone. Falling short by just one rank may be a bit disappointing, but I am sure you will get in another competition. </p>",
      "rawMarkdown": "Congratulations @fergusoci . I had my fingers crossed for you to get a solo gold in this competitive challenge with a lot of large teams in the gold zone. Falling short by just one rank may be a bit disappointing, but I am sure you will get in another competition. ",
      "votes": 2,
      "replies": [
        {
          "id": 1335946,
          "postDate": "2021-06-04T14:29:15.750Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    },
    {
      "id": 1344278,
      "postDate": "2021-06-10T19:08:44.997Z",
      "rawMarkdown": "",
      "votes": -3,
      "isDeleted": true,
      "replies": [
        {
          "id": 1344315,
          "postDate": "2021-06-10T19:54:21.487Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> </p>",
          "rawMarkdown": "Thanks @morizin ",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1335928,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2021-06-04T14:16:54.170000",
      "content": "<p>I am sure you are disappointed Ciara but you can be proud of an amazing performance in a very competitive competition as a solo competitor fighting against talent and hardware! Great job!</p>",
      "votes": 11,
      "replies": [
        {
          "id": 1335944,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2021-06-04T14:28:34.873000",
          "content": "<p>Cheers. I'm going to go sleep for a week. Might have the capacity for a brighter outlook on it after that!</p>",
          "votes": 12,
          "replies": []
        }
      ]
    },
    {
      "id": 1335663,
      "author_name": "Murat Cihan Sorkun",
      "author_url": "",
      "post_date": "2021-06-04T11:00:09.060000",
      "content": "<p>Amazing work! Not the gold medal but you gained the respect of me and many kagglers by having a such good solo score. Also, all other solo competitors below 1.00 did a great job. Congratulations all!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1335404,
      "author_name": "Jonathan Besomi",
      "author_url": "",
      "post_date": "2021-06-04T07:38:14.183000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> for your outstanding result, especially considering the high competition in this challenge (17 GMs in the top 11 teams). I was rooting for you throughout the full competition! I'm sure you will attain the solo-gold medal very soon!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1336319,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-06-04T19:30:25.930000",
      "content": "<blockquote>\n  <p>Uuuuff. Kaggle is tough! The last two months have felt like a really hard slog. I was trying to go for a solo gold but fell just short </p>\n</blockquote>\n<p>I've been there exactly one year ago in the tweet sentiment competition.  One rank short of gold felt like the worst possible outcome at the time…</p>\n<p>Reality is that even if not gold, this is a huge achievement.  As you said Kaggle competitions are tough, especially for solo competitors.  You can be proud of your result.</p>\n<p>I hope we will see you enter new competitions soon.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1336338,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2021-06-04T19:44:09.107000",
          "content": "<p>Well, fair play for continuing to make solo attempts after that. I'm not sure whether I have the capacity for another one! Life-ex-kaggle circumstances makes it pretty difficult…</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1336937,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-05T09:52:39.750000",
          "content": "<p>I'm sorry… ! :(</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1335732,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2021-06-04T12:03:50.877000",
      "content": "<p>Congrats for your strong finish. Single model with 0.63 is amazing!<br>\nAnd don't give up all hope, there is still a little chance, as you know🤗</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1335769,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2021-06-04T12:29:42.400000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a>. Congrats to you guys, great result. Fingers crossed we get a fair result, whatever that is.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1343650,
      "author_name": "Anton Semenistyy",
      "author_url": "",
      "post_date": "2021-06-10T10:49:48.803000",
      "content": "<p><a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> , congrats on such an outstanding result. If it’s possible, then could you please tell more detailed information about lstm model with 1.11 LB? How long did you train it and what parameters did you use?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1343753,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2021-06-10T12:12:50.393000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/antonymipt\" target=\"_blank\">@antonymipt</a>. Details for that one were as follows:</p>\n<ul>\n<li>Encoder backbone: resnest200e (timm)</li>\n<li>Optimizer: Adam</li>\n<li>Scheduler: cosine annealing</li>\n<li>Input size: 320 x 320</li>\n<li>Batch size: 32</li>\n<li>LR: 15e-5</li>\n<li>Attention dim: 512</li>\n<li>Embedding dim: 512</li>\n<li>Decoder dim: 2048</li>\n<li>mixed precision</li>\n<li># epochs: 20, each epoch contained all of the training data + additional randomly generated samples from <code>extra_approved_InChIs.csv</code>. Augmentation: Coarse dropout, random rotations</li>\n<li>5 additional epochs on pseudo labelled test data with a lower LR (5e-5)</li>\n</ul>\n<p>The choice of backbone was the biggest differentiator here. I experimented with many (efficientnets, efficientnet v2, different sizes of resnext, resnest, also different attention/embedding/decoder dims). Resnest200 took a long time to train, but the additional capacity meant it could achieve higher accuracy in the end.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1336385,
      "author_name": "Rob Mulla",
      "author_url": "",
      "post_date": "2021-06-04T20:56:30.663000",
      "content": "<p><a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> You should be incredibly proud of your performance in this competition.  I (like others) was keeping my eye on this competition rooting for you to get solo gold, it's a bummer that it didn't happen this time. I'm confident that with your skills you will get it soon enough!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1336400,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2021-06-04T21:32:06.060000",
          "content": "<p>Cheers, Rob</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1358556,
      "author_name": "Alexander Soare",
      "author_url": "",
      "post_date": "2021-06-20T15:11:18.873000",
      "content": "<p>I'm very impressed with your amazing solo result, especially given that each epoch took 24 hrs. Kaggle can be a lot about luck, with luck favoring better data scientists, but also those with fast hardware for more rolls of the dice. So… you must be an amazing data scientist!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1340791,
      "author_name": "Sohail Ahmed",
      "author_url": "",
      "post_date": "2021-06-08T09:09:14.690000",
      "content": "<p>hopefully, you will get what you deserved in near future.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1336502,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2021-06-05T01:53:11.107000",
      "content": "<p>Thx for sharing.  Also thanks  <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> here for his sharing, though I had no hardware available to reproduce. </p>\n<p>I tried a modified version of VIT, It ran timeout on Kaggle, giving up  However learned the methodology.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1336439,
      "author_name": "Charles",
      "author_url": "",
      "post_date": "2021-06-04T22:48:07.227000",
      "content": "<p>Probably one of the most difficult competitions to achieve a solo gold in, and you got so close! You should be proud of your efforts - and will surely achieve that solo gold in the near future. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1336402,
      "author_name": "Hasan N",
      "author_url": "",
      "post_date": "2021-06-04T21:32:35.027000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> , I participated in the competition early but gave up after three submissions since training one model was taking about 10 days and I didnt have enough computational resources in hand.<br>\nBut after that , I spectated the leaderboard and I was mainly watching you go up and hoping that you win it.<br>\nYour final placement is an inspiration for everyone ( What you have achieved is so hard, It was like hunger games ).<br>\nCongratulations, You really stole the spotlight in this competition .</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1335875,
      "author_name": "Rafi Hai",
      "author_url": "",
      "post_date": "2021-06-04T13:43:15.833000",
      "content": "<p>It was a tough competition, other top teams didn't achieve the score you did (with less than 24h per epoch). Congrats!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1335833,
      "author_name": "Salim Khazem",
      "author_url": "",
      "post_date": "2021-06-04T13:12:52.163000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a>  and thanks a lot for sharing you strategy, great job </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1335649,
      "author_name": "Mr_KnowNothing",
      "author_url": "",
      "post_date": "2021-06-04T10:45:34.613000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> you fought really hard , to finish 12th place solo in such a tough competition is an amazing feat.. <br>\nThanks for the writeup . Can you tell us the training strategy for your best ViT model (0.67 Lb is really great)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1335733,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2021-06-04T12:03:54.760000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a>. Details for ViT training as follows:</p>\n<p><strong>Model details:</strong></p>\n<p>encoder backbone = timm <code>vit_deit_base_patch16_384</code> pretrained<br>\ntransformer decoder = <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> TransformerDecode class. 8 attention heads; 6 layers; vocab size 300<br>\ninput size = (384, 384)<br>\noptimizer = Adam<br>\nlr = 15e-5<br>\nbatch size = 24<br>\nlr scheduler = cosine annealing</p>\n<p><strong>Training:</strong></p>\n<p>Mixed precision for 21 epochs (originally attempting 40 epochs). <br>\nAt this point I ran into NaN losses.<br>\nSwitched to full precision for 10 epochs, lowered batch size to 18.<br>\nCalculated pseudo labels for the test set based on ensemble of predictions from this + other models available at this point.<br>\nTrained for a further 5 epochs on the test set with a low lr.</p>\n<p>Things got interesting at this point. I still had time left, but not enough to train other models from scratch. The best option was to continue to train the ViT model. I decided to continue with the non-pseudo labelled version (the one that had been trained for 21 + 10 epochs). It went as follows:</p>\n<p>Train for 2 epochs. Power outage for a whole day. Attempt to train for another 8 epochs. Machine crash after epoch 2 (I assume load-related… At this point the whole thing looked like it was about to lift off). Train for another 3 epochs before time ran out.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1335766,
          "author_name": "Mr_KnowNothing",
          "author_url": "",
          "post_date": "2021-06-04T12:26:41.150000",
          "content": "<p>Great Job , We also tried with Pseudo labels and increased Img Size , but were not able to squeeze out much improvement  , perhaps we should not have given up on our model and let it train 😉</p>\n<p>Its good to know that we were on the right track</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1335418,
      "author_name": "Pascal Pfeiffer",
      "author_url": "",
      "post_date": "2021-06-04T07:47:50.227000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> . I had my fingers crossed for you to get a solo gold in this competitive challenge with a lot of large teams in the gold zone. Falling short by just one rank may be a bit disappointing, but I am sure you will get in another competition. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1335946,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2021-06-04T14:29:15.750000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1344278,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-06-10T19:08:44.997000",
      "content": "",
      "votes": -3,
      "replies": [
        {
          "id": 1344315,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2021-06-10T19:54:21.487000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1335307": "Uuuuff. Kaggle is tough! The last two months have felt like a really hard slog. I was trying to go for a solo gold but fell just short - congratulations to all the winners, and everyone in the medal zone, it was a pleasure to compete with you!\n\nSome recognition before I go into solution details: @hengck23 - delighted you made it to the gold zone. My strongest models were based on the solutions which you so generously opensourced early on. I learnt a huge amount from team mates in the last competition that I entered (@christofhenkel, @ilu000, @philippsinger, @robikscube) - I don't think I would have done as well had it not been for that. Cheers.\n\n**Solution strategy:**\n\nGiven the size of both the dataset and the highest performing models, it became clear early on that I would need to make a strategy call: train multiple smaller models and ensemble them, or train a single large model and try to push the limit of what that could do. \n\nI felt that this competition was similar to [Lyft](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles) in terms of data size. My main learning from that competition was that patience can be a virtue: there is value in simply training a model for far longer than you think you should. Improvements are slow, but they are steady. \n\nTo this end, I focussed on a single VIT transformer (based on @hengck23, 384 x 384 inputs). My best single checkpoint ultimately made it to 0.67 LB. Each epoch took c. 24 hours (!!!). Ensembled they reached 0.63.\n\nWhen experimenting early on I trained a LSTM+attention model based on @yasufuminakama 's code with a resnest200 backbone. This reached 1.11 LB so I included it in my final ensemble for a small amount of diversification. \n\n**Data:**\n\nBoth models were fit using all of the data provided, plus an additional 10% samples per epoch of randomly generated images from `extra_approved_InChIs.csv`. I based the image generation on @stainsby 's [notebook](https://www.kaggle.com/stainsby/improved-synthetic-data-for-bms-competition-v3). Augmentation comprised random rotation (plus coarse dropout in the case of the LSTM model). The final ensemble set also included models that had been fit on pseudo-labelled test data.\n\n**Ensembling strategy:**\n\nThroughout the competition I had been using @nofreewill's InChI normalisation as the basis for ensembling: each prediction was categorized as 1) normalized, where normalized result = prediction; 2) normalized, with modifications; 3) normalization fails. \n\nFor each image in the test set I isolated the predictions with the lowest normalization category. Within this group I used simple voting, with ties decided by model preference (i.e. model were ordered highest preference to lowest preference based on their LB scores).\n\nI ensembled 7 checkpoints for the VIT model + 2 for the Resnest200/LSTM model. Each sample was computed twice per checkpoint: once in its original form, once flipped L->R.\n\nCongrats, all.",
    "1335928": "I am sure you are disappointed Ciara but you can be proud of an amazing performance in a very competitive competition as a solo competitor fighting against talent and hardware! Great job!",
    "1335663": "Amazing work! Not the gold medal but you gained the respect of me and many kagglers by having a such good solo score. Also, all other solo competitors below 1.00 did a great job. Congratulations all!",
    "1335404": "Congratulations @fergusoci for your outstanding result, especially considering the high competition in this challenge (17 GMs in the top 11 teams). I was rooting for you throughout the full competition! I'm sure you will attain the solo-gold medal very soon!",
    "1336319": "> Uuuuff. Kaggle is tough! The last two months have felt like a really hard slog. I was trying to go for a solo gold but fell just short \n\nI've been there exactly one year ago in the tweet sentiment competition.  One rank short of gold felt like the worst possible outcome at the time...\n\nReality is that even if not gold, this is a huge achievement.  As you said Kaggle competitions are tough, especially for solo competitors.  You can be proud of your result.\n\nI hope we will see you enter new competitions soon.",
    "1335732": "Congrats for your strong finish. Single model with 0.63 is amazing!\nAnd don't give up all hope, there is still a little chance, as you know🤗",
    "1343650": "@fergusoci , congrats on such an outstanding result. If it’s possible, then could you please tell more detailed information about lstm model with 1.11 LB? How long did you train it and what parameters did you use?",
    "1336385": "@fergusoci You should be incredibly proud of your performance in this competition.  I (like others) was keeping my eye on this competition rooting for you to get solo gold, it's a bummer that it didn't happen this time. I'm confident that with your skills you will get it soon enough!",
    "1358556": "I'm very impressed with your amazing solo result, especially given that each epoch took 24 hrs. Kaggle can be a lot about luck, with luck favoring better data scientists, but also those with fast hardware for more rolls of the dice. So... you must be an amazing data scientist!",
    "1340791": "hopefully, you will get what you deserved in near future.",
    "1336502": "Thx for sharing.  Also thanks  @hengck23 here for his sharing, though I had no hardware available to reproduce. \n\nI tried a modified version of VIT, It ran timeout on Kaggle, giving up  However learned the methodology.",
    "1336439": "Probably one of the most difficult competitions to achieve a solo gold in, and you got so close! You should be proud of your efforts - and will surely achieve that solo gold in the near future. ",
    "1336402": "Hello @fergusoci , I participated in the competition early but gave up after three submissions since training one model was taking about 10 days and I didnt have enough computational resources in hand.\nBut after that , I spectated the leaderboard and I was mainly watching you go up and hoping that you win it.\nYour final placement is an inspiration for everyone ( What you have achieved is so hard, It was like hunger games ).\nCongratulations, You really stole the spotlight in this competition .\n",
    "1335875": "It was a tough competition, other top teams didn't achieve the score you did (with less than 24h per epoch). Congrats!",
    "1335833": "Congrats @fergusoci  and thanks a lot for sharing you strategy, great job ",
    "1335649": "Congrats @fergusoci you fought really hard , to finish 12th place solo in such a tough competition is an amazing feat.. \nThanks for the writeup . Can you tell us the training strategy for your best ViT model (0.67 Lb is really great)",
    "1335418": "Congratulations @fergusoci . I had my fingers crossed for you to get a solo gold in this competitive challenge with a lot of large teams in the gold zone. Falling short by just one rank may be a bit disappointing, but I am sure you will get in another competition. ",
    "1344278": ""
  }
}