{
  "id": 230887,
  "title": "we are watching a race of gpu power",
  "url": "/competitions/bms-molecular-translation/discussion/230887",
  "author_name": "hengck23",
  "post_date": "2021-04-06T02:45:01.546000",
  "votes": 46,
  "comment_count": 35,
  "views": 0,
  "content": "<p>the loss dynamics of LSTM (and other seq model) is like this</p>\n<ul>\n<li>loss decrease for the head part of the seq first in early training</li>\n<li>short sequences get corrected first</li>\n<li>as the training goes, sometimes you see the loss is stagnant, But actually, the tail part of the long sequences are being corrected. Long sequence are rare, that is why the improvement of loss are minor</li>\n<li>but as long as you keep your GPU running, lb score will increase</li>\n</ul>\n<p>eventually, everyone will reach  lb score of about 1 to 2 …if they have enough GPU resource and their seq model has \"enough capacity\"</p>\n<p>every one is racing now. final top lb score could be in the range of  0.5 to 0.8</p>",
  "messages": [
    {
      "id": 1264252,
      "postDate": "2021-04-06T02:45:01.547Z",
      "content": "<p>the loss dynamics of LSTM (and other seq model) is like this</p>\n<ul>\n<li>loss decrease for the head part of the seq first in early training</li>\n<li>short sequences get corrected first</li>\n<li>as the training goes, sometimes you see the loss is stagnant, But actually, the tail part of the long sequences are being corrected. Long sequence are rare, that is why the improvement of loss are minor</li>\n<li>but as long as you keep your GPU running, lb score will increase</li>\n</ul>\n<p>eventually, everyone will reach  lb score of about 1 to 2 …if they have enough GPU resource and their seq model has \"enough capacity\"</p>\n<p>every one is racing now. final top lb score could be in the range of  0.5 to 0.8</p>",
      "rawMarkdown": "the loss dynamics of LSTM (and other seq model) is like this\n- loss decrease for the head part of the seq first in early training\n- short sequences get corrected first\n- as the training goes, sometimes you see the loss is stagnant, But actually, the tail part of the long sequences are being corrected. Long sequence are rare, that is why the improvement of loss are minor\n- but as long as you keep your GPU running, lb score will increase\n\neventually, everyone will reach  lb score of about 1 to 2 ...if they have enough GPU resource and their seq model has \"enough capacity\"\n\nevery one is racing now. final top lb score could be in the range of  0.5 to 0.8",
      "votes": 45
    },
    {
      "id": 1265554,
      "postDate": "2021-04-07T02:36:31.630Z",
      "content": "<p>Just manually assigned structures for 200 random images from training set, using computational chemist's intuition. 7% are wrong (7/100 and 7/100)</p>\n<ul>\n<li>10 are due to <em>ghost</em> terminal methyl group (e.g., <code>4a4f6c36a791</code>)</li>\n<li>1 is due to fuzzy stereo depiction (e.g., <code>4080f41a9577</code>)</li>\n<li>2 are due to missing bonds in the complex multi-ring structure containing crossing bonds (e.g., <code>94a92ba4a8ae</code>)</li>\n<li>1 is due to extremely fuzzy fluorine depiction (e.g., <code>e159f8ad60e9</code>)</li>\n</ul>\n<p>Let's say 1 false prediction increases edit distance by 30 on average . According to this estimation, the winning models with LD &lt; 2 are probably overfitting to THE particular molecular sketcher &amp; noise generator, from a pure scientific perspective.  </p>\n<p><strong>We are watching a race of overfitting</strong></p>",
      "rawMarkdown": "Just manually assigned structures for 200 random images from training set, using computational chemist's intuition. 7% are wrong (7/100 and 7/100)\n\n* 10 are due to *ghost* terminal methyl group (e.g., `4a4f6c36a791`)\n* 1 is due to fuzzy stereo depiction (e.g., `4080f41a9577`)\n* 2 are due to missing bonds in the complex multi-ring structure containing crossing bonds (e.g., `94a92ba4a8ae`)\n* 1 is due to extremely fuzzy fluorine depiction (e.g., `e159f8ad60e9`)\n\nLet's say 1 false prediction increases edit distance by 30 on average . According to this estimation, the winning models with LD < 2 are probably overfitting to THE particular molecular sketcher & noise generator, from a pure scientific perspective.  \n\n**We are watching a race of overfitting**",
      "votes": 18,
      "replies": [
        {
          "id": 1271090,
          "postDate": "2021-04-12T09:45:07.413Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1276721,
          "postDate": "2021-04-17T22:06:02.283Z",
          "content": "<p>This is valuable insight. Hopefully the competition sponsors are cognisant of this. </p>",
          "rawMarkdown": "This is valuable insight. Hopefully the competition sponsors are cognisant of this. "
        }
      ]
    },
    {
      "id": 1264398,
      "postDate": "2021-04-06T06:04:23.313Z",
      "content": "<p>Kaggle should provide some GCP quota for this competition. <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> </p>",
      "rawMarkdown": "Kaggle should provide some GCP quota for this competition. @addisonhoward ",
      "votes": 14,
      "replies": [
        {
          "id": 1271793,
          "postDate": "2021-04-12T23:13:05.250Z",
          "content": "<p>Yeah there must be some.</p>",
          "rawMarkdown": "Yeah there must be some."
        }
      ]
    },
    {
      "id": 1276723,
      "postDate": "2021-04-17T22:08:37.163Z",
      "content": "<p>I am just glad that the weather here is getting colder as autumn sets in - so I am appreciative of the extra warmth in my house. =)</p>",
      "rawMarkdown": "I am just glad that the weather here is getting colder as autumn sets in - so I am appreciative of the extra warmth in my house. =)",
      "votes": 8
    },
    {
      "id": 1268679,
      "postDate": "2021-04-09T16:33:43.563Z",
      "content": "<p>Isn't it true for all image competitions? What is different from other image competitions here?</p>",
      "rawMarkdown": "Isn't it true for all image competitions? What is different from other image competitions here?",
      "votes": 5,
      "replies": [
        {
          "id": 1269167,
          "postDate": "2021-04-10T08:34:40.290Z",
          "content": "<p>This competition seems way more resource hungry than other competitions. The data is huge and training longer seems to help quite a bit from what I can gather from the discussions. There are many cv competitions where you dont need much HW to score high, recent examples: VinBigData, Rainforest, NFL</p>",
          "rawMarkdown": "This competition seems way more resource hungry than other competitions. The data is huge and training longer seems to help quite a bit from what I can gather from the discussions. There are many cv competitions where you dont need much HW to score high, recent examples: VinBigData, Rainforest, NFL",
          "votes": 7
        },
        {
          "id": 1269177,
          "postDate": "2021-04-10T09:00:26.870Z",
          "content": "<p>There is a long tail problem here. you really need the number of iterations!</p>",
          "rawMarkdown": "There is a long tail problem here. you really need the number of iterations!",
          "votes": 7
        },
        {
          "id": 1269189,
          "postDate": "2021-04-10T09:18:58.327Z",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Sure, TGS, Bengali were small data too.  But there are also large data CV competitions, this one is no the first one.</p>\n<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> I believe you, my point is that it is not the first compute intensive comp.</p>",
          "rawMarkdown": "@philippsinger Sure, TGS, Bengali were small data too.  But there are also large data CV competitions, this one is no the first one.\n\n@hengck23 I believe you, my point is that it is not the first compute intensive comp.",
          "votes": 2
        },
        {
          "id": 1269514,
          "postDate": "2021-04-10T15:49:59.660Z",
          "content": "<p>Right, there are other compute intensive ones.</p>",
          "rawMarkdown": "Right, there are other compute intensive ones.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1264763,
      "postDate": "2021-04-06T12:14:04.003Z",
      "content": "<p>This seems to be a tendency, it can be confirmed by the lack of \"end-to-end\" notebooks. Many competitions lack of notebooks with a training part, because if they train with Kaggle resources they can't achive a remarkable score. </p>",
      "rawMarkdown": "This seems to be a tendency, it can be confirmed by the lack of \"end-to-end\" notebooks. Many competitions lack of notebooks with a training part, because if they train with Kaggle resources they can't achive a remarkable score. ",
      "votes": 5
    },
    {
      "id": 1268773,
      "postDate": "2021-04-09T18:26:34.763Z",
      "content": "<p>We don't have access to 1024 A100 at NVIDIA, contrarily to what some team name could suggest…</p>\n<p>We don't even have access to A100 at this point, except for very limited test.  We work with V100 GPU, which is already quite good.</p>\n<p>Just to straighten expectations.</p>",
      "rawMarkdown": "We don't have access to 1024 A100 at NVIDIA, contrarily to what some team name could suggest...\n\nWe don't even have access to A100 at this point, except for very limited test.  We work with V100 GPU, which is already quite good.\n\nJust to straighten expectations.",
      "votes": 3
    },
    {
      "id": 1279462,
      "postDate": "2021-04-20T23:45:27.177Z",
      "content": "<p>I will likely switch to TF and use the full power of TPU + TFRecords. </p>\n<p>I love Pytorch but training is really slow on single V100 - 16Gb.  My current val score 2.20 (not submitted yet) took <br>\n more than 4 days of training. Let alone inference which is even slower (I don't use fairseq, but plain Pytorch transformer.)   It doesn't let me room to experiment quickly new ideas. </p>\n<p>I converted all my transformer code to run on Torch/XLA but we still have big i/o bottleneck.   </p>",
      "rawMarkdown": "I will likely switch to TF and use the full power of TPU + TFRecords. \n\nI love Pytorch but training is really slow on single V100 - 16Gb.  My current val score 2.20 (not submitted yet) took \n more than 4 days of training. Let alone inference which is even slower (I don't use fairseq, but plain Pytorch transformer.)   It doesn't let me room to experiment quickly new ideas. \n\nI converted all my transformer code to run on Torch/XLA but we still have big i/o bottleneck.   ",
      "votes": 1,
      "replies": [
        {
          "id": 1279477,
          "postDate": "2021-04-21T00:30:15.767Z",
          "content": "<p>you need yoshi</p>\n<p><img src=\"https://i.ibb.co/0m446GH/Selection-076.png\" alt=\"\"></p>",
          "rawMarkdown": "you need yoshi\n\n![](https://i.ibb.co/0m446GH/Selection-076.png)",
          "votes": 7
        }
      ]
    },
    {
      "id": 1271288,
      "postDate": "2021-04-12T13:00:45.650Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  , I observed some thing.<br>\nI trained a model on 1 epoch and the results were as follows.<br>\nNote : I am using same code as Used by Y Nakama and using some 3 lakhs images for validation .</p>\n<p>Train Loss : 0.53<br>\nVal Levhestein Distance : 32.X</p>\n<p>Later I trained for 1 more epoch based on above checkpoint , Results were as follows<br>\nTrain Loss = 0.29<br>\nVal Levhestein Distance = 95.X</p>\n<p>Is this normal has anyone observed such behaviour .</p>",
      "rawMarkdown": "@hengck23  , I observed some thing.\nI trained a model on 1 epoch and the results were as follows.\nNote : I am using same code as Used by Y Nakama and using some 3 lakhs images for validation .\n\nTrain Loss : 0.53\nVal Levhestein Distance : 32.X\n\nLater I trained for 1 more epoch based on above checkpoint , Results were as follows\nTrain Loss = 0.29\nVal Levhestein Distance = 95.X\n\nIs this normal has anyone observed such behaviour .",
      "votes": 1,
      "replies": [
        {
          "id": 1271416,
          "postDate": "2021-04-12T15:34:48.943Z",
          "content": "<p>I found that using the wrong mean &amp; std for normalization leads to unstable validation performance.<br>\nHere are the <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/232025\" target=\"_blank\">correct ones </a>for the data. Maybe that was your problem. But in this case it should also lead to different performances in the 1st epoch over several runs with the same parameters.   </p>\n<p>Other possible cause: Model is heavily overfitting to training set, but seems unlikely. Your parameter count would have to be really high.</p>",
          "rawMarkdown": "I found that using the wrong mean & std for normalization leads to unstable validation performance.\nHere are the [correct ones ](https://www.kaggle.com/c/bms-molecular-translation/discussion/232025)for the data. Maybe that was your problem. But in this case it should also lead to different performances in the 1st epoch over several runs with the same parameters.   \n\nOther possible cause: Model is heavily overfitting to training set, but seems unlikely. Your parameter count would have to be really high."
        },
        {
          "id": 1279215,
          "postDate": "2021-04-20T17:51:00.703Z",
          "content": "<p><a href=\"https://www.kaggle.com/sayedathar11\" target=\"_blank\">@sayedathar11</a> can I ask?</p>\n<ul>\n<li>while you were doing inferencing of validation set . Were you fixing the orientation? Someone had such a problem in the lb score first epoch it was 25.xx then I don't remember much 96.xx or 46.xx don't know clearly</li>\n</ul>",
          "rawMarkdown": "@sayedathar11 can I ask?\n- while you were doing inferencing of validation set . Were you fixing the orientation? Someone had such a problem in the lb score first epoch it was 25.xx then I don't remember much 96.xx or 46.xx don't know clearly",
          "votes": 1
        },
        {
          "id": 1279426,
          "postDate": "2021-04-20T22:33:59.117Z",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> can you specify what do you meant by fixing orientation ? I didn't get you? Can you please elaborate ?<br>\nIf by orientation you meant augmentation then I want to add that I was adding train augmentation as follows.</p>\n<p>A.Compose([<br>\n            A.Resize(CFG.size, CFG.size),<br>\n            A.OneOf([A.HorizontalFlip( p = 0.5) , # Horizontal Flip is basically rotation by an Angle , Hence its okay to use it <br>\n            A.VerticalFlip(p=0.5)] , p = 0.3 ),<br>\n            A.RandomRotate90(p = 0.4),<br>\n            Normalize(<br>\n                mean=[0.485, 0.456, 0.406],<br>\n                std=[0.229, 0.224, 0.225],<br>\n            ),<br>\n            ToTensorV2(),<br>\n        ])</p>\n<p>And on validation set I was applying augmentation as follows .<br>\nCompose([<br>\n            A.Resize(CFG.size, CFG.size),<br>\n            A.Normalize(<br>\n                mean=[0.485, 0.456, 0.406],<br>\n                std=[0.229, 0.224, 0.225],<br>\n            ),<br>\n            ToTensorV2(),<br>\n        ])</p>\n<p>Is there something problematic here ?</p>",
          "rawMarkdown": "@morizin can you specify what do you meant by fixing orientation ? I didn't get you? Can you please elaborate ?\nIf by orientation you meant augmentation then I want to add that I was adding train augmentation as follows.\n\nA.Compose([\n            A.Resize(CFG.size, CFG.size),\n            A.OneOf([A.HorizontalFlip( p = 0.5) , # Horizontal Flip is basically rotation by an Angle , Hence its okay to use it \n            A.VerticalFlip(p=0.5)] , p = 0.3 ),\n            A.RandomRotate90(p = 0.4),\n            Normalize(\n                mean=[0.485, 0.456, 0.406],\n                std=[0.229, 0.224, 0.225],\n            ),\n            ToTensorV2(),\n        ])\n\n\nAnd on validation set I was applying augmentation as follows .\nCompose([\n            A.Resize(CFG.size, CFG.size),\n            A.Normalize(\n                mean=[0.485, 0.456, 0.406],\n                std=[0.229, 0.224, 0.225],\n            ),\n            ToTensorV2(),\n        ])\n\nIs there something problematic here ?"
        },
        {
          "id": 1279439,
          "postDate": "2021-04-20T23:04:45.547Z",
          "content": "<p><code>Is there something problematic here ?</code><br>\nYes, you are using mean &amp; std from image net for normalization. As long as you don't use a model pretrained on image net your normalization is useless.<br>\nThat is most likely the cause for your instability, as already mentioned by me above.</p>",
          "rawMarkdown": "`Is there something problematic here ?`\nYes, you are using mean & std from image net for normalization. As long as you don't use a model pretrained on image net your normalization is useless.\nThat is most likely the cause for your instability, as already mentioned by me above.",
          "votes": -1
        },
        {
          "id": 1279445,
          "postDate": "2021-04-20T23:11:08.607Z",
          "content": "<p>I am ensuring that my model uses pretrained weights of imagenet and also others are doing same I guess there is most probably some other error can you please let <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> answer this .</p>",
          "rawMarkdown": "I am ensuring that my model uses pretrained weights of imagenet and also others are doing same I guess there is most probably some other error can you please let @morizin answer this .",
          "votes": -1
        },
        {
          "id": 1279449,
          "postDate": "2021-04-20T23:18:17.900Z",
          "content": "<p>Try the correct mean &amp; std anyway to check, because your normalized images should have mean of 0 and std 1 anyway. I argue it is better to use mean &amp; std of actual data if they are to far of from the original training data. <br>\nAnd <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> can still answer, no problem.<br>\nI mean it is your problem, I just want to help. And I had the <strong>exact</strong> same problems and was using the same normalization as you.</p>",
          "rawMarkdown": "Try the correct mean & std anyway to check, because your normalized images should have mean of 0 and std 1 anyway. I argue it is better to use mean & std of actual data if they are to far of from the original training data. \nAnd @morizin can still answer, no problem.\nI mean it is your problem, I just want to help. And I had the **exact** same problems and was using the same normalization as you.",
          "votes": -1
        },
        {
          "id": 1281217,
          "postDate": "2021-04-22T17:57:26.753Z",
          "content": "<p><a href=\"https://www.kaggle.com/sayedathar11\" target=\"_blank\">@sayedathar11</a>  sorry for late reply<br>\nactually the some images of test is rotated which our training model finds weird<br>\ni am not sure that your problem is same as this. but some experienced this weird behavior in LB but not in CV.<br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/224257#1237479\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/224257#1237479</a><br>\nI just started to explore the competition didnt try much till now.</p>\n<p><code>Is this normal has anyone observed such behaviour .</code><br>\ncan you check what happens in epoch 3 please. try only rotate as augmentation please for now</p>\n<p>because i think horizontal and vertical flip in train aug causes some problem</p>\n<p><img src=\"https://th.bing.com/th/id/R85095686b39fa2fe6648eaf0c70e3999?rik=H%2feB0Rv55fOCrA&amp;riu=http%3a%2f%2fwww.online-image-editor.com%2fhelp%2fimages%2frotate_flip_02.png\" alt=\"\"></p>",
          "rawMarkdown": "@sayedathar11  sorry for late reply\nactually the some images of test is rotated which our training model finds weird\ni am not sure that your problem is same as this. but some experienced this weird behavior in LB but not in CV.\nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/224257#1237479\nI just started to explore the competition didnt try much till now.\n\n`Is this normal has anyone observed such behaviour .`\ncan you check what happens in epoch 3 please. try only rotate as augmentation please for now\n\nbecause i think horizontal and vertical flip in train aug causes some problem\n\n![](https://th.bing.com/th/id/R85095686b39fa2fe6648eaf0c70e3999?rik=H%2feB0Rv55fOCrA&riu=http%3a%2f%2fwww.online-image-editor.com%2fhelp%2fimages%2frotate_flip_02.png)",
          "votes": 1
        },
        {
          "id": 1282024,
          "postDate": "2021-04-23T14:15:55.810Z",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> I guess the problem was arising due to using some augmentations which causes instability in training ,I have removed those augmentations and Trying Efficientnet B1 as encoder , still will inform if I see any changes in 3rd epoch.</p>",
          "rawMarkdown": "@morizin I guess the problem was arising due to using some augmentations which causes instability in training ,I have removed those augmentations and Trying Efficientnet B1 as encoder , still will inform if I see any changes in 3rd epoch."
        }
      ]
    },
    {
      "id": 1268334,
      "postDate": "2021-04-09T10:07:35.787Z",
      "content": "<p>i find an interesting results:</p>\n<p><img src=\"https://i.ibb.co/wJ9n3LV/Selection-054.png\" alt=\"\"><br>\n<a href=\"https://arxiv.org/pdf/2103.17239.pdf\" target=\"_blank\">https://arxiv.org/pdf/2103.17239.pdf</a></p>\n<p>data-mining : let GPU run and mine information from data</p>",
      "rawMarkdown": "i find an interesting results:\n\n![](https://i.ibb.co/wJ9n3LV/Selection-054.png)\nhttps://arxiv.org/pdf/2103.17239.pdf\n\ndata-mining : let GPU run and mine information from data",
      "votes": 1,
      "replies": [
        {
          "id": 1268376,
          "postDate": "2021-04-09T10:59:48.540Z",
          "content": "<p><img src=\"https://i.ibb.co/3NKgCHc/Selection-058.png\" alt=\"\"></p>\n<p>how to improve your image transformer (depth and distillation)</p>",
          "rawMarkdown": "![](https://i.ibb.co/3NKgCHc/Selection-058.png)\n\nhow to improve your image transformer (depth and distillation)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1281126,
      "postDate": "2021-04-22T16:30:55.207Z",
      "content": "<p>Just wanted to comment and post a photo showing results from my model on a holdout validation dataset. The model is an EfficientNetB7--&gt;Attn/LSTM run for around 6 epochs. Scores a public LB of 5.77. Note that although I say \"INCHI LENGTH\" below, I am referring to the number of tokens that make up the ground truth InChI string.</p>\n<p><img src=\"https://i.ibb.co/H7Ngq5J/Screen-Shot-2021-04-22-at-12-23-29-PM.png\" alt=\"validation\"></p>\n<p>I've noticed that if I train for more epochs the LSD continues to get better for the longer sequences. I don't have a photo showing this but as you pointed out above… this is obviously what's supposed to happen.</p>\n<p>I did wonder how much of an LSD drop we should expect for longer sequences when compared to shorter sequences (assuming an infinite amount of training)? Essentially, should I assume that this model is capable of scoring ~1.5 if trained for a LOoooong time? Or will there be some remaining negative effects as longer sequences may just be more difficult to predict?</p>\n<p>PS:<br>\nAs a random aside, my model currently only predicts ~1.5% valid inchi strings when submitting.</p>",
      "rawMarkdown": "Just wanted to comment and post a photo showing results from my model on a holdout validation dataset. The model is an EfficientNetB7-->Attn/LSTM run for around 6 epochs. Scores a public LB of 5.77. Note that although I say \"INCHI LENGTH\" below, I am referring to the number of tokens that make up the ground truth InChI string.\n\n![validation](https://i.ibb.co/H7Ngq5J/Screen-Shot-2021-04-22-at-12-23-29-PM.png)\n\nI've noticed that if I train for more epochs the LSD continues to get better for the longer sequences. I don't have a photo showing this but as you pointed out above... this is obviously what's supposed to happen.\n\nI did wonder how much of an LSD drop we should expect for longer sequences when compared to shorter sequences (assuming an infinite amount of training)? Essentially, should I assume that this model is capable of scoring ~1.5 if trained for a LOoooong time? Or will there be some remaining negative effects as longer sequences may just be more difficult to predict?\n\n\nPS:\nAs a random aside, my model currently only predicts ~1.5% valid inchi strings when submitting.\n",
      "votes": 2,
      "replies": [
        {
          "id": 1281318,
          "postDate": "2021-04-22T19:30:27.727Z",
          "content": "<p>I think LD~3 is totally achievable. Not sure about 1.5. But thanks for sharing valid InChI percent. </p>\n<p>Here is my baseline model on holdout set (N = 303k): <br>\n100% valid InChI. 78% perfect LD = 0 predictions. Mean LD = 7.2</p>\n<table>\n<thead>\n<tr>\n<th>mean LD</th>\n<th>perfect %</th>\n<th>N</th>\n<th>InChI length</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>6.333333</td>\n<td>67%</td>\n<td>3</td>\n<td>20</td>\n</tr>\n<tr>\n<td>3.860465</td>\n<td>78%</td>\n<td>1075</td>\n<td>40</td>\n</tr>\n<tr>\n<td>2.574591</td>\n<td>87%</td>\n<td>17710</td>\n<td>60</td>\n</tr>\n<tr>\n<td>2.416558</td>\n<td>90%</td>\n<td>59586</td>\n<td>80</td>\n</tr>\n<tr>\n<td>3.367635</td>\n<td>87%</td>\n<td>71533</td>\n<td>100</td>\n</tr>\n<tr>\n<td>5.563011</td>\n<td>80%</td>\n<td>58537</td>\n<td>120</td>\n</tr>\n<tr>\n<td>8.590690</td>\n<td>73%</td>\n<td>47304</td>\n<td>140</td>\n</tr>\n<tr>\n<td>13.629566</td>\n<td>62%</td>\n<td>26145</td>\n<td>160</td>\n</tr>\n<tr>\n<td>20.810648</td>\n<td>51%</td>\n<td>12585</td>\n<td>180</td>\n</tr>\n<tr>\n<td>30.792253</td>\n<td>39%</td>\n<td>5757</td>\n<td>200</td>\n</tr>\n<tr>\n<td>43.468519</td>\n<td>29%</td>\n<td>2160</td>\n<td>220</td>\n</tr>\n<tr>\n<td>65.289773</td>\n<td>19%</td>\n<td>704</td>\n<td>240</td>\n</tr>\n<tr>\n<td>106.171206</td>\n<td>9%</td>\n<td>257</td>\n<td>260</td>\n</tr>\n<tr>\n<td>149.389706</td>\n<td>4%</td>\n<td>136</td>\n<td>280</td>\n</tr>\n<tr>\n<td>182.104839</td>\n<td>0%</td>\n<td>124</td>\n<td>300</td>\n</tr>\n<tr>\n<td>194.864198</td>\n<td>0%</td>\n<td>81</td>\n<td>320</td>\n</tr>\n<tr>\n<td>207.985294</td>\n<td>1.5%</td>\n<td>68</td>\n<td>340</td>\n</tr>\n<tr>\n<td>234.826087</td>\n<td>0%</td>\n<td>23</td>\n<td>360</td>\n</tr>\n<tr>\n<td>247.500000</td>\n<td>0%</td>\n<td>6</td>\n<td>380</td>\n</tr>\n</tbody>\n</table>",
          "rawMarkdown": "I think LD~3 is totally achievable. Not sure about 1.5. But thanks for sharing valid InChI percent. \n\nHere is my baseline model on holdout set (N = 303k): \n100% valid InChI. 78% perfect LD = 0 predictions. Mean LD = 7.2\n\n|   mean LD  |  perfect %|   N |  InChI length |\n| --- | --- | --- | --- |\n|6.333333  |  67% | 3 | 20|\n|3.860465  |78% |1075 | 40|\n|2.574591| 87% |17710|  60|\n|2.416558 |90%| 59586 | 80|\n|3.367635| 87%| 71533| 100|\n|5.563011 |80%|58537| 120|\n|8.590690 |73%|47304| 140|\n|13.629566 |62%|26145| 160|\n|20.810648| 51%|12585| 180|\n|30.792253  |39%|5757| 200|\n|43.468519| 29%| 2160 |220|\n|65.289773   |19%|704| 240|\n|106.171206  | 9%|257 |260|\n|149.389706  | 4%|136| 280|\n|182.104839   |0%|124| 300|\n|194.864198  |  0%|81 |320|\n|207.985294   | 1.5%|68 |340|\n|234.826087  |  0%|23 |360|\n|247.500000    | 0%|6 |380|",
          "votes": 2
        }
      ]
    },
    {
      "id": 1274357,
      "postDate": "2021-04-15T08:19:26.297Z",
      "content": "<p>Is it worth it to keep training or use the GPU power to mine crypto then ? 😄</p>",
      "rawMarkdown": "Is it worth it to keep training or use the GPU power to mine crypto then ? 😄",
      "votes": 2
    },
    {
      "id": 1298495,
      "postDate": "2021-05-08T23:44:27.817Z",
      "content": "<p>Many thanks for this super valuable insight.</p>",
      "rawMarkdown": "Many thanks for this super valuable insight."
    },
    {
      "id": 1269444,
      "postDate": "2021-04-10T14:38:05.810Z",
      "content": "<p>Is there a way to sample more Long sequences to make training more efficient?</p>",
      "rawMarkdown": "Is there a way to sample more Long sequences to make training more efficient?",
      "replies": [
        {
          "id": 1270013,
          "postDate": "2021-04-11T07:06:34.627Z",
          "content": "<p>\"Is there a way to sample more Long sequences to make training more efficient?\"</p>\n<p>we should find out the reason why long sequence get worst results first.<br>\ne.g. because the are bigger in image and resizing them to a fixed size causes loss of information?<br>\ne.g. we need to model larger  range dependency because the atoms are far apart?<br>\ne.g. is there a correlation of sequence length and image size?</p>\n<p>when the reason become known, then we are decide how to sample, etc</p>",
          "rawMarkdown": "\"Is there a way to sample more Long sequences to make training more efficient?\"\n\nwe should find out the reason why long sequence get worst results first.\ne.g. because the are bigger in image and resizing them to a fixed size causes loss of information?\ne.g. we need to model larger  range dependency because the atoms are far apart?\ne.g. is there a correlation of sequence length and image size?\n\nwhen the reason become known, then we are decide how to sample, etc",
          "votes": 5
        }
      ]
    },
    {
      "id": 1265427,
      "postDate": "2021-04-06T21:44:38.773Z",
      "content": "<p>Thanks for the observation. This suggests that we should use whatever GPU resources we have as efficiently as we can.  If I understand correctly, it would seem that changing the mix of the training data in later epochs, to include more longer sequences (either by sampling or augmentation) would help to speed up the process of correcting the long sequences.</p>",
      "rawMarkdown": "Thanks for the observation. This suggests that we should use whatever GPU resources we have as efficiently as we can.  If I understand correctly, it would seem that changing the mix of the training data in later epochs, to include more longer sequences (either by sampling or augmentation) would help to speed up the process of correcting the long sequences.",
      "replies": [
        {
          "id": 1266750,
          "postDate": "2021-04-08T03:52:08.410Z",
          "content": "<p>note that there is smart batching in training transformer:</p>\n<p><a href=\"https://towardsdatascience.com/divide-hugging-face-transformers-training-time-by-2-or-more-21bf7129db9q-21bf7129db9e\" target=\"_blank\">https://towardsdatascience.com/divide-hugging-face-transformers-training-time-by-2-or-more-21bf7129db9q-21bf7129db9e</a></p>\n<p><a href=\"https://wandb.ai/pommedeterresautee/speed_training/reports/Train-HuggingFace-Models-Twice-As-Fast--VmlldzoxMDgzOTI\" target=\"_blank\">https://wandb.ai/pommedeterresautee/speed_training/reports/Train-HuggingFace-Models-Twice-As-Fast--VmlldzoxMDgzOTI</a></p>\n<p>\"smart batching (named uniform length batching in the article, because experiments have been logged with this stupid name, I will keep it for this report).\"</p>\n<hr>\n<p>you can check the following papers, etc:</p>\n<p>Training Tips for the Transformer Model<br>\n<a href=\"https://ufal.mff.cuni.cz/pbml/110/art-popel-bojar.pdf\" target=\"_blank\">https://ufal.mff.cuni.cz/pbml/110/art-popel-bojar.pdf</a></p>\n<p>Tips on max_length<br>\n\"So in our case, if the batch size is high enough, the max_length has almost no<br>\neffect on BLEU, but this should be checked for each new dataset.\"</p>\n<hr>\n<p><a href=\"https://www.aclweb.org/anthology/W17-5708.pdf\" target=\"_blank\">https://www.aclweb.org/anthology/W17-5708.pdf</a><br>\nA Bag of Useful Tricks for Practical Neural Machine Translation:<br>\nEmbedding Layer Initialization and Large Batch Size</p>\n<p><a href=\"https://github.com/neubig/nmt-tips\" target=\"_blank\">https://github.com/neubig/nmt-tips</a></p>",
          "rawMarkdown": "note that there is smart batching in training transformer:\n\nhttps://towardsdatascience.com/divide-hugging-face-transformers-training-time-by-2-or-more-21bf7129db9q-21bf7129db9e\n\nhttps://wandb.ai/pommedeterresautee/speed_training/reports/Train-HuggingFace-Models-Twice-As-Fast--VmlldzoxMDgzOTI\n\n\"smart batching (named uniform length batching in the article, because experiments have been logged with this stupid name, I will keep it for this report).\"\n\n----\n\nyou can check the following papers, etc:\n\nTraining Tips for the Transformer Model\nhttps://ufal.mff.cuni.cz/pbml/110/art-popel-bojar.pdf\n\nTips on max_length\n\"So in our case, if the batch size is high enough, the max_length has almost no\neffect on BLEU, but this should be checked for each new dataset.\"\n\n---\n\nhttps://www.aclweb.org/anthology/W17-5708.pdf\nA Bag of Useful Tricks for Practical Neural Machine Translation:\nEmbedding Layer Initialization and Large Batch Size\n\nhttps://github.com/neubig/nmt-tips",
          "votes": 6
        },
        {
          "id": 1271708,
          "postDate": "2021-04-12T19:59:42.010Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>,  thanks for your response and super helpful links… yes I can see a path to faster results!</p>",
          "rawMarkdown": "@hengck23,  thanks for your response and super helpful links... yes I can see a path to faster results!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1265554,
      "author_name": "human intelligence",
      "author_url": "",
      "post_date": "2021-04-07T02:36:31.630000",
      "content": "<p>Just manually assigned structures for 200 random images from training set, using computational chemist's intuition. 7% are wrong (7/100 and 7/100)</p>\n<ul>\n<li>10 are due to <em>ghost</em> terminal methyl group (e.g., <code>4a4f6c36a791</code>)</li>\n<li>1 is due to fuzzy stereo depiction (e.g., <code>4080f41a9577</code>)</li>\n<li>2 are due to missing bonds in the complex multi-ring structure containing crossing bonds (e.g., <code>94a92ba4a8ae</code>)</li>\n<li>1 is due to extremely fuzzy fluorine depiction (e.g., <code>e159f8ad60e9</code>)</li>\n</ul>\n<p>Let's say 1 false prediction increases edit distance by 30 on average . According to this estimation, the winning models with LD &lt; 2 are probably overfitting to THE particular molecular sketcher &amp; noise generator, from a pure scientific perspective.  </p>\n<p><strong>We are watching a race of overfitting</strong></p>",
      "votes": 18,
      "replies": [
        {
          "id": 1271090,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-12T09:45:07.413000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1276721,
          "author_name": "Charles",
          "author_url": "",
          "post_date": "2021-04-17T22:06:02.283000",
          "content": "<p>This is valuable insight. Hopefully the competition sponsors are cognisant of this. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1264398,
      "author_name": "Johnny Lee",
      "author_url": "",
      "post_date": "2021-04-06T06:04:23.313000",
      "content": "<p>Kaggle should provide some GCP quota for this competition. <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> </p>",
      "votes": 14,
      "replies": [
        {
          "id": 1271793,
          "author_name": "Santosh kumar",
          "author_url": "",
          "post_date": "2021-04-12T23:13:05.250000",
          "content": "<p>Yeah there must be some.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1276723,
      "author_name": "Charles",
      "author_url": "",
      "post_date": "2021-04-17T22:08:37.163000",
      "content": "<p>I am just glad that the weather here is getting colder as autumn sets in - so I am appreciative of the extra warmth in my house. =)</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 1268679,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-04-09T16:33:43.563000",
      "content": "<p>Isn't it true for all image competitions? What is different from other image competitions here?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1269167,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-04-10T08:34:40.290000",
          "content": "<p>This competition seems way more resource hungry than other competitions. The data is huge and training longer seems to help quite a bit from what I can gather from the discussions. There are many cv competitions where you dont need much HW to score high, recent examples: VinBigData, Rainforest, NFL</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1269177,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-10T09:00:26.870000",
          "content": "<p>There is a long tail problem here. you really need the number of iterations!</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1269189,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-04-10T09:18:58.327000",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Sure, TGS, Bengali were small data too.  But there are also large data CV competitions, this one is no the first one.</p>\n<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> I believe you, my point is that it is not the first compute intensive comp.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1269514,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-04-10T15:49:59.660000",
          "content": "<p>Right, there are other compute intensive ones.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1264763,
      "author_name": "Hiram Coria 🧬",
      "author_url": "",
      "post_date": "2021-04-06T12:14:04.003000",
      "content": "<p>This seems to be a tendency, it can be confirmed by the lack of \"end-to-end\" notebooks. Many competitions lack of notebooks with a training part, because if they train with Kaggle resources they can't achive a remarkable score. </p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1268773,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-04-09T18:26:34.763000",
      "content": "<p>We don't have access to 1024 A100 at NVIDIA, contrarily to what some team name could suggest…</p>\n<p>We don't even have access to A100 at this point, except for very limited test.  We work with V100 GPU, which is already quite good.</p>\n<p>Just to straighten expectations.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1279462,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2021-04-20T23:45:27.177000",
      "content": "<p>I will likely switch to TF and use the full power of TPU + TFRecords. </p>\n<p>I love Pytorch but training is really slow on single V100 - 16Gb.  My current val score 2.20 (not submitted yet) took <br>\n more than 4 days of training. Let alone inference which is even slower (I don't use fairseq, but plain Pytorch transformer.)   It doesn't let me room to experiment quickly new ideas. </p>\n<p>I converted all my transformer code to run on Torch/XLA but we still have big i/o bottleneck.   </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1279477,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-21T00:30:15.767000",
          "content": "<p>you need yoshi</p>\n<p><img src=\"https://i.ibb.co/0m446GH/Selection-076.png\" alt=\"\"></p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 1271288,
      "author_name": "Athar Sayed",
      "author_url": "",
      "post_date": "2021-04-12T13:00:45.650000",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  , I observed some thing.<br>\nI trained a model on 1 epoch and the results were as follows.<br>\nNote : I am using same code as Used by Y Nakama and using some 3 lakhs images for validation .</p>\n<p>Train Loss : 0.53<br>\nVal Levhestein Distance : 32.X</p>\n<p>Later I trained for 1 more epoch based on above checkpoint , Results were as follows<br>\nTrain Loss = 0.29<br>\nVal Levhestein Distance = 95.X</p>\n<p>Is this normal has anyone observed such behaviour .</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1271416,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-04-12T15:34:48.943000",
          "content": "<p>I found that using the wrong mean &amp; std for normalization leads to unstable validation performance.<br>\nHere are the <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/232025\" target=\"_blank\">correct ones </a>for the data. Maybe that was your problem. But in this case it should also lead to different performances in the 1st epoch over several runs with the same parameters.   </p>\n<p>Other possible cause: Model is heavily overfitting to training set, but seems unlikely. Your parameter count would have to be really high.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1279215,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-04-20T17:51:00.703000",
          "content": "<p><a href=\"https://www.kaggle.com/sayedathar11\" target=\"_blank\">@sayedathar11</a> can I ask?</p>\n<ul>\n<li>while you were doing inferencing of validation set . Were you fixing the orientation? Someone had such a problem in the lb score first epoch it was 25.xx then I don't remember much 96.xx or 46.xx don't know clearly</li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1279426,
          "author_name": "Athar Sayed",
          "author_url": "",
          "post_date": "2021-04-20T22:33:59.117000",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> can you specify what do you meant by fixing orientation ? I didn't get you? Can you please elaborate ?<br>\nIf by orientation you meant augmentation then I want to add that I was adding train augmentation as follows.</p>\n<p>A.Compose([<br>\n            A.Resize(CFG.size, CFG.size),<br>\n            A.OneOf([A.HorizontalFlip( p = 0.5) , # Horizontal Flip is basically rotation by an Angle , Hence its okay to use it <br>\n            A.VerticalFlip(p=0.5)] , p = 0.3 ),<br>\n            A.RandomRotate90(p = 0.4),<br>\n            Normalize(<br>\n                mean=[0.485, 0.456, 0.406],<br>\n                std=[0.229, 0.224, 0.225],<br>\n            ),<br>\n            ToTensorV2(),<br>\n        ])</p>\n<p>And on validation set I was applying augmentation as follows .<br>\nCompose([<br>\n            A.Resize(CFG.size, CFG.size),<br>\n            A.Normalize(<br>\n                mean=[0.485, 0.456, 0.406],<br>\n                std=[0.229, 0.224, 0.225],<br>\n            ),<br>\n            ToTensorV2(),<br>\n        ])</p>\n<p>Is there something problematic here ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1279439,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-04-20T23:04:45.547000",
          "content": "<p><code>Is there something problematic here ?</code><br>\nYes, you are using mean &amp; std from image net for normalization. As long as you don't use a model pretrained on image net your normalization is useless.<br>\nThat is most likely the cause for your instability, as already mentioned by me above.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1279445,
          "author_name": "Athar Sayed",
          "author_url": "",
          "post_date": "2021-04-20T23:11:08.607000",
          "content": "<p>I am ensuring that my model uses pretrained weights of imagenet and also others are doing same I guess there is most probably some other error can you please let <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> answer this .</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1279449,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-04-20T23:18:17.900000",
          "content": "<p>Try the correct mean &amp; std anyway to check, because your normalized images should have mean of 0 and std 1 anyway. I argue it is better to use mean &amp; std of actual data if they are to far of from the original training data. <br>\nAnd <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> can still answer, no problem.<br>\nI mean it is your problem, I just want to help. And I had the <strong>exact</strong> same problems and was using the same normalization as you.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1281217,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-04-22T17:57:26.753000",
          "content": "<p><a href=\"https://www.kaggle.com/sayedathar11\" target=\"_blank\">@sayedathar11</a>  sorry for late reply<br>\nactually the some images of test is rotated which our training model finds weird<br>\ni am not sure that your problem is same as this. but some experienced this weird behavior in LB but not in CV.<br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/224257#1237479\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/224257#1237479</a><br>\nI just started to explore the competition didnt try much till now.</p>\n<p><code>Is this normal has anyone observed such behaviour .</code><br>\ncan you check what happens in epoch 3 please. try only rotate as augmentation please for now</p>\n<p>because i think horizontal and vertical flip in train aug causes some problem</p>\n<p><img src=\"https://th.bing.com/th/id/R85095686b39fa2fe6648eaf0c70e3999?rik=H%2feB0Rv55fOCrA&amp;riu=http%3a%2f%2fwww.online-image-editor.com%2fhelp%2fimages%2frotate_flip_02.png\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1282024,
          "author_name": "Athar Sayed",
          "author_url": "",
          "post_date": "2021-04-23T14:15:55.810000",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> I guess the problem was arising due to using some augmentations which causes instability in training ,I have removed those augmentations and Trying Efficientnet B1 as encoder , still will inform if I see any changes in 3rd epoch.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1268334,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-09T10:07:35.787000",
      "content": "<p>i find an interesting results:</p>\n<p><img src=\"https://i.ibb.co/wJ9n3LV/Selection-054.png\" alt=\"\"><br>\n<a href=\"https://arxiv.org/pdf/2103.17239.pdf\" target=\"_blank\">https://arxiv.org/pdf/2103.17239.pdf</a></p>\n<p>data-mining : let GPU run and mine information from data</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1268376,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-09T10:59:48.540000",
          "content": "<p><img src=\"https://i.ibb.co/3NKgCHc/Selection-058.png\" alt=\"\"></p>\n<p>how to improve your image transformer (depth and distillation)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1281126,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-04-22T16:30:55.207000",
      "content": "<p>Just wanted to comment and post a photo showing results from my model on a holdout validation dataset. The model is an EfficientNetB7--&gt;Attn/LSTM run for around 6 epochs. Scores a public LB of 5.77. Note that although I say \"INCHI LENGTH\" below, I am referring to the number of tokens that make up the ground truth InChI string.</p>\n<p><img src=\"https://i.ibb.co/H7Ngq5J/Screen-Shot-2021-04-22-at-12-23-29-PM.png\" alt=\"validation\"></p>\n<p>I've noticed that if I train for more epochs the LSD continues to get better for the longer sequences. I don't have a photo showing this but as you pointed out above… this is obviously what's supposed to happen.</p>\n<p>I did wonder how much of an LSD drop we should expect for longer sequences when compared to shorter sequences (assuming an infinite amount of training)? Essentially, should I assume that this model is capable of scoring ~1.5 if trained for a LOoooong time? Or will there be some remaining negative effects as longer sequences may just be more difficult to predict?</p>\n<p>PS:<br>\nAs a random aside, my model currently only predicts ~1.5% valid inchi strings when submitting.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1281318,
          "author_name": "human intelligence",
          "author_url": "",
          "post_date": "2021-04-22T19:30:27.727000",
          "content": "<p>I think LD~3 is totally achievable. Not sure about 1.5. But thanks for sharing valid InChI percent. </p>\n<p>Here is my baseline model on holdout set (N = 303k): <br>\n100% valid InChI. 78% perfect LD = 0 predictions. Mean LD = 7.2</p>\n<table>\n<thead>\n<tr>\n<th>mean LD</th>\n<th>perfect %</th>\n<th>N</th>\n<th>InChI length</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>6.333333</td>\n<td>67%</td>\n<td>3</td>\n<td>20</td>\n</tr>\n<tr>\n<td>3.860465</td>\n<td>78%</td>\n<td>1075</td>\n<td>40</td>\n</tr>\n<tr>\n<td>2.574591</td>\n<td>87%</td>\n<td>17710</td>\n<td>60</td>\n</tr>\n<tr>\n<td>2.416558</td>\n<td>90%</td>\n<td>59586</td>\n<td>80</td>\n</tr>\n<tr>\n<td>3.367635</td>\n<td>87%</td>\n<td>71533</td>\n<td>100</td>\n</tr>\n<tr>\n<td>5.563011</td>\n<td>80%</td>\n<td>58537</td>\n<td>120</td>\n</tr>\n<tr>\n<td>8.590690</td>\n<td>73%</td>\n<td>47304</td>\n<td>140</td>\n</tr>\n<tr>\n<td>13.629566</td>\n<td>62%</td>\n<td>26145</td>\n<td>160</td>\n</tr>\n<tr>\n<td>20.810648</td>\n<td>51%</td>\n<td>12585</td>\n<td>180</td>\n</tr>\n<tr>\n<td>30.792253</td>\n<td>39%</td>\n<td>5757</td>\n<td>200</td>\n</tr>\n<tr>\n<td>43.468519</td>\n<td>29%</td>\n<td>2160</td>\n<td>220</td>\n</tr>\n<tr>\n<td>65.289773</td>\n<td>19%</td>\n<td>704</td>\n<td>240</td>\n</tr>\n<tr>\n<td>106.171206</td>\n<td>9%</td>\n<td>257</td>\n<td>260</td>\n</tr>\n<tr>\n<td>149.389706</td>\n<td>4%</td>\n<td>136</td>\n<td>280</td>\n</tr>\n<tr>\n<td>182.104839</td>\n<td>0%</td>\n<td>124</td>\n<td>300</td>\n</tr>\n<tr>\n<td>194.864198</td>\n<td>0%</td>\n<td>81</td>\n<td>320</td>\n</tr>\n<tr>\n<td>207.985294</td>\n<td>1.5%</td>\n<td>68</td>\n<td>340</td>\n</tr>\n<tr>\n<td>234.826087</td>\n<td>0%</td>\n<td>23</td>\n<td>360</td>\n</tr>\n<tr>\n<td>247.500000</td>\n<td>0%</td>\n<td>6</td>\n<td>380</td>\n</tr>\n</tbody>\n</table>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1274357,
      "author_name": "Yijie Xu",
      "author_url": "",
      "post_date": "2021-04-15T08:19:26.297000",
      "content": "<p>Is it worth it to keep training or use the GPU power to mine crypto then ? 😄</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1298495,
      "author_name": "HinePo",
      "author_url": "",
      "post_date": "2021-05-08T23:44:27.817000",
      "content": "<p>Many thanks for this super valuable insight.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1269444,
      "author_name": "ZavodRobotov",
      "author_url": "",
      "post_date": "2021-04-10T14:38:05.810000",
      "content": "<p>Is there a way to sample more Long sequences to make training more efficient?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1270013,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-11T07:06:34.627000",
          "content": "<p>\"Is there a way to sample more Long sequences to make training more efficient?\"</p>\n<p>we should find out the reason why long sequence get worst results first.<br>\ne.g. because the are bigger in image and resizing them to a fixed size causes loss of information?<br>\ne.g. we need to model larger  range dependency because the atoms are far apart?<br>\ne.g. is there a correlation of sequence length and image size?</p>\n<p>when the reason become known, then we are decide how to sample, etc</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1265427,
      "author_name": "Andy Penrose",
      "author_url": "",
      "post_date": "2021-04-06T21:44:38.773000",
      "content": "<p>Thanks for the observation. This suggests that we should use whatever GPU resources we have as efficiently as we can.  If I understand correctly, it would seem that changing the mix of the training data in later epochs, to include more longer sequences (either by sampling or augmentation) would help to speed up the process of correcting the long sequences.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1266750,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-08T03:52:08.410000",
          "content": "<p>note that there is smart batching in training transformer:</p>\n<p><a href=\"https://towardsdatascience.com/divide-hugging-face-transformers-training-time-by-2-or-more-21bf7129db9q-21bf7129db9e\" target=\"_blank\">https://towardsdatascience.com/divide-hugging-face-transformers-training-time-by-2-or-more-21bf7129db9q-21bf7129db9e</a></p>\n<p><a href=\"https://wandb.ai/pommedeterresautee/speed_training/reports/Train-HuggingFace-Models-Twice-As-Fast--VmlldzoxMDgzOTI\" target=\"_blank\">https://wandb.ai/pommedeterresautee/speed_training/reports/Train-HuggingFace-Models-Twice-As-Fast--VmlldzoxMDgzOTI</a></p>\n<p>\"smart batching (named uniform length batching in the article, because experiments have been logged with this stupid name, I will keep it for this report).\"</p>\n<hr>\n<p>you can check the following papers, etc:</p>\n<p>Training Tips for the Transformer Model<br>\n<a href=\"https://ufal.mff.cuni.cz/pbml/110/art-popel-bojar.pdf\" target=\"_blank\">https://ufal.mff.cuni.cz/pbml/110/art-popel-bojar.pdf</a></p>\n<p>Tips on max_length<br>\n\"So in our case, if the batch size is high enough, the max_length has almost no<br>\neffect on BLEU, but this should be checked for each new dataset.\"</p>\n<hr>\n<p><a href=\"https://www.aclweb.org/anthology/W17-5708.pdf\" target=\"_blank\">https://www.aclweb.org/anthology/W17-5708.pdf</a><br>\nA Bag of Useful Tricks for Practical Neural Machine Translation:<br>\nEmbedding Layer Initialization and Large Batch Size</p>\n<p><a href=\"https://github.com/neubig/nmt-tips\" target=\"_blank\">https://github.com/neubig/nmt-tips</a></p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1271708,
          "author_name": "Andy Penrose",
          "author_url": "",
          "post_date": "2021-04-12T19:59:42.010000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>,  thanks for your response and super helpful links… yes I can see a path to faster results!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1264252": "the loss dynamics of LSTM (and other seq model) is like this\n- loss decrease for the head part of the seq first in early training\n- short sequences get corrected first\n- as the training goes, sometimes you see the loss is stagnant, But actually, the tail part of the long sequences are being corrected. Long sequence are rare, that is why the improvement of loss are minor\n- but as long as you keep your GPU running, lb score will increase\n\neventually, everyone will reach  lb score of about 1 to 2 ...if they have enough GPU resource and their seq model has \"enough capacity\"\n\nevery one is racing now. final top lb score could be in the range of  0.5 to 0.8",
    "1265554": "Just manually assigned structures for 200 random images from training set, using computational chemist's intuition. 7% are wrong (7/100 and 7/100)\n\n* 10 are due to *ghost* terminal methyl group (e.g., `4a4f6c36a791`)\n* 1 is due to fuzzy stereo depiction (e.g., `4080f41a9577`)\n* 2 are due to missing bonds in the complex multi-ring structure containing crossing bonds (e.g., `94a92ba4a8ae`)\n* 1 is due to extremely fuzzy fluorine depiction (e.g., `e159f8ad60e9`)\n\nLet's say 1 false prediction increases edit distance by 30 on average . According to this estimation, the winning models with LD < 2 are probably overfitting to THE particular molecular sketcher & noise generator, from a pure scientific perspective.  \n\n**We are watching a race of overfitting**",
    "1264398": "Kaggle should provide some GCP quota for this competition. @addisonhoward ",
    "1276723": "I am just glad that the weather here is getting colder as autumn sets in - so I am appreciative of the extra warmth in my house. =)",
    "1268679": "Isn't it true for all image competitions? What is different from other image competitions here?",
    "1264763": "This seems to be a tendency, it can be confirmed by the lack of \"end-to-end\" notebooks. Many competitions lack of notebooks with a training part, because if they train with Kaggle resources they can't achive a remarkable score. ",
    "1268773": "We don't have access to 1024 A100 at NVIDIA, contrarily to what some team name could suggest...\n\nWe don't even have access to A100 at this point, except for very limited test.  We work with V100 GPU, which is already quite good.\n\nJust to straighten expectations.",
    "1279462": "I will likely switch to TF and use the full power of TPU + TFRecords. \n\nI love Pytorch but training is really slow on single V100 - 16Gb.  My current val score 2.20 (not submitted yet) took \n more than 4 days of training. Let alone inference which is even slower (I don't use fairseq, but plain Pytorch transformer.)   It doesn't let me room to experiment quickly new ideas. \n\nI converted all my transformer code to run on Torch/XLA but we still have big i/o bottleneck.   ",
    "1271288": "@hengck23  , I observed some thing.\nI trained a model on 1 epoch and the results were as follows.\nNote : I am using same code as Used by Y Nakama and using some 3 lakhs images for validation .\n\nTrain Loss : 0.53\nVal Levhestein Distance : 32.X\n\nLater I trained for 1 more epoch based on above checkpoint , Results were as follows\nTrain Loss = 0.29\nVal Levhestein Distance = 95.X\n\nIs this normal has anyone observed such behaviour .",
    "1268334": "i find an interesting results:\n\n![](https://i.ibb.co/wJ9n3LV/Selection-054.png)\nhttps://arxiv.org/pdf/2103.17239.pdf\n\ndata-mining : let GPU run and mine information from data",
    "1281126": "Just wanted to comment and post a photo showing results from my model on a holdout validation dataset. The model is an EfficientNetB7-->Attn/LSTM run for around 6 epochs. Scores a public LB of 5.77. Note that although I say \"INCHI LENGTH\" below, I am referring to the number of tokens that make up the ground truth InChI string.\n\n![validation](https://i.ibb.co/H7Ngq5J/Screen-Shot-2021-04-22-at-12-23-29-PM.png)\n\nI've noticed that if I train for more epochs the LSD continues to get better for the longer sequences. I don't have a photo showing this but as you pointed out above... this is obviously what's supposed to happen.\n\nI did wonder how much of an LSD drop we should expect for longer sequences when compared to shorter sequences (assuming an infinite amount of training)? Essentially, should I assume that this model is capable of scoring ~1.5 if trained for a LOoooong time? Or will there be some remaining negative effects as longer sequences may just be more difficult to predict?\n\n\nPS:\nAs a random aside, my model currently only predicts ~1.5% valid inchi strings when submitting.\n",
    "1274357": "Is it worth it to keep training or use the GPU power to mine crypto then ? 😄",
    "1298495": "Many thanks for this super valuable insight.",
    "1269444": "Is there a way to sample more Long sequences to make training more efficient?",
    "1265427": "Thanks for the observation. This suggests that we should use whatever GPU resources we have as efficiently as we can.  If I understand correctly, it would seem that changing the mix of the training data in later epochs, to include more longer sequences (either by sampling or augmentation) would help to speed up the process of correcting the long sequences."
  }
}