{
  "id": 224257,
  "title": "Best Single Model",
  "url": "/competitions/bms-molecular-translation/discussion/224257",
  "author_name": "DrHB",
  "post_date": "2021-03-07T15:10:01.531000",
  "votes": 108,
  "comment_count": 235,
  "views": 0,
  "content": "<p>Just starting a common topic =)</p>\n<pre><code>model: resnet34 \ntrain_samples: 800k\nvalid_samples: 100k\nepoch: 5\nlr: 1e-3\nschedule: cosine\naugmentation: None\nsize: 128\nCV score: 15.7\nLB score: 26.7\n</code></pre>",
  "messages": [
    {
      "id": 1229719,
      "postDate": "2021-03-07T15:10:01.533Z",
      "content": "<p>Just starting a common topic =)</p>\n<pre><code>model: resnet34 \ntrain_samples: 800k\nvalid_samples: 100k\nepoch: 5\nlr: 1e-3\nschedule: cosine\naugmentation: None\nsize: 128\nCV score: 15.7\nLB score: 26.7\n</code></pre>",
      "rawMarkdown": "Just starting a common topic =)\n\n```\nmodel: resnet34 \ntrain_samples: 800k\nvalid_samples: 100k\nepoch: 5\nlr: 1e-3\nschedule: cosine\naugmentation: None\nsize: 128\nCV score: 15.7\nLB score: 26.7\n```",
      "votes": 108
    },
    {
      "id": 1260670,
      "postDate": "2021-04-02T10:26:26.977Z",
      "content": "<p>power of big encoder.</p>\n<p>use CNN-attention-LSTM (1 layer) model, single fold</p>\n<p>train: input 224x224, no augmentation (i.e. rotation is not used)<br>\ntest: use YNakama w&lt;h rotate trick, greedy (argmax) decoder<br>\ntokenizer: YNakama's Tokenizer</p>\n<p>training: train as long as my HW can support … basically the below results are for epoch&gt;10. results are not optimal, but i estimate within +/-1 optimality</p>\n<p>resnet34d: ~LB 9.5/CV 8.7<br>\nresnet26d: ~LB 8.5/CV 6.8<br>\nresnet101d:~LB 4.9/CV 4.0<br>\nresnet200d:~LB 3.8/CV 3.2</p>\n<hr>\n<p>below are the results for training in progress. no submission has been made</p>\n<ul>\n<li>load  pretrained encoder  into transformer and finetune end-to-end:<ul>\n<li>resnet26d : train/valid cross-entropy loss seems better than CNN-attention-LSTM-resnet101d and resnet200d</li>\n<li>hence, transformer is clearly the winner</li></ul></li>\n</ul>",
      "rawMarkdown": "power of big encoder.\n\nuse CNN-attention-LSTM (1 layer) model, single fold\n\ntrain: input 224x224, no augmentation (i.e. rotation is not used)\ntest: use YNakama w<h rotate trick, greedy (argmax) decoder\ntokenizer: YNakama's Tokenizer\n\ntraining: train as long as my HW can support ... basically the below results are for epoch>10. results are not optimal, but i estimate within +/-1 optimality\n\nresnet34d: ~LB 9.5/CV 8.7\nresnet26d: ~LB 8.5/CV 6.8\nresnet101d:~LB 4.9/CV 4.0\nresnet200d:~LB 3.8/CV 3.2\n\n---\nbelow are the results for training in progress. no submission has been made\n\n- load  pretrained encoder  into transformer and finetune end-to-end:\n   - resnet26d : train/valid cross-entropy loss seems better than CNN-attention-LSTM-resnet101d and resnet200d\n   - hence, transformer is clearly the winner\n ",
      "votes": 15,
      "replies": [
        {
          "id": 1260682,
          "postDate": "2021-04-02T10:41:26.237Z",
          "content": "<p>i wonder did anyone did experiment on size?<br>\ne.g. 192,224,256,288,320 ….</p>\n<p>in my early experiment for CNN encoder resnet34 with stride=32, 288 seems to perform worse than 224. but i cannot confirm since training takes long long time which i cannot afford</p>",
          "rawMarkdown": "i wonder did anyone did experiment on size?\ne.g. 192,224,256,288,320 ....\n\nin my early experiment for CNN encoder resnet34 with stride=32, 288 seems to perform worse than 224. but i cannot confirm since training takes long long time which i cannot afford",
          "votes": 1
        },
        {
          "id": 1260761,
          "postDate": "2021-04-02T11:54:50.103Z",
          "content": "<p>I stopped resnet101 after two epochs, because I got worse than resnet50.  May be I should continue for few more epochs (but it's really time consuming for my hardware )</p>\n<p>My current submit is resnet50 trained for 12 epochs<br>\nI use transformer as decoder and not LSTM. </p>",
          "rawMarkdown": "I stopped resnet101 after two epochs, because I got worse than resnet50.  May be I should continue for few more epochs (but it's really time consuming for my hardware )\n\nMy current submit is resnet50 trained for 12 epochs\nI use transformer as decoder and not LSTM. ",
          "votes": 1
        },
        {
          "id": 1260774,
          "postDate": "2021-04-02T12:00:23.117Z",
          "content": "<p>\"I use transformer as decoder and not LSTM.\"</p>\n<p>my previous experience is that transformer should perform better</p>\n<p>i plan to change the decoder later (I can freeze my encoder for faster training when changing the decoder)<br>\nbut the problem with the transformer is inference. i haven't found a way to make it fast. having to spend 10 to 20 hrs to make a submission is frightening.</p>",
          "rawMarkdown": "\"I use transformer as decoder and not LSTM.\"\n\nmy previous experience is that transformer should perform better\n\ni plan to change the decoder later (I can freeze my encoder for faster training when changing the decoder)\nbut the problem with the transformer is inference. i haven't found a way to make it fast. having to spend 10 to 20 hrs to make a submission is frightening.",
          "votes": 3
        },
        {
          "id": 1260792,
          "postDate": "2021-04-02T12:13:41.627Z",
          "content": "<p>You are  right.  That's why I use TPU (Torch/XLA) for inference and train on GPU. </p>\n<p>Anyway I think we can  experiment a lot before running inference, given Validation and LB  seem pretty much aligned. </p>",
          "rawMarkdown": "You are  right.  That's why I use TPU (Torch/XLA) for inference and train on GPU. \n\nAnyway I think we can  experiment a lot before running inference, given Validation and LB  seem pretty much aligned. ",
          "votes": 2
        },
        {
          "id": 1260810,
          "postDate": "2021-04-02T12:25:58.240Z",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> I experimented with transformers but I didn't get better performance (other than faster training) out of them. I'm thinking my positional embedding was not optimal but I tried both sin/cos and simple nn.Embedding; none worked out good for me. I'm really unfamiliar with Transformers btw.</p>",
          "rawMarkdown": "@serigne I experimented with transformers but I didn't get better performance (other than faster training) out of them. I'm thinking my positional embedding was not optimal but I tried both sin/cos and simple nn.Embedding; none worked out good for me. I'm really unfamiliar with Transformers btw."
        },
        {
          "id": 1260821,
          "postDate": "2021-04-02T12:35:55.460Z",
          "content": "<p>there is a trick for faster inference</p>\n<ol>\n<li><p>based on my experiment results about 50% of the prediction has 0  Levenshtein distance (i.e. perfect results)</p></li>\n<li><p>you can find a way to identify correct test prediction, either by:</p></li>\n</ol>\n<ul>\n<li>consistent results of different model and probing, etc</li>\n<li>use rdkit to generate image from predicted Inchi string and you also need to train a verifier to verify if kaggle image is same as rdkit image or not</li>\n</ul>\n<p>for new models, you only need to submit results for difficult test samples, while keeping the easy one fixed.</p>\n<p>some goes for training. some of the train samples are really easy. a good sampling method can reduce number of train samples</p>",
          "rawMarkdown": "there is a trick for faster inference\n1. based on my experiment results about 50% of the prediction has 0  Levenshtein distance (i.e. perfect results)\n\n2. you can find a way to identify correct test prediction, either by:\n- consistent results of different model and probing, etc\n- use rdkit to generate image from predicted Inchi string and you also need to train a verifier to verify if kaggle image is same as rdkit image or not\n\nfor new models, you only need to submit results for difficult test samples, while keeping the easy one fixed.\n\n\nsome goes for training. some of the train samples are really easy. a good sampling method can reduce number of train samples",
          "votes": 7
        },
        {
          "id": 1260823,
          "postDate": "2021-04-02T12:36:58.167Z",
          "content": "<p>\"I'm really unfamiliar with Transformers btw.\"</p>\n<p>the rule is if LSTM can do it, transformer will do better</p>",
          "rawMarkdown": "\"I'm really unfamiliar with Transformers btw.\"\n\nthe rule is if LSTM can do it, transformer will do better",
          "votes": 4
        },
        {
          "id": 1260825,
          "postDate": "2021-04-02T12:37:52.063Z",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Is TPU much faster than GPU on Pytorch? I've experiemented with TPU but got some issues and now I'm thinking to go deeper with this problem but does it worth it?</p>",
          "rawMarkdown": "@serigne Is TPU much faster than GPU on Pytorch? I've experiemented with TPU but got some issues and now I'm thinking to go deeper with this problem but does it worth it?",
          "votes": 1
        },
        {
          "id": 1260842,
          "postDate": "2021-04-02T12:55:20.093Z",
          "content": "<p>When the model gets better, the % of 0 distance prediction will go up too. I think the top team should have &gt; 85% 0 distance predictions now. </p>\n<p>A quick way to tell if the prediction is likely 0 distance or not is simply to use rdkit to test whether the prediction is a valid inchi string, which should give &gt; 90% precision. </p>\n<p>So, pseudo labels would probably play some role in this game now. And that thought alone makes me want to abandon this competition…</p>",
          "rawMarkdown": "When the model gets better, the % of 0 distance prediction will go up too. I think the top team should have > 85% 0 distance predictions now. \n\nA quick way to tell if the prediction is likely 0 distance or not is simply to use rdkit to test whether the prediction is a valid inchi string, which should give > 90% precision. \n\nSo, pseudo labels would probably play some role in this game now. And that thought alone makes me want to abandon this competition...",
          "votes": 4
        },
        {
          "id": 1260843,
          "postDate": "2021-04-02T12:55:55.270Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> I've tried sequence bucketing based on the length of the InChI string but it showed results worse than I had without this sampler. I had an idea that this kind of grouping could be helpful for updating the gradients but probably it didn't work out because of the poor predictive power of my decoder. </p>",
          "rawMarkdown": "@hengck23 I've tried sequence bucketing based on the length of the InChI string but it showed results worse than I had without this sampler. I had an idea that this kind of grouping could be helpful for updating the gradients but probably it didn't work out because of the poor predictive power of my decoder. "
        },
        {
          "id": 1260876,
          "postDate": "2021-04-02T13:36:17.640Z",
          "content": "<p><a href=\"https://www.kaggle.com/alexvishnevskiy\" target=\"_blank\">@alexvishnevskiy</a> <br>\nThere is still I/O and CPU bottlenecks , compared to TensorFlow TPU.<br>\nBut we can still run 8 inferences in parallel </p>",
          "rawMarkdown": "@alexvishnevskiy \nThere is still I/O and CPU bottlenecks , compared to TensorFlow TPU.\nBut we can still run 8 inferences in parallel ",
          "votes": 1
        },
        {
          "id": 1261038,
          "postDate": "2021-04-02T16:34:55.027Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> For resnet34, is it okay to share what kind of adjustments you made based on the public kernel? ‘cause if I use the default setting, it became overfit even at epoch 5, single fold. Thanks!</p>",
          "rawMarkdown": "@hengck23 For resnet34, is it okay to share what kind of adjustments you made based on the public kernel? ‘cause if I use the default setting, it became overfit even at epoch 5, single fold. Thanks!",
          "replies": [
            {
              "id": 1286849,
              "postDate": "2021-04-28T13:15:41.837Z",
              "content": "<p>how can it be?there are millions of samples with long labels.&gt; <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> For resnet34, is it okay to share what kind of adjustments you made based on the public kernel? ‘cause if I use the default setting, it became overfit even at epoch 5, single fold. Thanks!</p>",
              "rawMarkdown": "how can it be?there are millions of samples with long labels.> @hengck23 For resnet34, is it okay to share what kind of adjustments you made based on the public kernel? ‘cause if I use the default setting, it became overfit even at epoch 5, single fold. Thanks!\n\n",
              "votes": 1
            }
          ]
        },
        {
          "id": 1261381,
          "postDate": "2021-04-03T02:57:41.180Z",
          "content": "<p><a href=\"https://www.kaggle.com/xm2203\" target=\"_blank\">@xm2203</a> </p>\n<p>i rewrite the code from the public kernel. but there is not much difference</p>\n<p>i don't think it will overfit if you are using all the train samples (80/20 train/validation split)<br>\ntry to use learning rate 1e-3 and train end-to-end, then drop to 1e-4,1e-5 only when there is no more improvement</p>\n<p>i will publish my train log files and loss curve curves later after organization</p>",
          "rawMarkdown": "@xm2203 \n\ni rewrite the code from the public kernel. but there is not much difference\n\ni don't think it will overfit if you are using all the train samples (80/20 train/validation split)\ntry to use learning rate 1e-3 and train end-to-end, then drop to 1e-4,1e-5 only when there is no more improvement\n\ni will publish my train log files and loss curve curves later after organization",
          "votes": 2
        },
        {
          "id": 1261508,
          "postDate": "2021-04-03T06:04:30.357Z",
          "content": "<p><a href=\"https://www.kaggle.com/xm2203\" target=\"_blank\">@xm2203</a> <br>\nIf you are using the public kernel you are not overfitting because the LR scheduler used is Cosine annealing and it has T_max of 5 so the LR resets to higher value which makes your score higher. If you continue training you should see good improvements till epoch-10. </p>",
          "rawMarkdown": "@xm2203 \nIf you are using the public kernel you are not overfitting because the LR scheduler used is Cosine annealing and it has T_max of 5 so the LR resets to higher value which makes your score higher. If you continue training you should see good improvements till epoch-10. ",
          "votes": 2
        },
        {
          "id": 1285021,
          "postDate": "2021-04-26T13:59:11.133Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> \"the rule is if LSTM can do it, transformer will do better\"… \"if you have a LOT of data\" :)</p>",
          "rawMarkdown": "@hengck23 \"the rule is if LSTM can do it, transformer will do better\"... \"if you have a LOT of data\" :)"
        },
        {
          "id": 1286847,
          "postDate": "2021-04-28T13:14:02.827Z",
          "content": "<p>how can it be?there are millions of samples with long labels.</p>",
          "rawMarkdown": "how can it be?there are millions of samples with long labels."
        }
      ]
    },
    {
      "id": 1236521,
      "postDate": "2021-03-13T08:08:21.683Z",
      "content": "<pre><code>encoder: resnet34 \ndecoder: LSTM with attention\ntrain_samples: 80%\nvalid_samples: 20%\nepoch: 1\naugmentation: None\nsize: 224\nCV: 24.1\n</code></pre>\n<p>Training is done with kaggle notebooks, I will continue training and share notebooks as starter code.</p>",
      "rawMarkdown": "```\nencoder: resnet34 \ndecoder: LSTM with attention\ntrain_samples: 80%\nvalid_samples: 20%\nepoch: 1\naugmentation: None\nsize: 224\nCV: 24.1\n```\nTraining is done with kaggle notebooks, I will continue training and share notebooks as starter code.",
      "votes": 12,
      "replies": [
        {
          "id": 1236664,
          "postDate": "2021-03-13T10:58:30.830Z",
          "content": "<p>if you didn't use augmentation, the LB will be about 30.</p>",
          "rawMarkdown": "if you didn't use augmentation, the LB will be about 30.",
          "votes": 2
        },
        {
          "id": 1237433,
          "postDate": "2021-03-14T06:22:03.810Z",
          "content": "<p>Hmm… I trained 2 epochs and CV is 18.7, LB is 46.9, maybe something is wrong… I will try to fix it.</p>",
          "rawMarkdown": "Hmm... I trained 2 epochs and CV is 18.7, LB is 46.9, maybe something is wrong... I will try to fix it.",
          "votes": 2
        },
        {
          "id": 1237454,
          "postDate": "2021-03-14T06:46:44.630Z",
          "content": "<p>Is it possible that you did not take into account that the training set only has horizontal compounds, but the test has 90° rotated ones, too?</p>",
          "rawMarkdown": "Is it possible that you did not take into account that the training set only has horizontal compounds, but the test has 90° rotated ones, too?",
          "votes": 7
        },
        {
          "id": 1237479,
          "postDate": "2021-03-14T07:26:30.600Z",
          "content": "<p>Thanks for pointing out! It seems that's the reason of large gap between my CV and LB. </p>\n<p>===== Edit1 =====<br>\nApplied below code for inference and LB changed from 46.9 to 21.9</p>\n<pre><code>h, w, _ = image.shape\nif h &gt; w:\n    image = image.transpose(1, 0, 2)\n</code></pre>\n<p>===== Edit2 =====<br>\nThis is correct one. LB changed from 21.9 to 20.3</p>\n<pre><code>import albumentations as A\n\ntransform = A.Compose([A.Transpose(p=1), A.VerticalFlip(p=1)])\n\nh, w, _ = image.shape\nif h &gt; w:\n    image = transform(image=image)['image']\n</code></pre>",
          "rawMarkdown": "Thanks for pointing out! It seems that's the reason of large gap between my CV and LB. \n\n===== Edit1 =====\nApplied below code for inference and LB changed from 46.9 to 21.9\n```\nh, w, _ = image.shape\nif h > w:\n    image = image.transpose(1, 0, 2)\n```\n===== Edit2 =====\nThis is correct one. LB changed from 21.9 to 20.3\n```\nimport albumentations as A\n\ntransform = A.Compose([A.Transpose(p=1), A.VerticalFlip(p=1)])\n\nh, w, _ = image.shape\nif h > w:\n    image = transform(image=image)['image']\n```",
          "votes": 19
        },
        {
          "id": 1237692,
          "postDate": "2021-03-14T11:28:23.743Z",
          "content": "<p>I'm glad I could help! :)</p>",
          "rawMarkdown": "I'm glad I could help! :)",
          "votes": 2
        },
        {
          "id": 1237974,
          "postDate": "2021-03-14T14:44:19.980Z",
          "content": "<p>Isn't that flip via diagnoal…..?</p>",
          "rawMarkdown": "Isn't that flip via diagnoal.....?",
          "votes": 4
        },
        {
          "id": 1238045,
          "postDate": "2021-03-14T15:36:18.040Z",
          "content": "<p>Oh, thanks for pointing out! I added correct one.</p>",
          "rawMarkdown": "Oh, thanks for pointing out! I added correct one.",
          "votes": 4
        },
        {
          "id": 1238055,
          "postDate": "2021-03-14T15:43:13.193Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1240629,
          "postDate": "2021-03-16T14:41:09.247Z",
          "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> your public kerne is greatl. We have slightly similar set up but with few tweaks.</p>\n<p>Its possible to achieve <code>CV-5.4</code> and <code>LB-6.9</code> with just <code>resnet34</code> trained for only <code>10</code> epoch.  In fact our best epoch is <code>9</code> =) </p>",
          "rawMarkdown": "@yasufuminakama your public kerne is greatl. We have slightly similar set up but with few tweaks.\n\nIts possible to achieve `CV-5.4` and `LB-6.9` with just `resnet34` trained for only `10` epoch.  In fact our best epoch is `9` =) ",
          "votes": 8
        },
        {
          "id": 1241092,
          "postDate": "2021-03-16T21:49:47.353Z",
          "content": "<p><a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> Thanks for information! <br>\nI have already reached CV-4.x with resnet34 encoder :D<br>\nSo then your current LB-3.8 is from other encoder? :)</p>",
          "rawMarkdown": "@drhabib Thanks for information! \nI have already reached CV-4.x with resnet34 encoder :D\nSo then your current LB-3.8 is from other encoder? :)",
          "votes": 6
        },
        {
          "id": 1241120,
          "postDate": "2021-03-16T22:44:56.903Z",
          "content": "<p>yes current score can be achieved with <code>resnet50</code> xD But also if I try hard can be reached with <code>resnet34</code> =) </p>",
          "rawMarkdown": "yes current score can be achieved with `resnet50` xD But also if I try hard can be reached with `resnet34` =) ",
          "votes": 5
        },
        {
          "id": 1243165,
          "postDate": "2021-03-18T04:11:06.780Z",
          "content": "<p>your score is impressive,how long to train for your current LB?</p>",
          "rawMarkdown": "your score is impressive,how long to train for your current LB?",
          "votes": 1
        },
        {
          "id": 1243194,
          "postDate": "2021-03-18T04:38:57.803Z",
          "content": "<p>~2hours per epoch for 80% train images on TITAN RTX.</p>",
          "rawMarkdown": "~2hours per epoch for 80% train images on TITAN RTX.",
          "votes": 9
        },
        {
          "id": 1246869,
          "postDate": "2021-03-21T07:37:34.793Z",
          "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> Have you updated the public notebook with the flip code</p>\n<pre><code>import albumentations as A\n\ntransform = A.Compose([A.Transpose(p=1), A.VerticalFlip(p=1)])\n\nh, w, _ = image.shape\nif h &gt; w:\n    image = transform(image=image)['image']\n</code></pre>",
          "rawMarkdown": "@yasufuminakama Have you updated the public notebook with the flip code\n```\nimport albumentations as A\n\ntransform = A.Compose([A.Transpose(p=1), A.VerticalFlip(p=1)])\n\nh, w, _ = image.shape\nif h > w:\n    image = transform(image=image)['image']\n```"
        },
        {
          "id": 1246874,
          "postDate": "2021-03-21T07:41:00.720Z",
          "content": "<p>Yes, it's used in latest version and scored LB:20.3</p>",
          "rawMarkdown": "Yes, it's used in latest version and scored LB:20.3",
          "votes": 1
        },
        {
          "id": 1246886,
          "postDate": "2021-03-21T07:54:05.403Z",
          "content": "<p>if we apply this augmentation every image will be of a different shape right? will it be any problem while training</p>",
          "rawMarkdown": "if we apply this augmentation every image will be of a different shape right? will it be any problem while training"
        },
        {
          "id": 1246895,
          "postDate": "2021-03-21T08:11:04.977Z",
          "content": "<p>Yes, I don't use it for training.</p>",
          "rawMarkdown": "Yes, I don't use it for training.",
          "votes": 1
        },
        {
          "id": 1246920,
          "postDate": "2021-03-21T08:50:20Z",
          "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> i think you didnt changed in train notebook</p>",
          "rawMarkdown": "@yasufuminakama i think you didnt changed in train notebook"
        },
        {
          "id": 1246934,
          "postDate": "2021-03-21T09:06:37.440Z",
          "content": "<p>It's always h &lt; w for train images so you don't need it.</p>",
          "rawMarkdown": "It's always h < w for train images so you don't need it.",
          "votes": 1
        },
        {
          "id": 1246936,
          "postDate": "2021-03-21T09:07:15.163Z",
          "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>  then where do you use it?</p>",
          "rawMarkdown": "@yasufuminakama  then where do you use it?"
        },
        {
          "id": 1246938,
          "postDate": "2021-03-21T09:09:03.953Z",
          "content": "<p>For inference.</p>",
          "rawMarkdown": "For inference.",
          "votes": 2
        },
        {
          "id": 1258149,
          "postDate": "2021-03-31T11:45:08.617Z",
          "content": "<p>end to end training for CNN-LSTM is very slow. what i did in other competition is:</p>\n<ol>\n<li><p>pretrain CNN only using classification (this is competition, you can probably do a few tasks like predicting number of C atoms, H atoms, etc (regression), classify if the image contains some functional group (or partial InChi substring, or if a word is present or not from the token dictionary, aka multi binary label), segmentation, clustering, contrastive learning, etc …. just think of some simple way to generate ground truths from the InChI)</p></li>\n<li><p>with the CNN layers froze, train your LSTM only</p></li>\n<li><p>finally finetune CNN+LSTM with low learning rate for the target image-to-text translation task.</p></li>\n</ol>\n<p>The trick is to think of simple tasks in step (1)  that is closely related to the target task (3)</p>",
          "rawMarkdown": "end to end training for CNN-LSTM is very slow. what i did in other competition is:\n\n1. pretrain CNN only using classification (this is competition, you can probably do a few tasks like predicting number of C atoms, H atoms, etc (regression), classify if the image contains some functional group (or partial InChi substring, or if a word is present or not from the token dictionary, aka multi binary label), segmentation, clustering, contrastive learning, etc .... just think of some simple way to generate ground truths from the InChI)\n\n2. with the CNN layers froze, train your LSTM only\n\n3. finally finetune CNN+LSTM with low learning rate for the target image-to-text translation task.\n\nThe trick is to think of simple tasks in step (1)  that is closely related to the target task (3)\n\n",
          "votes": 5
        },
        {
          "id": 1259826,
          "postDate": "2021-04-01T17:24:49.933Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>. When I saw this yesterday, I immediately started a notebook go through these steps. I got good results with doing Regression on the number of atoms of each element. I just shared my notebook publicly here: <a href=\"https://www.kaggle.com/moeinshariatnia/cnn-rnn-cnn-pretraining-w-regression-train\" target=\"_blank\">[CNN+RNN]CNN pretraining w/ regression -TRAIN</a></p>\n<p>I'll be really glad if you and others tell me what you think about it.</p>\n<p><img src=\"https://i.ibb.co/WPDk5d7/Presentation1.jpg\" alt=\"\"></p>",
          "rawMarkdown": "Thanks @hengck23. When I saw this yesterday, I immediately started a notebook go through these steps. I got good results with doing Regression on the number of atoms of each element. I just shared my notebook publicly here: [[CNN+RNN]CNN pretraining w/ regression -TRAIN](https://www.kaggle.com/moeinshariatnia/cnn-rnn-cnn-pretraining-w-regression-train)\n\nI'll be really glad if you and others tell me what you think about it.\n\n![](https://i.ibb.co/WPDk5d7/Presentation1.jpg)"
        }
      ]
    },
    {
      "id": 1268777,
      "postDate": "2021-04-09T18:37:30.067Z",
      "content": "<p>I completed CNN+LSTM experiments and now are playing with vision transformer(essentially a transformer encoder)+transformer decoder. <br>\n(also refer to my post for code and details: <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231190\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/231190</a>)</p>\n<p>here are the first results:</p>\n<p>transformer in transformer:    <br>\n     - TNT-S-224-16: LB 3.11/CV 1.96 (inference = 80 min on 4xti1080)<br>\n     - TNT-S-320-16: LB 2.27/CV 1.55 (inference = 140 min on 4xti1080)</p>\n<p>(do you have others to suggest?)</p>\n<hr>\n<p>train: input 224x224, no augmentation (i.e. rotation is not used)<br>\ntest: use YNakama w&lt;h rotate trick, greedy (argmax) decoder, aka. no beam search yet<br>\ntokenizer: YNakama's Tokenizer</p>\n<hr>\n<p>a few quick observations:</p>\n<ul>\n<li>transformer train faster, 2 days is enough</li>\n<li>you actually don't need many epoch (7 to 10 is ok), but you do need many iterations and low learning rate at the end.<br>\n(many iterations can mean you use a smaller batch size. i use 64)</li>\n<li>cv/lb gap is larger (which i think if you use teacher distillation with CNN+LSTM, maybe you can get interesting results)</li>\n</ul>\n<p>for more ideas: read this <a href=\"https://github.com/lucidrains/vit-pytorch\" target=\"_blank\">https://github.com/lucidrains/vit-pytorch</a></p>",
      "rawMarkdown": "I completed CNN+LSTM experiments and now are playing with vision transformer(essentially a transformer encoder)+transformer decoder. \n(also refer to my post for code and details: https://www.kaggle.com/c/bms-molecular-translation/discussion/231190)\n\nhere are the first results:\n\ntransformer in transformer:    \n     - TNT-S-224-16: LB 3.11/CV 1.96 (inference = 80 min on 4xti1080)\n     - TNT-S-320-16: LB 2.27/CV 1.55 (inference = 140 min on 4xti1080)\n\n \n\n(do you have others to suggest?)\n\n---\n\ntrain: input 224x224, no augmentation (i.e. rotation is not used)\ntest: use YNakama w<h rotate trick, greedy (argmax) decoder, aka. no beam search yet\ntokenizer: YNakama's Tokenizer\n\n---\n\na few quick observations:\n- transformer train faster, 2 days is enough\n- you actually don't need many epoch (7 to 10 is ok), but you do need many iterations and low learning rate at the end.\n(many iterations can mean you use a smaller batch size. i use 64)\n- cv/lb gap is larger (which i think if you use teacher distillation with CNN+LSTM, maybe you can get interesting results)\n\nfor more ideas: read this https://github.com/lucidrains/vit-pytorch",
      "votes": 9,
      "replies": [
        {
          "id": 1268813,
          "postDate": "2021-04-09T19:33:15.647Z",
          "content": "<p>I have also the same problem with the batch size :) bs=64 somehow always perform better than bigger batch sizes. Seems to be problem of multi GPU training?</p>",
          "rawMarkdown": "I have also the same problem with the batch size :) bs=64 somehow always perform better than bigger batch sizes. Seems to be problem of multi GPU training?",
          "votes": 3
        },
        {
          "id": 1268892,
          "postDate": "2021-04-09T23:01:44.633Z",
          "content": "<p>no. my GPU has 48 GB. </p>",
          "rawMarkdown": "no. my GPU has 48 GB. ",
          "votes": 2
        },
        {
          "id": 1271054,
          "postDate": "2021-04-12T08:50:32.463Z",
          "content": "<p>For the TNT-S-320-16 model, have you used a pretrained model ?</p>",
          "rawMarkdown": "For the TNT-S-320-16 model, have you used a pretrained model ?"
        },
        {
          "id": 1272367,
          "postDate": "2021-04-13T12:27:19.690Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> do we have a pretrained model for TNT-S-320-16 model</p>",
          "rawMarkdown": "@hengck23 do we have a pretrained model for TNT-S-320-16 model"
        },
        {
          "id": 1287622,
          "postDate": "2021-04-29T08:35:11.817Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<blockquote>\n  <p>transformer train faster, 2 days is enough</p>\n</blockquote>\n<p>I don't agree with that in my case. it takes ages for me</p>",
          "rawMarkdown": "@hengck23 \n> transformer train faster, 2 days is enough\n\nI don't agree with that in my case. it takes ages for me"
        },
        {
          "id": 1287813,
          "postDate": "2021-04-29T12:29:01.647Z",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> I think he has many GPUs :)</p>",
          "rawMarkdown": "@morizin I think he has many GPUs :)",
          "votes": 2
        }
      ]
    },
    {
      "id": 1303020,
      "postDate": "2021-05-11T21:19:50.243Z",
      "content": "<p>Hi there. We thought we would share some of our benchmarks as others have been so forthcoming. It's been a wonderful competition so far.</p>\n<hr>\n<p><strong>SETUP:</strong><br>\n    - EfficientNetB5 <em>(headless… no extra layers)</em><br>\n    - 256/768 (attn/lstm-units) LSTM<br>\n    - 192x384 cropped/stretched/rotation-fixed images<br>\n    - TFRecords and TPU<br>\n    - Everything done on Kaggle in 1-2 sessions.</p>\n<hr>\n<p>12 EPOCHS (5-6 hours of training) - LIMIT MAX LENGTH TO 140 TOKENS<br>\n    - <strong>1.95 CV</strong> - No Beam Search - No Postprocessing<br>\n    - <strong>3.75 LB</strong> - No Beam Search - No Postprocessing</p>\n<p>10 ADDITIONAL EPOCHS (5-6 hours of training) - FINE TUNE ON FULL LENGTH (277 TOKENS)<br>\n    - <strong>1.89 CV</strong> - No Beam Search - No Postprocessing<br>\n    - <strong>2.68 LB</strong> - No Beam Search - No Postprocessing<br>\n    - <strong>2.58 LB</strong> - No Beam Search - RDKit Postprocessing</p>\n<hr>\n<p>We are implementing Beam Search and swapping in Transformers shortly and hopefully, that will help. At some point in the near future, we will scale up image size and model size and increase training time and try to squeeze below 1.</p>",
      "rawMarkdown": "Hi there. We thought we would share some of our benchmarks as others have been so forthcoming. It's been a wonderful competition so far.\n\n---\n\n**SETUP:**\n    - EfficientNetB5 *(headless... no extra layers)*\n    - 256/768 (attn/lstm-units) LSTM\n    - 192x384 cropped/stretched/rotation-fixed images\n    - TFRecords and TPU\n    - Everything done on Kaggle in 1-2 sessions.\n\n---\n\n12 EPOCHS (5-6 hours of training) - LIMIT MAX LENGTH TO 140 TOKENS\n    - **1.95 CV** - No Beam Search - No Postprocessing\n    - **3.75 LB** - No Beam Search - No Postprocessing\n\n10 ADDITIONAL EPOCHS (5-6 hours of training) - FINE TUNE ON FULL LENGTH (277 TOKENS)\n    - **1.89 CV** - No Beam Search - No Postprocessing\n    - **2.68 LB** - No Beam Search - No Postprocessing\n    - **2.58 LB** - No Beam Search - RDKit Postprocessing\n\n---\n\nWe are implementing Beam Search and swapping in Transformers shortly and hopefully, that will help. At some point in the near future, we will scale up image size and model size and increase training time and try to squeeze below 1.",
      "votes": 9,
      "replies": [
        {
          "id": 1303255,
          "postDate": "2021-05-12T01:31:22.723Z",
          "content": "<p>\"2.58 LB - No Beam Search - RDKit Postprocessing\"</p>\n<p>if inchi validation is the key, you might as well train an \"inchi correction model\"</p>\n<p>how to get corrupted inchi? </p>\n<ul>\n<li>use prediction of your model in earlier iterations (i.e. before it converges)</li>\n<li>add noise to input image</li>\n</ul>\n<p>this is seq-to-seq model : input invalid inchi, output best guess of valid inchi</p>\n<hr>\n<p>better still, train a refinement model and this can be repeated many rounds<br>\n(e.g. alternating between lstm and transformer)</p>\n<p>model : input previous model predicted inchi+ image, output best guess of true inchi</p>",
          "rawMarkdown": "\"2.58 LB - No Beam Search - RDKit Postprocessing\"\n\nif inchi validation is the key, you might as well train an \"inchi correction model\"\n\nhow to get corrupted inchi? \n- use prediction of your model in earlier iterations (i.e. before it converges)\n- add noise to input image\n\nthis is seq-to-seq model : input invalid inchi, output best guess of valid inchi\n\n---\n\nbetter still, train a refinement model and this can be repeated many rounds\n(e.g. alternating between lstm and transformer)\n\nmodel : input previous model predicted inchi+ image, output best guess of true inchi",
          "votes": 1
        },
        {
          "id": 1303543,
          "postDate": "2021-05-12T06:06:31.603Z",
          "content": "<p>This is a great result, thank you for sharing all this in such a readable way <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>!</p>\n<p>Training on a GPU  for ~9 epochs, using LSTM + attention (1024 decoder dim), image size of 384x384 (rotated during training by 90 deg clockwise with a p of 0.5) I get 4.71 local CV.</p>\n<p>I wonder - would you be willing to share what augmentations you are using? My reasoning is this - there is so much data in train, one should be good without using augmentations? Really curious if people are seeing augmentations improve their results and what augmentations are helping.</p>",
          "rawMarkdown": "This is a great result, thank you for sharing all this in such a readable way @dschettler8845!\n\nTraining on a GPU  for ~9 epochs, using LSTM + attention (1024 decoder dim), image size of 384x384 (rotated during training by 90 deg clockwise with a p of 0.5) I get 4.71 local CV.\n\nI wonder - would you be willing to share what augmentations you are using? My reasoning is this - there is so much data in train, one should be good without using augmentations? Really curious if people are seeing augmentations improve their results and what augmentations are helping.",
          "votes": 1
        },
        {
          "id": 1303892,
          "postDate": "2021-05-12T10:21:13.063Z",
          "content": "<p>Thank you! </p>\n<p>The only augmentation I’m performing is rotating when the cropped molecule aspect ratio is smaller than 1 (ie h&gt;w).</p>\n<p>I plan to iron out the kinks in this approach (instead of just h v. W) in the future (ie a model to predict if an image needs to be rotated).</p>\n<p>The only augmentation I can think of is increasing the amount of noise in the image (or decreasing)… I haven’t implemented it yet though. I believe others may have. </p>\n<p>I hope this helps!</p>",
          "rawMarkdown": "Thank you! \n\nThe only augmentation I’m performing is rotating when the cropped molecule aspect ratio is smaller than 1 (ie h>w).\n\nI plan to iron out the kinks in this approach (instead of just h v. W) in the future (ie a model to predict if an image needs to be rotated).\n\nThe only augmentation I can think of is increasing the amount of noise in the image (or decreasing)... I haven’t implemented it yet though. I believe others may have. \n\nI hope this helps!",
          "votes": 1
        },
        {
          "id": 1305188,
          "postDate": "2021-05-13T06:20:31.723Z",
          "content": "<p>This was very helpful indeed! 😊 Thank you very much for sharing your thoughts!</p>\n<p>Seems that if I want to improve my results, my time would be better spent elsewhere than messing around with the augmentations. Thank you!</p>",
          "rawMarkdown": "This was very helpful indeed! 😊 Thank you very much for sharing your thoughts!\n\nSeems that if I want to improve my results, my time would be better spent elsewhere than messing around with the augmentations. Thank you!"
        },
        {
          "id": 1306824,
          "postDate": "2021-05-14T05:33:44.987Z",
          "content": "<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> apologies, one more question if I may please. I am trying to get an intuition for just how fast TPUs are.</p>\n<p>When you mention a single epoch, how many images do you mean? The entire train set of 2.4 mln images? You are using all 8 TPU cores I would imagine?</p>",
          "rawMarkdown": "@dschettler8845 apologies, one more question if I may please. I am trying to get an intuition for just how fast TPUs are.\n\nWhen you mention a single epoch, how many images do you mean? The entire train set of 2.4 mln images? You are using all 8 TPU cores I would imagine?"
        },
        {
          "id": 1307614,
          "postDate": "2021-05-14T14:54:34.293Z",
          "content": "<p>Yup. When I say a single epoch I mean all of the training data (- 80,000 validation images)… so approximately 2.3-2.4 million images.</p>\n<p>TPU is incredibly useful in this competition.</p>\n<p>I am distributing the training across all 8 cores.</p>\n<p>I use a batch size of 1024 (128 for each core).</p>\n<p>Hope this helps!</p>",
          "rawMarkdown": "Yup. When I say a single epoch I mean all of the training data (- 80,000 validation images)... so approximately 2.3-2.4 million images.\n\nTPU is incredibly useful in this competition.\n\nI am distributing the training across all 8 cores.\n\nI use a batch size of 1024 (128 for each core).\n\nHope this helps!",
          "votes": 1
        },
        {
          "id": 1307652,
          "postDate": "2021-05-14T15:26:07.710Z",
          "content": "<p>TF distributes efficiently both the model and the data across all cores.   That's why you can train really fast even with huge models on TPU. <br>\nMixed Precision with bfloat16 reduces the training time by half. </p>\n<p>Gradient Accumulation helps to even speed up the training again.  But it's a bit trickier to implement on TF while it's really straightfoward on Pytorch. </p>\n<p>And of course with the TFRecords dataset in GCS near the TPU VM, there is no I/O bottleneck</p>",
          "rawMarkdown": "TF distributes efficiently both the model and the data across all cores.   That's why you can train really fast even with huge models on TPU. \nMixed Precision with bfloat16 reduces the training time by half. \n\nGradient Accumulation helps to even speed up the training again.  But it's a bit trickier to implement on TF while it's really straightfoward on Pytorch. \n\nAnd of course with the TFRecords dataset in GCS near the TPU VM, there is no I/O bottleneck",
          "votes": 3
        },
        {
          "id": 1308484,
          "postDate": "2021-05-15T08:50:19.517Z",
          "content": "<p>Thank you very much for these additional details <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>! 2.4 million images in ~30 minutes, especially with a sequential model such as RNN, that is super impressive 🙂</p>",
          "rawMarkdown": "Thank you very much for these additional details @dschettler8845! 2.4 million images in ~30 minutes, especially with a sequential model such as RNN, that is super impressive 🙂",
          "votes": 2
        },
        {
          "id": 1308517,
          "postDate": "2021-05-15T09:20:39.817Z",
          "content": "<p>Thank you for sharing this information <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a>! Appreciate it</p>\n<p>Could I please ask about why you believe gradient accumulation can be helpful? </p>",
          "rawMarkdown": "Thank you for sharing this information @serigne! Appreciate it\n\nCould I please ask about why you believe gradient accumulation can be helpful? ",
          "votes": 1
        },
        {
          "id": 1308591,
          "postDate": "2021-05-15T10:18:16.543Z",
          "content": "<p>Sure, </p>\n<p>Because when you apply gradient accumulation, The gradient update (<code>optimizer.apply_gradients</code> on TF  and <code>optimizer.step()</code> on Pytorch) is done only  every K steps intead of every step. </p>\n<p>Without Gradient Accumulation, the distributed training would compute loss reduce for all cores at every step . Therefore, with gradient accumulation you will save <code>num_cores*(K-1)</code> loss reduce time at every gradient update.  Then, the more you have number of cores, the more you will accelerate training with GA. <br>\nBut you may need to tune the learning rate accordingly. </p>\n<p>On Colab( which has TPU v2-8), you can manage to make the training almost as fast as on kaggle (which has TPU v3-8) with GA.   That's why I don't bother to use Kaggle TPU. </p>\n<p>You can see <a href=\"https://www.tensorflow.org/xla\" target=\"_blank\">here</a> a benchmark w/ and w/o GA for distributed training </p>",
          "rawMarkdown": "Sure, \n\nBecause when you apply gradient accumulation, The gradient update (`optimizer.apply_gradients` on TF  and `optimizer.step() ` on Pytorch) is done only  every K steps intead of every step. \n\nWithout Gradient Accumulation, the distributed training would compute loss reduce for all cores at every step . Therefore, with gradient accumulation you will save `num_cores*(K-1)` loss reduce time at every gradient update.  Then, the more you have number of cores, the more you will accelerate training with GA. \nBut you may need to tune the learning rate accordingly. \n\nOn Colab( which has TPU v2-8), you can manage to make the training almost as fast as on kaggle (which has TPU v3-8) with GA.   That's why I don't bother to use Kaggle TPU. \n\nYou can see [here](https://www.tensorflow.org/xla) a benchmark w/ and w/o GA for distributed training ",
          "votes": 2
        },
        {
          "id": 1308663,
          "postDate": "2021-05-15T11:15:49.523Z",
          "content": "<p>Thank you very much for your answer <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a>!!!</p>",
          "rawMarkdown": "Thank you very much for your answer @serigne!!!"
        }
      ]
    },
    {
      "id": 1267440,
      "postDate": "2021-04-08T14:13:17.970Z",
      "content": "<p>With the update of new rules, i want to update our score. Our current LB standing doesn’t break any rules. Best model has CV of 1.29 and LB of 1.34. I hope this will encourage participants to build model with high scores:) </p>",
      "rawMarkdown": "With the update of new rules, i want to update our score. Our current LB standing doesn’t break any rules. Best model has CV of 1.29 and LB of 1.34. I hope this will encourage participants to build model with high scores:) ",
      "votes": 8
    },
    {
      "id": 1230093,
      "postDate": "2021-03-07T19:16:32.597Z",
      "content": "<p>Just trained 1 epoch to verify my code works… e0+lstm. got 24 in my holdout set and 44 on the public. Quite a large gap.</p>\n<p>Mar-10<br>\nFixed some issue, and I am moving again. Finished another epoch and I got 19 local and 32 on the public. </p>\n<table>\n<thead>\n<tr>\n<th>local 5% hold out</th>\n<th>public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>24</td>\n<td>44</td>\n</tr>\n<tr>\n<td>19</td>\n<td>32</td>\n</tr>\n<tr>\n<td>16</td>\n<td>29</td>\n</tr>\n<tr>\n<td>13</td>\n<td>27</td>\n</tr>\n</tbody>\n</table>\n<p>Above table does not handle the different orient in test set. </p>\n<p>Mar-17 Had some significant code changes and trained new model end to end from scratch, still e0 encoder lstm decoder. </p>\n<table>\n<thead>\n<tr>\n<th>epoch</th>\n<th>local 20% hold out</th>\n<th>public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>3</td>\n<td>5.4</td>\n<td>5.9</td>\n</tr>\n<tr>\n<td>4</td>\n<td>4.9</td>\n<td>5.6</td>\n</tr>\n<tr>\n<td>5</td>\n<td>4.6</td>\n<td>5.2</td>\n</tr>\n</tbody>\n</table>\n<p>ensemble 5 fold(all similar validation scores) of the above setup -&gt; 4.11 LB. <br>\nIt seems ensemble is much better than average of cv</p>\n<p>Used rdkit to normalize the 4.11 predictions, score improved to 4.07. It seems there is some juice worth squeezing. </p>",
      "rawMarkdown": "Just trained 1 epoch to verify my code works... e0+lstm. got 24 in my holdout set and 44 on the public. Quite a large gap.\n\nMar-10\nFixed some issue, and I am moving again. Finished another epoch and I got 19 local and 32 on the public. \n\n| local 5% hold out | public |\n| --- | --- |\n| 24 | 44 |\n| 19 | 32 |\n| 16 | 29 |\n| 13 | 27 |\n\nAbove table does not handle the different orient in test set. \n\nMar-17 Had some significant code changes and trained new model end to end from scratch, still e0 encoder lstm decoder. \n| epoch | local 20% hold out | public |\n| --- | --- | --- |\n| 3 | 5.4 | 5.9 |\n| 4 | 4.9 | 5.6 | \n| 5 | 4.6 | 5.2 |\n\nensemble 5 fold(all similar validation scores) of the above setup -> 4.11 LB. \nIt seems ensemble is much better than average of cv\n\nUsed rdkit to normalize the 4.11 predictions, score improved to 4.07. It seems there is some juice worth squeezing. ",
      "votes": 7,
      "replies": [
        {
          "id": 1230124,
          "postDate": "2021-03-07T19:47:55.717Z",
          "content": "<p>Hopefully you are continuing training =) Score increases greatly after first epoch (in my case) … Yes the gap thing is concerning, it will be important to analyze the data and the output of models ..</p>",
          "rawMarkdown": "Hopefully you are continuing training =) Score increases greatly after first epoch (in my case) ... Yes the gap thing is concerning, it will be important to analyze the data and the output of models ..",
          "votes": 3
        },
        {
          "id": 1230163,
          "postDate": "2021-03-07T20:35:05.877Z",
          "content": "<p>Are you using an autoregressive model? If so, while evaluating do you generate the output token by token or you generate it all in parallel by giving it the real previous tokens?</p>",
          "rawMarkdown": "Are you using an autoregressive model? If so, while evaluating do you generate the output token by token or you generate it all in parallel by giving it the real previous tokens?",
          "votes": 3
        },
        {
          "id": 1230165,
          "postDate": "2021-03-07T20:39:34.717Z",
          "content": "<p>Yes, autoregressive model. When validating, I am not using real tokens. I generate until <code>&lt;EOS&gt;</code>. After that I decode the generated text to calculate the score. Thus I don't think I have leak in my local validation code. </p>",
          "rawMarkdown": "Yes, autoregressive model. When validating, I am not using real tokens. I generate until `<EOS>`. After that I decode the generated text to calculate the score. Thus I don't think I have leak in my local validation code. ",
          "votes": 2
        },
        {
          "id": 1230174,
          "postDate": "2021-03-07T20:48:38.050Z",
          "content": "<p>Ok, just checking, because even if I haven't achieved your results yet, my public score has been pretty consistent with my validation so far, maybe as the model gets better the gap increases.</p>",
          "rawMarkdown": "Ok, just checking, because even if I haven't achieved your results yet, my public score has been pretty consistent with my validation so far, maybe as the model gets better the gap increases.",
          "votes": 4
        },
        {
          "id": 1232345,
          "postDate": "2021-03-09T17:17:32.573Z",
          "content": "<p>Did you look at the train and test images?<br>\nTraining images are horizontal only, but in the test, there are horizontal and 90° (clockwise) rotated ones.<br>\nDid you take care of this difference?</p>",
          "rawMarkdown": "Did you look at the train and test images?\nTraining images are horizontal only, but in the test, there are horizontal and 90° (clockwise) rotated ones.\nDid you take care of this difference?",
          "votes": 2
        },
        {
          "id": 1232356,
          "postDate": "2021-03-09T17:30:56.097Z",
          "content": "<p>I haven't looked at any images at all. I was primarily testing my code. Thanks for the info <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a>.  I think I will try to add rotation as an augmentation. </p>\n<p><a href=\"https://www.kaggle.com/DrHB\" target=\"_blank\">@DrHB</a> mind share what batch size are you using? I stopped it after one epoch and loaded it up in a larger GPU to train in a larger batch size(128 in this case) and my training became very unstable, had a few blow-ups. Or, are you clipping gradients?</p>",
          "rawMarkdown": "I haven't looked at any images at all. I was primarily testing my code. Thanks for the info @nofreewill.  I think I will try to add rotation as an augmentation. \n\n@DrHB mind share what batch size are you using? I stopped it after one epoch and loaded it up in a larger GPU to train in a larger batch size(128 in this case) and my training became very unstable, had a few blow-ups. Or, are you clipping gradients?",
          "votes": 1
        },
        {
          "id": 1232362,
          "postDate": "2021-03-09T17:37:24.990Z",
          "content": "<p>I can train on <code>fp16</code> with batch size <code>256</code> and my <code>gpu</code> still have 3-4 gb free left from <code>16</code> =)  </p>",
          "rawMarkdown": "I can train on `fp16` with batch size `256` and my `gpu` still have 3-4 gb free left from `16` =)  ",
          "votes": 5
        },
        {
          "id": 1232366,
          "postDate": "2021-03-09T17:39:17.607Z",
          "content": "<p>btw with <code>fp16</code> several time I was getting <code>nan</code> loss … adding warm up helped with <code>gradient clipping</code>. </p>",
          "rawMarkdown": "btw with `fp16` several time I was getting `nan` loss ... adding warm up helped with `gradient clipping`. ",
          "votes": 2
        },
        {
          "id": 1232374,
          "postDate": "2021-03-09T17:43:43.410Z",
          "content": "<p>Ah, thanks a lot! </p>",
          "rawMarkdown": "Ah, thanks a lot! ",
          "votes": 1
        },
        {
          "id": 1232381,
          "postDate": "2021-03-09T17:47:44.267Z",
          "content": "<p>this worked for me… I think <code>max_norm</code> can go to <code>4.</code> you have to play around…  </p>\n<pre><code>norm_type = 2.0\nmax_norm  = 1.\nnn.utils.clip_grad_norm_(model.parameters(), max_norm, norm_type)\n</code></pre>\n<p>good luck!</p>",
          "rawMarkdown": "this worked for me... I think `max_norm` can go to `4.` you have to play around...  \n```\nnorm_type = 2.0\nmax_norm  = 1.\nnn.utils.clip_grad_norm_(model.parameters(), max_norm, norm_type)\n```\n\ngood luck!",
          "votes": 3
        },
        {
          "id": 1255883,
          "postDate": "2021-03-29T10:08:59.043Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ryanzhang\" target=\"_blank\">@ryanzhang</a> , <br>\nIf you dont mind can you give a brief idea of how you are creating the ensemble….</p>",
          "rawMarkdown": "Hi @ryanzhang , \nIf you dont mind can you give a brief idea of how you are creating the ensemble...."
        },
        {
          "id": 1255894,
          "postDate": "2021-03-29T10:17:53.217Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1253530,
      "postDate": "2021-03-26T20:48:24.033Z",
      "content": "<p>resnet34, 10 epochs -&gt; single fold validation: 8.27, lb: 8.87</p>",
      "rawMarkdown": "resnet34, 10 epochs -> single fold validation: 8.27, lb: 8.87",
      "votes": 5,
      "replies": [
        {
          "id": 1253534,
          "postDate": "2021-03-26T20:54:02.083Z",
          "content": "<p>Is this the public kernel?</p>",
          "rawMarkdown": "Is this the public kernel?"
        },
        {
          "id": 1253894,
          "postDate": "2021-03-27T06:37:15.570Z",
          "content": "<p>I started with this one as basis, yes.</p>",
          "rawMarkdown": "I started with this one as basis, yes.",
          "votes": 1
        },
        {
          "id": 1254279,
          "postDate": "2021-03-27T13:39:09.027Z",
          "content": "<p>Have you added beam search to get such result ?</p>",
          "rawMarkdown": "Have you added beam search to get such result ?"
        },
        {
          "id": 1254284,
          "postDate": "2021-03-27T13:44:35.547Z",
          "content": "<p>Not yet, no.</p>",
          "rawMarkdown": "Not yet, no."
        }
      ]
    },
    {
      "id": 1265529,
      "postDate": "2021-04-07T01:47:33.157Z",
      "content": "<ul>\n<li>Encoder: Pruned EfficientNetB3</li>\n<li>Decoder: GRU</li>\n<li>Image size: <code>224</code></li>\n<li>Augmentation: None</li>\n<li>CV: <code>2.79</code></li>\n<li>CV with beam search (k=5): <code>2.65</code></li>\n<li>LB without beam search: <code>3.57</code></li>\n<li>LB with beam search (k=5): <code>3.22</code></li>\n</ul>",
      "rawMarkdown": "- Encoder: Pruned EfficientNetB3\n- Decoder: GRU\n- Image size: `224`\n- Augmentation: None\n- CV: `2.79`\n- CV with beam search (k=5): `2.65`\n- LB without beam search: `3.57`\n- LB with beam search (k=5): `3.22`\n\n",
      "votes": 6,
      "replies": [
        {
          "id": 1265567,
          "postDate": "2021-04-07T03:12:13.547Z",
          "content": "<p>thanks!.</p>\n<p>i forget there is such a thing called pruned model though i used it in work.<br>\nby the way, seems that pruned model uses less watts (saving in electricity bill)</p>\n<pre><code>|===============================+======================+=|\n|   0  TITAN X (Pascal)    Off  | 00000000:05:00.0  On |                  N/A |\n| 59%   84C    P2   105W / 250W |  10302MiB / 12192MiB |    100%      Default |\n+-------------------------------+----------------------+----------------------+\n</code></pre>",
          "rawMarkdown": "thanks!.\n\ni forget there is such a thing called pruned model though i used it in work.\nby the way, seems that pruned model uses less watts (saving in electricity bill)\n\n```\n|===============================+======================+=|\n|   0  TITAN X (Pascal)    Off  | 00000000:05:00.0  On |                  N/A |\n| 59%   84C    P2   105W / 250W |  10302MiB / 12192MiB |    100%      Default |\n+-------------------------------+----------------------+----------------------+\n  ``` ",
          "votes": 1
        },
        {
          "id": 1265724,
          "postDate": "2021-04-07T06:55:03.383Z",
          "content": "<p>Are you using attention with GRU ?</p>",
          "rawMarkdown": "Are you using attention with GRU ?",
          "votes": 1
        },
        {
          "id": 1265747,
          "postDate": "2021-04-07T07:22:34.857Z",
          "content": "<p>Yes, GRU with attention</p>",
          "rawMarkdown": "Yes, GRU with attention"
        },
        {
          "id": 1267343,
          "postDate": "2021-04-08T13:29:49.410Z",
          "content": "<p><a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> just one layer of GRU or you stack some?</p>",
          "rawMarkdown": "@tuckerarrants just one layer of GRU or you stack some?",
          "votes": 2
        },
        {
          "id": 1267395,
          "postDate": "2021-04-08T13:45:20.860Z",
          "content": "<p>Just one GRU layer</p>",
          "rawMarkdown": "Just one GRU layer",
          "votes": 1
        },
        {
          "id": 1308750,
          "postDate": "2021-05-15T12:41:28.323Z",
          "content": "<p>I wonder if using a pruned EffNet makes sense here. It was pruned for the purposed of the task it was trained for right? Here the task is rather different. What are your thoughts?</p>",
          "rawMarkdown": "I wonder if using a pruned EffNet makes sense here. It was pruned for the purposed of the task it was trained for right? Here the task is rather different. What are your thoughts?",
          "votes": 1
        }
      ]
    },
    {
      "id": 1230112,
      "postDate": "2021-03-07T19:33:57.840Z",
      "content": "<p>Great job on pushing the leaderboard forward. I'm still working on my first submission, hopefully I'll have the basis for my model setup in a couple days.</p>",
      "rawMarkdown": "Great job on pushing the leaderboard forward. I'm still working on my first submission, hopefully I'll have the basis for my model setup in a couple days.",
      "votes": 4
    },
    {
      "id": 1267645,
      "postDate": "2021-04-08T17:08:15.797Z",
      "content": "<p>I just used some public kernels and got CV 3.xx with B0 without any change in the decoder.<br>\nPersonally, I think finding a good training strategy with suitable batch size and LR scheduler is essential.<br>\n P/S: My hardware resource is limited.  </p>",
      "rawMarkdown": "I just used some public kernels and got CV 3.xx with B0 without any change in the decoder.\nPersonally, I think finding a good training strategy with suitable batch size and LR scheduler is essential.\n P/S: My hardware resource is limited.  ",
      "votes": 3
    },
    {
      "id": 1267208,
      "postDate": "2021-04-08T11:40:30.710Z",
      "content": "<p>Encoder: eca_nfnet_l0<br>\nDecoder: LSTM+GRU<br>\nImage size: 224<br>\nBS=128<br>\nscheduler='CosineAnnealingLR'<br>\nAugmentation: None<br>\nCV: 4.61<br>\nLB: 5.85</p>\n<p>1 epoch takes 5h on my local machine.</p>",
      "rawMarkdown": "Encoder: eca_nfnet_l0\nDecoder: LSTM+GRU\nImage size: 224\nBS=128\nscheduler='CosineAnnealingLR'\nAugmentation: None\nCV: 4.61\nLB: 5.85\n\n1 epoch takes 5h on my local machine.",
      "votes": 3
    },
    {
      "id": 1291931,
      "postDate": "2021-05-03T13:04:27.760Z",
      "content": "<p>Let's re-activate cv/lb thread.</p>\n<p>CV:1.17 -&gt; LB:1.49<br>\nCV:1.10 -&gt; LB:1.42<br>\nSeems stable for me, but a little big discrepancy.</p>\n<p>Do you have any updates, <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> ??<br>\nYou should;)</p>",
      "rawMarkdown": "Let's re-activate cv/lb thread.\n\nCV:1.17 -> LB:1.49\nCV:1.10 -> LB:1.42\nSeems stable for me, but a little big discrepancy.\n\nDo you have any updates, @drhabib ??\nYou should;)",
      "votes": 4,
      "replies": [
        {
          "id": 1291934,
          "postDate": "2021-05-03T13:11:41.140Z",
          "content": "<p>=) We have single models below 1.0. With CV and LB close. Cant say anything more at this time =)  But you are doing impressive progress for 2 submission, good luck to you  =) </p>",
          "rawMarkdown": "=) We have single models below 1.0. With CV and LB close. Cant say anything more at this time =)  But you are doing impressive progress for 2 submission, good luck to you  =) ",
          "votes": 6
        },
        {
          "id": 1291936,
          "postDate": "2021-05-03T13:12:34.870Z",
          "content": "<p>my current best gap is 0.15, but it depends on your fold.</p>",
          "rawMarkdown": "my current best gap is 0.15, but it depends on your fold.",
          "votes": 4
        },
        {
          "id": 1291947,
          "postDate": "2021-05-03T13:24:08.103Z",
          "content": "<p>is these raw results without k-beam search or inchi post-processing?</p>",
          "rawMarkdown": "is these raw results without k-beam search or inchi post-processing?",
          "votes": 2
        },
        {
          "id": 1291952,
          "postDate": "2021-05-03T13:28:40.683Z",
          "content": "<p>We don't use K - beam.. and we use <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> normalize script… </p>",
          "rawMarkdown": "We don't use K - beam.. and we use @nofreewill normalize script... ",
          "votes": 4
        },
        {
          "id": 1291978,
          "postDate": "2021-05-03T13:51:25.317Z",
          "content": "<p>Glad to see so many to use my script. Thanks for mentioning me! : )</p>",
          "rawMarkdown": "Glad to see so many to use my script. Thanks for mentioning me! : )",
          "votes": 8
        },
        {
          "id": 1291981,
          "postDate": "2021-05-03T13:55:13.430Z",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> normalize script beats the best network</p>",
          "rawMarkdown": "@nofreewill normalize script beats the best network",
          "votes": 4
        },
        {
          "id": 1291992,
          "postDate": "2021-05-03T14:05:31.867Z",
          "content": "<p><a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> <a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> Wow, single models below 1.0 are impressive! <br>\nAnd also thanks for letting me know you don't use k-beam, it was my next priority but now I decide to focus on something different;)<br>\nI'm trying hard to catch up you guys, good luck!</p>",
          "rawMarkdown": "@drhabib @tugstugi Wow, single models below 1.0 are impressive! \nAnd also thanks for letting me know you don't use k-beam, it was my next priority but now I decide to focus on something different;)\nI'm trying hard to catch up you guys, good luck!",
          "votes": 4
        },
        {
          "id": 1298890,
          "postDate": "2021-05-09T09:48:23.293Z",
          "content": "<p><a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a> nice score! Is that with a single model, without beam search?</p>",
          "rawMarkdown": "@bamps53 nice score! Is that with a single model, without beam search?"
        },
        {
          "id": 1298906,
          "postDate": "2021-05-09T10:04:20.257Z",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> Thanks, you should have better model than this since you haven't submitted for this 10 days:)<br>\nYes, I am too lazy to implement beam search, the score is of single model, without beam search, but with your rdkit normalization!</p>",
          "rawMarkdown": "@nofreewill Thanks, you should have better model than this since you haven't submitted for this 10 days:)\nYes, I am too lazy to implement beam search, the score is of single model, without beam search, but with your rdkit normalization!",
          "votes": 1
        },
        {
          "id": 1298914,
          "postDate": "2021-05-09T10:11:13.880Z",
          "content": "<p>I may have. But I also worked on it for much longer time than you. You made that in about only a few weeks. Hats off!</p>\n<p>I only implemented the normalization. The idea is from <a href=\"https://www.kaggle.com/stassl\" target=\"_blank\">@stassl</a> and <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> <br>\nJust wanted to emphasize it :)</p>",
          "rawMarkdown": "I may have. But I also worked on it for much longer time than you. You made that in about only a few weeks. Hats off!\n\nI only implemented the normalization. The idea is from @stassl and @yasufuminakama \nJust wanted to emphasize it :)",
          "votes": 1
        },
        {
          "id": 1298926,
          "postDate": "2021-05-09T10:27:57.163Z",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> Actually I've worked on this competition more than a month but I just haven't submitted until my cv score got good enough:)<br>\nBy the way you and <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> both say not using beam search, does it mean it doesn't help or it works but you don't use it due to longer inference time?</p>",
          "rawMarkdown": "@nofreewill Actually I've worked on this competition more than a month but I just haven't submitted until my cv score got good enough:)\nBy the way you and @drhabib both say not using beam search, does it mean it doesn't help or it works but you don't use it due to longer inference time?",
          "votes": 2
        },
        {
          "id": 1298933,
          "postDate": "2021-05-09T10:38:31.020Z",
          "content": "<p><code>I've worked on this competition more than a month</code><br>\nNow that's a bit comforting information. :D</p>\n<p>Beam search definitely helps my CV.<br>\nBut normalization had no effect on my CV (didn't change any prediction), so I don't know how much beam search helps together with normalization. I have a bad feeling that the effect of normalization will degrade a lot. I hope I'm wrong about it.</p>\n<p>But yeah, I would say that not using beam search is a bad idea. Especially because there is a lot of time to implement it while model is being trained.</p>",
          "rawMarkdown": "`I've worked on this competition more than a month`\nNow that's a bit comforting information. :D\n\nBeam search definitely helps my CV.\nBut normalization had no effect on my CV (didn't change any prediction), so I don't know how much beam search helps together with normalization. I have a bad feeling that the effect of normalization will degrade a lot. I hope I'm wrong about it.\n\nBut yeah, I would say that not using beam search is a bad idea. Especially because there is a lot of time to implement it while model is being trained.",
          "replies": [
            {
              "id": 1298936,
              "postDate": "2021-05-09T10:41:03.093Z",
              "content": "<blockquote>\n  <p>Beam search definitely helps my CV.</p>\n</blockquote>\n<p>How much does the beam search on transformer help? Haven't implemented yet :)</p>",
              "rawMarkdown": "> Beam search definitely helps my CV.\n\nHow much does the beam search on transformer help? Haven't implemented yet :)\n"
            }
          ]
        },
        {
          "id": 1298946,
          "postDate": "2021-05-09T10:55:28.030Z",
          "content": "<p>The best I've got with only 2 beams is 3.7% LD reduction.</p>",
          "rawMarkdown": "The best I've got with only 2 beams is 3.7% LD reduction.",
          "votes": 1
        },
        {
          "id": 1298987,
          "postDate": "2021-05-09T11:45:34.557Z",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> Really, I always get 0.07 improvement by rdkit normalization. Maybe my original predictions are too naive🤔</p>\n<blockquote>\n  <p>Especially because there is a lot of time to implement it while model is being trained.</p>\n</blockquote>\n<p>Definitely true;)</p>",
          "rawMarkdown": "@nofreewill Really, I always get 0.07 improvement by rdkit normalization. Maybe my original predictions are too naive🤔\n> Especially because there is a lot of time to implement it while model is being trained.\n\nDefinitely true;)",
          "votes": 2
        },
        {
          "id": 1298994,
          "postDate": "2021-05-09T11:49:30.343Z",
          "content": "<p>I see diminishing effects as my predictions get better. But could be just coincidence.</p>",
          "rawMarkdown": "I see diminishing effects as my predictions get better. But could be just coincidence.",
          "votes": 1
        },
        {
          "id": 1299014,
          "postDate": "2021-05-09T12:15:12.680Z",
          "content": "<p>I too see diminishing effects as my predictions get better. </p>\n<p>For my models, beam search + normalization yields better CV / leaderboard than just beam search or normalization by itself. </p>",
          "rawMarkdown": "I too see diminishing effects as my predictions get better. \n\nFor my models, beam search + normalization yields better CV / leaderboard than just beam search or normalization by itself. ",
          "votes": 2
        },
        {
          "id": 1299064,
          "postDate": "2021-05-09T12:50:17.763Z",
          "content": "<p>Then <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a> may secretly be doing something different :D</p>",
          "rawMarkdown": "Then @bamps53 may secretly be doing something different :D",
          "votes": 2
        }
      ]
    },
    {
      "id": 1257282,
      "postDate": "2021-03-30T17:19:03.957Z",
      "content": "<p>efficientnet_b0 as encoder, 5 epochs, CV: 7.21, LB: 7.76 using around 50% of the train data and no augmentations. For now wondering what's going to make difference, because usually model really slows down learning after that point…</p>",
      "rawMarkdown": "efficientnet_b0 as encoder, 5 epochs, CV: 7.21, LB: 7.76 using around 50% of the train data and no augmentations. For now wondering what's going to make difference, because usually model really slows down learning after that point...",
      "votes": 3,
      "replies": [
        {
          "id": 1258330,
          "postDate": "2021-03-31T14:37:01.567Z",
          "content": "<p>Why no augmentation? Several participants have reported better LB scores using random Rotation (90-deg) augmentation.</p>",
          "rawMarkdown": "Why no augmentation? Several participants have reported better LB scores using random Rotation (90-deg) augmentation."
        },
        {
          "id": 1258605,
          "postDate": "2021-03-31T18:44:17.747Z",
          "content": "<p>I think it depends on test data you using, (the 90 degree one)</p>",
          "rawMarkdown": "I think it depends on test data you using, (the 90 degree one)"
        }
      ]
    },
    {
      "id": 1253878,
      "postDate": "2021-03-27T06:01:28.540Z",
      "content": "<p>Anyone using <strong>Beam Search</strong> for prediction? Does it make a huge difference in your predictions instead of simple argmax at each timestep?<br>\nI wrote a beam search function but the problem is that it is not batch-wise and takes a lot of time. I also did not find an efficient batch-wise beam search online.</p>",
      "rawMarkdown": "Anyone using **Beam Search** for prediction? Does it make a huge difference in your predictions instead of simple argmax at each timestep?\nI wrote a beam search function but the problem is that it is not batch-wise and takes a lot of time. I also did not find an efficient batch-wise beam search online.",
      "votes": 3,
      "replies": [
        {
          "id": 1253961,
          "postDate": "2021-03-27T07:55:32.240Z",
          "content": "<p>You can test what difference it makes on a small holdout set.<br>\nI also did not find one online and it wasn't easy task to make it efficient.<br>\nFor me it does make a difference, though:<br>\nCV went ~4.0 -&gt; 3.15<br>\nLB went 7.57 -&gt; 6.59</p>",
          "rawMarkdown": "You can test what difference it makes on a small holdout set.\nI also did not find one online and it wasn't easy task to make it efficient.\nFor me it does make a difference, though:\nCV went ~4.0 -> 3.15\nLB went 7.57 -> 6.59",
          "votes": 3
        },
        {
          "id": 1254013,
          "postDate": "2021-03-27T09:00:27.807Z",
          "content": "<p>Thanks for mentioning that. So I think I'm gonna have a hard time implementing it 😅</p>",
          "rawMarkdown": "Thanks for mentioning that. So I think I'm gonna have a hard time implementing it 😅"
        },
        {
          "id": 1254037,
          "postDate": "2021-03-27T09:36:19.337Z",
          "content": "<p>It's not easy if you implement it the first time, but was absolutely enjoyable.</p>",
          "rawMarkdown": "It's not easy if you implement it the first time, but was absolutely enjoyable."
        },
        {
          "id": 1254077,
          "postDate": "2021-03-27T10:22:20.393Z",
          "content": "<p>You can't parallelize beam-search. </p>\n<p>You may skip it for (quick) early experiments. </p>",
          "rawMarkdown": "You can't parallelize beam-search. \n\nYou may skip it for (quick) early experiments. "
        },
        {
          "id": 1254079,
          "postDate": "2021-03-27T10:27:31.430Z",
          "content": "<p>It may depend on the architecture, because I did parallelize my beam search.</p>",
          "rawMarkdown": "It may depend on the architecture, because I did parallelize my beam search."
        },
        {
          "id": 1254085,
          "postDate": "2021-03-27T10:37:01.517Z",
          "content": "<p>I don't know much about RNNs, but I'm skeptical that those cannot be beam search parallelized. You just need to cache all the hidden states, but again, I have zero practical experience with RNNs.</p>",
          "rawMarkdown": "I don't know much about RNNs, but I'm skeptical that those cannot be beam search parallelized. You just need to cache all the hidden states, but again, I have zero practical experience with RNNs."
        },
        {
          "id": 1254175,
          "postDate": "2021-03-27T12:00:16.257Z",
          "content": "<p>You need to know top_k sequences with highest overall probability.  I don't see how you can compute this with single forward pass. </p>",
          "rawMarkdown": "You need to know top_k sequences with highest overall probability.  I don't see how you can compute this with single forward pass. "
        },
        {
          "id": 1254194,
          "postDate": "2021-03-27T12:17:56.143Z",
          "content": "<p>You have 2 beams: beam1 with probability 0.9 and beam2 with probability 0.85. You <strong>forward pass</strong> the two and get 2×(vocabulary size) -&gt; flatten it and get topk. ([[vocab_size],[vocab_size]]-&gt;[vocab_size<strong>;</strong>vocab_size])</p>\n<ul>\n<li>If you now have […<strong>0.87</strong>,…0.81,… <strong>;</strong> …<strong>0.83</strong>,…0.82,…] then you select the 0.87 from beam1 and 0.83 from beam2, that is straightforward.</li>\n<li>If you have […<strong>0.87</strong>,…<strong>0.84</strong>,… <strong>;</strong> …0.83,…0.82,…] however, then you select 0.87 and 0.84 from beam1 and copy the model from beam1 to beam2.</li>\n</ul>\n<p>These <strong>forward passes</strong> are the target of parallelization.</p>\n<hr>\n<p>With batch size of 2 and beam width of 2:<br>\n<strong>Forward pass</strong>:<br>\n[[vocab_size],<br>\n[vocab_size],<br>\n[vocab_size],<br>\n[vocab_size],]<br>\n<strong>Flattened for topk</strong>:<br>\n[[vocab_size, vocab_size],<br>\n[vocab_size, vocab_size],]</p>\n<hr>\n<p>So the main idea is that you handle the two beams separately while you forward pass, but you consider them together for selecting topk, and take the corresponding models in accordance to topk.</p>",
          "rawMarkdown": "You have 2 beams: beam1 with probability 0.9 and beam2 with probability 0.85. You **forward pass** the two and get 2×(vocabulary size) -> flatten it and get topk. ([[vocab_size],[vocab_size]]->[vocab_size**;**vocab_size])\n - If you now have [...**0.87**,...0.81,... **;** ...**0.83**,...0.82,...] then you select the 0.87 from beam1 and 0.83 from beam2, that is straightforward.\n - If you have [...**0.87**,...**0.84**,... **;** ...0.83,...0.82,...] however, then you select 0.87 and 0.84 from beam1 and copy the model from beam1 to beam2.\n\nThese **forward passes** are the target of parallelization.\n\n---\n\nWith batch size of 2 and beam width of 2:\n**Forward pass**:\n[[vocab_size],\n[vocab_size],\n[vocab_size],\n[vocab_size],]\n**Flattened for topk**:\n[[vocab_size, vocab_size],\n[vocab_size, vocab_size],]\n\n---\n\nSo the main idea is that you handle the two beams separately while you forward pass, but you consider them together for selecting topk, and take the corresponding models in accordance to topk."
        }
      ]
    },
    {
      "id": 1236612,
      "postDate": "2021-03-13T09:56:39.843Z",
      "content": "<p>Anyone successful to use transformer as a decoder?</p>",
      "rawMarkdown": "Anyone successful to use transformer as a decoder?",
      "votes": 3,
      "replies": [
        {
          "id": 1252085,
          "postDate": "2021-03-25T12:17:44.460Z",
          "content": "<p>I've tried it. It produces true tokens but the ordering is way wrong. I'm working on the positional embeddings and post-processing to make it work</p>",
          "rawMarkdown": "I've tried it. It produces true tokens but the ordering is way wrong. I'm working on the positional embeddings and post-processing to make it work",
          "votes": 3
        },
        {
          "id": 1278133,
          "postDate": "2021-04-19T15:30:11.857Z",
          "content": "<p>my transformer also didn’t achieve a satisfying results. is really the problem of a wrong positional emdedding in your case？</p>",
          "rawMarkdown": "my transformer also didn’t achieve a satisfying results. is really the problem of a wrong positional emdedding in your case？",
          "votes": 1
        }
      ]
    },
    {
      "id": 1234172,
      "postDate": "2021-03-11T03:07:09.093Z",
      "content": "<p>Surprising that baseline model can perform so well.</p>",
      "rawMarkdown": "Surprising that baseline model can perform so well.",
      "votes": 3,
      "replies": [
        {
          "id": 1243809,
          "postDate": "2021-03-18T14:00:38.220Z",
          "content": "<p>I am not a chemist but I don't think that a model with a Levenshtein value of ~20 would have any value for scientific use. At the longest InChi lables, that's almost 10% of the string that is wrong. I suspect a value &lt; 1 is needed.</p>",
          "rawMarkdown": "I am not a chemist but I don't think that a model with a Levenshtein value of ~20 would have any value for scientific use. At the longest InChi lables, that's almost 10% of the string that is wrong. I suspect a value < 1 is needed.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1232308,
      "postDate": "2021-03-09T16:35:14.690Z",
      "content": "<p>Interesting! Are you using attention during the decoding? I'm curious how useful that would be.</p>",
      "rawMarkdown": "Interesting! Are you using attention during the decoding? I'm curious how useful that would be.",
      "votes": 3,
      "replies": [
        {
          "id": 1232334,
          "postDate": "2021-03-09T17:01:19.743Z",
          "content": "<p>Yes attention. And I have plotted some attention and it seems like it recognizes edges and letters… but also a lot of background empty space ..  </p>",
          "rawMarkdown": "Yes attention. And I have plotted some attention and it seems like it recognizes edges and letters... but also a lot of background empty space ..  ",
          "votes": 3
        }
      ]
    },
    {
      "id": 1230221,
      "postDate": "2021-03-07T22:36:08.857Z",
      "content": "<p>I'm waiting an end to end transformer model in this discussion :)</p>",
      "rawMarkdown": "I'm waiting an end to end transformer model in this discussion :)",
      "votes": 3
    },
    {
      "id": 1260240,
      "postDate": "2021-04-02T00:43:12.320Z",
      "content": "<p>use batchwise beam search?</p>\n<p><a href=\"https://stackoverflow.com/questions/64356953/batch-wise-beam-search-in-pytorch\" target=\"_blank\">https://stackoverflow.com/questions/64356953/batch-wise-beam-search-in-pytorch</a><br>\n<a href=\"https://zhuanlan.zhihu.com/p/167072494\" target=\"_blank\">https://zhuanlan.zhihu.com/p/167072494</a></p>\n<p><a href=\"https://medium.com/the-artificial-impostor/implementing-beam-search-part-1-4f53482daabe\" target=\"_blank\">https://medium.com/the-artificial-impostor/implementing-beam-search-part-1-4f53482daabe</a><br>\nHow to Do Beam Search Efficiently<br>\n\"The smarter way is to put these N nodes into a batch and feed it to the model. It’ll be much faster since we can parallelize the computation, especially when using a CUDA backend.\"</p>\n<hr>\n<p>if you train multiple models (e.g. improved models as you make more submissions or for ensemble), these models can provide a hypothesis to reduce the time for beam search (i.e. they provide some clue on the top-k decoding path)</p>",
      "rawMarkdown": "use batchwise beam search?\n\nhttps://stackoverflow.com/questions/64356953/batch-wise-beam-search-in-pytorch\nhttps://zhuanlan.zhihu.com/p/167072494\n\nhttps://medium.com/the-artificial-impostor/implementing-beam-search-part-1-4f53482daabe\nHow to Do Beam Search Efficiently\n\"The smarter way is to put these N nodes into a batch and feed it to the model. It’ll be much faster since we can parallelize the computation, especially when using a CUDA backend.\"\n\n---\n\nif you train multiple models (e.g. improved models as you make more submissions or for ensemble), these models can provide a hypothesis to reduce the time for beam search (i.e. they provide some clue on the top-k decoding path)",
      "votes": 4
    },
    {
      "id": 1237838,
      "postDate": "2021-03-14T13:22:17.290Z",
      "content": "<p>end to end model: CNN + RNN<br>\nattention: None (don't know how to implement yet)<br>\ntrain / valid: 80% / 20%<br>\nepoch: 18<br>\nCV: 10.8<br>\nLB: 11.6</p>",
      "rawMarkdown": "end to end model: CNN + RNN\nattention: None (don't know how to implement yet)\ntrain / valid: 80% / 20%\nepoch: 18\nCV: 10.8\nLB: 11.6",
      "votes": 4,
      "replies": [
        {
          "id": 1237844,
          "postDate": "2021-03-14T13:27:14.047Z",
          "content": "<p>Wow 18 epochs… How long does it take and 18th epoch is the best CV?</p>",
          "rawMarkdown": "Wow 18 epochs... How long does it take and 18th epoch is the best CV?",
          "votes": 4
        },
        {
          "id": 1237915,
          "postDate": "2021-03-14T13:55:48.640Z",
          "content": "<p>It takes about 15 hours with TPU and continue training for now. My model needs to be optimized. It's very slow to train.</p>",
          "rawMarkdown": "It takes about 15 hours with TPU and continue training for now. My model needs to be optimized. It's very slow to train.",
          "votes": 6
        },
        {
          "id": 1237954,
          "postDate": "2021-03-14T14:23:23.187Z",
          "content": "<p>Thanks for information, TPU is very fast…</p>",
          "rawMarkdown": "Thanks for information, TPU is very fast...",
          "votes": 4
        },
        {
          "id": 1239746,
          "postDate": "2021-03-16T01:17:37.967Z",
          "content": "<p>any hint or paper on how you use the RNN on top of the CNN?</p>",
          "rawMarkdown": "any hint or paper on how you use the RNN on top of the CNN?",
          "votes": 1
        },
        {
          "id": 1239750,
          "postDate": "2021-03-16T01:28:47.563Z",
          "content": "<p><a href=\"https://arxiv.org/pdf/1502.03044.pdf\" target=\"_blank\">https://arxiv.org/pdf/1502.03044.pdf</a></p>",
          "rawMarkdown": "https://arxiv.org/pdf/1502.03044.pdf",
          "votes": 2
        },
        {
          "id": 1239754,
          "postDate": "2021-03-16T01:34:24.563Z",
          "content": "<p><a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> </p>\n<p>Thank you very much! the paper has 6700 citations, i have a lot to read</p>",
          "rawMarkdown": "@drhabib \n\nThank you very much! the paper has 6700 citations, i have a lot to read",
          "votes": 2
        },
        {
          "id": 1240662,
          "postDate": "2021-03-16T15:05:17.963Z",
          "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> I was surprised that you can achieve 24.1 with only one epoch. Now I think I know the reason. <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> I use some different approach. I don't predict the character one by one. so my model need more epoch to train.</p>",
          "rawMarkdown": "@yasufuminakama I was surprised that you can achieve 24.1 with only one epoch. Now I think I know the reason. @hengck23 I use some different approach. I don't predict the character one by one. so my model need more epoch to train.",
          "votes": 3
        },
        {
          "id": 1243805,
          "postDate": "2021-03-18T13:57:44.780Z",
          "content": "<p><a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a> </p>\n<blockquote>\n  <p>It takes about 15 hours with TPU and continue training for now. My model needs to be optimized. It's very slow to train.</p>\n</blockquote>\n<p>Are you training on TPU available via Kaggle notebooks? Or somewhere else (where?)?</p>",
          "rawMarkdown": "@wuliaokaola \n\n> It takes about 15 hours with TPU and continue training for now. My model needs to be optimized. It's very slow to train.\n\nAre you training on TPU available via Kaggle notebooks? Or somewhere else (where?)?"
        },
        {
          "id": 1243818,
          "postDate": "2021-03-18T14:10:31.873Z",
          "content": "<p><a href=\"https://www.kaggle.com/marketneutral\" target=\"_blank\">@marketneutral</a> You can train 3 epochs, and then load it into another notebook to train next 3 epochs.</p>",
          "rawMarkdown": "@marketneutral You can train 3 epochs, and then load it into another notebook to train next 3 epochs.",
          "votes": 4
        },
        {
          "id": 1243884,
          "postDate": "2021-03-18T14:56:25.250Z",
          "content": "<p>So you are training in Kaggle TPU for 3 epochs, saving the model, and then restarting training in a new TPU session?</p>",
          "rawMarkdown": "So you are training in Kaggle TPU for 3 epochs, saving the model, and then restarting training in a new TPU session?"
        },
        {
          "id": 1287894,
          "postDate": "2021-04-29T13:58:16.633Z",
          "content": "<p><a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a> would you like to share how to use RNN without outputting one by one?</p>",
          "rawMarkdown": "@wuliaokaola would you like to share how to use RNN without outputting one by one?"
        }
      ]
    },
    {
      "id": 1229749,
      "postDate": "2021-03-07T15:27:07.327Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> thanks for sharing your baseline results and welcome to the competition. Are you training a Image Captioning with resnet as backbone ? I am trying to understand how you have trained resnet <br>\nThanks</p>",
      "rawMarkdown": "Hi @drhabib thanks for sharing your baseline results and welcome to the competition. Are you training a Image Captioning with resnet as backbone ? I am trying to understand how you have trained resnet \nThanks",
      "votes": 4,
      "replies": [
        {
          "id": 1229754,
          "postDate": "2021-03-07T15:29:57.097Z",
          "content": "<p>Same question…</p>",
          "rawMarkdown": "Same question...",
          "votes": 1
        },
        {
          "id": 1229762,
          "postDate": "2021-03-07T15:32:42.140Z",
          "content": "<p>yes its basic captioning model. To be honest I did nothing special so far.. just a simple baseline =). Maybe one thing to mention that my tokenization take cares of chemical elements e.g <code>Br</code> is  one <code>token</code> and not to <code>B</code> and <code>r</code>. I am not sure how important this is for now…. perhaps not so important because sequence model these days can learn this relationship pretty easy =) </p>",
          "rawMarkdown": "yes its basic captioning model. To be honest I did nothing special so far.. just a simple baseline =). Maybe one thing to mention that my tokenization take cares of chemical elements e.g `Br` is  one `token` and not to `B` and `r`. I am not sure how important this is for now.... perhaps not so important because sequence model these days can learn this relationship pretty easy =) ",
          "votes": 11
        },
        {
          "id": 1229784,
          "postDate": "2021-03-07T15:44:41.733Z",
          "content": "<p>Thanks for answer. Two more questions if you don't mind, first, did you use PyTorch? second, is there any good tutorial with PyTorch that you recommends?</p>",
          "rawMarkdown": "Thanks for answer. Two more questions if you don't mind, first, did you use PyTorch? second, is there any good tutorial with PyTorch that you recommends?",
          "votes": 3
        },
        {
          "id": 1229812,
          "postDate": "2021-03-07T15:57:58.360Z",
          "content": "<p>Yes I use PyTorch (pretty sure this can be done also in TF) and training done in fp16. </p>\n<p><a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> posted several good resources:<br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223419\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/223419</a><br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223218\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/223218</a><br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223223\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/223223</a></p>\n<p>I would start from them =)</p>",
          "rawMarkdown": "Yes I use PyTorch (pretty sure this can be done also in TF) and training done in fp16. \n\n@usharengaraju posted several good resources:\nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/223419\nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/223218\nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/223223\n\nI would start from them =)",
          "votes": 7
        },
        {
          "id": 1230351,
          "postDate": "2021-03-08T03:47:11.637Z",
          "content": "<p>Thank you for sharing <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a>, at least I still have some hope with my model later after training<br>\nBut in your conclusion, is image captioning is the best approach to this competition?</p>",
          "rawMarkdown": "Thank you for sharing @drhabib, at least I still have some hope with my model later after training\nBut in your conclusion, is image captioning is the best approach to this competition?",
          "votes": 4
        },
        {
          "id": 1230367,
          "postDate": "2021-03-08T04:34:22.113Z",
          "content": "<blockquote>\n  <p>But in your conclusion, is image captioning is the best approach to this competition?</p>\n</blockquote>\n<p>I think we'll find out in a few weeks :)</p>",
          "rawMarkdown": "> But in your conclusion, is image captioning is the best approach to this competition?\n\nI think we'll find out in a few weeks :)",
          "votes": 3
        },
        {
          "id": 1231080,
          "postDate": "2021-03-08T16:50:15.027Z",
          "content": "<p><a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> Congratulations on being first in LB and thank you for the mention . Glad you found my posts helpful</p>",
          "rawMarkdown": "@drhabib Congratulations on being first in LB and thank you for the mention . Glad you found my posts helpful",
          "votes": 4
        }
      ]
    },
    {
      "id": 1293314,
      "postDate": "2021-05-04T18:30:25.663Z",
      "content": "<p>Hey everyone! What's your lowest score with the GRU decoder? How do you think, is it possible to get the gold with that?</p>",
      "rawMarkdown": "Hey everyone! What's your lowest score with the GRU decoder? How do you think, is it possible to get the gold with that?",
      "votes": 1,
      "replies": [
        {
          "id": 1293405,
          "postDate": "2021-05-04T20:04:25.430Z",
          "content": "<p>Heng's theorem : \"Whatever LSTM(GRU) can do , Transformer can do it even better\" :)</p>",
          "rawMarkdown": "Heng's theorem : \"Whatever LSTM(GRU) can do , Transformer can do it even better\" :)",
          "votes": 5
        },
        {
          "id": 1293815,
          "postDate": "2021-05-05T07:20:57.907Z",
          "content": "<p>I still have not switched to a Transformer decoder and currently use LSTM. My impression is that it's definitely possible to get between 1.5 and 2.0 LB with LSTM, but going below 1.0, which will probably be a requirement for Gold, might not be feasible.</p>",
          "rawMarkdown": "I still have not switched to a Transformer decoder and currently use LSTM. My impression is that it's definitely possible to get between 1.5 and 2.0 LB with LSTM, but going below 1.0, which will probably be a requirement for Gold, might not be feasible.",
          "votes": 3
        },
        {
          "id": 1294786,
          "postDate": "2021-05-05T23:09:34.487Z",
          "content": "<p>I am switching back and forth between a Pytorch Transformer and TF LSTM and will submit the one that win at the end ^^</p>\n<p>For now, I have better CV with Pytorch while using  smaller encoder and smaller resolution. </p>\n<p>But TF run incredibly fast on TPU even with huge model, letting room to explore new ideas. <br>\nUnfortunately some operations are not supported by TPU during the training loop and wrapping numpy functions doesn't work either. I had very hard time diving into TF documentations and code and hack around a bit to make them work. </p>",
          "rawMarkdown": "I am switching back and forth between a Pytorch Transformer and TF LSTM and will submit the one that win at the end ^^\n\nFor now, I have better CV with Pytorch while using  smaller encoder and smaller resolution. \n\n But TF run incredibly fast on TPU even with huge model, letting room to explore new ideas. \nUnfortunately some operations are not supported by TPU during the training loop and wrapping numpy functions doesn't work either. I had very hard time diving into TF documentations and code and hack around a bit to make them work. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1242789,
      "postDate": "2021-03-17T19:54:39.973Z",
      "content": "<p>Apart from the scores; can you guys share what type of augmentations you have used ( or can be used)<br>\nI think augmentation will be a vital part since we are working with Chemical molecules.<br>\nThanks</p>",
      "rawMarkdown": "Apart from the scores; can you guys share what type of augmentations you have used ( or can be used)\nI think augmentation will be a vital part since we are working with Chemical molecules.\nThanks",
      "votes": 1,
      "replies": [
        {
          "id": 1243804,
          "postDate": "2021-03-18T13:57:39.663Z",
          "content": "<p>for the moment I am not using anything special… you can score 3-4 using just this <code>augs</code> below</p>\n<pre><code>    train_transforms  = A.Compose([\n            A.RandomRotate90(p=.5),\n            A.Normalize(mean=[0.485, 0.456, 0.406],\n                        std=[0.229, 0.224, 0.225],\n                        always_apply=True),\n            AtoTensor()\n        ])\n</code></pre>",
          "rawMarkdown": "for the moment I am not using anything special... you can score 3-4 using just this `augs` below\n```\n    train_transforms  = A.Compose([\n            A.RandomRotate90(p=.5),\n            A.Normalize(mean=[0.485, 0.456, 0.406],\n                        std=[0.229, 0.224, 0.225],\n                        always_apply=True),\n            AtoTensor()\n        ])\n```",
          "votes": 13
        },
        {
          "id": 1243930,
          "postDate": "2021-03-18T15:34:55.080Z",
          "content": "<p>Thanks For replying;<br>\nFor normalizing should we use the mean, std for the given images, as these are very different from the imagenet data</p>",
          "rawMarkdown": "Thanks For replying;\nFor normalizing should we use the mean, std for the given images, as these are very different from the imagenet data\n",
          "votes": 1
        },
        {
          "id": 1243941,
          "postDate": "2021-03-18T15:43:55.840Z",
          "content": "<p>Based on my past experience (I might be wrong) its best to use <code>imagenet</code> if you are using pertained network.  If you are using  network from scratch then you have to use dataset statistics. </p>",
          "rawMarkdown": "Based on my past experience (I might be wrong) its best to use `imagenet` if you are using pertained network.  If you are using  network from scratch then you have to use dataset statistics. ",
          "votes": 2
        },
        {
          "id": 1243942,
          "postDate": "2021-03-18T15:48:09.313Z",
          "content": "<p>After a lot of experimentation I found 3 total types of augmentations can be helpful for this comp…</p>\n<ol>\n<li>Adding noise -- test images contain noise so maybe this has a good effect</li>\n<li>Improving image quality -- the letters are not clearly visible and also the lines used to create bonds are broken and  these types of augmentations help with this.  </li>\n<li>Rotating augmentations </li>\n</ol>",
          "rawMarkdown": "After a lot of experimentation I found 3 total types of augmentations can be helpful for this comp...\n1. Adding noise -- test images contain noise so maybe this has a good effect\n2. Improving image quality -- the letters are not clearly visible and also the lines used to create bonds are broken and  these types of augmentations help with this.  \n3. Rotating augmentations ",
          "votes": 9
        },
        {
          "id": 1243943,
          "postDate": "2021-03-18T15:50:36.527Z",
          "content": "<p><a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> <code>you can score 3-4</code> by this you mean scoring on LB right??</p>",
          "rawMarkdown": "@drhabib `you can score 3-4 ` by this you mean scoring on LB right??",
          "votes": 5
        },
        {
          "id": 1243946,
          "postDate": "2021-03-18T15:52:23.680Z",
          "content": "<p>yes <code>LB</code> =)</p>",
          "rawMarkdown": "yes `LB` =)",
          "votes": 4
        }
      ]
    },
    {
      "id": 1236024,
      "postDate": "2021-03-12T17:19:53.393Z",
      "content": "<p>Looks like you are making good progress on the LB. 😃</p>\n<p>Out of interest, are you fine-tuning the ResNet, or have you frozen the ResNet layers? My hunch would be that the challenge dataset images are very different to pretraining images, and hence models might benefit significantly from the fine-tuning.</p>",
      "rawMarkdown": "Looks like you are making good progress on the LB. 😃\n\nOut of interest, are you fine-tuning the ResNet, or have you frozen the ResNet layers? My hunch would be that the challenge dataset images are very different to pretraining images, and hence models might benefit significantly from the fine-tuning.",
      "votes": 1,
      "replies": [
        {
          "id": 1236117,
          "postDate": "2021-03-12T19:26:19.977Z",
          "content": "<p>Your <code>hunch</code> is correct. Its good idea to also finetune your image <code>encoder</code> =)</p>",
          "rawMarkdown": "Your `hunch` is correct. Its good idea to also finetune your image `encoder` =)",
          "votes": 1
        },
        {
          "id": 1237076,
          "postDate": "2021-03-13T18:01:50.357Z",
          "content": "<p>Actually I gave wrong answer… I thought I was fine-tuning <code>encoder</code> as well .. but my teammate found a bug in my code where I was only training last layer of my <code>encoder</code>. I am rerunning now my training script. </p>\n<p>Coming back to your question. Our current score is still from<code>resnet34</code> but with trained last layer. I will update once I finish training full model. </p>",
          "rawMarkdown": "Actually I gave wrong answer... I thought I was fine-tuning `encoder` as well .. but my teammate found a bug in my code where I was only training last layer of my `encoder`. I am rerunning now my training script. \n\nComing back to your question. Our current score is still from` resnet34` but with trained last layer. I will update once I finish training full model. ",
          "votes": 5
        },
        {
          "id": 1237086,
          "postDate": "2021-03-13T18:25:04.227Z",
          "content": "<p>The last layer be the Conv layer before pooling? Or are you projecting with a linear layer?</p>\n<p>I am training end to end. And last night I froze my CNN encoder and swapped on a new decoder, it trains really fast. After 1 epoch I am getting 11 on my local validation set. It seems the decoder part is likely much more important. </p>\n<p>Edit: probably not. I retrained end-to-end with some tweaks, 1 epoch can get validation distance to 11. </p>",
          "rawMarkdown": "The last layer be the Conv layer before pooling? Or are you projecting with a linear layer?\n\nI am training end to end. And last night I froze my CNN encoder and swapped on a new decoder, it trains really fast. After 1 epoch I am getting 11 on my local validation set. It seems the decoder part is likely much more important. \n\nEdit: probably not. I retrained end-to-end with some tweaks, 1 epoch can get validation distance to 11. ",
          "votes": 2
        },
        {
          "id": 1237090,
          "postDate": "2021-03-13T18:31:27.420Z",
          "content": "<p>Thanks for the update. Interesting to hear that you've only been fine-tuning the last layer. Can't wait to see what scores you will get with full fine-tuning!</p>",
          "rawMarkdown": "Thanks for the update. Interesting to hear that you've only been fine-tuning the last layer. Can't wait to see what scores you will get with full fine-tuning!",
          "votes": 2
        },
        {
          "id": 1237108,
          "postDate": "2021-03-13T18:57:01.210Z",
          "content": "<p>Yes correct the last <code>conv_block</code> before <code>pooling</code>. It will be interesting to see the difference … I assume the change might be small since early layers of model trained on imagenett are already really good at recognizing shapes (e.g. <code>lines</code>, 'circles`) </p>\n<p><a href=\"https://arxiv.org/pdf/1311.2901.pdf\" target=\"_blank\">https://arxiv.org/pdf/1311.2901.pdf</a><br>\n<a href=\"https://distill.pub/2020/circuits/zoom-in/\" target=\"_blank\">https://distill.pub/2020/circuits/zoom-in/</a></p>\n<p>experiments will show =)</p>",
          "rawMarkdown": "Yes correct the last `conv_block` before `pooling`. It will be interesting to see the difference ... I assume the change might be small since early layers of model trained on imagenett are already really good at recognizing shapes (e.g. `lines`, 'circles`) \n\nhttps://arxiv.org/pdf/1311.2901.pdf\nhttps://distill.pub/2020/circuits/zoom-in/\n\nexperiments will show =)",
          "votes": 3
        }
      ]
    },
    {
      "id": 1230768,
      "postDate": "2021-03-08T12:51:12.797Z",
      "content": "<p>Hello guys; I am also trying to solve it as an image captioning model type task, but after getting the image features from a backbone network I like to train through a transformer network;<br>\nSo can someone share any links/resources where I can train a Transformer from scratch? <br>\n(will fine-tuning any other language models like BERT, DistillBERT be of any use/ since they  are trained and have tokens for the English language)</p>",
      "rawMarkdown": "Hello guys; I am also trying to solve it as an image captioning model type task, but after getting the image features from a backbone network I like to train through a transformer network;\nSo can someone share any links/resources where I can train a Transformer from scratch? \n(will fine-tuning any other language models like BERT, DistillBERT be of any use/ since they  are trained and have tokens for the English language)",
      "votes": 1,
      "replies": [
        {
          "id": 1230843,
          "postDate": "2021-03-08T14:03:13.217Z",
          "content": "<p>I think there is people who share the transformer model for this on discussion forum, I think it was tensor girl</p>",
          "rawMarkdown": "I think there is people who share the transformer model for this on discussion forum, I think it was tensor girl",
          "votes": 2
        },
        {
          "id": 1231079,
          "postDate": "2021-03-08T16:49:14.587Z",
          "content": "<p><a href=\"https://www.kaggle.com/houoinkyoma\" target=\"_blank\">@houoinkyoma</a> Thank you so much for the mention Uruha</p>",
          "rawMarkdown": "@houoinkyoma Thank you so much for the mention Uruha",
          "votes": 2
        },
        {
          "id": 1232259,
          "postDate": "2021-03-09T15:41:22.117Z",
          "content": "<p>This is great if you are using PyTorch: <a href=\"https://github.com/bentrevett/pytorch-seq2seq/blob/master/6%20-%20Attention%20is%20All%20You%20Need.ipynb\" target=\"_blank\">https://github.com/bentrevett/pytorch-seq2seq/blob/master/6%20-%20Attention%20is%20All%20You%20Need.ipynb</a></p>",
          "rawMarkdown": "This is great if you are using PyTorch: https://github.com/bentrevett/pytorch-seq2seq/blob/master/6%20-%20Attention%20is%20All%20You%20Need.ipynb",
          "votes": 2
        }
      ]
    },
    {
      "id": 1292463,
      "postDate": "2021-05-04T01:47:40.530Z",
      "content": "<p>i am still waiting for reports of non end-to-end solution.<br>\nThe LG SMILE (Dacon.ai from) top solution uses direct graph parsing from the image. (lstm-attention was ranked third)<br>\nI wonder has anyone tried methods like that?</p>",
      "rawMarkdown": "i am still waiting for reports of non end-to-end solution.\nThe LG SMILE (Dacon.ai from) top solution uses direct graph parsing from the image. (lstm-attention was ranked third)\nI wonder has anyone tried methods like that?",
      "votes": 2,
      "replies": [
        {
          "id": 1292482,
          "postDate": "2021-05-04T02:20:19.730Z",
          "content": "<p>Have you tried??</p>",
          "rawMarkdown": "Have you tried??",
          "votes": 1
        },
        {
          "id": 1292594,
          "postDate": "2021-05-04T06:11:51.577Z",
          "content": "<p>trying now</p>",
          "rawMarkdown": "trying now",
          "votes": 3
        },
        {
          "id": 1292847,
          "postDate": "2021-05-04T10:42:59.540Z",
          "content": "<p>I tried but for the moment I'm not able to get a Levenshtein distance under 0.4 on formula while the transformer encoder - transformer decoder goes under 0.1</p>",
          "rawMarkdown": "I tried but for the moment I'm not able to get a Levenshtein distance under 0.4 on formula while the transformer encoder - transformer decoder goes under 0.1",
          "votes": 6
        }
      ]
    },
    {
      "id": 1270959,
      "postDate": "2021-04-12T07:26:37.330Z",
      "content": "<p>anyone manages to make cnn-attention-lstm obtain LB score below 2.5 (without beam search)?</p>",
      "rawMarkdown": "anyone manages to make cnn-attention-lstm obtain LB score below 2.5 (without beam search)?",
      "votes": 2,
      "replies": [
        {
          "id": 1270972,
          "postDate": "2021-04-12T07:38:00.123Z",
          "content": "<p>Do you get a LB score below 3.0 with cnn-attention-lstm (I just started the competition) ?</p>",
          "rawMarkdown": "Do you get a LB score below 3.0 with cnn-attention-lstm (I just started the competition) ?",
          "votes": 1
        },
        {
          "id": 1271017,
          "postDate": "2021-04-12T08:06:26.490Z",
          "content": "<p>I got CV around 2.5 but didn't submit. I am moving now to transformer :) Increasing the hidden size dimension increases atleast CV.</p>",
          "rawMarkdown": "I got CV around 2.5 but didn't submit. I am moving now to transformer :) Increasing the hidden size dimension increases atleast CV.",
          "votes": 2
        },
        {
          "id": 1277705,
          "postDate": "2021-04-19T05:46:36.690Z",
          "content": "<p>effnet-b4-attention-lstm, input_size=380, cv=1.86,lb=2.5(without beam search).<br>\nThank you for your <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231190\" target=\"_blank\">share</a>,it help me a lot.</p>\n<p><a href=\"https://www.kaggle.com/yingpengchen/pl-bms-train\" target=\"_blank\">Training notebook</a><br>\n<a href=\"https://www.kaggle.com/yingpengchen/pl-bms-molecular-translation/edit/run/59116148\" target=\"_blank\">Inference notebook</a></p>",
          "rawMarkdown": "effnet-b4-attention-lstm, input_size=380, cv=1.86,lb=2.5(without beam search).\nThank you for your [share](https://www.kaggle.com/c/bms-molecular-translation/discussion/231190),it help me a lot.\n\n[Training notebook](https://www.kaggle.com/yingpengchen/pl-bms-train)\n[Inference notebook](https://www.kaggle.com/yingpengchen/pl-bms-molecular-translation/edit/run/59116148)",
          "votes": 4,
          "replies": [
            {
              "id": 1279613,
              "postDate": "2021-04-21T05:15:43.950Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 1277714,
          "postDate": "2021-04-19T06:04:58.390Z",
          "content": "<p>thanks! I am surprised that attention-lstm can get below 2.5</p>\n<p>i also plugin effcientnet (and resnet) into trasnformer encoder. the results is better than lstm but worse than vision transformer encoder (which i don't know why).</p>\n<p>It seems that image size is very important</p>",
          "rawMarkdown": "thanks! I am surprised that attention-lstm can get below 2.5\n\ni also plugin effcientnet (and resnet) into trasnformer encoder. the results is better than lstm but worse than vision transformer encoder (which i don't know why).\n\nIt seems that image size is very important",
          "votes": 1
        },
        {
          "id": 1279324,
          "postDate": "2021-04-20T19:34:16.470Z",
          "content": "<p><a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> may I ask, how big that \"big gap…\" is?</p>",
          "rawMarkdown": "@tugstugi may I ask, how big that \"big gap...\" is?"
        },
        {
          "id": 1279336,
          "postDate": "2021-04-20T19:38:35.820Z",
          "content": "<p>was* I guess</p>",
          "rawMarkdown": "was* I guess",
          "votes": 2
        },
        {
          "id": 1279580,
          "postDate": "2021-04-21T04:32:35.250Z",
          "content": "<p>B4 + input 380…</p>",
          "rawMarkdown": "B4 + input 380..."
        },
        {
          "id": 1279612,
          "postDate": "2021-04-21T05:13:20.560Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1260248,
      "postDate": "2021-04-02T00:54:22.753Z",
      "content": "<p>fyi:</p>\n<p>it is possible to predict the image orientation (rotate +/- 90 clockwise or 180) with 98%~99% using resnet34 on validation set</p>",
      "rawMarkdown": "fyi:\n\nit is possible to predict the image orientation (rotate +/- 90 clockwise or 180) with 98%~99% using resnet34 on validation set",
      "votes": 2,
      "replies": [
        {
          "id": 1260580,
          "postDate": "2021-04-02T08:36:10.970Z",
          "content": "<p>The 1-2% \"error\" could contain images that are actually rotated ones.</p>",
          "rawMarkdown": "The 1-2% \"error\" could contain images that are actually rotated ones.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1258104,
      "postDate": "2021-03-31T10:59:37.003Z",
      "content": "<p>Has anyone tried label smoothing? I know that in machine translation it matters.</p>",
      "rawMarkdown": "Has anyone tried label smoothing? I know that in machine translation it matters.",
      "votes": 2
    },
    {
      "id": 1253185,
      "postDate": "2021-03-26T12:58:53.030Z",
      "content": "<p>My architecture is different form public kernels.</p>\n<p>Single fold validation: 5.50<br>\nLB : 6.84</p>\n<p>I don't have descent hardware and patience to wait for 5 folds for now ^^</p>\n<p>Anyway I think there is still room for improvement in this fold. I stopped  training  just to check  my  submission</p>",
      "rawMarkdown": "My architecture is different form public kernels.\n\nSingle fold validation: 5.50\nLB : 6.84\n\nI don't have descent hardware and patience to wait for 5 folds for now ^^\n\nAnyway I think there is still room for improvement in this fold. I stopped  training  just to check  my  submission",
      "votes": 2
    },
    {
      "id": 1231908,
      "postDate": "2021-03-09T11:17:01.910Z",
      "content": "<p>My first CV-LB: 18.44-23.7</p>",
      "rawMarkdown": "My first CV-LB: 18.44-23.7",
      "votes": 2,
      "replies": [
        {
          "id": 1233297,
          "postDate": "2021-03-10T09:34:22.270Z",
          "content": "<p>2nd: 8.66-11.3</p>",
          "rawMarkdown": "2nd: 8.66-11.3",
          "votes": 1
        },
        {
          "id": 1233449,
          "postDate": "2021-03-10T12:15:05.777Z",
          "content": "<p>How many samples are you taking for training and validation?</p>",
          "rawMarkdown": "How many samples are you taking for training and validation?",
          "votes": 1
        },
        {
          "id": 1233459,
          "postDate": "2021-03-10T12:41:08.850Z",
          "content": "<p>I use all the images; ~50k of them are for validation.</p>\n<p>I like your profpic.</p>",
          "rawMarkdown": "I use all the images; ~50k of them are for validation.\n\nI like your profpic.",
          "votes": 4
        },
        {
          "id": 1233855,
          "postDate": "2021-03-10T18:20:30.507Z",
          "content": "<p>Thanks, u seem to be another Breaking Bad fan !!<br>\nCan you share how much time it is taking for the training and what hardware are you using?<br>\nThanks again!!</p>",
          "rawMarkdown": "Thanks, u seem to be another Breaking Bad fan !!\nCan you share how much time it is taking for the training and what hardware are you using?\nThanks again!!",
          "votes": 1
        },
        {
          "id": 1233865,
          "postDate": "2021-03-10T18:33:18.687Z",
          "content": "<p>That's the best series ever!:D</p>\n<p>1 epoch is ~2hours on a 1080Ti.</p>",
          "rawMarkdown": "That's the best series ever!:D\n\n1 epoch is ~2hours on a 1080Ti.",
          "votes": 2
        },
        {
          "id": 1236107,
          "postDate": "2021-03-12T19:11:31.067Z",
          "content": "<p>Hi, your 1 epoch is about ~2 million images? </p>",
          "rawMarkdown": "Hi, your 1 epoch is about ~2 million images? ",
          "votes": 1
        },
        {
          "id": 1236108,
          "postDate": "2021-03-12T19:14:33.237Z",
          "content": "<p>Yes, 1 epoch on all (~2million) the training images is ~2 hours on my 1080 Ti.</p>",
          "rawMarkdown": "Yes, 1 epoch on all (~2million) the training images is ~2 hours on my 1080 Ti.",
          "votes": 1
        },
        {
          "id": 1236114,
          "postDate": "2021-03-12T19:20:47.363Z",
          "content": "<p>Thanks!                       </p>",
          "rawMarkdown": "Thanks!                       ",
          "votes": 2
        },
        {
          "id": 1248043,
          "postDate": "2021-03-22T09:40:30.953Z",
          "content": "<p>I've changed my validation and now my CV-LB gap is huge (3.15-6.59). Has anyone experienced such a gap?</p>",
          "rawMarkdown": "I've changed my validation and now my CV-LB gap is huge (3.15-6.59). Has anyone experienced such a gap?"
        },
        {
          "id": 1248100,
          "postDate": "2021-03-22T10:47:13.777Z",
          "content": "<p>My current gap is 2.80 - 3.06.</p>",
          "rawMarkdown": "My current gap is 2.80 - 3.06.",
          "votes": 1
        },
        {
          "id": 1248121,
          "postDate": "2021-03-22T11:10:37.780Z",
          "content": "<p><a href=\"https://www.kaggle.com/anokas\" target=\"_blank\">@anokas</a> had made a 5.99 LB submission yesterday and now a 3.17 today. That made me think.. That's about the x2 gap that I have. Also looking at the table of <a href=\"https://www.kaggle.com/ryanzhang\" target=\"_blank\">@ryanzhang</a> , he also had twice the LB as CV when he did not rotate the test images. In the hope that I do something wrong there, too, I looked at my code and yes I did mess it up. Now I'm creating my next submission and I hope this time the gap disappears.</p>",
          "rawMarkdown": "@anokas had made a 5.99 LB submission yesterday and now a 3.17 today. That made me think.. That's about the x2 gap that I have. Also looking at the table of @ryanzhang , he also had twice the LB as CV when he did not rotate the test images. In the hope that I do something wrong there, too, I looked at my code and yes I did mess it up. Now I'm creating my next submission and I hope this time the gap disappears."
        },
        {
          "id": 1248162,
          "postDate": "2021-03-22T11:49:30.793Z",
          "content": "<p>It seems I did not mess up.. shoot. I know now nothing :D</p>",
          "rawMarkdown": "It seems I did not mess up.. shoot. I know now nothing :D"
        },
        {
          "id": 1252830,
          "postDate": "2021-03-26T05:10:43.993Z",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> 2hrs on 1080Ti. Is that using CNN encoder / RNN decoder? I cant get anywhere close to that for 2M images. </p>",
          "rawMarkdown": "@nofreewill 2hrs on 1080Ti. Is that using CNN encoder / RNN decoder? I cant get anywhere close to that for 2M images. "
        },
        {
          "id": 1253034,
          "postDate": "2021-03-26T10:08:51.227Z",
          "content": "<p><a href=\"https://www.kaggle.com/trushk\" target=\"_blank\">@trushk</a> y me either… I'm getting more like a little over 3 hrs per epoch on my titan rtx for 2M images… I've tried preprocessing (resizing images outside of train loop) but that doesn't put a dent in the train time… <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> any tips on reducing train time?  are you using fp16 yet?  Thanks…  </p>",
          "rawMarkdown": "@trushk y me either... I'm getting more like a little over 3 hrs per epoch on my titan rtx for 2M images... I've tried preprocessing (resizing images outside of train loop) but that doesn't put a dent in the train time... @nofreewill any tips on reducing train time?  are you using fp16 yet?  Thanks...  "
        },
        {
          "id": 1253145,
          "postDate": "2021-03-26T12:01:08.140Z",
          "content": "<p>I think it is likely due to the difference in decoders between people. Most of the training time is spent on the autoregressive step-by-step decoding if your setup is similar to public kernels.  I just started testing another decoder setup and was able to get 2hrs per 2m datapoints. </p>",
          "rawMarkdown": "I think it is likely due to the difference in decoders between people. Most of the training time is spent on the autoregressive step-by-step decoding if your setup is similar to public kernels.  I just started testing another decoder setup and was able to get 2hrs per 2m datapoints. ",
          "votes": 6
        },
        {
          "id": 1253152,
          "postDate": "2021-03-26T12:12:56.770Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/ryanzhang\" target=\"_blank\">@ryanzhang</a> <br>\nI'm using transformer to decode the features of the ResNet but the predictions are not good and there are a lot of repetitive tokens out of the decoder and also the ordering of them has problems. I've tried different positional embedding methods but none worked out.<br>\nI'm really sorry to ask but can you give me a gist that what I am probably missing (if you are using a transformer to decode)? It is completely okay if you do not want to answer this.</p>",
          "rawMarkdown": "Hey @ryanzhang \nI'm using transformer to decode the features of the ResNet but the predictions are not good and there are a lot of repetitive tokens out of the decoder and also the ordering of them has problems. I've tried different positional embedding methods but none worked out.\nI'm really sorry to ask but can you give me a gist that what I am probably missing (if you are using a transformer to decode)? It is completely okay if you do not want to answer this."
        },
        {
          "id": 1253182,
          "postDate": "2021-03-26T12:53:51.880Z",
          "content": "<p>I don't use the public kernel, that's all I'm willing to say. :)</p>",
          "rawMarkdown": "I don't use the public kernel, that's all I'm willing to say. :)",
          "votes": 3
        },
        {
          "id": 1253265,
          "postDate": "2021-03-26T14:40:51.920Z",
          "content": "<p>I see the opposite. I hooked on a transformer decoder to my frozen e0 encoder, and just tuning the decoer. It quickly outperformed my old setup. </p>\n<p>My issue now is the inference time is very slow, I probably have some problems with my code. </p>",
          "rawMarkdown": "I see the opposite. I hooked on a transformer decoder to my frozen e0 encoder, and just tuning the decoer. It quickly outperformed my old setup. \n\nMy issue now is the inference time is very slow, I probably have some problems with my code. ",
          "votes": 3
        },
        {
          "id": 1253273,
          "postDate": "2021-03-26T14:47:57.880Z",
          "content": "<p>I don't think its your code… Transformers inference is in general super slow and its very noticeable when you run prediction on millions of images </p>",
          "rawMarkdown": "I don't think its your code... Transformers inference is in general super slow and its very noticeable when you run prediction on millions of images ",
          "votes": 1
        },
        {
          "id": 1253275,
          "postDate": "2021-03-26T14:51:08.073Z",
          "content": "<p><a href=\"https://www.kaggle.com/moeinshariatnia\" target=\"_blank\">@moeinshariatnia</a> In order to avoid this issue you have to add more regularization. e.g by penalizing attention… or having  more dropouts, skip connections or this approach <a href=\"https://arxiv.org/pdf/2004.13342.pdf\" target=\"_blank\">https://arxiv.org/pdf/2004.13342.pdf</a></p>",
          "rawMarkdown": "@moeinshariatnia In order to avoid this issue you have to add more regularization. e.g by penalizing attention... or having  more dropouts, skip connections or this approach https://arxiv.org/pdf/2004.13342.pdf",
          "votes": 1
        },
        {
          "id": 1253282,
          "postDate": "2021-03-26T14:57:55.290Z",
          "content": "<p><a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> Thank you very much. I'll start experimenting right now</p>",
          "rawMarkdown": "@drhabib Thank you very much. I'll start experimenting right now",
          "votes": 1
        },
        {
          "id": 1253375,
          "postDate": "2021-03-26T16:51:24.513Z",
          "content": "<p>Thanks, I coded up logic to try to speed it up. Something like what huggingface did: <img src=\"https://huggingface.co/blog/assets/09_accelerated_inference/optimized_graph.png\" alt=\"this\">,  by the problem in my code, I mean this part. </p>",
          "rawMarkdown": "Thanks, I coded up logic to try to speed it up. Something like what huggingface did: ![this](https://huggingface.co/blog/assets/09_accelerated_inference/optimized_graph.png),  by the problem in my code, I mean this part. "
        },
        {
          "id": 1253599,
          "postDate": "2021-03-26T22:03:59.450Z",
          "content": "<p><a href=\"https://www.kaggle.com/ryanzhang\" target=\"_blank\">@ryanzhang</a>: What is the e0 encoder ? Is it EfficientNet B0 ?</p>",
          "rawMarkdown": "@ryanzhang: What is the e0 encoder ? Is it EfficientNet B0 ?"
        },
        {
          "id": 1253643,
          "postDate": "2021-03-26T22:17:17.147Z",
          "content": "<p>yes                                  </p>",
          "rawMarkdown": "yes                                  "
        },
        {
          "id": 1287811,
          "postDate": "2021-04-29T12:25:18.267Z",
          "content": "<p>Just a note: I've upgraded my gpu.</p>",
          "rawMarkdown": "Just a note: I've upgraded my gpu."
        }
      ]
    },
    {
      "id": 1231037,
      "postDate": "2021-03-08T16:11:23.380Z",
      "content": "<p>Has your CV-LB gap changed since then?</p>",
      "rawMarkdown": "Has your CV-LB gap changed since then?",
      "votes": 2,
      "replies": [
        {
          "id": 1231085,
          "postDate": "2021-03-08T16:56:42.283Z",
          "content": "<p>I incorporated some augmentations as expected the score improved <code>CV: 5.03</code> LB <code>11.2</code>…</p>",
          "rawMarkdown": "I incorporated some augmentations as expected the score improved `CV: 5.03` LB `11.2`...",
          "votes": 4
        },
        {
          "id": 1231115,
          "postDate": "2021-03-08T17:17:43.830Z",
          "content": "<p>Thank you for the reply!<br>\nLB/CV 26.7/15.7~1.7 -&gt; 11.2/5.03~2.22<br>\nIn relative value that seems to be an increase, though.</p>",
          "rawMarkdown": "Thank you for the reply!\nLB/CV 26.7/15.7~1.7 -> 11.2/5.03~2.22\nIn relative value that seems to be an increase, though.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1230777,
      "postDate": "2021-03-08T12:57:35.847Z",
      "content": "<p>Amazing accuracy! This model performance has surprised many cheminformaticians! Just curious, if you run prediction on the rest of training set, what's the distribution of the levenshtein distances, and what's the percentage of perfect predictions (distance = 0)?</p>",
      "rawMarkdown": "Amazing accuracy! This model performance has surprised many cheminformaticians! Just curious, if you run prediction on the rest of training set, what's the distribution of the levenshtein distances, and what's the percentage of perfect predictions (distance = 0)?",
      "votes": 2,
      "replies": [
        {
          "id": 1230858,
          "postDate": "2021-03-08T14:17:49.253Z",
          "content": "<p>I haven't done for full dataset prediction yet.. but for hold out (200k - separate from training and validation) I got something like this:</p>\n<pre><code>Levenshtein_distance\n0                       137603\n2                        11653\n1                        11431\n3                         8382\n4                         8135\n6                         6070\n5                         4921\n7                         4346\n8                         4261\n9                         2965\n10                        2807\n11                        2353\n</code></pre>\n<p>this means that from <code>200k</code>, <code>137603</code> have <code>Levenshtein_distance</code> of <code>0</code> and so on… Again this still vanilla model , I am pretty sure kagglers will push this competition to limit =) </p>",
          "rawMarkdown": "I haven't done for full dataset prediction yet.. but for hold out (200k - separate from training and validation) I got something like this:\n\n```\nLevenshtein_distance\n0                       137603\n2                        11653\n1                        11431\n3                         8382\n4                         8135\n6                         6070\n5                         4921\n7                         4346\n8                         4261\n9                         2965\n10                        2807\n11                        2353\n\n```\nthis means that from `200k`, `137603` have `Levenshtein_distance` of `0` and so on... Again this still vanilla model , I am pretty sure kagglers will push this competition to limit =) ",
          "votes": 10
        },
        {
          "id": 1231086,
          "postDate": "2021-03-08T16:57:54.590Z",
          "content": "<p>Was 11 the highest distance you had with this hold out set? Great work and thanks for sharing!</p>",
          "rawMarkdown": "Was 11 the highest distance you had with this hold out set? Great work and thanks for sharing!"
        },
        {
          "id": 1231098,
          "postDate": "2021-03-08T17:05:19.280Z",
          "content": "<p>Ah I just printed top 11 … =) Here is full <a href=\"https://ibb.co/r4jdH7N\" target=\"_blank\">distribution</a>. </p>\n<p>Edit: for some reason I can't load images from my computer .. </p>",
          "rawMarkdown": "Ah I just printed top 11 ... =) Here is full [distribution](https://ibb.co/r4jdH7N). \n\nEdit: for some reason I can't load images from my computer .. ",
          "votes": 3
        },
        {
          "id": 1231110,
          "postDate": "2021-03-08T17:15:14.300Z",
          "content": "<p>Ah okay, that makes sense, thanks for sharing it :)</p>",
          "rawMarkdown": "Ah okay, that makes sense, thanks for sharing it :)",
          "votes": 2
        }
      ]
    },
    {
      "id": 1307632,
      "postDate": "2021-05-14T15:09:30.087Z",
      "content": "<p>Can anypne explain what is rdkit post-processing.</p>",
      "rawMarkdown": "Can anypne explain what is rdkit post-processing.",
      "replies": [
        {
          "id": 1307823,
          "postDate": "2021-05-14T17:27:25.497Z",
          "content": "<p>See this notebook from <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> </p>\n<p><a href=\"https://www.kaggle.com/nofreewill/normalize-your-predictions\" target=\"_blank\">https://www.kaggle.com/nofreewill/normalize-your-predictions</a></p>\n<hr>\n<p>It's essentially trying to pass the generated InChI into RDKit, and, if it's valid, capturing the returned, normalized, InChI string. It has been seen that this has a positive impact (sometimes) on test scores. That being said, it causes a segmentation fault to try and do this operation when passing certain InChI's, hence, the wonderful notebook provided by <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> linked to above.</p>",
          "rawMarkdown": "See this notebook from @nofreewill \n\nhttps://www.kaggle.com/nofreewill/normalize-your-predictions\n\n---\n\nIt's essentially trying to pass the generated InChI into RDKit, and, if it's valid, capturing the returned, normalized, InChI string. It has been seen that this has a positive impact (sometimes) on test scores. That being said, it causes a segmentation fault to try and do this operation when passing certain InChI's, hence, the wonderful notebook provided by @nofreewill linked to above.",
          "votes": 4
        },
        {
          "id": 1308403,
          "postDate": "2021-05-15T07:15:43.793Z",
          "content": "<p>is this allowed by host?<br>\nI've read in other post that fixing inchi post prediction is not allowed, i'm not sure if validating inchi with RDKit is allowed or not.</p>",
          "rawMarkdown": "is this allowed by host?\nI've read in other post that fixing inchi post prediction is not allowed, i'm not sure if validating inchi with RDKit is allowed or not."
        },
        {
          "id": 1308415,
          "postDate": "2021-05-15T07:37:45.357Z",
          "content": "<p><a href=\"https://www.kaggle.com/lahninethou\" target=\"_blank\">@lahninethou</a> I've asked same question to the host and they said it's valid.<br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231318#1292491\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/231318#1292491</a></p>",
          "rawMarkdown": "@lahninethou I've asked same question to the host and they said it's valid.\nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/231318#1292491",
          "votes": 3
        }
      ]
    },
    {
      "id": 1260793,
      "postDate": "2021-04-02T12:16:24.657Z",
      "content": "<p>I used Efficientnet B0 as encoder.<br>\nUsing about half of the train images<br>\nI went for 6 epochs and I'm at 10.4 on validation set<br>\nI have not submitted it yet because I'm out of GPU time :)</p>\n<p>But the main problem is that the loss does not seem to decrease anymore and training has literally stopped<br>\nI think next week I'm gonna try another encoder out!!<br>\nEdit: 10.4 Valid/ 11.78 LB</p>",
      "rawMarkdown": "I used Efficientnet B0 as encoder.\nUsing about half of the train images\nI went for 6 epochs and I'm at 10.4 on validation set\nI have not submitted it yet because I'm out of GPU time :)\n\nBut the main problem is that the loss does not seem to decrease anymore and training has literally stopped\nI think next week I'm gonna try another encoder out!!\nEdit: 10.4 Valid/ 11.78 LB",
      "replies": [
        {
          "id": 1260798,
          "postDate": "2021-04-02T12:19:12.100Z",
          "content": "<p>Btw, has anyone tried a more powerful RNN decoder? Like multi-layer LSTM ? Or maybe with a bigger hidden size? <br>\nBecause now I'm thinking maybe the encoder is alright and decoder cannot learn further and needs be stronger!</p>",
          "rawMarkdown": "Btw, has anyone tried a more powerful RNN decoder? Like multi-layer LSTM ? Or maybe with a bigger hidden size? \nBecause now I'm thinking maybe the encoder is alright and decoder cannot learn further and needs be stronger!"
        },
        {
          "id": 1260834,
          "postDate": "2021-04-02T12:48:43.637Z",
          "content": "<p>\"Because now I'm thinking maybe the encoder is alright and decoder cannot learn further and needs be stronger!\"</p>\n<p>my strategy is to think of a model that can train fast and infer fast … and i can increase its capacity</p>\n<p>as for accuracy, generate more data with rdkit will do the trick</p>",
          "rawMarkdown": "\"Because now I'm thinking maybe the encoder is alright and decoder cannot learn further and needs be stronger!\"\n\n\nmy strategy is to think of a model that can train fast and infer fast ... and i can increase its capacity\n\nas for accuracy, generate more data with rdkit will do the trick",
          "votes": 1
        }
      ]
    },
    {
      "id": 1258474,
      "postDate": "2021-03-31T16:30:44.640Z",
      "content": "<p>transformer-based model. CV: 4.37 and LB: 6.17. I think the model is still underfitted even I trained for 50 epochs.</p>",
      "rawMarkdown": "transformer-based model. CV: 4.37 and LB: 6.17. I think the model is still underfitted even I trained for 50 epochs.",
      "replies": [
        {
          "id": 1260701,
          "postDate": "2021-04-02T10:55:40.123Z",
          "content": "<p>does transformer train faster than LSTM?</p>",
          "rawMarkdown": "does transformer train faster than LSTM?"
        },
        {
          "id": 1260718,
          "postDate": "2021-04-02T11:21:15.343Z",
          "content": "<p>Well, I don't know because I haven't train LSTM model in this competition.</p>",
          "rawMarkdown": "Well, I don't know because I haven't train LSTM model in this competition."
        }
      ]
    },
    {
      "id": 1287808,
      "postDate": "2021-04-29T12:23:52.047Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1279639,
      "postDate": "2021-04-21T05:53:27.157Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1260670,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-02T10:26:26.977000",
      "content": "<p>power of big encoder.</p>\n<p>use CNN-attention-LSTM (1 layer) model, single fold</p>\n<p>train: input 224x224, no augmentation (i.e. rotation is not used)<br>\ntest: use YNakama w&lt;h rotate trick, greedy (argmax) decoder<br>\ntokenizer: YNakama's Tokenizer</p>\n<p>training: train as long as my HW can support … basically the below results are for epoch&gt;10. results are not optimal, but i estimate within +/-1 optimality</p>\n<p>resnet34d: ~LB 9.5/CV 8.7<br>\nresnet26d: ~LB 8.5/CV 6.8<br>\nresnet101d:~LB 4.9/CV 4.0<br>\nresnet200d:~LB 3.8/CV 3.2</p>\n<hr>\n<p>below are the results for training in progress. no submission has been made</p>\n<ul>\n<li>load  pretrained encoder  into transformer and finetune end-to-end:<ul>\n<li>resnet26d : train/valid cross-entropy loss seems better than CNN-attention-LSTM-resnet101d and resnet200d</li>\n<li>hence, transformer is clearly the winner</li></ul></li>\n</ul>",
      "votes": 15,
      "replies": [
        {
          "id": 1260682,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-02T10:41:26.237000",
          "content": "<p>i wonder did anyone did experiment on size?<br>\ne.g. 192,224,256,288,320 ….</p>\n<p>in my early experiment for CNN encoder resnet34 with stride=32, 288 seems to perform worse than 224. but i cannot confirm since training takes long long time which i cannot afford</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1260761,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-04-02T11:54:50.103000",
          "content": "<p>I stopped resnet101 after two epochs, because I got worse than resnet50.  May be I should continue for few more epochs (but it's really time consuming for my hardware )</p>\n<p>My current submit is resnet50 trained for 12 epochs<br>\nI use transformer as decoder and not LSTM. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1260774,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-02T12:00:23.117000",
          "content": "<p>\"I use transformer as decoder and not LSTM.\"</p>\n<p>my previous experience is that transformer should perform better</p>\n<p>i plan to change the decoder later (I can freeze my encoder for faster training when changing the decoder)<br>\nbut the problem with the transformer is inference. i haven't found a way to make it fast. having to spend 10 to 20 hrs to make a submission is frightening.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1260792,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-04-02T12:13:41.627000",
          "content": "<p>You are  right.  That's why I use TPU (Torch/XLA) for inference and train on GPU. </p>\n<p>Anyway I think we can  experiment a lot before running inference, given Validation and LB  seem pretty much aligned. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1260810,
          "author_name": "Moein",
          "author_url": "",
          "post_date": "2021-04-02T12:25:58.240000",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> I experimented with transformers but I didn't get better performance (other than faster training) out of them. I'm thinking my positional embedding was not optimal but I tried both sin/cos and simple nn.Embedding; none worked out good for me. I'm really unfamiliar with Transformers btw.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1260821,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-02T12:35:55.460000",
          "content": "<p>there is a trick for faster inference</p>\n<ol>\n<li><p>based on my experiment results about 50% of the prediction has 0  Levenshtein distance (i.e. perfect results)</p></li>\n<li><p>you can find a way to identify correct test prediction, either by:</p></li>\n</ol>\n<ul>\n<li>consistent results of different model and probing, etc</li>\n<li>use rdkit to generate image from predicted Inchi string and you also need to train a verifier to verify if kaggle image is same as rdkit image or not</li>\n</ul>\n<p>for new models, you only need to submit results for difficult test samples, while keeping the easy one fixed.</p>\n<p>some goes for training. some of the train samples are really easy. a good sampling method can reduce number of train samples</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1260823,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-02T12:36:58.167000",
          "content": "<p>\"I'm really unfamiliar with Transformers btw.\"</p>\n<p>the rule is if LSTM can do it, transformer will do better</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1260825,
          "author_name": "Alex Vishnevskiy",
          "author_url": "",
          "post_date": "2021-04-02T12:37:52.063000",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Is TPU much faster than GPU on Pytorch? I've experiemented with TPU but got some issues and now I'm thinking to go deeper with this problem but does it worth it?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1260842,
          "author_name": "Ren",
          "author_url": "",
          "post_date": "2021-04-02T12:55:20.093000",
          "content": "<p>When the model gets better, the % of 0 distance prediction will go up too. I think the top team should have &gt; 85% 0 distance predictions now. </p>\n<p>A quick way to tell if the prediction is likely 0 distance or not is simply to use rdkit to test whether the prediction is a valid inchi string, which should give &gt; 90% precision. </p>\n<p>So, pseudo labels would probably play some role in this game now. And that thought alone makes me want to abandon this competition…</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1260843,
          "author_name": "Alex Vishnevskiy",
          "author_url": "",
          "post_date": "2021-04-02T12:55:55.270000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> I've tried sequence bucketing based on the length of the InChI string but it showed results worse than I had without this sampler. I had an idea that this kind of grouping could be helpful for updating the gradients but probably it didn't work out because of the poor predictive power of my decoder. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1260876,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-04-02T13:36:17.640000",
          "content": "<p><a href=\"https://www.kaggle.com/alexvishnevskiy\" target=\"_blank\">@alexvishnevskiy</a> <br>\nThere is still I/O and CPU bottlenecks , compared to TensorFlow TPU.<br>\nBut we can still run 8 inferences in parallel </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1261038,
          "author_name": "will_try",
          "author_url": "",
          "post_date": "2021-04-02T16:34:55.027000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> For resnet34, is it okay to share what kind of adjustments you made based on the public kernel? ‘cause if I use the default setting, it became overfit even at epoch 5, single fold. Thanks!</p>",
          "votes": 0,
          "replies": [
            {
              "id": 1286849,
              "author_name": "yanzhengxxj",
              "author_url": "",
              "post_date": "2021-04-28T13:15:41.837000",
              "content": "<p>how can it be?there are millions of samples with long labels.&gt; <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> For resnet34, is it okay to share what kind of adjustments you made based on the public kernel? ‘cause if I use the default setting, it became overfit even at epoch 5, single fold. Thanks!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 1261381,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-03T02:57:41.180000",
          "content": "<p><a href=\"https://www.kaggle.com/xm2203\" target=\"_blank\">@xm2203</a> </p>\n<p>i rewrite the code from the public kernel. but there is not much difference</p>\n<p>i don't think it will overfit if you are using all the train samples (80/20 train/validation split)<br>\ntry to use learning rate 1e-3 and train end-to-end, then drop to 1e-4,1e-5 only when there is no more improvement</p>\n<p>i will publish my train log files and loss curve curves later after organization</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1261508,
          "author_name": "Nitin Datta",
          "author_url": "",
          "post_date": "2021-04-03T06:04:30.357000",
          "content": "<p><a href=\"https://www.kaggle.com/xm2203\" target=\"_blank\">@xm2203</a> <br>\nIf you are using the public kernel you are not overfitting because the LR scheduler used is Cosine annealing and it has T_max of 5 so the LR resets to higher value which makes your score higher. If you continue training you should see good improvements till epoch-10. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1285021,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2021-04-26T13:59:11.133000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> \"the rule is if LSTM can do it, transformer will do better\"… \"if you have a LOT of data\" :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1286847,
          "author_name": "yanzhengxxj",
          "author_url": "",
          "post_date": "2021-04-28T13:14:02.827000",
          "content": "<p>how can it be?there are millions of samples with long labels.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1236521,
      "author_name": "Y.Nakama",
      "author_url": "",
      "post_date": "2021-03-13T08:08:21.683000",
      "content": "<pre><code>encoder: resnet34 \ndecoder: LSTM with attention\ntrain_samples: 80%\nvalid_samples: 20%\nepoch: 1\naugmentation: None\nsize: 224\nCV: 24.1\n</code></pre>\n<p>Training is done with kaggle notebooks, I will continue training and share notebooks as starter code.</p>",
      "votes": 12,
      "replies": [
        {
          "id": 1236664,
          "author_name": "Johnny Lee",
          "author_url": "",
          "post_date": "2021-03-13T10:58:30.830000",
          "content": "<p>if you didn't use augmentation, the LB will be about 30.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1237433,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2021-03-14T06:22:03.810000",
          "content": "<p>Hmm… I trained 2 epochs and CV is 18.7, LB is 46.9, maybe something is wrong… I will try to fix it.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1237454,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-03-14T06:46:44.630000",
          "content": "<p>Is it possible that you did not take into account that the training set only has horizontal compounds, but the test has 90° rotated ones, too?</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1237479,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2021-03-14T07:26:30.600000",
          "content": "<p>Thanks for pointing out! It seems that's the reason of large gap between my CV and LB. </p>\n<p>===== Edit1 =====<br>\nApplied below code for inference and LB changed from 46.9 to 21.9</p>\n<pre><code>h, w, _ = image.shape\nif h &gt; w:\n    image = image.transpose(1, 0, 2)\n</code></pre>\n<p>===== Edit2 =====<br>\nThis is correct one. LB changed from 21.9 to 20.3</p>\n<pre><code>import albumentations as A\n\ntransform = A.Compose([A.Transpose(p=1), A.VerticalFlip(p=1)])\n\nh, w, _ = image.shape\nif h &gt; w:\n    image = transform(image=image)['image']\n</code></pre>",
          "votes": 19,
          "replies": []
        },
        {
          "id": 1237692,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-03-14T11:28:23.743000",
          "content": "<p>I'm glad I could help! :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1237974,
          "author_name": "Ren",
          "author_url": "",
          "post_date": "2021-03-14T14:44:19.980000",
          "content": "<p>Isn't that flip via diagnoal…..?</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1238045,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2021-03-14T15:36:18.040000",
          "content": "<p>Oh, thanks for pointing out! I added correct one.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1238055,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-14T15:43:13.193000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1240629,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-03-16T14:41:09.247000",
          "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> your public kerne is greatl. We have slightly similar set up but with few tweaks.</p>\n<p>Its possible to achieve <code>CV-5.4</code> and <code>LB-6.9</code> with just <code>resnet34</code> trained for only <code>10</code> epoch.  In fact our best epoch is <code>9</code> =) </p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 1241092,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2021-03-16T21:49:47.353000",
          "content": "<p><a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> Thanks for information! <br>\nI have already reached CV-4.x with resnet34 encoder :D<br>\nSo then your current LB-3.8 is from other encoder? :)</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1241120,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-03-16T22:44:56.903000",
          "content": "<p>yes current score can be achieved with <code>resnet50</code> xD But also if I try hard can be reached with <code>resnet34</code> =) </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1243165,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "2021-03-18T04:11:06.780000",
          "content": "<p>your score is impressive,how long to train for your current LB?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1243194,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2021-03-18T04:38:57.803000",
          "content": "<p>~2hours per epoch for 80% train images on TITAN RTX.</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 1246869,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-03-21T07:37:34.793000",
          "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> Have you updated the public notebook with the flip code</p>\n<pre><code>import albumentations as A\n\ntransform = A.Compose([A.Transpose(p=1), A.VerticalFlip(p=1)])\n\nh, w, _ = image.shape\nif h &gt; w:\n    image = transform(image=image)['image']\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1246874,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2021-03-21T07:41:00.720000",
          "content": "<p>Yes, it's used in latest version and scored LB:20.3</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1246886,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2021-03-21T07:54:05.403000",
          "content": "<p>if we apply this augmentation every image will be of a different shape right? will it be any problem while training</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1246895,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2021-03-21T08:11:04.977000",
          "content": "<p>Yes, I don't use it for training.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1246920,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-03-21T08:50:20",
          "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> i think you didnt changed in train notebook</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1246934,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2021-03-21T09:06:37.440000",
          "content": "<p>It's always h &lt; w for train images so you don't need it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1246936,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2021-03-21T09:07:15.163000",
          "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>  then where do you use it?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1246938,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2021-03-21T09:09:03.953000",
          "content": "<p>For inference.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1258149,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-03-31T11:45:08.617000",
          "content": "<p>end to end training for CNN-LSTM is very slow. what i did in other competition is:</p>\n<ol>\n<li><p>pretrain CNN only using classification (this is competition, you can probably do a few tasks like predicting number of C atoms, H atoms, etc (regression), classify if the image contains some functional group (or partial InChi substring, or if a word is present or not from the token dictionary, aka multi binary label), segmentation, clustering, contrastive learning, etc …. just think of some simple way to generate ground truths from the InChI)</p></li>\n<li><p>with the CNN layers froze, train your LSTM only</p></li>\n<li><p>finally finetune CNN+LSTM with low learning rate for the target image-to-text translation task.</p></li>\n</ol>\n<p>The trick is to think of simple tasks in step (1)  that is closely related to the target task (3)</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1259826,
          "author_name": "Moein",
          "author_url": "",
          "post_date": "2021-04-01T17:24:49.933000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>. When I saw this yesterday, I immediately started a notebook go through these steps. I got good results with doing Regression on the number of atoms of each element. I just shared my notebook publicly here: <a href=\"https://www.kaggle.com/moeinshariatnia/cnn-rnn-cnn-pretraining-w-regression-train\" target=\"_blank\">[CNN+RNN]CNN pretraining w/ regression -TRAIN</a></p>\n<p>I'll be really glad if you and others tell me what you think about it.</p>\n<p><img src=\"https://i.ibb.co/WPDk5d7/Presentation1.jpg\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1268777,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-09T18:37:30.067000",
      "content": "<p>I completed CNN+LSTM experiments and now are playing with vision transformer(essentially a transformer encoder)+transformer decoder. <br>\n(also refer to my post for code and details: <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231190\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/231190</a>)</p>\n<p>here are the first results:</p>\n<p>transformer in transformer:    <br>\n     - TNT-S-224-16: LB 3.11/CV 1.96 (inference = 80 min on 4xti1080)<br>\n     - TNT-S-320-16: LB 2.27/CV 1.55 (inference = 140 min on 4xti1080)</p>\n<p>(do you have others to suggest?)</p>\n<hr>\n<p>train: input 224x224, no augmentation (i.e. rotation is not used)<br>\ntest: use YNakama w&lt;h rotate trick, greedy (argmax) decoder, aka. no beam search yet<br>\ntokenizer: YNakama's Tokenizer</p>\n<hr>\n<p>a few quick observations:</p>\n<ul>\n<li>transformer train faster, 2 days is enough</li>\n<li>you actually don't need many epoch (7 to 10 is ok), but you do need many iterations and low learning rate at the end.<br>\n(many iterations can mean you use a smaller batch size. i use 64)</li>\n<li>cv/lb gap is larger (which i think if you use teacher distillation with CNN+LSTM, maybe you can get interesting results)</li>\n</ul>\n<p>for more ideas: read this <a href=\"https://github.com/lucidrains/vit-pytorch\" target=\"_blank\">https://github.com/lucidrains/vit-pytorch</a></p>",
      "votes": 9,
      "replies": [
        {
          "id": 1268813,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "2021-04-09T19:33:15.647000",
          "content": "<p>I have also the same problem with the batch size :) bs=64 somehow always perform better than bigger batch sizes. Seems to be problem of multi GPU training?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1268892,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-09T23:01:44.633000",
          "content": "<p>no. my GPU has 48 GB. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1271054,
          "author_name": "Thomas SELECK",
          "author_url": "",
          "post_date": "2021-04-12T08:50:32.463000",
          "content": "<p>For the TNT-S-320-16 model, have you used a pretrained model ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1272367,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-04-13T12:27:19.690000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> do we have a pretrained model for TNT-S-320-16 model</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1287622,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-04-29T08:35:11.817000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<blockquote>\n  <p>transformer train faster, 2 days is enough</p>\n</blockquote>\n<p>I don't agree with that in my case. it takes ages for me</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1287813,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "2021-04-29T12:29:01.647000",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> I think he has many GPUs :)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1303020,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-05-11T21:19:50.243000",
      "content": "<p>Hi there. We thought we would share some of our benchmarks as others have been so forthcoming. It's been a wonderful competition so far.</p>\n<hr>\n<p><strong>SETUP:</strong><br>\n    - EfficientNetB5 <em>(headless… no extra layers)</em><br>\n    - 256/768 (attn/lstm-units) LSTM<br>\n    - 192x384 cropped/stretched/rotation-fixed images<br>\n    - TFRecords and TPU<br>\n    - Everything done on Kaggle in 1-2 sessions.</p>\n<hr>\n<p>12 EPOCHS (5-6 hours of training) - LIMIT MAX LENGTH TO 140 TOKENS<br>\n    - <strong>1.95 CV</strong> - No Beam Search - No Postprocessing<br>\n    - <strong>3.75 LB</strong> - No Beam Search - No Postprocessing</p>\n<p>10 ADDITIONAL EPOCHS (5-6 hours of training) - FINE TUNE ON FULL LENGTH (277 TOKENS)<br>\n    - <strong>1.89 CV</strong> - No Beam Search - No Postprocessing<br>\n    - <strong>2.68 LB</strong> - No Beam Search - No Postprocessing<br>\n    - <strong>2.58 LB</strong> - No Beam Search - RDKit Postprocessing</p>\n<hr>\n<p>We are implementing Beam Search and swapping in Transformers shortly and hopefully, that will help. At some point in the near future, we will scale up image size and model size and increase training time and try to squeeze below 1.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 1303255,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-05-12T01:31:22.723000",
          "content": "<p>\"2.58 LB - No Beam Search - RDKit Postprocessing\"</p>\n<p>if inchi validation is the key, you might as well train an \"inchi correction model\"</p>\n<p>how to get corrupted inchi? </p>\n<ul>\n<li>use prediction of your model in earlier iterations (i.e. before it converges)</li>\n<li>add noise to input image</li>\n</ul>\n<p>this is seq-to-seq model : input invalid inchi, output best guess of valid inchi</p>\n<hr>\n<p>better still, train a refinement model and this can be repeated many rounds<br>\n(e.g. alternating between lstm and transformer)</p>\n<p>model : input previous model predicted inchi+ image, output best guess of true inchi</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1303543,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2021-05-12T06:06:31.603000",
          "content": "<p>This is a great result, thank you for sharing all this in such a readable way <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>!</p>\n<p>Training on a GPU  for ~9 epochs, using LSTM + attention (1024 decoder dim), image size of 384x384 (rotated during training by 90 deg clockwise with a p of 0.5) I get 4.71 local CV.</p>\n<p>I wonder - would you be willing to share what augmentations you are using? My reasoning is this - there is so much data in train, one should be good without using augmentations? Really curious if people are seeing augmentations improve their results and what augmentations are helping.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1303892,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-05-12T10:21:13.063000",
          "content": "<p>Thank you! </p>\n<p>The only augmentation I’m performing is rotating when the cropped molecule aspect ratio is smaller than 1 (ie h&gt;w).</p>\n<p>I plan to iron out the kinks in this approach (instead of just h v. W) in the future (ie a model to predict if an image needs to be rotated).</p>\n<p>The only augmentation I can think of is increasing the amount of noise in the image (or decreasing)… I haven’t implemented it yet though. I believe others may have. </p>\n<p>I hope this helps!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1305188,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2021-05-13T06:20:31.723000",
          "content": "<p>This was very helpful indeed! 😊 Thank you very much for sharing your thoughts!</p>\n<p>Seems that if I want to improve my results, my time would be better spent elsewhere than messing around with the augmentations. Thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1306824,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2021-05-14T05:33:44.987000",
          "content": "<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> apologies, one more question if I may please. I am trying to get an intuition for just how fast TPUs are.</p>\n<p>When you mention a single epoch, how many images do you mean? The entire train set of 2.4 mln images? You are using all 8 TPU cores I would imagine?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1307614,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-05-14T14:54:34.293000",
          "content": "<p>Yup. When I say a single epoch I mean all of the training data (- 80,000 validation images)… so approximately 2.3-2.4 million images.</p>\n<p>TPU is incredibly useful in this competition.</p>\n<p>I am distributing the training across all 8 cores.</p>\n<p>I use a batch size of 1024 (128 for each core).</p>\n<p>Hope this helps!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1307652,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-05-14T15:26:07.710000",
          "content": "<p>TF distributes efficiently both the model and the data across all cores.   That's why you can train really fast even with huge models on TPU. <br>\nMixed Precision with bfloat16 reduces the training time by half. </p>\n<p>Gradient Accumulation helps to even speed up the training again.  But it's a bit trickier to implement on TF while it's really straightfoward on Pytorch. </p>\n<p>And of course with the TFRecords dataset in GCS near the TPU VM, there is no I/O bottleneck</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1308484,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2021-05-15T08:50:19.517000",
          "content": "<p>Thank you very much for these additional details <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>! 2.4 million images in ~30 minutes, especially with a sequential model such as RNN, that is super impressive 🙂</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1308517,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2021-05-15T09:20:39.817000",
          "content": "<p>Thank you for sharing this information <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a>! Appreciate it</p>\n<p>Could I please ask about why you believe gradient accumulation can be helpful? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1308591,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-05-15T10:18:16.543000",
          "content": "<p>Sure, </p>\n<p>Because when you apply gradient accumulation, The gradient update (<code>optimizer.apply_gradients</code> on TF  and <code>optimizer.step()</code> on Pytorch) is done only  every K steps intead of every step. </p>\n<p>Without Gradient Accumulation, the distributed training would compute loss reduce for all cores at every step . Therefore, with gradient accumulation you will save <code>num_cores*(K-1)</code> loss reduce time at every gradient update.  Then, the more you have number of cores, the more you will accelerate training with GA. <br>\nBut you may need to tune the learning rate accordingly. </p>\n<p>On Colab( which has TPU v2-8), you can manage to make the training almost as fast as on kaggle (which has TPU v3-8) with GA.   That's why I don't bother to use Kaggle TPU. </p>\n<p>You can see <a href=\"https://www.tensorflow.org/xla\" target=\"_blank\">here</a> a benchmark w/ and w/o GA for distributed training </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1308663,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2021-05-15T11:15:49.523000",
          "content": "<p>Thank you very much for your answer <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a>!!!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1267440,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2021-04-08T14:13:17.970000",
      "content": "<p>With the update of new rules, i want to update our score. Our current LB standing doesn’t break any rules. Best model has CV of 1.29 and LB of 1.34. I hope this will encourage participants to build model with high scores:) </p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 1230093,
      "author_name": "Ren",
      "author_url": "",
      "post_date": "2021-03-07T19:16:32.597000",
      "content": "<p>Just trained 1 epoch to verify my code works… e0+lstm. got 24 in my holdout set and 44 on the public. Quite a large gap.</p>\n<p>Mar-10<br>\nFixed some issue, and I am moving again. Finished another epoch and I got 19 local and 32 on the public. </p>\n<table>\n<thead>\n<tr>\n<th>local 5% hold out</th>\n<th>public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>24</td>\n<td>44</td>\n</tr>\n<tr>\n<td>19</td>\n<td>32</td>\n</tr>\n<tr>\n<td>16</td>\n<td>29</td>\n</tr>\n<tr>\n<td>13</td>\n<td>27</td>\n</tr>\n</tbody>\n</table>\n<p>Above table does not handle the different orient in test set. </p>\n<p>Mar-17 Had some significant code changes and trained new model end to end from scratch, still e0 encoder lstm decoder. </p>\n<table>\n<thead>\n<tr>\n<th>epoch</th>\n<th>local 20% hold out</th>\n<th>public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>3</td>\n<td>5.4</td>\n<td>5.9</td>\n</tr>\n<tr>\n<td>4</td>\n<td>4.9</td>\n<td>5.6</td>\n</tr>\n<tr>\n<td>5</td>\n<td>4.6</td>\n<td>5.2</td>\n</tr>\n</tbody>\n</table>\n<p>ensemble 5 fold(all similar validation scores) of the above setup -&gt; 4.11 LB. <br>\nIt seems ensemble is much better than average of cv</p>\n<p>Used rdkit to normalize the 4.11 predictions, score improved to 4.07. It seems there is some juice worth squeezing. </p>",
      "votes": 7,
      "replies": [
        {
          "id": 1230124,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-03-07T19:47:55.717000",
          "content": "<p>Hopefully you are continuing training =) Score increases greatly after first epoch (in my case) … Yes the gap thing is concerning, it will be important to analyze the data and the output of models ..</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1230163,
          "author_name": "Pablo Pernías",
          "author_url": "",
          "post_date": "2021-03-07T20:35:05.877000",
          "content": "<p>Are you using an autoregressive model? If so, while evaluating do you generate the output token by token or you generate it all in parallel by giving it the real previous tokens?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1230165,
          "author_name": "Ren",
          "author_url": "",
          "post_date": "2021-03-07T20:39:34.717000",
          "content": "<p>Yes, autoregressive model. When validating, I am not using real tokens. I generate until <code>&lt;EOS&gt;</code>. After that I decode the generated text to calculate the score. Thus I don't think I have leak in my local validation code. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1230174,
          "author_name": "Pablo Pernías",
          "author_url": "",
          "post_date": "2021-03-07T20:48:38.050000",
          "content": "<p>Ok, just checking, because even if I haven't achieved your results yet, my public score has been pretty consistent with my validation so far, maybe as the model gets better the gap increases.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1232345,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-03-09T17:17:32.573000",
          "content": "<p>Did you look at the train and test images?<br>\nTraining images are horizontal only, but in the test, there are horizontal and 90° (clockwise) rotated ones.<br>\nDid you take care of this difference?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1232356,
          "author_name": "Ren",
          "author_url": "",
          "post_date": "2021-03-09T17:30:56.097000",
          "content": "<p>I haven't looked at any images at all. I was primarily testing my code. Thanks for the info <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a>.  I think I will try to add rotation as an augmentation. </p>\n<p><a href=\"https://www.kaggle.com/DrHB\" target=\"_blank\">@DrHB</a> mind share what batch size are you using? I stopped it after one epoch and loaded it up in a larger GPU to train in a larger batch size(128 in this case) and my training became very unstable, had a few blow-ups. Or, are you clipping gradients?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1232362,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-03-09T17:37:24.990000",
          "content": "<p>I can train on <code>fp16</code> with batch size <code>256</code> and my <code>gpu</code> still have 3-4 gb free left from <code>16</code> =)  </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1232366,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-03-09T17:39:17.607000",
          "content": "<p>btw with <code>fp16</code> several time I was getting <code>nan</code> loss … adding warm up helped with <code>gradient clipping</code>. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1232374,
          "author_name": "Ren",
          "author_url": "",
          "post_date": "2021-03-09T17:43:43.410000",
          "content": "<p>Ah, thanks a lot! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1232381,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-03-09T17:47:44.267000",
          "content": "<p>this worked for me… I think <code>max_norm</code> can go to <code>4.</code> you have to play around…  </p>\n<pre><code>norm_type = 2.0\nmax_norm  = 1.\nnn.utils.clip_grad_norm_(model.parameters(), max_norm, norm_type)\n</code></pre>\n<p>good luck!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1255883,
          "author_name": "Nitin Datta",
          "author_url": "",
          "post_date": "2021-03-29T10:08:59.043000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ryanzhang\" target=\"_blank\">@ryanzhang</a> , <br>\nIf you dont mind can you give a brief idea of how you are creating the ensemble….</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1255894,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-29T10:17:53.217000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1253530,
      "author_name": "Konrad Banachewicz",
      "author_url": "",
      "post_date": "2021-03-26T20:48:24.033000",
      "content": "<p>resnet34, 10 epochs -&gt; single fold validation: 8.27, lb: 8.87</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1253534,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-03-26T20:54:02.083000",
          "content": "<p>Is this the public kernel?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1253894,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2021-03-27T06:37:15.570000",
          "content": "<p>I started with this one as basis, yes.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1254279,
          "author_name": "Thomas SELECK",
          "author_url": "",
          "post_date": "2021-03-27T13:39:09.027000",
          "content": "<p>Have you added beam search to get such result ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1254284,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2021-03-27T13:44:35.547000",
          "content": "<p>Not yet, no.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1265529,
      "author_name": "Tucker Arrants",
      "author_url": "",
      "post_date": "2021-04-07T01:47:33.157000",
      "content": "<ul>\n<li>Encoder: Pruned EfficientNetB3</li>\n<li>Decoder: GRU</li>\n<li>Image size: <code>224</code></li>\n<li>Augmentation: None</li>\n<li>CV: <code>2.79</code></li>\n<li>CV with beam search (k=5): <code>2.65</code></li>\n<li>LB without beam search: <code>3.57</code></li>\n<li>LB with beam search (k=5): <code>3.22</code></li>\n</ul>",
      "votes": 6,
      "replies": [
        {
          "id": 1265567,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-07T03:12:13.547000",
          "content": "<p>thanks!.</p>\n<p>i forget there is such a thing called pruned model though i used it in work.<br>\nby the way, seems that pruned model uses less watts (saving in electricity bill)</p>\n<pre><code>|===============================+======================+=|\n|   0  TITAN X (Pascal)    Off  | 00000000:05:00.0  On |                  N/A |\n| 59%   84C    P2   105W / 250W |  10302MiB / 12192MiB |    100%      Default |\n+-------------------------------+----------------------+----------------------+\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1265724,
          "author_name": "Thomas SELECK",
          "author_url": "",
          "post_date": "2021-04-07T06:55:03.383000",
          "content": "<p>Are you using attention with GRU ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1265747,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2021-04-07T07:22:34.857000",
          "content": "<p>Yes, GRU with attention</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1267343,
          "author_name": "Ertuğrul Demir",
          "author_url": "",
          "post_date": "2021-04-08T13:29:49.410000",
          "content": "<p><a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> just one layer of GRU or you stack some?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1267395,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2021-04-08T13:45:20.860000",
          "content": "<p>Just one GRU layer</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1308750,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2021-05-15T12:41:28.323000",
          "content": "<p>I wonder if using a pruned EffNet makes sense here. It was pruned for the purposed of the task it was trained for right? Here the task is rather different. What are your thoughts?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1230112,
      "author_name": "Matt",
      "author_url": "",
      "post_date": "2021-03-07T19:33:57.840000",
      "content": "<p>Great job on pushing the leaderboard forward. I'm still working on my first submission, hopefully I'll have the basis for my model setup in a couple days.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1267645,
      "author_name": "Hai Nam Nguyen",
      "author_url": "",
      "post_date": "2021-04-08T17:08:15.797000",
      "content": "<p>I just used some public kernels and got CV 3.xx with B0 without any change in the decoder.<br>\nPersonally, I think finding a good training strategy with suitable batch size and LR scheduler is essential.<br>\n P/S: My hardware resource is limited.  </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1267208,
      "author_name": "atfujita",
      "author_url": "",
      "post_date": "2021-04-08T11:40:30.710000",
      "content": "<p>Encoder: eca_nfnet_l0<br>\nDecoder: LSTM+GRU<br>\nImage size: 224<br>\nBS=128<br>\nscheduler='CosineAnnealingLR'<br>\nAugmentation: None<br>\nCV: 4.61<br>\nLB: 5.85</p>\n<p>1 epoch takes 5h on my local machine.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1291931,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2021-05-03T13:04:27.760000",
      "content": "<p>Let's re-activate cv/lb thread.</p>\n<p>CV:1.17 -&gt; LB:1.49<br>\nCV:1.10 -&gt; LB:1.42<br>\nSeems stable for me, but a little big discrepancy.</p>\n<p>Do you have any updates, <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> ??<br>\nYou should;)</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1291934,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-05-03T13:11:41.140000",
          "content": "<p>=) We have single models below 1.0. With CV and LB close. Cant say anything more at this time =)  But you are doing impressive progress for 2 submission, good luck to you  =) </p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1291936,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "2021-05-03T13:12:34.870000",
          "content": "<p>my current best gap is 0.15, but it depends on your fold.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1291947,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-05-03T13:24:08.103000",
          "content": "<p>is these raw results without k-beam search or inchi post-processing?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1291952,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-05-03T13:28:40.683000",
          "content": "<p>We don't use K - beam.. and we use <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> normalize script… </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1291978,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-05-03T13:51:25.317000",
          "content": "<p>Glad to see so many to use my script. Thanks for mentioning me! : )</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 1291981,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-05-03T13:55:13.430000",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> normalize script beats the best network</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1291992,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2021-05-03T14:05:31.867000",
          "content": "<p><a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> <a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> Wow, single models below 1.0 are impressive! <br>\nAnd also thanks for letting me know you don't use k-beam, it was my next priority but now I decide to focus on something different;)<br>\nI'm trying hard to catch up you guys, good luck!</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1298890,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-05-09T09:48:23.293000",
          "content": "<p><a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a> nice score! Is that with a single model, without beam search?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1298906,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2021-05-09T10:04:20.257000",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> Thanks, you should have better model than this since you haven't submitted for this 10 days:)<br>\nYes, I am too lazy to implement beam search, the score is of single model, without beam search, but with your rdkit normalization!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1298914,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-05-09T10:11:13.880000",
          "content": "<p>I may have. But I also worked on it for much longer time than you. You made that in about only a few weeks. Hats off!</p>\n<p>I only implemented the normalization. The idea is from <a href=\"https://www.kaggle.com/stassl\" target=\"_blank\">@stassl</a> and <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> <br>\nJust wanted to emphasize it :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1298926,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2021-05-09T10:27:57.163000",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> Actually I've worked on this competition more than a month but I just haven't submitted until my cv score got good enough:)<br>\nBy the way you and <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> both say not using beam search, does it mean it doesn't help or it works but you don't use it due to longer inference time?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1298933,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-05-09T10:38:31.020000",
          "content": "<p><code>I've worked on this competition more than a month</code><br>\nNow that's a bit comforting information. :D</p>\n<p>Beam search definitely helps my CV.<br>\nBut normalization had no effect on my CV (didn't change any prediction), so I don't know how much beam search helps together with normalization. I have a bad feeling that the effect of normalization will degrade a lot. I hope I'm wrong about it.</p>\n<p>But yeah, I would say that not using beam search is a bad idea. Especially because there is a lot of time to implement it while model is being trained.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 1298936,
              "author_name": "tugstugi",
              "author_url": "",
              "post_date": "2021-05-09T10:41:03.093000",
              "content": "<blockquote>\n  <p>Beam search definitely helps my CV.</p>\n</blockquote>\n<p>How much does the beam search on transformer help? Haven't implemented yet :)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1298946,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-05-09T10:55:28.030000",
          "content": "<p>The best I've got with only 2 beams is 3.7% LD reduction.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1298987,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2021-05-09T11:45:34.557000",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> Really, I always get 0.07 improvement by rdkit normalization. Maybe my original predictions are too naive🤔</p>\n<blockquote>\n  <p>Especially because there is a lot of time to implement it while model is being trained.</p>\n</blockquote>\n<p>Definitely true;)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1298994,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-05-09T11:49:30.343000",
          "content": "<p>I see diminishing effects as my predictions get better. But could be just coincidence.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1299014,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2021-05-09T12:15:12.680000",
          "content": "<p>I too see diminishing effects as my predictions get better. </p>\n<p>For my models, beam search + normalization yields better CV / leaderboard than just beam search or normalization by itself. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1299064,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-05-09T12:50:17.763000",
          "content": "<p>Then <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a> may secretly be doing something different :D</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1257282,
      "author_name": "Ertuğrul Demir",
      "author_url": "",
      "post_date": "2021-03-30T17:19:03.957000",
      "content": "<p>efficientnet_b0 as encoder, 5 epochs, CV: 7.21, LB: 7.76 using around 50% of the train data and no augmentations. For now wondering what's going to make difference, because usually model really slows down learning after that point…</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1258330,
          "author_name": "Sirish Somanchi",
          "author_url": "",
          "post_date": "2021-03-31T14:37:01.567000",
          "content": "<p>Why no augmentation? Several participants have reported better LB scores using random Rotation (90-deg) augmentation.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1258605,
          "author_name": "Ertuğrul Demir",
          "author_url": "",
          "post_date": "2021-03-31T18:44:17.747000",
          "content": "<p>I think it depends on test data you using, (the 90 degree one)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1253878,
      "author_name": "Moein",
      "author_url": "",
      "post_date": "2021-03-27T06:01:28.540000",
      "content": "<p>Anyone using <strong>Beam Search</strong> for prediction? Does it make a huge difference in your predictions instead of simple argmax at each timestep?<br>\nI wrote a beam search function but the problem is that it is not batch-wise and takes a lot of time. I also did not find an efficient batch-wise beam search online.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1253961,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-03-27T07:55:32.240000",
          "content": "<p>You can test what difference it makes on a small holdout set.<br>\nI also did not find one online and it wasn't easy task to make it efficient.<br>\nFor me it does make a difference, though:<br>\nCV went ~4.0 -&gt; 3.15<br>\nLB went 7.57 -&gt; 6.59</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1254013,
          "author_name": "Moein",
          "author_url": "",
          "post_date": "2021-03-27T09:00:27.807000",
          "content": "<p>Thanks for mentioning that. So I think I'm gonna have a hard time implementing it 😅</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1254037,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-03-27T09:36:19.337000",
          "content": "<p>It's not easy if you implement it the first time, but was absolutely enjoyable.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1254077,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-03-27T10:22:20.393000",
          "content": "<p>You can't parallelize beam-search. </p>\n<p>You may skip it for (quick) early experiments. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1254079,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-03-27T10:27:31.430000",
          "content": "<p>It may depend on the architecture, because I did parallelize my beam search.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1254085,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-03-27T10:37:01.517000",
          "content": "<p>I don't know much about RNNs, but I'm skeptical that those cannot be beam search parallelized. You just need to cache all the hidden states, but again, I have zero practical experience with RNNs.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1254175,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-03-27T12:00:16.257000",
          "content": "<p>You need to know top_k sequences with highest overall probability.  I don't see how you can compute this with single forward pass. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1254194,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-03-27T12:17:56.143000",
          "content": "<p>You have 2 beams: beam1 with probability 0.9 and beam2 with probability 0.85. You <strong>forward pass</strong> the two and get 2×(vocabulary size) -&gt; flatten it and get topk. ([[vocab_size],[vocab_size]]-&gt;[vocab_size<strong>;</strong>vocab_size])</p>\n<ul>\n<li>If you now have […<strong>0.87</strong>,…0.81,… <strong>;</strong> …<strong>0.83</strong>,…0.82,…] then you select the 0.87 from beam1 and 0.83 from beam2, that is straightforward.</li>\n<li>If you have […<strong>0.87</strong>,…<strong>0.84</strong>,… <strong>;</strong> …0.83,…0.82,…] however, then you select 0.87 and 0.84 from beam1 and copy the model from beam1 to beam2.</li>\n</ul>\n<p>These <strong>forward passes</strong> are the target of parallelization.</p>\n<hr>\n<p>With batch size of 2 and beam width of 2:<br>\n<strong>Forward pass</strong>:<br>\n[[vocab_size],<br>\n[vocab_size],<br>\n[vocab_size],<br>\n[vocab_size],]<br>\n<strong>Flattened for topk</strong>:<br>\n[[vocab_size, vocab_size],<br>\n[vocab_size, vocab_size],]</p>\n<hr>\n<p>So the main idea is that you handle the two beams separately while you forward pass, but you consider them together for selecting topk, and take the corresponding models in accordance to topk.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1236612,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2021-03-13T09:56:39.843000",
      "content": "<p>Anyone successful to use transformer as a decoder?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1252085,
          "author_name": "Moein",
          "author_url": "",
          "post_date": "2021-03-25T12:17:44.460000",
          "content": "<p>I've tried it. It produces true tokens but the ordering is way wrong. I'm working on the positional embeddings and post-processing to make it work</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1278133,
          "author_name": "Matrix",
          "author_url": "",
          "post_date": "2021-04-19T15:30:11.857000",
          "content": "<p>my transformer also didn’t achieve a satisfying results. is really the problem of a wrong positional emdedding in your case？</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1234172,
      "author_name": "Matrix",
      "author_url": "",
      "post_date": "2021-03-11T03:07:09.093000",
      "content": "<p>Surprising that baseline model can perform so well.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1243809,
          "author_name": "marketneutral",
          "author_url": "",
          "post_date": "2021-03-18T14:00:38.220000",
          "content": "<p>I am not a chemist but I don't think that a model with a Levenshtein value of ~20 would have any value for scientific use. At the longest InChi lables, that's almost 10% of the string that is wrong. I suspect a value &lt; 1 is needed.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1232308,
      "author_name": "xhlulu",
      "author_url": "",
      "post_date": "2021-03-09T16:35:14.690000",
      "content": "<p>Interesting! Are you using attention during the decoding? I'm curious how useful that would be.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1232334,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-03-09T17:01:19.743000",
          "content": "<p>Yes attention. And I have plotted some attention and it seems like it recognizes edges and letters… but also a lot of background empty space ..  </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1230221,
      "author_name": "Sinan Calisir",
      "author_url": "",
      "post_date": "2021-03-07T22:36:08.857000",
      "content": "<p>I'm waiting an end to end transformer model in this discussion :)</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1260240,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-02T00:43:12.320000",
      "content": "<p>use batchwise beam search?</p>\n<p><a href=\"https://stackoverflow.com/questions/64356953/batch-wise-beam-search-in-pytorch\" target=\"_blank\">https://stackoverflow.com/questions/64356953/batch-wise-beam-search-in-pytorch</a><br>\n<a href=\"https://zhuanlan.zhihu.com/p/167072494\" target=\"_blank\">https://zhuanlan.zhihu.com/p/167072494</a></p>\n<p><a href=\"https://medium.com/the-artificial-impostor/implementing-beam-search-part-1-4f53482daabe\" target=\"_blank\">https://medium.com/the-artificial-impostor/implementing-beam-search-part-1-4f53482daabe</a><br>\nHow to Do Beam Search Efficiently<br>\n\"The smarter way is to put these N nodes into a batch and feed it to the model. It’ll be much faster since we can parallelize the computation, especially when using a CUDA backend.\"</p>\n<hr>\n<p>if you train multiple models (e.g. improved models as you make more submissions or for ensemble), these models can provide a hypothesis to reduce the time for beam search (i.e. they provide some clue on the top-k decoding path)</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1237838,
      "author_name": "Johnny Lee",
      "author_url": "",
      "post_date": "2021-03-14T13:22:17.290000",
      "content": "<p>end to end model: CNN + RNN<br>\nattention: None (don't know how to implement yet)<br>\ntrain / valid: 80% / 20%<br>\nepoch: 18<br>\nCV: 10.8<br>\nLB: 11.6</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1237844,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2021-03-14T13:27:14.047000",
          "content": "<p>Wow 18 epochs… How long does it take and 18th epoch is the best CV?</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1237915,
          "author_name": "Johnny Lee",
          "author_url": "",
          "post_date": "2021-03-14T13:55:48.640000",
          "content": "<p>It takes about 15 hours with TPU and continue training for now. My model needs to be optimized. It's very slow to train.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1237954,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2021-03-14T14:23:23.187000",
          "content": "<p>Thanks for information, TPU is very fast…</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1239746,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-03-16T01:17:37.967000",
          "content": "<p>any hint or paper on how you use the RNN on top of the CNN?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1239750,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-03-16T01:28:47.563000",
          "content": "<p><a href=\"https://arxiv.org/pdf/1502.03044.pdf\" target=\"_blank\">https://arxiv.org/pdf/1502.03044.pdf</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1239754,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-03-16T01:34:24.563000",
          "content": "<p><a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> </p>\n<p>Thank you very much! the paper has 6700 citations, i have a lot to read</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1240662,
          "author_name": "Johnny Lee",
          "author_url": "",
          "post_date": "2021-03-16T15:05:17.963000",
          "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> I was surprised that you can achieve 24.1 with only one epoch. Now I think I know the reason. <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> I use some different approach. I don't predict the character one by one. so my model need more epoch to train.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1243805,
          "author_name": "marketneutral",
          "author_url": "",
          "post_date": "2021-03-18T13:57:44.780000",
          "content": "<p><a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a> </p>\n<blockquote>\n  <p>It takes about 15 hours with TPU and continue training for now. My model needs to be optimized. It's very slow to train.</p>\n</blockquote>\n<p>Are you training on TPU available via Kaggle notebooks? Or somewhere else (where?)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1243818,
          "author_name": "Johnny Lee",
          "author_url": "",
          "post_date": "2021-03-18T14:10:31.873000",
          "content": "<p><a href=\"https://www.kaggle.com/marketneutral\" target=\"_blank\">@marketneutral</a> You can train 3 epochs, and then load it into another notebook to train next 3 epochs.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1243884,
          "author_name": "marketneutral",
          "author_url": "",
          "post_date": "2021-03-18T14:56:25.250000",
          "content": "<p>So you are training in Kaggle TPU for 3 epochs, saving the model, and then restarting training in a new TPU session?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1287894,
          "author_name": "yanzhengxxj",
          "author_url": "",
          "post_date": "2021-04-29T13:58:16.633000",
          "content": "<p><a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a> would you like to share how to use RNN without outputting one by one?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1229749,
      "author_name": "Mr_KnowNothing",
      "author_url": "",
      "post_date": "2021-03-07T15:27:07.327000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> thanks for sharing your baseline results and welcome to the competition. Are you training a Image Captioning with resnet as backbone ? I am trying to understand how you have trained resnet <br>\nThanks</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1229754,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2021-03-07T15:29:57.097000",
          "content": "<p>Same question…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1229762,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-03-07T15:32:42.140000",
          "content": "<p>yes its basic captioning model. To be honest I did nothing special so far.. just a simple baseline =). Maybe one thing to mention that my tokenization take cares of chemical elements e.g <code>Br</code> is  one <code>token</code> and not to <code>B</code> and <code>r</code>. I am not sure how important this is for now…. perhaps not so important because sequence model these days can learn this relationship pretty easy =) </p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 1229784,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2021-03-07T15:44:41.733000",
          "content": "<p>Thanks for answer. Two more questions if you don't mind, first, did you use PyTorch? second, is there any good tutorial with PyTorch that you recommends?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1229812,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-03-07T15:57:58.360000",
          "content": "<p>Yes I use PyTorch (pretty sure this can be done also in TF) and training done in fp16. </p>\n<p><a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> posted several good resources:<br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223419\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/223419</a><br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223218\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/223218</a><br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223223\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/223223</a></p>\n<p>I would start from them =)</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1230351,
          "author_name": "Yanfei SIMP",
          "author_url": "",
          "post_date": "2021-03-08T03:47:11.637000",
          "content": "<p>Thank you for sharing <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a>, at least I still have some hope with my model later after training<br>\nBut in your conclusion, is image captioning is the best approach to this competition?</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1230367,
          "author_name": "Sanyam Bhutani",
          "author_url": "",
          "post_date": "2021-03-08T04:34:22.113000",
          "content": "<blockquote>\n  <p>But in your conclusion, is image captioning is the best approach to this competition?</p>\n</blockquote>\n<p>I think we'll find out in a few weeks :)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1231080,
          "author_name": "Tensor Girl",
          "author_url": "",
          "post_date": "2021-03-08T16:50:15.027000",
          "content": "<p><a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> Congratulations on being first in LB and thank you for the mention . Glad you found my posts helpful</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1293314,
      "author_name": "Vadim Timakin",
      "author_url": "",
      "post_date": "2021-05-04T18:30:25.663000",
      "content": "<p>Hey everyone! What's your lowest score with the GRU decoder? How do you think, is it possible to get the gold with that?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1293405,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-05-04T20:04:25.430000",
          "content": "<p>Heng's theorem : \"Whatever LSTM(GRU) can do , Transformer can do it even better\" :)</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1293815,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-05-05T07:20:57.907000",
          "content": "<p>I still have not switched to a Transformer decoder and currently use LSTM. My impression is that it's definitely possible to get between 1.5 and 2.0 LB with LSTM, but going below 1.0, which will probably be a requirement for Gold, might not be feasible.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1294786,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-05-05T23:09:34.487000",
          "content": "<p>I am switching back and forth between a Pytorch Transformer and TF LSTM and will submit the one that win at the end ^^</p>\n<p>For now, I have better CV with Pytorch while using  smaller encoder and smaller resolution. </p>\n<p>But TF run incredibly fast on TPU even with huge model, letting room to explore new ideas. <br>\nUnfortunately some operations are not supported by TPU during the training loop and wrapping numpy functions doesn't work either. I had very hard time diving into TF documentations and code and hack around a bit to make them work. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1242789,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2021-03-17T19:54:39.973000",
      "content": "<p>Apart from the scores; can you guys share what type of augmentations you have used ( or can be used)<br>\nI think augmentation will be a vital part since we are working with Chemical molecules.<br>\nThanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1243804,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-03-18T13:57:39.663000",
          "content": "<p>for the moment I am not using anything special… you can score 3-4 using just this <code>augs</code> below</p>\n<pre><code>    train_transforms  = A.Compose([\n            A.RandomRotate90(p=.5),\n            A.Normalize(mean=[0.485, 0.456, 0.406],\n                        std=[0.229, 0.224, 0.225],\n                        always_apply=True),\n            AtoTensor()\n        ])\n</code></pre>",
          "votes": 13,
          "replies": []
        },
        {
          "id": 1243930,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2021-03-18T15:34:55.080000",
          "content": "<p>Thanks For replying;<br>\nFor normalizing should we use the mean, std for the given images, as these are very different from the imagenet data</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1243941,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-03-18T15:43:55.840000",
          "content": "<p>Based on my past experience (I might be wrong) its best to use <code>imagenet</code> if you are using pertained network.  If you are using  network from scratch then you have to use dataset statistics. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1243942,
          "author_name": "Nitin Datta",
          "author_url": "",
          "post_date": "2021-03-18T15:48:09.313000",
          "content": "<p>After a lot of experimentation I found 3 total types of augmentations can be helpful for this comp…</p>\n<ol>\n<li>Adding noise -- test images contain noise so maybe this has a good effect</li>\n<li>Improving image quality -- the letters are not clearly visible and also the lines used to create bonds are broken and  these types of augmentations help with this.  </li>\n<li>Rotating augmentations </li>\n</ol>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 1243943,
          "author_name": "Nitin Datta",
          "author_url": "",
          "post_date": "2021-03-18T15:50:36.527000",
          "content": "<p><a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> <code>you can score 3-4</code> by this you mean scoring on LB right??</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1243946,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-03-18T15:52:23.680000",
          "content": "<p>yes <code>LB</code> =)</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1236024,
      "author_name": "Andy Wilkinson",
      "author_url": "",
      "post_date": "2021-03-12T17:19:53.393000",
      "content": "<p>Looks like you are making good progress on the LB. 😃</p>\n<p>Out of interest, are you fine-tuning the ResNet, or have you frozen the ResNet layers? My hunch would be that the challenge dataset images are very different to pretraining images, and hence models might benefit significantly from the fine-tuning.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1236117,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-03-12T19:26:19.977000",
          "content": "<p>Your <code>hunch</code> is correct. Its good idea to also finetune your image <code>encoder</code> =)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1237076,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-03-13T18:01:50.357000",
          "content": "<p>Actually I gave wrong answer… I thought I was fine-tuning <code>encoder</code> as well .. but my teammate found a bug in my code where I was only training last layer of my <code>encoder</code>. I am rerunning now my training script. </p>\n<p>Coming back to your question. Our current score is still from<code>resnet34</code> but with trained last layer. I will update once I finish training full model. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1237086,
          "author_name": "Ren",
          "author_url": "",
          "post_date": "2021-03-13T18:25:04.227000",
          "content": "<p>The last layer be the Conv layer before pooling? Or are you projecting with a linear layer?</p>\n<p>I am training end to end. And last night I froze my CNN encoder and swapped on a new decoder, it trains really fast. After 1 epoch I am getting 11 on my local validation set. It seems the decoder part is likely much more important. </p>\n<p>Edit: probably not. I retrained end-to-end with some tweaks, 1 epoch can get validation distance to 11. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1237090,
          "author_name": "Andy Wilkinson",
          "author_url": "",
          "post_date": "2021-03-13T18:31:27.420000",
          "content": "<p>Thanks for the update. Interesting to hear that you've only been fine-tuning the last layer. Can't wait to see what scores you will get with full fine-tuning!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1237108,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-03-13T18:57:01.210000",
          "content": "<p>Yes correct the last <code>conv_block</code> before <code>pooling</code>. It will be interesting to see the difference … I assume the change might be small since early layers of model trained on imagenett are already really good at recognizing shapes (e.g. <code>lines</code>, 'circles`) </p>\n<p><a href=\"https://arxiv.org/pdf/1311.2901.pdf\" target=\"_blank\">https://arxiv.org/pdf/1311.2901.pdf</a><br>\n<a href=\"https://distill.pub/2020/circuits/zoom-in/\" target=\"_blank\">https://distill.pub/2020/circuits/zoom-in/</a></p>\n<p>experiments will show =)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1230768,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2021-03-08T12:51:12.797000",
      "content": "<p>Hello guys; I am also trying to solve it as an image captioning model type task, but after getting the image features from a backbone network I like to train through a transformer network;<br>\nSo can someone share any links/resources where I can train a Transformer from scratch? <br>\n(will fine-tuning any other language models like BERT, DistillBERT be of any use/ since they  are trained and have tokens for the English language)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1230843,
          "author_name": "Yanfei SIMP",
          "author_url": "",
          "post_date": "2021-03-08T14:03:13.217000",
          "content": "<p>I think there is people who share the transformer model for this on discussion forum, I think it was tensor girl</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1231079,
          "author_name": "Tensor Girl",
          "author_url": "",
          "post_date": "2021-03-08T16:49:14.587000",
          "content": "<p><a href=\"https://www.kaggle.com/houoinkyoma\" target=\"_blank\">@houoinkyoma</a> Thank you so much for the mention Uruha</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1232259,
          "author_name": "Kaushal Shah",
          "author_url": "",
          "post_date": "2021-03-09T15:41:22.117000",
          "content": "<p>This is great if you are using PyTorch: <a href=\"https://github.com/bentrevett/pytorch-seq2seq/blob/master/6%20-%20Attention%20is%20All%20You%20Need.ipynb\" target=\"_blank\">https://github.com/bentrevett/pytorch-seq2seq/blob/master/6%20-%20Attention%20is%20All%20You%20Need.ipynb</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1292463,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-05-04T01:47:40.530000",
      "content": "<p>i am still waiting for reports of non end-to-end solution.<br>\nThe LG SMILE (Dacon.ai from) top solution uses direct graph parsing from the image. (lstm-attention was ranked third)<br>\nI wonder has anyone tried methods like that?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1292482,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2021-05-04T02:20:19.730000",
          "content": "<p>Have you tried??</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1292594,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-05-04T06:11:51.577000",
          "content": "<p>trying now</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1292847,
          "author_name": "Matthieu Planté",
          "author_url": "",
          "post_date": "2021-05-04T10:42:59.540000",
          "content": "<p>I tried but for the moment I'm not able to get a Levenshtein distance under 0.4 on formula while the transformer encoder - transformer decoder goes under 0.1</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 1270959,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-12T07:26:37.330000",
      "content": "<p>anyone manages to make cnn-attention-lstm obtain LB score below 2.5 (without beam search)?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1270972,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-04-12T07:38:00.123000",
          "content": "<p>Do you get a LB score below 3.0 with cnn-attention-lstm (I just started the competition) ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1271017,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "2021-04-12T08:06:26.490000",
          "content": "<p>I got CV around 2.5 but didn't submit. I am moving now to transformer :) Increasing the hidden size dimension increases atleast CV.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1277705,
          "author_name": "Roc",
          "author_url": "",
          "post_date": "2021-04-19T05:46:36.690000",
          "content": "<p>effnet-b4-attention-lstm, input_size=380, cv=1.86,lb=2.5(without beam search).<br>\nThank you for your <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231190\" target=\"_blank\">share</a>,it help me a lot.</p>\n<p><a href=\"https://www.kaggle.com/yingpengchen/pl-bms-train\" target=\"_blank\">Training notebook</a><br>\n<a href=\"https://www.kaggle.com/yingpengchen/pl-bms-molecular-translation/edit/run/59116148\" target=\"_blank\">Inference notebook</a></p>",
          "votes": 4,
          "replies": [
            {
              "id": 1279613,
              "author_name": "",
              "author_url": "",
              "post_date": "2021-04-21T05:15:43.950000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1277714,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-19T06:04:58.390000",
          "content": "<p>thanks! I am surprised that attention-lstm can get below 2.5</p>\n<p>i also plugin effcientnet (and resnet) into trasnformer encoder. the results is better than lstm but worse than vision transformer encoder (which i don't know why).</p>\n<p>It seems that image size is very important</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1279324,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-04-20T19:34:16.470000",
          "content": "<p><a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> may I ask, how big that \"big gap…\" is?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1279336,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-04-20T19:38:35.820000",
          "content": "<p>was* I guess</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1279580,
          "author_name": "Tian",
          "author_url": "",
          "post_date": "2021-04-21T04:32:35.250000",
          "content": "<p>B4 + input 380…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1279612,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-21T05:13:20.560000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1260248,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-02T00:54:22.753000",
      "content": "<p>fyi:</p>\n<p>it is possible to predict the image orientation (rotate +/- 90 clockwise or 180) with 98%~99% using resnet34 on validation set</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1260580,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-04-02T08:36:10.970000",
          "content": "<p>The 1-2% \"error\" could contain images that are actually rotated ones.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1258104,
      "author_name": "Alex Vishnevskiy",
      "author_url": "",
      "post_date": "2021-03-31T10:59:37.003000",
      "content": "<p>Has anyone tried label smoothing? I know that in machine translation it matters.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1253185,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-26T12:58:53.030000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1231908,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-09T11:17:01.910000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1233297,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-10T09:34:22.270000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1233449,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-10T12:15:05.777000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1233459,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-10T12:41:08.850000",
          "content": "",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1233855,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-10T18:20:30.507000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1233865,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-10T18:33:18.687000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1236107,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-12T19:11:31.067000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1236108,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-12T19:14:33.237000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1236114,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-12T19:20:47.363000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1248043,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-22T09:40:30.953000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1248100,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-22T10:47:13.777000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1248121,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-22T11:10:37.780000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1248162,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-22T11:49:30.793000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1252830,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-26T05:10:43.993000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1253034,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-26T10:08:51.227000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1253145,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-26T12:01:08.140000",
          "content": "",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1253152,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-26T12:12:56.770000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1253182,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-26T12:53:51.880000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1253265,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-26T14:40:51.920000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1253273,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-26T14:47:57.880000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1253275,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-26T14:51:08.073000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1253282,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-26T14:57:55.290000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1253375,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-26T16:51:24.513000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1253599,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-26T22:03:59.450000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1253643,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-26T22:17:17.147000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1287811,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-29T12:25:18.267000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1231037,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-08T16:11:23.380000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1231085,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-08T16:56:42.283000",
          "content": "",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1231115,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-08T17:17:43.830000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1230777,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-08T12:57:35.847000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1230858,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-08T14:17:49.253000",
          "content": "",
          "votes": 10,
          "replies": []
        },
        {
          "id": 1231086,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-08T16:57:54.590000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1231098,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-08T17:05:19.280000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1231110,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-08T17:15:14.300000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1307632,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-14T15:09:30.087000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1307823,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-14T17:27:25.497000",
          "content": "",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1308403,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-15T07:15:43.793000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1308415,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-15T07:37:45.357000",
          "content": "",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1260793,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-02T12:16:24.657000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1260798,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-02T12:19:12.100000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1260834,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-02T12:48:43.637000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1258474,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-31T16:30:44.640000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1260701,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-02T10:55:40.123000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1260718,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-02T11:21:15.343000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1287808,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-29T12:23:52.047000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1279639,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-21T05:53:27.157000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1229719": "Just starting a common topic =)\n\n```\nmodel: resnet34 \ntrain_samples: 800k\nvalid_samples: 100k\nepoch: 5\nlr: 1e-3\nschedule: cosine\naugmentation: None\nsize: 128\nCV score: 15.7\nLB score: 26.7\n```",
    "1260670": "power of big encoder.\n\nuse CNN-attention-LSTM (1 layer) model, single fold\n\ntrain: input 224x224, no augmentation (i.e. rotation is not used)\ntest: use YNakama w<h rotate trick, greedy (argmax) decoder\ntokenizer: YNakama's Tokenizer\n\ntraining: train as long as my HW can support ... basically the below results are for epoch>10. results are not optimal, but i estimate within +/-1 optimality\n\nresnet34d: ~LB 9.5/CV 8.7\nresnet26d: ~LB 8.5/CV 6.8\nresnet101d:~LB 4.9/CV 4.0\nresnet200d:~LB 3.8/CV 3.2\n\n---\nbelow are the results for training in progress. no submission has been made\n\n- load  pretrained encoder  into transformer and finetune end-to-end:\n   - resnet26d : train/valid cross-entropy loss seems better than CNN-attention-LSTM-resnet101d and resnet200d\n   - hence, transformer is clearly the winner\n ",
    "1236521": "```\nencoder: resnet34 \ndecoder: LSTM with attention\ntrain_samples: 80%\nvalid_samples: 20%\nepoch: 1\naugmentation: None\nsize: 224\nCV: 24.1\n```\nTraining is done with kaggle notebooks, I will continue training and share notebooks as starter code.",
    "1268777": "I completed CNN+LSTM experiments and now are playing with vision transformer(essentially a transformer encoder)+transformer decoder. \n(also refer to my post for code and details: https://www.kaggle.com/c/bms-molecular-translation/discussion/231190)\n\nhere are the first results:\n\ntransformer in transformer:    \n     - TNT-S-224-16: LB 3.11/CV 1.96 (inference = 80 min on 4xti1080)\n     - TNT-S-320-16: LB 2.27/CV 1.55 (inference = 140 min on 4xti1080)\n\n \n\n(do you have others to suggest?)\n\n---\n\ntrain: input 224x224, no augmentation (i.e. rotation is not used)\ntest: use YNakama w<h rotate trick, greedy (argmax) decoder, aka. no beam search yet\ntokenizer: YNakama's Tokenizer\n\n---\n\na few quick observations:\n- transformer train faster, 2 days is enough\n- you actually don't need many epoch (7 to 10 is ok), but you do need many iterations and low learning rate at the end.\n(many iterations can mean you use a smaller batch size. i use 64)\n- cv/lb gap is larger (which i think if you use teacher distillation with CNN+LSTM, maybe you can get interesting results)\n\nfor more ideas: read this https://github.com/lucidrains/vit-pytorch",
    "1303020": "Hi there. We thought we would share some of our benchmarks as others have been so forthcoming. It's been a wonderful competition so far.\n\n---\n\n**SETUP:**\n    - EfficientNetB5 *(headless... no extra layers)*\n    - 256/768 (attn/lstm-units) LSTM\n    - 192x384 cropped/stretched/rotation-fixed images\n    - TFRecords and TPU\n    - Everything done on Kaggle in 1-2 sessions.\n\n---\n\n12 EPOCHS (5-6 hours of training) - LIMIT MAX LENGTH TO 140 TOKENS\n    - **1.95 CV** - No Beam Search - No Postprocessing\n    - **3.75 LB** - No Beam Search - No Postprocessing\n\n10 ADDITIONAL EPOCHS (5-6 hours of training) - FINE TUNE ON FULL LENGTH (277 TOKENS)\n    - **1.89 CV** - No Beam Search - No Postprocessing\n    - **2.68 LB** - No Beam Search - No Postprocessing\n    - **2.58 LB** - No Beam Search - RDKit Postprocessing\n\n---\n\nWe are implementing Beam Search and swapping in Transformers shortly and hopefully, that will help. At some point in the near future, we will scale up image size and model size and increase training time and try to squeeze below 1.",
    "1267440": "With the update of new rules, i want to update our score. Our current LB standing doesn’t break any rules. Best model has CV of 1.29 and LB of 1.34. I hope this will encourage participants to build model with high scores:) ",
    "1230093": "Just trained 1 epoch to verify my code works... e0+lstm. got 24 in my holdout set and 44 on the public. Quite a large gap.\n\nMar-10\nFixed some issue, and I am moving again. Finished another epoch and I got 19 local and 32 on the public. \n\n| local 5% hold out | public |\n| --- | --- |\n| 24 | 44 |\n| 19 | 32 |\n| 16 | 29 |\n| 13 | 27 |\n\nAbove table does not handle the different orient in test set. \n\nMar-17 Had some significant code changes and trained new model end to end from scratch, still e0 encoder lstm decoder. \n| epoch | local 20% hold out | public |\n| --- | --- | --- |\n| 3 | 5.4 | 5.9 |\n| 4 | 4.9 | 5.6 | \n| 5 | 4.6 | 5.2 |\n\nensemble 5 fold(all similar validation scores) of the above setup -> 4.11 LB. \nIt seems ensemble is much better than average of cv\n\nUsed rdkit to normalize the 4.11 predictions, score improved to 4.07. It seems there is some juice worth squeezing. ",
    "1253530": "resnet34, 10 epochs -> single fold validation: 8.27, lb: 8.87",
    "1265529": "- Encoder: Pruned EfficientNetB3\n- Decoder: GRU\n- Image size: `224`\n- Augmentation: None\n- CV: `2.79`\n- CV with beam search (k=5): `2.65`\n- LB without beam search: `3.57`\n- LB with beam search (k=5): `3.22`\n\n",
    "1230112": "Great job on pushing the leaderboard forward. I'm still working on my first submission, hopefully I'll have the basis for my model setup in a couple days.",
    "1267645": "I just used some public kernels and got CV 3.xx with B0 without any change in the decoder.\nPersonally, I think finding a good training strategy with suitable batch size and LR scheduler is essential.\n P/S: My hardware resource is limited.  ",
    "1267208": "Encoder: eca_nfnet_l0\nDecoder: LSTM+GRU\nImage size: 224\nBS=128\nscheduler='CosineAnnealingLR'\nAugmentation: None\nCV: 4.61\nLB: 5.85\n\n1 epoch takes 5h on my local machine.",
    "1291931": "Let's re-activate cv/lb thread.\n\nCV:1.17 -> LB:1.49\nCV:1.10 -> LB:1.42\nSeems stable for me, but a little big discrepancy.\n\nDo you have any updates, @drhabib ??\nYou should;)",
    "1257282": "efficientnet_b0 as encoder, 5 epochs, CV: 7.21, LB: 7.76 using around 50% of the train data and no augmentations. For now wondering what's going to make difference, because usually model really slows down learning after that point...",
    "1253878": "Anyone using **Beam Search** for prediction? Does it make a huge difference in your predictions instead of simple argmax at each timestep?\nI wrote a beam search function but the problem is that it is not batch-wise and takes a lot of time. I also did not find an efficient batch-wise beam search online.",
    "1236612": "Anyone successful to use transformer as a decoder?",
    "1234172": "Surprising that baseline model can perform so well.",
    "1232308": "Interesting! Are you using attention during the decoding? I'm curious how useful that would be.",
    "1230221": "I'm waiting an end to end transformer model in this discussion :)",
    "1260240": "use batchwise beam search?\n\nhttps://stackoverflow.com/questions/64356953/batch-wise-beam-search-in-pytorch\nhttps://zhuanlan.zhihu.com/p/167072494\n\nhttps://medium.com/the-artificial-impostor/implementing-beam-search-part-1-4f53482daabe\nHow to Do Beam Search Efficiently\n\"The smarter way is to put these N nodes into a batch and feed it to the model. It’ll be much faster since we can parallelize the computation, especially when using a CUDA backend.\"\n\n---\n\nif you train multiple models (e.g. improved models as you make more submissions or for ensemble), these models can provide a hypothesis to reduce the time for beam search (i.e. they provide some clue on the top-k decoding path)",
    "1237838": "end to end model: CNN + RNN\nattention: None (don't know how to implement yet)\ntrain / valid: 80% / 20%\nepoch: 18\nCV: 10.8\nLB: 11.6",
    "1229749": "Hi @drhabib thanks for sharing your baseline results and welcome to the competition. Are you training a Image Captioning with resnet as backbone ? I am trying to understand how you have trained resnet \nThanks",
    "1293314": "Hey everyone! What's your lowest score with the GRU decoder? How do you think, is it possible to get the gold with that?",
    "1242789": "Apart from the scores; can you guys share what type of augmentations you have used ( or can be used)\nI think augmentation will be a vital part since we are working with Chemical molecules.\nThanks",
    "1236024": "Looks like you are making good progress on the LB. 😃\n\nOut of interest, are you fine-tuning the ResNet, or have you frozen the ResNet layers? My hunch would be that the challenge dataset images are very different to pretraining images, and hence models might benefit significantly from the fine-tuning.",
    "1230768": "Hello guys; I am also trying to solve it as an image captioning model type task, but after getting the image features from a backbone network I like to train through a transformer network;\nSo can someone share any links/resources where I can train a Transformer from scratch? \n(will fine-tuning any other language models like BERT, DistillBERT be of any use/ since they  are trained and have tokens for the English language)",
    "1292463": "i am still waiting for reports of non end-to-end solution.\nThe LG SMILE (Dacon.ai from) top solution uses direct graph parsing from the image. (lstm-attention was ranked third)\nI wonder has anyone tried methods like that?",
    "1270959": "anyone manages to make cnn-attention-lstm obtain LB score below 2.5 (without beam search)?",
    "1260248": "fyi:\n\nit is possible to predict the image orientation (rotate +/- 90 clockwise or 180) with 98%~99% using resnet34 on validation set",
    "1258104": "Has anyone tried label smoothing? I know that in machine translation it matters.",
    "1253185": "My architecture is different form public kernels.\n\nSingle fold validation: 5.50\nLB : 6.84\n\nI don't have descent hardware and patience to wait for 5 folds for now ^^\n\nAnyway I think there is still room for improvement in this fold. I stopped  training  just to check  my  submission",
    "1231908": "My first CV-LB: 18.44-23.7",
    "1231037": "Has your CV-LB gap changed since then?",
    "1230777": "Amazing accuracy! This model performance has surprised many cheminformaticians! Just curious, if you run prediction on the rest of training set, what's the distribution of the levenshtein distances, and what's the percentage of perfect predictions (distance = 0)?",
    "1307632": "Can anypne explain what is rdkit post-processing.",
    "1260793": "I used Efficientnet B0 as encoder.\nUsing about half of the train images\nI went for 6 epochs and I'm at 10.4 on validation set\nI have not submitted it yet because I'm out of GPU time :)\n\nBut the main problem is that the loss does not seem to decrease anymore and training has literally stopped\nI think next week I'm gonna try another encoder out!!\nEdit: 10.4 Valid/ 11.78 LB",
    "1258474": "transformer-based model. CV: 4.37 and LB: 6.17. I think the model is still underfitted even I trained for 50 epochs.",
    "1287808": "",
    "1279639": ""
  }
}