{
  "id": 231190,
  "title": "[completed] transformer starter kit ... in pytorch",
  "url": "/competitions/bms-molecular-translation/discussion/231190",
  "author_name": "hengck23",
  "post_date": "2021-04-07T11:55:28.548000",
  "votes": 179,
  "comment_count": 216,
  "views": 0,
  "content": "<p></p>\n<p></p>\n<p>speed now is  \"using fairseq+jit\", you can do inference at 2 min for 10_000 images with single GPU.</p>\n<ul>\n<li>224x224 images</li>\n<li>pre-norm activated trnasformer are used. transformer layer are coded from scrtach.</li>\n<li>resnet101d+transformer  : performance: 3.3 CV with 40_000 images , in 78 min </li>\n<li>resnet26d+transformer  :  performance: ~3.7 CV  estimated</li>\n<li>[2] resnet26d+attention+LSTM and [1]resnet26d+LSTM are also included for your study and experiment</li>\n<li>train/validation log files, loss curves and intermediate trained models are included</li>\n<li>i first trained the image encoder using attention+lstm. the pretained model is then used in the transformer to save time. i not sure if this affect performance</li>\n<li>no augmentation in training.</li>\n<li>rotation prediction in inference</li>\n</ul>\n<p>[1] Show and Tell: A Neural Image Caption Generator<br>\n<a href=\"https://arxiv.org/abs/1411.4555\" target=\"_blank\">https://arxiv.org/abs/1411.4555</a></p>\n<p>[2] Show, Attend and Tell: Neural Image Caption Generation with Visual Attention<br>\n<a href=\"https://arxiv.org/abs/1502.03044\" target=\"_blank\">https://arxiv.org/abs/1502.03044</a></p>\n<hr>\n<p>all files are at google drive: <a href=\"https://drive.google.com/drive/folders/1dTfmZxDkDkrnRzz5DOkPOYr9oDIBcrq5?usp=sharing\" target=\"_blank\">https://drive.google.com/drive/folders/1dTfmZxDkDkrnRzz5DOkPOYr9oDIBcrq5?usp=sharing</a></p>\n<p>[2021-apr-06]<br>\n <img src=\"https://i.ibb.co/wzjSzWf/Selection-046.png\" alt=\"\"></p>\n<ul>\n<li>please refer to readme.ppt to run the software and experiments</li>\n</ul>\n<p>[2021-apr-06a]</p>\n<ul>\n<li>dirty code for torch.jit.script model</li>\n<li>if i sort the validation samples by length, i can complete 40_000 test samples in 18 min</li>\n</ul>\n<p>[2021-apr-07a]</p>\n<ul>\n<li>code to migrate from my transformer to fairseq transformer.</li>\n<li>using fairseq+jit, you can do interence at 2 min for 10_000 images.</li>\n<li>fairseq uses cache for key, values.</li>\n</ul>\n<p>[2021-apr-07b]</p>\n<ul>\n<li>code to to train TNT(transformer in transformer) image encoder+ token transformer decoder</li>\n<li>80 min to do inference for 1.6 millions image with 4xTi1080. CV =1.9, LB =3.1</li>\n<li>use fairseq API</li>\n<li>you can modify 224 input to 320 to get CV=1.4, LB2.0 !!!</li>\n</ul>\n<p>[2021-apr-24]</p>\n<ul>\n<li>implement patch-based input for transformer encoder. Please referto th PPT for deails.</li>\n<li>you learn how to prepare and store image as patches</li>\n<li>how to modify image transformer to accept variable-length input of patches. How to modify position encoding for patch input.</li>\n<li>how to set the mask values for transformer encoder and decoder</li>\n<li>a small trained model (using 0.8 input scale) and training log is provided. the results of patched based image transformer is the similar to image one, but running much faster (CV-teacher forcing: 1.27, CV-without teacher forcing : about 1.35 for 40000 validation set). time: 5.5 min for 40000 validation set, 55 min for 400000 test images (25% of all). LB = 2.01</li>\n<li>there is no submission code or jit inference. This is left as an exercise for you. It is easy to modify the code</li>\n</ul>\n<p>[bug] as i was running the submission code, i note that the some test images have larger num of patch than the train, this affects the positional encoding for the input patch.</p>\n<p>[2021-apr-25]</p>\n<ul>\n<li>model py file for the original VIT vision transformer. you can use this to replace the TNT. There is no pretrain model for deep TNT. But there are larger and deeper pretrain model for VIT</li>\n</ul>\n<hr>\n<p>Note:  in some way, it is similar to this keras code:<br>\n<a href=\"https://www.kaggle.com/aditya08/imagecaptioning-show-attnd-tell-w-transformer\" target=\"_blank\">https://www.kaggle.com/aditya08/imagecaptioning-show-attnd-tell-w-transformer</a></p>",
  "messages": [
    {
      "id": 1265992,
      "postDate": "2021-04-07T11:55:28.550Z",
      "content": "<p></p>\n<p></p>\n<p>speed now is  \"using fairseq+jit\", you can do inference at 2 min for 10_000 images with single GPU.</p>\n<ul>\n<li>224x224 images</li>\n<li>pre-norm activated trnasformer are used. transformer layer are coded from scrtach.</li>\n<li>resnet101d+transformer  : performance: 3.3 CV with 40_000 images , in 78 min </li>\n<li>resnet26d+transformer  :  performance: ~3.7 CV  estimated</li>\n<li>[2] resnet26d+attention+LSTM and [1]resnet26d+LSTM are also included for your study and experiment</li>\n<li>train/validation log files, loss curves and intermediate trained models are included</li>\n<li>i first trained the image encoder using attention+lstm. the pretained model is then used in the transformer to save time. i not sure if this affect performance</li>\n<li>no augmentation in training.</li>\n<li>rotation prediction in inference</li>\n</ul>\n<p>[1] Show and Tell: A Neural Image Caption Generator<br>\n<a href=\"https://arxiv.org/abs/1411.4555\" target=\"_blank\">https://arxiv.org/abs/1411.4555</a></p>\n<p>[2] Show, Attend and Tell: Neural Image Caption Generation with Visual Attention<br>\n<a href=\"https://arxiv.org/abs/1502.03044\" target=\"_blank\">https://arxiv.org/abs/1502.03044</a></p>\n<hr>\n<p>all files are at google drive: <a href=\"https://drive.google.com/drive/folders/1dTfmZxDkDkrnRzz5DOkPOYr9oDIBcrq5?usp=sharing\" target=\"_blank\">https://drive.google.com/drive/folders/1dTfmZxDkDkrnRzz5DOkPOYr9oDIBcrq5?usp=sharing</a></p>\n<p>[2021-apr-06]<br>\n <img src=\"https://i.ibb.co/wzjSzWf/Selection-046.png\" alt=\"\"></p>\n<ul>\n<li>please refer to readme.ppt to run the software and experiments</li>\n</ul>\n<p>[2021-apr-06a]</p>\n<ul>\n<li>dirty code for torch.jit.script model</li>\n<li>if i sort the validation samples by length, i can complete 40_000 test samples in 18 min</li>\n</ul>\n<p>[2021-apr-07a]</p>\n<ul>\n<li>code to migrate from my transformer to fairseq transformer.</li>\n<li>using fairseq+jit, you can do interence at 2 min for 10_000 images.</li>\n<li>fairseq uses cache for key, values.</li>\n</ul>\n<p>[2021-apr-07b]</p>\n<ul>\n<li>code to to train TNT(transformer in transformer) image encoder+ token transformer decoder</li>\n<li>80 min to do inference for 1.6 millions image with 4xTi1080. CV =1.9, LB =3.1</li>\n<li>use fairseq API</li>\n<li>you can modify 224 input to 320 to get CV=1.4, LB2.0 !!!</li>\n</ul>\n<p>[2021-apr-24]</p>\n<ul>\n<li>implement patch-based input for transformer encoder. Please referto th PPT for deails.</li>\n<li>you learn how to prepare and store image as patches</li>\n<li>how to modify image transformer to accept variable-length input of patches. How to modify position encoding for patch input.</li>\n<li>how to set the mask values for transformer encoder and decoder</li>\n<li>a small trained model (using 0.8 input scale) and training log is provided. the results of patched based image transformer is the similar to image one, but running much faster (CV-teacher forcing: 1.27, CV-without teacher forcing : about 1.35 for 40000 validation set). time: 5.5 min for 40000 validation set, 55 min for 400000 test images (25% of all). LB = 2.01</li>\n<li>there is no submission code or jit inference. This is left as an exercise for you. It is easy to modify the code</li>\n</ul>\n<p>[bug] as i was running the submission code, i note that the some test images have larger num of patch than the train, this affects the positional encoding for the input patch.</p>\n<p>[2021-apr-25]</p>\n<ul>\n<li>model py file for the original VIT vision transformer. you can use this to replace the TNT. There is no pretrain model for deep TNT. But there are larger and deeper pretrain model for VIT</li>\n</ul>\n<hr>\n<p>Note:  in some way, it is similar to this keras code:<br>\n<a href=\"https://www.kaggle.com/aditya08/imagecaptioning-show-attnd-tell-w-transformer\" target=\"_blank\">https://www.kaggle.com/aditya08/imagecaptioning-show-attnd-tell-w-transformer</a></p>",
      "rawMarkdown": "~~this is my terribly slow implementation, i couldn't make a submission yet. Anyway, it shows how to use transformer for a beginner ...~~\n\n\n~~i manage to speed up.  resnet101d+transformer  achieves LB of 3.92, taking 3h to predict 1.6 million test images with 4x Ti1080. The local CV is 3.19 (i train more iterations than the intermediate models in google drive)~~\n\nspeed now is  \"using fairseq+jit\", you can do inference at 2 min for 10_000 images with single GPU.\n\n- 224x224 images\n- pre-norm activated trnasformer are used. transformer layer are coded from scrtach.\n- resnet101d+transformer  : performance: 3.3 CV with 40_000 images , in 78 min \n- resnet26d+transformer  :  performance: ~3.7 CV  estimated\n- [2] resnet26d+attention+LSTM and [1]resnet26d+LSTM are also included for your study and experiment\n- train/validation log files, loss curves and intermediate trained models are included\n- i first trained the image encoder using attention+lstm. the pretained model is then used in the transformer to save time. i not sure if this affect performance\n- no augmentation in training.\n- rotation prediction in inference\n\n[1] Show and Tell: A Neural Image Caption Generator\nhttps://arxiv.org/abs/1411.4555\n\n[2] Show, Attend and Tell: Neural Image Caption Generation with Visual Attention\nhttps://arxiv.org/abs/1502.03044\n\n---\nall files are at google drive: https://drive.google.com/drive/folders/1dTfmZxDkDkrnRzz5DOkPOYr9oDIBcrq5?usp=sharing\n\n\n[2021-apr-06]\n ![](https://i.ibb.co/wzjSzWf/Selection-046.png)\n\n- please refer to readme.ppt to run the software and experiments\n\n\n\n[2021-apr-06a]\n- dirty code for torch.jit.script model\n- if i sort the validation samples by length, i can complete 40_000 test samples in 18 min\n\n\n[2021-apr-07a]\n- code to migrate from my transformer to fairseq transformer.\n- using fairseq+jit, you can do interence at 2 min for 10_000 images.\n- fairseq uses cache for key, values.\n\n\n\n[2021-apr-07b]\n- code to to train TNT(transformer in transformer) image encoder+ token transformer decoder\n- 80 min to do inference for 1.6 millions image with 4xTi1080. CV =1.9, LB =3.1\n- use fairseq API\n- you can modify 224 input to 320 to get CV=1.4, LB2.0 !!!\n\n\n\n[2021-apr-24]\n- implement patch-based input for transformer encoder. Please referto th PPT for deails.\n- you learn how to prepare and store image as patches\n- how to modify image transformer to accept variable-length input of patches. How to modify position encoding for patch input.\n- how to set the mask values for transformer encoder and decoder\n- a small trained model (using 0.8 input scale) and training log is provided. the results of patched based image transformer is the similar to image one, but running much faster (CV-teacher forcing: 1.27, CV-without teacher forcing : about 1.35 for 40000 validation set). time: 5.5 min for 40000 validation set, 55 min for 400000 test images (25% of all). LB = 2.01\n- there is no submission code or jit inference. This is left as an exercise for you. It is easy to modify the code\n\n[bug] as i was running the submission code, i note that the some test images have larger num of patch than the train, this affects the positional encoding for the input patch.\n\n\n[2021-apr-25]\n- model py file for the original VIT vision transformer. you can use this to replace the TNT. There is no pretrain model for deep TNT. But there are larger and deeper pretrain model for VIT\n\n---\n\n\nNote:  in some way, it is similar to this keras code:\nhttps://www.kaggle.com/aditya08/imagecaptioning-show-attnd-tell-w-transformer",
      "votes": 177
    },
    {
      "id": 1280094,
      "postDate": "2021-04-21T14:54:52.107Z",
      "content": "<p><br>\nalready implemented (version 2021-apr-24)</p>\n<p>forget about image size!<br>\nwe use non-empty patch as input (i.e. variable input length). empty image patch are discarded</p>\n<p><img src=\"https://i.ibb.co/nQqfV0S/Selection-104.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/NjVv60V/Selection-106.png\" alt=\"\"></p>",
      "rawMarkdown": "~~next version is coming soon!~~\nalready implemented (version 2021-apr-24)\n\n\nforget about image size!\nwe use non-empty patch as input (i.e. variable input length). empty image patch are discarded\n\n![](https://i.ibb.co/nQqfV0S/Selection-104.png)\n![](https://i.ibb.co/NjVv60V/Selection-106.png)",
      "votes": 14,
      "replies": [
        {
          "id": 1280103,
          "postDate": "2021-04-21T15:06:25.563Z",
          "content": "<p>maybe next,next version?<br>\n<img src=\"https://i.ibb.co/g7d5nMh/Selection-107.png\" alt=\"\"></p>\n<p>anyone can recommend bi-translation paper? (e.g. a model that can translate FR to EN and also EN to FR)</p>\n<p><a href=\"https://arxiv.org/pdf/1805.11213.pdf\" target=\"_blank\">https://arxiv.org/pdf/1805.11213.pdf</a><br>\n\"Our technique trains a single model for both directions of a language pair, allowing us<br>\nto back-translate source or target monolingual data without requiring an auxiliary<br>\nmodel. \"</p>\n<hr>\n<p>back translation is a data augmentation method: <a href=\"https://amitness.com/back-translation/\" target=\"_blank\">https://amitness.com/back-translation/</a><br>\nback translation can also be used as self-supervised or semi-supervsied learning</p>",
          "rawMarkdown": "maybe next,next version?\n![](https://i.ibb.co/g7d5nMh/Selection-107.png)\n\nanyone can recommend bi-translation paper? (e.g. a model that can translate FR to EN and also EN to FR)\n\nhttps://arxiv.org/pdf/1805.11213.pdf\n\"Our technique trains a single model for both directions of a language pair, allowing us\nto back-translate source or target monolingual data without requiring an auxiliary\nmodel. \"\n\n---\n\nback translation is a data augmentation method: https://amitness.com/back-translation/\nback translation can also be used as self-supervised or semi-supervsied learning",
          "votes": 2
        },
        {
          "id": 1280118,
          "postDate": "2021-04-21T15:25:59.227Z",
          "content": "<p>Can you then please post your CV/LB performance for this technique? Would be interested if it can keep up. And your iteration speed improvements, as that was your main goal as you have stated in a past comment.</p>\n<p>Interesting fact: The length distribution over the image patches is quite similar to the length distribution of my InChI tokenization. Only the long tail is quite longer and the mean is off by 5.<br>\nSo we have direct, decent correlation between effective molecule size and InChI length. Makes sense.<br>\nAs you have the data you could plot this correlation, if your are interested.</p>",
          "rawMarkdown": "Can you then please post your CV/LB performance for this technique? Would be interested if it can keep up. And your iteration speed improvements, as that was your main goal as you have stated in a past comment.\n\nInteresting fact: The length distribution over the image patches is quite similar to the length distribution of my InChI tokenization. Only the long tail is quite longer and the mean is off by 5.\nSo we have direct, decent correlation between effective molecule size and InChI length. Makes sense.\nAs you have the data you could plot this correlation, if your are interested.",
          "votes": 1
        },
        {
          "id": 1280382,
          "postDate": "2021-04-21T21:18:39.973Z",
          "content": "<p>very interesting idea <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, thanks!<br>\nI am interested in what order are you going to feed patches into encoder?<br>\nor this does not much matter?</p>",
          "rawMarkdown": "very interesting idea @hengck23, thanks!\nI am interested in what order are you going to feed patches into encoder?\nor this does not much matter?"
        },
        {
          "id": 1280415,
          "postDate": "2021-04-21T23:32:22.590Z",
          "content": "<p>it doesn't matter for transformer as it uses positional encoding.<br>\nin fact, it isn't feed in one by one. it is feed in all at once</p>",
          "rawMarkdown": "it doesn't matter for transformer as it uses positional encoding.\nin fact, it isn't feed in one by one. it is feed in all at once",
          "votes": 1
        },
        {
          "id": 1283267,
          "postDate": "2021-04-24T18:46:49.060Z",
          "content": "<p>updated!<br>\n[2021-apr-24]<br>\nimplement patch-based input for transformer encoder. Please refer to the PPT for details.</p>\n<p>please refer to the comments in the main top window.</p>",
          "rawMarkdown": "updated!\n[2021-apr-24]\nimplement patch-based input for transformer encoder. Please refer to the PPT for details.\n\nplease refer to the comments in the main top window.",
          "votes": 5
        },
        {
          "id": 1283327,
          "postDate": "2021-04-24T20:19:10.793Z",
          "content": "<p>Can you tell us what is the LB score for the [2021-apr-24] patch based model ?</p>",
          "rawMarkdown": "Can you tell us what is the LB score for the [2021-apr-24] patch based model ?"
        },
        {
          "id": 1283606,
          "postDate": "2021-04-25T04:59:24.790Z",
          "content": "<p>i update the results. LB = 2.01</p>",
          "rawMarkdown": "i update the results. LB = 2.01"
        }
      ]
    },
    {
      "id": 1272273,
      "postDate": "2021-04-13T11:25:54.260Z",
      "content": "<p>experiment update:<br>\n<img src=\"https://i.ibb.co/9qcKLQy/Selection-029.png\" alt=\"\"></p>\n<p>updated with new results (LB=2.06 for latest patch+coord input, version 2021-apr-24)</p>\n<p><img src=\"https://i.ibb.co/4dZmVVj/Selection-129.png\" alt=\"\"></p>",
      "rawMarkdown": "experiment update:\n![](https://i.ibb.co/9qcKLQy/Selection-029.png)\n\nupdated with new results (LB=2.06 for latest patch+coord input, version 2021-apr-24)\n\n![](https://i.ibb.co/4dZmVVj/Selection-129.png)",
      "votes": 10,
      "replies": [
        {
          "id": 1272484,
          "postDate": "2021-04-13T14:11:02.707Z",
          "content": "<p>Can I ask <br>\nchanging img_size out in the below code to 320 will make it train with image size from 224 to 320.</p>\n<pre><code>class TNT(nn.Module):\n    \"\"\" Transformer in Transformer - https://arxiv.org/abs/2103.00112\n    \"\"\"\n\n    def __init__(self, img_size=224, patch_size=16, in_chans=3, num_classes=1000, embed_dim=768, in_dim=48, depth=12,\n                 num_heads=12, in_num_head=4, mlp_ratio=4., qkv_bias=False, drop_rate=0., attn_drop_rate=0.,\n                 drop_path_rate=0., norm_layer=nn.LayerNorm, first_stride=4):\n</code></pre>",
          "rawMarkdown": "Can I ask \nchanging img_size out in the below code to 320 will make it train with image size from 224 to 320.\n\n```\nclass TNT(nn.Module):\n    \"\"\" Transformer in Transformer - https://arxiv.org/abs/2103.00112\n    \"\"\"\n\n    def __init__(self, img_size=224, patch_size=16, in_chans=3, num_classes=1000, embed_dim=768, in_dim=48, depth=12,\n                 num_heads=12, in_num_head=4, mlp_ratio=4., qkv_bias=False, drop_rate=0., attn_drop_rate=0.,\n                 drop_path_rate=0., norm_layer=nn.LayerNorm, first_stride=4):\n```"
        },
        {
          "id": 1273217,
          "postDate": "2021-04-14T07:06:59.550Z",
          "content": "<p><img src=\"https://i.ibb.co/vzLgnpk/Selection-028.png\" alt=\"\"></p>",
          "rawMarkdown": "![](https://i.ibb.co/vzLgnpk/Selection-028.png)",
          "votes": 4
        },
        {
          "id": 1274122,
          "postDate": "2021-04-15T02:37:00.163Z",
          "content": "<p>That's why \"a picture is worth a thousand words\"……thank you Heng!</p>",
          "rawMarkdown": "That's why \"a picture is worth a thousand words\"......thank you Heng!",
          "votes": 1
        },
        {
          "id": 1278732,
          "postDate": "2021-04-20T08:30:01.263Z",
          "content": "<p>Great work. Could you please share how many epochs averagely need to be run to reach those scores?</p>",
          "rawMarkdown": "Great work. Could you please share how many epochs averagely need to be run to reach those scores?"
        },
        {
          "id": 1279999,
          "postDate": "2021-04-21T13:07:35.157Z",
          "content": "<p>deleted, :D</p>",
          "rawMarkdown": "deleted, :D"
        },
        {
          "id": 1306975,
          "postDate": "2021-05-14T07:18:45.823Z",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  was that the only line needed to get diff image sizes to work? </p>",
          "rawMarkdown": "@morizin @hengck23  was that the only line needed to get diff image sizes to work? "
        },
        {
          "id": 1307464,
          "postDate": "2021-05-14T13:03:15.563Z",
          "content": "<p>i think yes <a href=\"https://www.kaggle.com/trushk\" target=\"_blank\">@trushk</a> </p>",
          "rawMarkdown": "i think yes @trushk "
        },
        {
          "id": 1308673,
          "postDate": "2021-05-15T11:25:42.990Z",
          "content": "<p>what's the df_fold.fine.csv? thank Heng!</p>",
          "rawMarkdown": "what's the df_fold.fine.csv? thank Heng!"
        }
      ]
    },
    {
      "id": 1268242,
      "postDate": "2021-04-09T08:28:44.973Z",
      "content": "<p>What do you want to know about transformer application in this competition?</p>\n<p>I see some posts that newbies want to know more about transformers.<br>\nif you can specifically post your question here, maybe i can organize a talk on it (e.g. zoom video) or include some of the problems and solution in my starter kit</p>\n<p>please share your thoughts here!</p>",
      "rawMarkdown": "What do you want to know about transformer application in this competition?\n\nI see some posts that newbies want to know more about transformers.\nif you can specifically post your question here, maybe i can organize a talk on it (e.g. zoom video) or include some of the problems and solution in my starter kit\n\nplease share your thoughts here!",
      "votes": 10,
      "replies": [
        {
          "id": 1285213,
          "postDate": "2021-04-26T17:31:59.080Z",
          "content": "<p>I would definitely be interested in attending a talk from you. For me though, I'm okay with transformers. I'm more interested in your overall knowledge of toolkits and optimisation. This is the first time I've seen torch jit and the first time I've seen fairseq. If you want a more concrete request, I'd be happy to watch a talk solely on jit. What is it in general? How does it apply to torch? How do you use it?</p>",
          "rawMarkdown": "I would definitely be interested in attending a talk from you. For me though, I'm okay with transformers. I'm more interested in your overall knowledge of toolkits and optimisation. This is the first time I've seen torch jit and the first time I've seen fairseq. If you want a more concrete request, I'd be happy to watch a talk solely on jit. What is it in general? How does it apply to torch? How do you use it?",
          "votes": 1
        },
        {
          "id": 1285260,
          "postDate": "2021-04-26T18:08:23.747Z",
          "content": "<p>Another concrete thing I'd like to learn about is how you organise your code. With many newer Kagglers used to using a notebook, your code base can look quite daunting. I'm not really a \"new\" Kaggler and I'm bewildered by it. In my head it's because \"that's how the real pros do it\".</p>",
          "rawMarkdown": "Another concrete thing I'd like to learn about is how you organise your code. With many newer Kagglers used to using a notebook, your code base can look quite daunting. I'm not really a \"new\" Kaggler and I'm bewildered by it. In my head it's because \"that's how the real pros do it\"."
        },
        {
          "id": 1285523,
          "postDate": "2021-04-27T03:38:55.357Z",
          "content": "<p><a href=\"https://www.kaggle.com/alexandersoare\" target=\"_blank\">@alexandersoare</a>,<br>\nTotally agree with you. If more people are interested we can conduct a zoom call/ Google meet and learn a lot of techniques from him. </p>",
          "rawMarkdown": "@alexandersoare,\nTotally agree with you. If more people are interested we can conduct a zoom call/ Google meet and learn a lot of techniques from him. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1284482,
      "postDate": "2021-04-26T01:26:21.667Z",
      "content": "<p>a few good open source to generate new train data by using random perturbation</p>\n<p>ICLR-2020 poster paper \"Augmenting Genetic Algorithms with Deep Neural Networks for exploring the chemical space\"<br>\n<a href=\"https://www.youtube.com/watch?v=9VilhlEXm9w&amp;t=16s\" target=\"_blank\">https://www.youtube.com/watch?v=9VilhlEXm9w&amp;t=16s</a></p>\n<p>EvoMol: a flexible and interpretable evolutionary algorithm for unbiased de novo molecular generation. J Cheminform 12, 55 (2020)<br>\n<a href=\"https://jcheminf.biomedcentral.com/articles/10.1186/s13321-020-00458-z\" target=\"_blank\">https://jcheminf.biomedcentral.com/articles/10.1186/s13321-020-00458-z</a><br>\n<a href=\"https://github.com/jules-leguy/EvoMol\" target=\"_blank\">https://github.com/jules-leguy/EvoMol</a></p>\n<p><img src=\"https://github.com/jules-leguy/EvoMol/raw/master/examples/figures/detailed_expl_tree.png\" alt=\"\"></p>\n<p>using GAN:<br>\n<a href=\"https://github.com/ardigen/mol-cycle-gan\" target=\"_blank\">https://github.com/ardigen/mol-cycle-gan</a></p>",
      "rawMarkdown": "a few good open source to generate new train data by using random perturbation\n\n ICLR-2020 poster paper \"Augmenting Genetic Algorithms with Deep Neural Networks for exploring the chemical space\"\nhttps://www.youtube.com/watch?v=9VilhlEXm9w&t=16s\n\nEvoMol: a flexible and interpretable evolutionary algorithm for unbiased de novo molecular generation. J Cheminform 12, 55 (2020)\nhttps://jcheminf.biomedcentral.com/articles/10.1186/s13321-020-00458-z\nhttps://github.com/jules-leguy/EvoMol\n\n\n\n![](https://github.com/jules-leguy/EvoMol/raw/master/examples/figures/detailed_expl_tree.png)\n\nusing GAN:\nhttps://github.com/ardigen/mol-cycle-gan",
      "votes": 8,
      "replies": [
        {
          "id": 1291396,
          "postDate": "2021-05-03T02:34:54.820Z",
          "content": "<p>for the adventurous<br>\nOptimization of Molecules via Deep Reinforcement Learning<br>\n<a href=\"https://www.nature.com/articles/s41598-019-47148-x\" target=\"_blank\">https://www.nature.com/articles/s41598-019-47148-x</a></p>\n<p><img src=\"https://www.researchgate.net/publication/334663607/figure/fig2/AS:784360231948290@1564017460710/Sample-molecules-in-the-property-optimization-task-a-Optimization-of-penalized-logP.png\" alt=\"\"></p>",
          "rawMarkdown": "for the adventurous\nOptimization of Molecules via Deep Reinforcement Learning\nhttps://www.nature.com/articles/s41598-019-47148-x\n\n![](https://www.researchgate.net/publication/334663607/figure/fig2/AS:784360231948290@1564017460710/Sample-molecules-in-the-property-optimization-task-a-Optimization-of-penalized-logP.png)"
        },
        {
          "id": 1291397,
          "postDate": "2021-05-03T02:36:31.560Z",
          "content": "<p>the simple competition may end up me learning transformer, image caption, attention-lstm, … and even RL</p>",
          "rawMarkdown": "the simple competition may end up me learning transformer, image caption, attention-lstm, ... and even RL"
        },
        {
          "id": 1295067,
          "postDate": "2021-05-06T07:15:11.090Z",
          "content": "<p>Wait, so is this allowed? Are we not meant to only stick to the provided InChIs?</p>",
          "rawMarkdown": "Wait, so is this allowed? Are we not meant to only stick to the provided InChIs?"
        },
        {
          "id": 1295275,
          "postDate": "2021-05-06T10:52:35.910Z",
          "content": "<p>i am using provided InChIs to create more InchIs</p>",
          "rawMarkdown": "i am using provided InChIs to create more InchIs"
        },
        {
          "id": 1295365,
          "postDate": "2021-05-06T11:57:27.910Z",
          "content": "<p>Yes, that's what I'm concerned about. Do you know for sure that is allowed?</p>",
          "rawMarkdown": "Yes, that's what I'm concerned about. Do you know for sure that is allowed?",
          "votes": 1
        },
        {
          "id": 1295407,
          "postDate": "2021-05-06T12:41:12.127Z",
          "content": "<p>Here, check this out <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231318#1267279\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/231318#1267279</a></p>\n<p>BTW not trying to police you. I just want to do this too and checking to see if you've got a clear green light. If not, I suppose I can ask in that thread.</p>",
          "rawMarkdown": "Here, check this out https://www.kaggle.com/c/bms-molecular-translation/discussion/231318#1267279\n\nBTW not trying to police you. I just want to do this too and checking to see if you've got a clear green light. If not, I suppose I can ask in that thread."
        }
      ]
    },
    {
      "id": 1270712,
      "postDate": "2021-04-11T23:10:13.607Z",
      "content": "<p>there is an obvious bug in all validation code during training:</p>\n<pre><code>correct version\ndef do_valid(net, tokenizer, valid_loader):\n\n    if 1:\n        score = []\n        for i, (p, t) in enumerate(zip(predict, truth)):\n            t = truth[i][1:length[i]-1]     # in the buggy version, i have used 1 instead of i\n            p = predict[i][1:length[i]-1]\n            t = tokenizer.one_predict_to_inchi(t)\n            p = tokenizer.one_predict_to_inchi(p)\n            s = Levenshtein.distance(p, t)\n            score.append(s)\n        lb_score = np.mean(score)\n</code></pre>",
      "rawMarkdown": "there is an obvious bug in all validation code during training:\n\n```\ncorrect version\ndef do_valid(net, tokenizer, valid_loader):\n\n    if 1:\n        score = []\n        for i, (p, t) in enumerate(zip(predict, truth)):\n            t = truth[i][1:length[i]-1]     # in the buggy version, i have used 1 instead of i\n            p = predict[i][1:length[i]-1]\n            t = tokenizer.one_predict_to_inchi(t)\n            p = tokenizer.one_predict_to_inchi(p)\n            s = Levenshtein.distance(p, t)\n            score.append(s)\n        lb_score = np.mean(score)\n\n\n\n```",
      "votes": 7,
      "replies": [
        {
          "id": 1278624,
          "postDate": "2021-04-20T05:53:28.510Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<p>This script is correct?<br>\nI have a score gap between this script and compute_lb_score().</p>\n<p>e.g.<br>\nLB: 1.4(this script) vs LB: 1.9(compute_lb_score)<br>\nBoth valid_df are same fold number and using all data.</p>\n<p>Now, I investigate this reason.<br>\nthanks.</p>",
          "rawMarkdown": "@hengck23 \n\nThis script is correct?\nI have a score gap between this script and compute_lb_score().\n\ne.g.\nLB: 1.4(this script) vs LB: 1.9(compute_lb_score)\nBoth valid_df are same fold number and using all data.\n\nNow, I investigate this reason.\nthanks."
        },
        {
          "id": 1278676,
          "postDate": "2021-04-20T06:55:35.233Z",
          "content": "<p>i haven't check in details, but should be correct. just note that this validation is under teaching forcing (using forward function).</p>\n<p>the submit one is using autoregressive, forward with argmax function (i.e. decoder, without teacher forcing)</p>\n<hr>\n<p>also check the sampler and df file of the data loader. it could be different</p>",
          "rawMarkdown": "i haven't check in details, but should be correct. just note that this validation is under teaching forcing (using forward function).\n\nthe submit one is using autoregressive, forward with argmax function (i.e. decoder, without teacher forcing)\n\n---\n\nalso check the sampler and df file of the data loader. it could be different",
          "votes": 1
        },
        {
          "id": 1278886,
          "postDate": "2021-04-20T11:23:35.093Z",
          "content": "<p>Thanks!</p>\n<blockquote>\n  <p>the submit one is using autoregressive, forward with argmax function (i.e. decoder, without teacher forcing)</p>\n</blockquote>\n<p>I think that this is the cause of that.<br>\nAnd I understood using autoregressive takes long time.</p>",
          "rawMarkdown": "Thanks!\n\n> the submit one is using autoregressive, forward with argmax function (i.e. decoder, without teacher forcing)\n\nI think that this is the cause of that.\nAnd I understood using autoregressive takes long time."
        },
        {
          "id": 1302551,
          "postDate": "2021-05-11T15:36:53.587Z",
          "content": "<p>I think it shoule be p = predict[i][0:length[i]-2]</p>",
          "rawMarkdown": "I think it shoule be p = predict[i][0:length[i]-2]",
          "votes": 1
        }
      ]
    },
    {
      "id": 1269047,
      "postDate": "2021-04-10T05:03:00.400Z",
      "content": "<p>here is the latest [2021-apr-07b] version:</p>\n<p><img src=\"https://i.ibb.co/rb0GNHB/Selection-082.png\" alt=\"\"></p>\n<p>if you want to help in this development, you can:</p>\n<ul>\n<li>develop and share code on how to use fairseq batch k beam search</li>\n<li>experiment with more transformer decoder layers, or larger input image size, or use a larger image transformer encoder.  </li>\n<li>share results for training with augmentation</li>\n</ul>\n<p>Note that you do not need to train from scratch as you can also use your previous trained model (e.g. 224 image, 3 layers decoder) as initial checkpoint for your new experiment (e.g. 320 image, 6 layers decoder)</p>",
      "rawMarkdown": "here is the latest [2021-apr-07b] version:\n\n![](https://i.ibb.co/rb0GNHB/Selection-082.png)\n\nif you want to help in this development, you can:\n- develop and share code on how to use fairseq batch k beam search\n- experiment with more transformer decoder layers, or larger input image size, or use a larger image transformer encoder.  \n- share results for training with augmentation\n\nNote that you do not need to train from scratch as you can also use your previous trained model (e.g. 224 image, 3 layers decoder) as initial checkpoint for your new experiment (e.g. 320 image, 6 layers decoder)\n\n\n",
      "votes": 5,
      "replies": [
        {
          "id": 1270460,
          "postDate": "2021-04-11T17:16:14.053Z",
          "content": "<p>Why do you need 3-layers-decoder for 224x224 images and 6-layers-decoder for 320x320 images?</p>",
          "rawMarkdown": "Why do you need 3-layers-decoder for 224x224 images and 6-layers-decoder for 320x320 images?"
        },
        {
          "id": 1270468,
          "postDate": "2021-04-11T17:22:51.130Z",
          "content": "<p><a href=\"https://www.kaggle.com/thomasseleck\" target=\"_blank\">@thomasseleck</a> it's not a must it's just some configs to test. You can use 320x320 images with a 3 layers decoder ;) </p>",
          "rawMarkdown": "@thomasseleck it's not a must it's just some configs to test. You can use 320x320 images with a 3 layers decoder ;) ",
          "votes": 1
        },
        {
          "id": 1273124,
          "postDate": "2021-04-14T05:34:50.497Z",
          "content": "<p>Are you also inputing the embedding of the CLS token into the decoder? </p>",
          "rawMarkdown": "Are you also inputing the embedding of the CLS token into the decoder? "
        },
        {
          "id": 1273432,
          "postDate": "2021-04-14T10:46:34.710Z",
          "content": "<p>\"Are you also inputing the embedding of the CLS token into the decoder?\"</p>\n<p>yes, but i don't think it is important</p>",
          "rawMarkdown": "\"Are you also inputing the embedding of the CLS token into the decoder?\"\n\nyes, but i don't think it is important"
        },
        {
          "id": 1276234,
          "postDate": "2021-04-17T09:35:25.697Z",
          "content": "<p>\"Note that you do not need to train from scratch as you can also use your previous trained model (e.g. 224 image, 3 layers decoder) as initial checkpoint for your new experiment (e.g. 320 image, 6 layers decoder)\"</p>\n<p>In this scenario, when you perform the state_dict load, how do you address the size mismatch between patch_pos? Because the number of patches changes, do you increase the patch size, or approach it in some other way? </p>",
          "rawMarkdown": "\"Note that you do not need to train from scratch as you can also use your previous trained model (e.g. 224 image, 3 layers decoder) as initial checkpoint for your new experiment (e.g. 320 image, 6 layers decoder)\"\n\nIn this scenario, when you perform the state_dict load, how do you address the size mismatch between patch_pos? Because the number of patches changes, do you increase the patch size, or approach it in some other way? "
        },
        {
          "id": 1276290,
          "postDate": "2021-04-17T10:58:35.043Z",
          "content": "<p>delete this key in the state dict and let this part be learned from scratch.</p>\n<pre><code>    if initial_checkpoint is not None:\n        f = torch.load(initial_checkpoint, map_location=lambda storage, loc: storage)\n        start_iteration = f['iteration']\n        start_epoch     = f['epoch']\n        state_dict = f['state_dict']\n        #del state_dict['cnn.e.patch_pos']\n        #del state_dict['text_pos.pos']\n        #net.load_state_dict(state_dict, strict=False)  # True\n        net.load_state_dict(state_dict, strict=True)  # True\n</code></pre>",
          "rawMarkdown": "delete this key in the state dict and let this part be learned from scratch.\n\n```\n\n    if initial_checkpoint is not None:\n        f = torch.load(initial_checkpoint, map_location=lambda storage, loc: storage)\n        start_iteration = f['iteration']\n        start_epoch     = f['epoch']\n        state_dict = f['state_dict']\n        #del state_dict['cnn.e.patch_pos']\n        #del state_dict['text_pos.pos']\n        #net.load_state_dict(state_dict, strict=False)  # True\n        net.load_state_dict(state_dict, strict=True)  # True\n```",
          "votes": 1
        },
        {
          "id": 1276726,
          "postDate": "2021-04-17T22:16:46.910Z",
          "content": "<p>Thank you, I will see what I can contribute in the next weeks. </p>",
          "rawMarkdown": "Thank you, I will see what I can contribute in the next weeks. ",
          "votes": 1
        },
        {
          "id": 1307451,
          "postDate": "2021-05-14T12:55:35.700Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1266823,
      "postDate": "2021-04-08T05:18:15.290Z",
      "content": "<p>i have decided to use pytorch/fairseq for my next development</p>\n<ul>\n<li>it is actively developed and has lots of pretain model (including attention-LSTM image caption)</li>\n<li>It is part of pytorch ecosystem</li>\n<li>it can export to onnx (tensorRT) and TPU</li>\n<li>it has pre-norm transformer decoder and encoder (i will try to copy my trained weights to its modules)</li>\n<li>it supports caching of keys and values, incremental decoding</li>\n<li>it supports ensemble of seq model</li>\n</ul>\n<p><a href=\"https://github.com/pytorch/fairseq\" target=\"_blank\">https://github.com/pytorch/fairseq</a></p>\n<p>stay tuned on how to use fairseq for this challenge!</p>",
      "rawMarkdown": "i have decided to use pytorch/fairseq for my next development\n\n- it is actively developed and has lots of pretain model (including attention-LSTM image caption)\n- It is part of pytorch ecosystem\n- it can export to onnx (tensorRT) and TPU\n- it has pre-norm transformer decoder and encoder (i will try to copy my trained weights to its modules)\n- it supports caching of keys and values, incremental decoding\n- it supports ensemble of seq model\n\nhttps://github.com/pytorch/fairseq\n\nstay tuned on how to use fairseq for this challenge!",
      "votes": 5,
      "replies": [
        {
          "id": 1268248,
          "postDate": "2021-04-09T08:32:21.760Z",
          "content": "<p>just a quick update. I am preparing the next software version, which should be up in 24hrs.<br>\nsome very good news:</p>\n<ol>\n<li><p>my transformer from scratch is almost similar to fairseq version (after all, all are based on the \"what you need is attention\" paper). Hence i can convert my model to fairseq easily. there are little numercial differences, but can be get rid of in fine tunning some iterations</p></li>\n<li><p>fairseq key,value caching speedup inference. now i can run resnet101d at 13 to 14 min at 40_000 test images, reduction of 4 min from the [2021-apr-06a] version</p></li>\n<li><p>there are several transformer based image encoder, e.g. Vit. But I find TNT is magic (transformer in transformer). It can get local CV in the range of 2.0. Unlike resnet, efficient, etc, there is not image striding, hence transformer-based classifier can with with small image size like 224</p></li>\n<li><p>number of transformer decoder layers can improve results (i think) … but more experiments needs to confirm this </p></li>\n</ol>",
          "rawMarkdown": "just a quick update. I am preparing the next software version, which should be up in 24hrs.\nsome very good news:\n1. my transformer from scratch is almost similar to fairseq version (after all, all are based on the \"what you need is attention\" paper). Hence i can convert my model to fairseq easily. there are little numercial differences, but can be get rid of in fine tunning some iterations\n\n2. fairseq key,value caching speedup inference. now i can run resnet101d at 13 to 14 min at 40_000 test images, reduction of 4 min from the [2021-apr-06a] version\n\n3. there are several transformer based image encoder, e.g. Vit. But I find TNT is magic (transformer in transformer). It can get local CV in the range of 2.0. Unlike resnet, efficient, etc, there is not image striding, hence transformer-based classifier can with with small image size like 224\n\n4. number of transformer decoder layers can improve results (i think) ... but more experiments needs to confirm this ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1323657,
      "postDate": "2021-05-26T11:31:39.413Z",
      "content": "<ul>\n<li>2021-apr-25 software version  (i,e,  sparse patch method: keep only patches with black pixels, empty patches are discarded) + modifications</li>\n<li>add augmentation(e.g. noise, rotate, scale, resize artifacts, …)  to mimick test image </li>\n<li>remove scale and rotation predictor in pre-processing </li>\n<li>in testing i revert to use w&gt;h condition to unrotate the image</li>\n</ul>\n<p>the important point to note is that local CV and LB can be very smiliar. 40_000 validations images seem good enough</p>\n<p>training score is about 0.65 (without rdkit inchi validation check). This means that the top kaggler has reached the limit and modeled train=test very well ?</p>\n<p><img src=\"https://i.ibb.co/y4Gchs8/Selection-083.png\" alt=\"\"></p>",
      "rawMarkdown": "- 2021-apr-25 software version  (i,e,  sparse patch method: keep only patches with black pixels, empty patches are discarded) + modifications\n- add augmentation(e.g. noise, rotate, scale, resize artifacts, ...)  to mimick test image \n- remove scale and rotation predictor in pre-processing \n- in testing i revert to use w>h condition to unrotate the image\n\nthe important point to note is that local CV and LB can be very smiliar. 40\\_000 validations images seem good enough\n\ntraining score is about 0.65 (without rdkit inchi validation check). This means that the top kaggler has reached the limit and modeled train=test very well ?\n\n![](https://i.ibb.co/y4Gchs8/Selection-083.png)",
      "votes": 5,
      "replies": [
        {
          "id": 1323717,
          "postDate": "2021-05-26T12:08:53.863Z",
          "content": "<p>Thanks for your awesome work !</p>\n<p>Is it possible to get the 01495000_model.pth file to try to reproduce the results ?</p>",
          "rawMarkdown": "Thanks for your awesome work !\n\nIs it possible to get the 01495000_model.pth file to try to reproduce the results ?",
          "votes": -1
        },
        {
          "id": 1324018,
          "postDate": "2021-05-26T15:00:54.273Z",
          "content": "<p>there is no further plan to release code or model. if you want to get the similiar results, your training log should be close to the below:</p>\n<pre><code>   batch_size = 128\n   experiment = ['vit-s0.8-p16-06b4r6', 'run_train.py']\n                      |----- VALID ---|---- TRAIN/BATCH --------------\nrate     iter   epoch | loss  lb(lev) | loss0  loss1  | time          \n----------------------------------------------------------------------\n0.00000  155.0000* 53.51  | 0.005   0.18  | 0.000  0.000  0.000  |  0 hr 00 min\n0.00003  155.1000  53.54  | 0.005   0.18  | 0.005  0.001  0.000  |  0 hr 19 min\n0.00003  155.1806  53.56  | 0.005   0.18  | 0.003  0.000  0.000  |  0 hr 35 min\n0.00003  155.2000  53.56  | 0.005   0.17  | 0.004  0.000  0.000  |  0 hr 39 min\n0.00003  155.3000  53.59  | 0.005   0.18  | 0.005  0.000  0.000  |  0 hr 58 min\n</code></pre>",
          "rawMarkdown": "there is no further plan to release code or model. if you want to get the similiar results, your training log should be close to the below:\n\n```\n   batch_size = 128\n   experiment = ['vit-s0.8-p16-06b4r6', 'run_train.py']\n                      |----- VALID ---|---- TRAIN/BATCH --------------\nrate     iter   epoch | loss  lb(lev) | loss0  loss1  | time          \n----------------------------------------------------------------------\n0.00000  155.0000* 53.51  | 0.005   0.18  | 0.000  0.000  0.000  |  0 hr 00 min\n0.00003  155.1000  53.54  | 0.005   0.18  | 0.005  0.001  0.000  |  0 hr 19 min\n0.00003  155.1806  53.56  | 0.005   0.18  | 0.003  0.000  0.000  |  0 hr 35 min\n0.00003  155.2000  53.56  | 0.005   0.17  | 0.004  0.000  0.000  |  0 hr 39 min\n0.00003  155.3000  53.59  | 0.005   0.18  | 0.005  0.000  0.000  |  0 hr 58 min\n\n```",
          "votes": 4
        },
        {
          "id": 1327398,
          "postDate": "2021-05-29T08:30:07.587Z",
          "content": "<p>doesn't this patch-based method require image resize?</p>\n<p>I'm trying to figure out how to solve this without resizing the images</p>",
          "rawMarkdown": "doesn't this patch-based method require image resize?\n\nI'm trying to figure out how to solve this without resizing the images"
        },
        {
          "id": 1327405,
          "postDate": "2021-05-29T08:35:13.030Z",
          "content": "<p>i build one model with image resize (scale = 0.8) and another one without (i.e. scale=1.0 or original size).</p>\n<p>if you resize image, you have less patches and the model will train faster and uses less memory.</p>\n<p>i haven't validate and test the performance for model without image resize (training is still in progress)</p>\n<p>but the both model have smiliar training loss.<br>\ni don't expect huge improvement for a single model.</p>\n<p>maybe there will be some gain from ensembling, test time augmentation, post-processing etc</p>",
          "rawMarkdown": "i build one model with image resize (scale = 0.8) and another one without (i.e. scale=1.0 or original size).\n\nif you resize image, you have less patches and the model will train faster and uses less memory.\n\ni haven't validate and test the performance for model without image resize (training is still in progress)\n\nbut the both model have smiliar training loss.\ni don't expect huge improvement for a single model.\n\nmaybe there will be some gain from ensembling, test time augmentation, post-processing etc",
          "votes": 1
        },
        {
          "id": 1339725,
          "postDate": "2021-06-07T12:29:05.423Z",
          "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> .This gave me a lot of intuition and ideas.</p>",
          "rawMarkdown": "Thank you so much @hengck23 .This gave me a lot of intuition and ideas."
        }
      ]
    },
    {
      "id": 1324800,
      "postDate": "2021-05-27T08:23:52.037Z",
      "content": "<p>The day <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> uses github (or any other git hosting solution)… 😄</p>\n<p>Joke aside, thanks for sharing this starter kit!</p>",
      "rawMarkdown": "The day @hengck23 uses github (or any other git hosting solution)... 😄\n\nJoke aside, thanks for sharing this starter kit!",
      "votes": 3
    },
    {
      "id": 1304058,
      "postDate": "2021-05-12T11:54:52.813Z",
      "content": "<p>some symmetry mages that may cause havoc to your algorithm</p>\n<p><img src=\"https://i.ibb.co/Jp9HsVt/Selection-062.png\" alt=\"\"></p>",
      "rawMarkdown": "some symmetry mages that may cause havoc to your algorithm\n\n![](https://i.ibb.co/Jp9HsVt/Selection-062.png)",
      "votes": 3,
      "replies": [
        {
          "id": 1304871,
          "postDate": "2021-05-13T00:34:57.167Z",
          "content": "<p>it is easy to detect such symmetric images : just rotate and do a comparison e.g. l2 loss, see if black pixels align, etc</p>\n<p>you can rotate such image as augmentation in training (to create different noise) </p>",
          "rawMarkdown": "it is easy to detect such symmetric images : just rotate and do a comparison e.g. l2 loss, see if black pixels align, etc\n\nyou can rotate such image as augmentation in training (to create different noise) ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1300309,
      "postDate": "2021-05-10T11:39:29.253Z",
      "content": "<p>transformer set decoding, aka object detection</p>\n<p><img src=\"https://i.ibb.co/dQVWGKX/Selection-031.png\" alt=\"\"></p>\n<p><img src=\"https://i.ibb.co/qmhkk1Y/Selection-037.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/xDZFJSm/Selection-038.png\" alt=\"\"></p>\n<p>i had this idea in the past, but stuck at the bond assignment. e.g. a bond has attribute X-Y where X,Y are the numbering of the atoms. I was wondering how to set the ground truth as the numbering could be unknown. Then I realize that it is not a problem if you use dynamic set assignment of ground truth as it is used in object detection</p>",
      "rawMarkdown": "transformer set decoding, aka object detection\n\n![](https://i.ibb.co/dQVWGKX/Selection-031.png)\n\n![](https://i.ibb.co/qmhkk1Y/Selection-037.png)\n![](https://i.ibb.co/xDZFJSm/Selection-038.png)\n\n\ni had this idea in the past, but stuck at the bond assignment. e.g. a bond has attribute X-Y where X,Y are the numbering of the atoms. I was wondering how to set the ground truth as the numbering could be unknown. Then I realize that it is not a problem if you use dynamic set assignment of ground truth as it is used in object detection",
      "votes": 3,
      "replies": [
        {
          "id": 1300819,
          "postDate": "2021-05-10T18:34:05.297Z",
          "content": "<p>instead of having a seq decoder transformer, i am thinking of changing the strategy. there is only one seq in our case (unlike the language model, where there can be many possible valid sequences)</p>\n<p>since there is one and only one ground truth seq (due to only one canonical atom numbering), we can directly predict the atom at the atomic number location</p>\n<p>the query object will have position encoding = atomic numbering. this ensures we would have most of the atom with atomic numbering correct, hence lowering the LD distance</p>",
          "rawMarkdown": "instead of having a seq decoder transformer, i am thinking of changing the strategy. there is only one seq in our case (unlike the language model, where there can be many possible valid sequences)\n\nsince there is one and only one ground truth seq (due to only one canonical atom numbering), we can directly predict the atom at the atomic number location\n\nthe query object will have position encoding = atomic numbering. this ensures we would have most of the atom with atomic numbering correct, hence lowering the LD distance",
          "votes": 2
        },
        {
          "id": 1300829,
          "postDate": "2021-05-10T18:48:14.800Z",
          "content": "<p>That sounds phenomenal!<br>\nI perhaps would go for something like this if there was enough time left to train.</p>\n<p>You do have great ideas, that's for sure!</p>",
          "rawMarkdown": "That sounds phenomenal!\nI perhaps would go for something like this if there was enough time left to train.\n\nYou do have great ideas, that's for sure!",
          "votes": 2
        }
      ]
    },
    {
      "id": 1296482,
      "postDate": "2021-05-07T09:35:41.723Z",
      "content": "<p>Training curve with manually decreasing LR</p>\n<p><a href=\"https://ibb.co/GTHXWcv\"><img src=\"https://i.ibb.co/k1mp4Qg/train-plot.png\" alt=\"train-plot\"></a></p>",
      "rawMarkdown": "Training curve with manually decreasing LR\n\n<a href=\"https://ibb.co/GTHXWcv\"><img src=\"https://i.ibb.co/k1mp4Qg/train-plot.png\" alt=\"train-plot\" border=\"0\"></a>",
      "votes": 3,
      "replies": [
        {
          "id": 1296485,
          "postDate": "2021-05-07T09:39:32.553Z",
          "content": "<p>Thanks for sharing! Is this after every epoch?</p>",
          "rawMarkdown": "Thanks for sharing! Is this after every epoch?"
        },
        {
          "id": 1296496,
          "postDate": "2021-05-07T09:47:23.413Z",
          "content": "<p>Nope, the total graph is about one epoch (warm restart from another pretrained model with different size). </p>",
          "rawMarkdown": "Nope, the total graph is about one epoch (warm restart from another pretrained model with different size). ",
          "votes": 1
        },
        {
          "id": 1296510,
          "postDate": "2021-05-07T10:01:34.087Z",
          "content": "<p>Oh I see now, I'm dumb ^^'</p>",
          "rawMarkdown": "Oh I see now, I'm dumb ^^'"
        },
        {
          "id": 1296530,
          "postDate": "2021-05-07T10:12:42.917Z",
          "content": "<p>There's no such thing as a dumb question! </p>\n<p>Also, I wish I had the resources to train for so many epochs! 😂</p>",
          "rawMarkdown": "There's no such thing as a dumb question! \n\nAlso, I wish I had the resources to train for so many epochs! 😂",
          "votes": 1
        },
        {
          "id": 1296556,
          "postDate": "2021-05-07T10:36:23.540Z",
          "content": "<p>Now I wish I was a question. :(</p>",
          "rawMarkdown": "Now I wish I was a question. :(",
          "votes": 1
        },
        {
          "id": 1296588,
          "postDate": "2021-05-07T11:05:51.673Z",
          "content": "<p>you always restart training, e.g cyclic training rate, to see if you can get better result.</p>\n<p>you can compare with my train log if you are using the same fold and net parameters</p>",
          "rawMarkdown": "you always restart training, e.g cyclic training rate, to see if you can get better result.\n\nyou can compare with my train log if you are using the same fold and net parameters"
        },
        {
          "id": 1296984,
          "postDate": "2021-05-07T16:17:05.283Z",
          "content": "<blockquote>\n  <p>Now I wish I was a question. :(</p>\n</blockquote>\n<p>Don't worry, questions can't think - this way they also can't be dumb. So you have a vast advantage there ;)</p>",
          "rawMarkdown": "> Now I wish I was a question. :(\n\nDon't worry, questions can't think - this way they also can't be dumb. So you have a vast advantage there ;)"
        },
        {
          "id": 1297705,
          "postDate": "2021-05-08T08:31:01.107Z",
          "content": "<p>never mind :D</p>",
          "rawMarkdown": "never mind :D"
        },
        {
          "id": 1298816,
          "postDate": "2021-05-09T08:29:10.523Z",
          "content": "<p>When I do manual LR restarts quite a few starting iterations have a higher cv than the checkpoint model. Any ideas why this might be happening or what I can do to fix? - Possible reason is that I deleted the trained pos embeddings even though I was using same image size. I should not do that. </p>",
          "rawMarkdown": "When I do manual LR restarts quite a few starting iterations have a higher cv than the checkpoint model. Any ideas why this might be happening or what I can do to fix? - Possible reason is that I deleted the trained pos embeddings even though I was using same image size. I should not do that. "
        },
        {
          "id": 1300324,
          "postDate": "2021-05-10T11:52:45.840Z",
          "content": "<p>if you restart training, you should load your previous trained model, including your trained pos encoding. you should modify the load previous state dict code</p>",
          "rawMarkdown": "if you restart training, you should load your previous trained model, including your trained pos encoding. you should modify the load previous state dict code",
          "votes": 2
        },
        {
          "id": 1300420,
          "postDate": "2021-05-10T12:52:53.720Z",
          "content": "<p>Yep, seeing the desired behavior now. </p>",
          "rawMarkdown": "Yep, seeing the desired behavior now. ",
          "votes": 1
        },
        {
          "id": 1307044,
          "postDate": "2021-05-14T08:06:45.750Z",
          "content": "<p>This is a grossly underrated string of internet bits among the 45 zettabytes in existence</p>\n<blockquote>\n  <p>Now I wish I was a question. :(</p>\n</blockquote>\n<p>cc <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> </p>",
          "rawMarkdown": "This is a grossly underrated string of internet bits among the 45 zettabytes in existence\n\n> Now I wish I was a question. :(\n\ncc @nofreewill ",
          "votes": 1
        },
        {
          "id": 1307186,
          "postDate": "2021-05-14T09:58:43.877Z",
          "content": "<p><a href=\"https://www.kaggle.com/alexandersoare\" target=\"_blank\">@alexandersoare</a> thanks for the appreciation! I'm not gonna lie, I had a pretty good laugh when I wrote that. :D</p>",
          "rawMarkdown": "@alexandersoare thanks for the appreciation! I'm not gonna lie, I had a pretty good laugh when I wrote that. :D",
          "votes": 1
        }
      ]
    },
    {
      "id": 1288190,
      "postDate": "2021-04-29T18:33:16.290Z",
      "content": "<p><img src=\"https://www.memecreator.org/static/images/memes/4307297.jpg\" alt=\"\"><br>\nAttention is all you need</p>",
      "rawMarkdown": "![](https://www.memecreator.org/static/images/memes/4307297.jpg)\nAttention is all you need",
      "votes": 4
    },
    {
      "id": 1273297,
      "postDate": "2021-04-14T08:21:51.083Z",
      "content": "<p>list of tricks to try:</p>\n<ol>\n<li><p><a href=\"https://arxiv.org/pdf/2006.12000.pdf\" target=\"_blank\">https://arxiv.org/pdf/2006.12000.pdf</a><br>\nSelf-Knowledge Distillation: A Simple Way for Better Generalization</p></li>\n<li><p>Sharpness-Aware Minimization for Efficiently Improving Generalization</p></li>\n<li><p><a href=\"https://cs.nju.edu.cn/wujx/paper/AAAI2021_Tricks.pdf\" target=\"_blank\">https://cs.nju.edu.cn/wujx/paper/AAAI2021_Tricks.pdf</a><br>\nBag of Tricks for Long-Tailed Visual Recognition with Deep Convolutional Neural Networks</p></li>\n</ol>\n<hr>\n<p><a href=\"https://neptune.ai/blog/text-classification-tips-and-tricks-kaggle-competitions\" target=\"_blank\">https://neptune.ai/blog/text-classification-tips-and-tricks-kaggle-competitions</a></p>",
      "rawMarkdown": "list of tricks to try:\n1. https://arxiv.org/pdf/2006.12000.pdf\nSelf-Knowledge Distillation: A Simple Way for Better Generalization\n\n2. Sharpness-Aware Minimization for Efficiently Improving Generalization\n\n3. https://cs.nju.edu.cn/wujx/paper/AAAI2021_Tricks.pdf\nBag of Tricks for Long-Tailed Visual Recognition with Deep Convolutional Neural Networks\n\n---\nhttps://neptune.ai/blog/text-classification-tips-and-tricks-kaggle-competitions",
      "votes": 4
    },
    {
      "id": 1327626,
      "postDate": "2021-05-29T13:23:36.343Z",
      "content": "<p>One minor typo here: pre-norm activated <strong>trnasformer</strong> are used. </p>",
      "rawMarkdown": "One minor typo here: pre-norm activated **trnasformer** are used. ",
      "votes": 1
    },
    {
      "id": 1314190,
      "postDate": "2021-05-19T02:51:12.303Z",
      "content": "<p>'advertisement' …<br>\n<a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240233\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/240233</a></p>",
      "rawMarkdown": "'advertisement' ...\nhttps://www.kaggle.com/c/siim-covid19-detection/discussion/240233",
      "votes": 1
    },
    {
      "id": 1301657,
      "postDate": "2021-05-11T07:20:15.203Z",
      "content": "<p>I read somewhere that we can do data aug by reversing the Inchi strings. Has anyone tried this as a data aug or is it a good idea even? Looking for thoughts. </p>",
      "rawMarkdown": "I read somewhere that we can do data aug by reversing the Inchi strings. Has anyone tried this as a data aug or is it a good idea even? Looking for thoughts. ",
      "votes": 1
    },
    {
      "id": 1296091,
      "postDate": "2021-05-07T01:12:45.840Z",
      "content": "<p>I want to ask such a question. How to use 224 image dimension pretrained weight to train higher dimension image.</p>",
      "rawMarkdown": "I want to ask such a question. How to use 224 image dimension pretrained weight to train higher dimension image.",
      "votes": 1,
      "replies": [
        {
          "id": 1297290,
          "postDate": "2021-05-07T22:08:53.193Z",
          "content": "<p>It has been found that CNN can actually be <a href=\"https://arxiv.org/pdf/1906.06423.pdf\" target=\"_blank\">finetuned for higher resolutions</a>. So it shouldn't be a problem to just load the pretrained weights and use the models as is <br>\n(The images look totally different anyway, all that pretraining might be useful for are general inductive biases about images. I have never tried a pretrained model so far, but I doubt that it makes more than a really tiny difference after 10+ epochs).</p>\n<p>Although according to the competition rules the pretrained weights should allow for commercial usage. Something few people here seem to take into consideration.</p>",
          "rawMarkdown": "It has been found that CNN can actually be [finetuned for higher resolutions](https://arxiv.org/pdf/1906.06423.pdf). So it shouldn't be a problem to just load the pretrained weights and use the models as is \n(The images look totally different anyway, all that pretraining might be useful for are general inductive biases about images. I have never tried a pretrained model so far, but I doubt that it makes more than a really tiny difference after 10+ epochs).\n\nAlthough according to the competition rules the pretrained weights should allow for commercial usage. Something few people here seem to take into consideration."
        },
        {
          "id": 1297386,
          "postDate": "2021-05-08T01:17:44.160Z",
          "content": "<p>Thank you for your answer.</p>",
          "rawMarkdown": "Thank you for your answer."
        }
      ]
    },
    {
      "id": 1280273,
      "postDate": "2021-04-21T19:15:29.767Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> How do you train the rotation detector please?</p>",
      "rawMarkdown": "@hengck23 How do you train the rotation detector please?",
      "votes": 1,
      "replies": [
        {
          "id": 1280416,
          "postDate": "2021-04-21T23:33:02.343Z",
          "content": "<p>rotate the train images. the rotation used is your ground truth label.</p>",
          "rawMarkdown": "rotate the train images. the rotation used is your ground truth label.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1272986,
      "postDate": "2021-04-14T01:32:54.670Z",
      "content": "<p>Hi! Just dowloaded the last version [2021-apr-07b], on <code>run_train.py</code> I seem to be missing some imports:</p>\n<pre><code>from common import *\nfrom bms import *\n\nfrom lib.net.lookahead import *\nfrom lib.net.radam import *\n</code></pre>\n<p>Where could I find those? Thanks!</p>",
      "rawMarkdown": "Hi! Just dowloaded the last version [2021-apr-07b], on `run_train.py` I seem to be missing some imports:\n```\nfrom common import *\nfrom bms import *\n\nfrom lib.net.lookahead import *\nfrom lib.net.radam import *\n```\n\nWhere could I find those? Thanks!",
      "votes": 1,
      "replies": [
        {
          "id": 1272999,
          "postDate": "2021-04-14T02:02:50.707Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/martinbeyerdeicas\" target=\"_blank\">@martinbeyerdeicas</a> <br>\nPlease check  [2021-apr-06] to find those files</p>",
          "rawMarkdown": "Hi @martinbeyerdeicas \nPlease check  [2021-apr-06] to find those files"
        }
      ]
    },
    {
      "id": 1270199,
      "postDate": "2021-04-11T12:16:37.390Z",
      "content": "<p>Wow! This is pure gold!</p>\n<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>: I have a few questions for you: How many GPUs did you use to train the \"2021-apr-07b\" version? How much time did it take ?</p>\n<p>Have you done some preprocessing on images before training like noise removal or something else ?</p>",
      "rawMarkdown": "Wow! This is pure gold!\n\n@hengck23: I have a few questions for you: How many GPUs did you use to train the \"2021-apr-07b\" version? How much time did it take ?\n\nHave you done some preprocessing on images before training like noise removal or something else ?",
      "votes": 1,
      "replies": [
        {
          "id": 1270205,
          "postDate": "2021-04-11T12:25:41.393Z",
          "content": "<p>\"Have you done some preprocessing on images before training like noise removal or something else ?\"<br>\nnot at this moment. But in the future, some preprocessing (e.g. size normalization) may improve results</p>\n<p>\"How many GPUs\"<br>\nI have a good GPU card. I will talk about that later.</p>\n<p>For a normal user, if you start with 224-TNT-S vision transformer, you just need to train with batch=64 for about 10 epoch. see log files for details. decrease your learning rate from 0.001 to 0.00005</p>\n<p>\"This is pure gold!\"<br>\nwe are still in the early part of the competition. At the end of the challenge, gold should be below 1.0,<br>\nsilver probably below 2 ~2.3. I haven't done anything special except train with long iterations with bigger models and better GPU. These results can be easily replicated and hence it is easy for others to catch up over time, using their different model </p>",
          "rawMarkdown": "\"Have you done some preprocessing on images before training like noise removal or something else ?\"\nnot at this moment. But in the future, some preprocessing (e.g. size normalization) may improve results\n\n\n\"How many GPUs\"\nI have a good GPU card. I will talk about that later.\n\nFor a normal user, if you start with 224-TNT-S vision transformer, you just need to train with batch=64 for about 10 epoch. see log files for details. decrease your learning rate from 0.001 to 0.00005\n\n\"This is pure gold!\"\nwe are still in the early part of the competition. At the end of the challenge, gold should be below 1.0,\nsilver probably below 2 ~2.3. I haven't done anything special except train with long iterations with bigger models and better GPU. These results can be easily replicated and hence it is easy for others to catch up over time, using their different model \n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1270197,
      "postDate": "2021-04-11T12:07:06.803Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , <br>\nI am encountering the following error when I try to run your code on Windows. I tried searching solutions online but could not find one…<br>\n<code>\nNameError: name 'Tensor' is not defined\n</code><br>\nDo you have any idea how to solve this? </p>",
      "rawMarkdown": "Hi @hengck23 , \nI am encountering the following error when I try to run your code on Windows. I tried searching solutions online but could not find one...\n`\nNameError: name 'Tensor' is not defined\n`\nDo you have any idea how to solve this? ",
      "votes": 1,
      "replies": [
        {
          "id": 1270198,
          "postDate": "2021-04-11T12:12:53.617Z",
          "content": "<p><a href=\"https://pytorch.org/docs/stable/jit.html\" target=\"_blank\">https://pytorch.org/docs/stable/jit.html</a></p>\n<p>try torch.Tensor</p>",
          "rawMarkdown": "https://pytorch.org/docs/stable/jit.html\n\ntry torch.Tensor",
          "votes": 3
        }
      ]
    },
    {
      "id": 1266756,
      "postDate": "2021-04-08T04:06:10.427Z",
      "content": "<p>you should implement this:</p>\n<p><a href=\"https://wandb.ai/pommedeterresautee/speed_training/reports/Train-HuggingFace-Models-Twice-As-Fast--VmlldzoxMDgzOTI\" target=\"_blank\">https://wandb.ai/pommedeterresautee/speed_training/reports/Train-HuggingFace-Models-Twice-As-Fast--VmlldzoxMDgzOTI</a><br>\n<a href=\"https://towardsdatascience.com/divide-hugging-face-transformers-training-time-by-2-or-more-21bf7129db9q-21bf7129db9e\" target=\"_blank\">https://towardsdatascience.com/divide-hugging-face-transformers-training-time-by-2-or-more-21bf7129db9q-21bf7129db9e</a></p>\n<p>Divide Hugging Face Transformers training time by 2 or more with dynamic padding and uniform length batching</p>",
      "rawMarkdown": "you should implement this:\n\nhttps://wandb.ai/pommedeterresautee/speed_training/reports/Train-HuggingFace-Models-Twice-As-Fast--VmlldzoxMDgzOTI\nhttps://towardsdatascience.com/divide-hugging-face-transformers-training-time-by-2-or-more-21bf7129db9q-21bf7129db9e\n\nDivide Hugging Face Transformers training time by 2 or more with dynamic padding and uniform length batching",
      "votes": 1,
      "replies": [
        {
          "id": 1267248,
          "postDate": "2021-04-08T12:14:53.733Z",
          "content": "<p>here is sample code for key,value caching:<br>\n<a href=\"https://github.com/tunz/transformer-pytorch/blob/master/model/fast_transformer.py\" target=\"_blank\">https://github.com/tunz/transformer-pytorch/blob/master/model/fast_transformer.py</a> <br>\n<a href=\"https://tunz.kr/post/4\" target=\"_blank\">https://tunz.kr/post/4</a></p>\n<p>(i am not implementing this since i am migrating to fairseq)</p>",
          "rawMarkdown": "here is sample code for key,value caching:\nhttps://github.com/tunz/transformer-pytorch/blob/master/model/fast_transformer.py \nhttps://tunz.kr/post/4\n\n\n(i am not implementing this since i am migrating to fairseq)"
        }
      ]
    },
    {
      "id": 1314510,
      "postDate": "2021-05-19T07:49:54.657Z",
      "content": "<p>hi  , I create two notebook base on your version apri 24</p>\n<ol>\n<li>preprocess:  <a href=\"https://www.kaggle.com/drzhuzhe/bms-preprocess-data-parallel\" target=\"_blank\">https://www.kaggle.com/drzhuzhe/bms-preprocess-data-parallel</a></li>\n<li>training: <a href=\"https://www.kaggle.com/drzhuzhe/training-on-gpu-bms/\" target=\"_blank\">https://www.kaggle.com/drzhuzhe/training-on-gpu-bms/</a></li>\n</ol>\n<p>During training I notice CPU memory keeping increase as iterate, result in out of memory crash </p>\n<p>I checkout all network modules <br>\ndelete all cv image and leak variable <br>\nbut memory still keeping increase, </p>\n<p>hardly locate what make this problem happen</p>\n<p>may someone have some tips?</p>",
      "rawMarkdown": "hi  , I create two notebook base on your version apri 24\n\n1. preprocess:  https://www.kaggle.com/drzhuzhe/bms-preprocess-data-parallel\n2. training: https://www.kaggle.com/drzhuzhe/training-on-gpu-bms/\n\nDuring training I notice CPU memory keeping increase as iterate, result in out of memory crash \n\nI checkout all network modules \ndelete all cv image and leak variable \nbut memory still keeping increase, \n\nhardly locate what make this problem happen\n\nmay someone have some tips?",
      "votes": 2,
      "replies": [
        {
          "id": 1327845,
          "postDate": "2021-05-29T17:03:46.413Z",
          "content": "<p>your datasets are private.</p>",
          "rawMarkdown": "your datasets are private."
        },
        {
          "id": 1328248,
          "postDate": "2021-05-30T04:34:00.507Z",
          "content": "<p>I made it public, <br>\nwatch out the BMS-train-full dataset has too much files, will be very slow for kaggle notebook to load it <br>\nIf you download it as a zip file , it will be much faster </p>",
          "rawMarkdown": "I made it public, \nwatch out the BMS-train-full dataset has too much files, will be very slow for kaggle notebook to load it \nIf you download it as a zip file , it will be much faster "
        },
        {
          "id": 1329265,
          "postDate": "2021-05-31T02:50:57.727Z",
          "content": "<p>yes, It crashed.  probably torch issue.  </p>",
          "rawMarkdown": "yes, It crashed.  probably torch issue.  "
        },
        {
          "id": 1329315,
          "postDate": "2021-05-31T03:49:37.257Z",
          "content": "<p>no no no ,</p>\n<ol>\n<li><p>to the out of memory problem</p>\n<p>this code is function well before Out of memory after 20k epoch<br>\nthat may due to image dataset load by both cpu and gpu<br>\nthis is definitely not torch issue, maybe hengk is trainning with larger RAM  </p></li>\n<li><p>to very slow to load input data problem </p>\n<p>it will take half a hour to load full dataset<br>\nif you get input files as a zip , it will be much faster<br>\nI zip all file and upload to kaggle dataset,<br>\nit is kaggle dataset automatic unzip my uploaded file and return to over 4000k small files</p></li>\n</ol>",
          "rawMarkdown": "no no no ,\n\n1. to the out of memory problem\n\n this code is function well before Out of memory after 20k epoch\n that may due to image dataset load by both cpu and gpu\n this is definitely not torch issue, maybe hengk is trainning with larger RAM  \n\n2. to very slow to load input data problem \n\n it will take half a hour to load full dataset\n if you get input files as a zip , it will be much faster\n I zip all file and upload to kaggle dataset,\n it is kaggle dataset automatic unzip my uploaded file and return to over 4000k small files",
          "votes": 1
        },
        {
          "id": 1329408,
          "postDate": "2021-05-31T06:07:41.153Z",
          "content": "<p>I loaded the pretrained weights, It crashed within 2000 epochs for resuming training.  Of course big enough RAM would not have this problem.  I have experiences that Pytorch consumes more ram.  Usually same algorithm if implemented in Tensorflow,  OOM happens less.</p>",
          "rawMarkdown": "I loaded the pretrained weights, It crashed within 2000 epochs for resuming training.  Of course big enough RAM would not have this problem.  I have experiences that Pytorch consumes more ram.  Usually same algorithm if implemented in Tensorflow,  OOM happens less."
        },
        {
          "id": 1329421,
          "postDate": "2021-05-31T06:13:41.323Z",
          "content": "<p>memory for each batch maybe different.<br>\ni think i allocate for the longest sequence length in the batch.</p>\n<p>you can sort your df_train by decreasing length and use sequential sampler to test the max batch size you system can time.</p>\n<p>once you determine the max batch size (and/or seq length), you can revert to random sampler</p>",
          "rawMarkdown": "memory for each batch maybe different.\ni think i allocate for the longest sequence length in the batch.\n\nyou can sort your df\\_train by decreasing length and use sequential sampler to test the max batch size you system can time.\n\nonce you determine the max batch size (and/or seq length), you can revert to random sampler"
        },
        {
          "id": 1329852,
          "postDate": "2021-05-31T12:26:13.110Z",
          "content": "<p>I reduce batchsize, It runs so slowly, No crash however Kaggle 9 hours timeout. I don't remember the previous crashed  iteration number. Mostly Hardware RAM  not big enough.  </p>",
          "rawMarkdown": "I reduce batchsize, It runs so slowly, No crash however Kaggle 9 hours timeout. I don't remember the previous crashed  iteration number. Mostly Hardware RAM  not big enough.  "
        }
      ]
    },
    {
      "id": 1290251,
      "postDate": "2021-05-01T18:39:43.463Z",
      "content": "<p>to my horror, many of the train images are rotated! (possibly 1 %)</p>\n<p>just to list a few example<br>\n1766b78e6eca<br>\n1b9063aabec5<br>\n1f60964241fd<br>\n1a56532cfc19<br>\n1fbedd22a0b4<br>\n1af43b2be6d3<br>\n1c9477a191a7<br>\n…</p>\n<p>so both test and train have rotated images. Just that the test has much more</p>\n<p>after preparing the images from some software like rdkit?, the organizer simply rotates the image if w&lt;h.</p>",
      "rawMarkdown": "to my horror, many of the train images are rotated! (possibly 1 %)\n\njust to list a few example\n1766b78e6eca\n1b9063aabec5\n1f60964241fd\n1a56532cfc19\n1fbedd22a0b4\n1af43b2be6d3\n1c9477a191a7\n...\n\nso both test and train have rotated images. Just that the test has much more\n\nafter preparing the images from some software like rdkit?, the organizer simply rotates the image if w<h.",
      "votes": 2
    },
    {
      "id": 1283914,
      "postDate": "2021-04-25T11:53:35.933Z",
      "content": "<p>i have been using Lookahead+Radam optimizer. I only save trained model parameters and not the optimizer.</p>\n<p>i notice that when i restart training, the optimizer would need to restart of zero (because i didn't save the previous state)<br>\nAnd there is always a drop and improvement in validation loss for the restart.</p>\n<p>i wonder what is the reason for this?</p>",
      "rawMarkdown": "i have been using Lookahead+Radam optimizer. I only save trained model parameters and not the optimizer.\n\ni notice that when i restart training, the optimizer would need to restart of zero (because i didn't save the previous state)\nAnd there is always a drop and improvement in validation loss for the restart.\n\ni wonder what is the reason for this?",
      "votes": 2
    },
    {
      "id": 1283843,
      "postDate": "2021-04-25T10:20:26.333Z",
      "content": "<p>for timm vision transformer models, there is \"def resize_pos_embed(posemb, posemb_new, num_tokens=1)\" function that can resize your positional embedding when you finetune your transformer from small size to a bigger image size</p>",
      "rawMarkdown": "for timm vision transformer models, there is \"def resize_pos_embed(posemb, posemb_new, num_tokens=1)\" function that can resize your positional embedding when you finetune your transformer from small size to a bigger image size",
      "votes": 2
    },
    {
      "id": 1277621,
      "postDate": "2021-04-19T02:11:46.217Z",
      "content": "<p>there is JAX implementation for TNT<br>\n<a href=\"https://github.com/NZ99/transformer_in_transformer_flax\" target=\"_blank\">https://github.com/NZ99/transformer_in_transformer_flax</a></p>\n<p>anyone want to take this code to TPU?<br>\n(in my latest experience image size 448 is better than 384 …. i think ultimately, maybe we need to go to 640 for a constant input image size. if not, we need to break the input image into patch tokens ….)</p>",
      "rawMarkdown": "there is JAX implementation for TNT\nhttps://github.com/NZ99/transformer_in_transformer_flax\n\nanyone want to take this code to TPU?\n(in my latest experience image size 448 is better than 384 .... i think ultimately, maybe we need to go to 640 for a constant input image size. if not, we need to break the input image into patch tokens ....)\n",
      "votes": 2
    },
    {
      "id": 1275234,
      "postDate": "2021-04-16T05:40:59.503Z",
      "content": "<p>This is great work, very educational. Thanks for your contribution <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>. </p>\n<p>PS: inference on the test set with 224x224 images takes me approximately 5 hours with a single 1080 Ti. </p>",
      "rawMarkdown": "This is great work, very educational. Thanks for your contribution @hengck23. \n\nPS: inference on the test set with 224x224 images takes me approximately 5 hours with a single 1080 Ti. ",
      "votes": 2
    },
    {
      "id": 1272950,
      "postDate": "2021-04-13T23:04:15.513Z",
      "content": "<p>i have completed basic implementation and experiments. Now it is time to go to the next stage … uncovering the magic of the data (i.e. look at the images, errors, etc)</p>\n<p>a quick study reveals</p>\n<ol>\n<li>at first I thought the images are randomly scaled. No, it isn't. there are only 2 scales. There is fixed (and different) padding of the image for the chemical structure for each scale. </li>\n</ol>\n<p>you can double your train images by inter-converting images from one scale to another …</p>\n<p>it seems that there much more to exploit. You can try it yourself  </p>",
      "rawMarkdown": "i have completed basic implementation and experiments. Now it is time to go to the next stage ... uncovering the magic of the data (i.e. look at the images, errors, etc)\n\na quick study reveals\n1. at first I thought the images are randomly scaled. No, it isn't. there are only 2 scales. There is fixed (and different) padding of the image for the chemical structure for each scale. \n\nyou can double your train images by inter-converting images from one scale to another ...\n\nit seems that there much more to exploit. You can try it yourself  ",
      "votes": 2
    },
    {
      "id": 1266193,
      "postDate": "2021-04-07T14:49:39.163Z",
      "content": "<p>these files are for my early development and benchmarking performance. i need to change my strategy for future work:</p>\n<ul>\n<li>use TIMM models and open-source transformer like huggingface, etc</li>\n<li>the reason is because TIMM models provide support for torch.jit (and onnx for tensorRT)</li>\n<li>i need to find a transformer accelerator that can run huggingface transformer model, e.g. Nvidia fast transformers, Tecent Turbo Transformers or Bytance LightSeq. Such accelerators provide speedup and also fast beam search etc.</li>\n</ul>\n<p>(Note, i will have to check license and copyright issue, but i will take care of these after i get my performance.)</p>",
      "rawMarkdown": "these files are for my early development and benchmarking performance. i need to change my strategy for future work:\n- use TIMM models and open-source transformer like huggingface, etc\n- the reason is because TIMM models provide support for torch.jit (and onnx for tensorRT)\n- i need to find a transformer accelerator that can run huggingface transformer model, e.g. Nvidia fast transformers, Tecent Turbo Transformers or Bytance LightSeq. Such accelerators provide speedup and also fast beam search etc.\n\n(Note, i will have to check license and copyright issue, but i will take care of these after i get my performance.)\n",
      "votes": 2,
      "replies": [
        {
          "id": 1266210,
          "postDate": "2021-04-07T15:01:52.493Z",
          "content": "<p>Thanks a lot for the share. Can you please tell how can inference fast which you told by using torch.jit you did faster inference!  I suffer a lot for resources!     </p>",
          "rawMarkdown": "Thanks a lot for the share. Can you please tell how can inference fast which you told by using torch.jit you did faster inference!  I suffer a lot for resources!     "
        },
        {
          "id": 1266757,
          "postDate": "2021-04-08T04:07:05.047Z",
          "content": "<p>check log file at new update at google drive: [2021-apr-06a]</p>",
          "rawMarkdown": "check log file at new update at google drive: [2021-apr-06a]",
          "votes": 1
        },
        {
          "id": 1266762,
          "postDate": "2021-04-08T04:13:22.227Z",
          "content": "<p>For the lack of knowledge still not figure it out by this portion : JIT : net = torch.jit.script(net) 😥</p>",
          "rawMarkdown": "For the lack of knowledge still not figure it out by this portion : JIT : net = torch.jit.script(net) 😥"
        }
      ]
    },
    {
      "id": 2765514,
      "postDate": "2024-04-21T07:05:38.717Z",
      "content": "<p>Thank you for your contributions. I have 1 question. ViT is pre-trained with size 224*224, dimension 768. How to continue training from pretraining effectively (when changing size and dimension)?</p>",
      "rawMarkdown": "Thank you for your contributions. I have 1 question. ViT is pre-trained with size 224*224, dimension 768. How to continue training from pretraining effectively (when changing size and dimension)?"
    },
    {
      "id": 1329882,
      "postDate": "2021-05-31T12:49:47.550Z",
      "content": "<p>Haven't tried it yet. For your second question, I would go for mixing the training set and the pseudo-labelled data rather than just the pseudo-labelled. As soon as you do pseudo-labelling you're at risk of overfitting to your pseudo-labels, and to the particular idiosyncrasies of whatever model produced them. To keep that risk lower, it's better to train with more data (and more diverse data).</p>\n<p>Would be interested to hear if anyone thinks otherwise.</p>",
      "rawMarkdown": "Haven't tried it yet. For your second question, I would go for mixing the training set and the pseudo-labelled data rather than just the pseudo-labelled. As soon as you do pseudo-labelling you're at risk of overfitting to your pseudo-labels, and to the particular idiosyncrasies of whatever model produced them. To keep that risk lower, it's better to train with more data (and more diverse data).\n\nWould be interested to hear if anyone thinks otherwise.",
      "replies": [
        {
          "id": 1330069,
          "postDate": "2021-05-31T15:12:25.087Z",
          "content": "<p>I notice you point out the problem of \"it's probably because you fed in normalized predictions.\" a few days ago. But I cant understand which this problem refer to. Can you explain this more detail? Thanks a lot. :)<br>\nBy the way, using pseudo-labelled data is very high risking things, and mixing them with train data is OK, but set this set of data less contribute to total loss may be more important?</p>",
          "rawMarkdown": "I notice you point out the problem of \"it's probably because you fed in normalized predictions.\" a few days ago. But I cant understand which this problem refer to. Can you explain this more detail? Thanks a lot. :)\nBy the way, using pseudo-labelled data is very high risking things, and mixing them with train data is OK, but set this set of data less contribute to total loss may be more important?"
        }
      ]
    },
    {
      "id": 1316269,
      "postDate": "2021-05-20T12:27:01.833Z",
      "content": "<p>Thank you for your excellent work. I retrained the [2021-apr-24] code, but found that the dev LB (Lev) is 1.29, so I think your shared pre-training model has converged, I added some code and predicted the test patch dataset, but After i submitted the csv, the LB score is above 5(worse than 07a(4.1)). I want to know why there is a huge gap between dev and test, maybe some of my prediction codes are wrong or I missed some details about your code. The only model code I modified is patch_pos, I use 9999 instead of any number exceed max_length in the patch matrix.</p>",
      "rawMarkdown": "Thank you for your excellent work. I retrained the [2021-apr-24] code, but found that the dev LB (Lev) is 1.29, so I think your shared pre-training model has converged, I added some code and predicted the test patch dataset, but After i submitted the csv, the LB score is above 5(worse than 07a(4.1)). I want to know why there is a huge gap between dev and test, maybe some of my prediction codes are wrong or I missed some details about your code. The only model code I modified is patch_pos, I use 9999 instead of any number exceed max_length in the patch matrix."
    },
    {
      "id": 1299749,
      "postDate": "2021-05-10T02:26:24.857Z",
      "content": "<p>i note that there is something wrong with the YNakamaTokenizer that i am using.<br>\nthe Tokenizer is learned from train data.</p>\n<p>there are missing dictionary items, e.g</p>\n<pre><code>key: value\n163: 100\n165: 101  #key164 is missing because there is not train data with such key\n166: 102\n</code></pre>\n<p>this exposes another problem. if there is no train data with atom numbering 164, the model cannot make this prediction in test.</p>\n<p>I have a feeling that treating this as a image caption problem seems to be wrong. we should decode the image to graph directly (e.g. molfile)</p>",
      "rawMarkdown": "i note that there is something wrong with the YNakamaTokenizer that i am using.\nthe Tokenizer is learned from train data.\n\nthere are missing dictionary items, e.g\n\n```\nkey: value\n163: 100\n165: 101  #key164 is missing because there is not train data with such key\n166: 102\n\n```\n\nthis exposes another problem. if there is no train data with atom numbering 164, the model cannot make this prediction in test.\n\nI have a feeling that treating this as a image caption problem seems to be wrong. we should decode the image to graph directly (e.g. molfile)"
    },
    {
      "id": 1299748,
      "postDate": "2021-05-10T02:26:04.310Z",
      "content": "<p><img src=\"https://i.ibb.co/VL3CY55/1620612915-1.png\" alt=\"\">.  Thanks you for your excellent code, but there may be a bug in do_valid. Since the predict[i] should start from 0 instead of 1. It means that the tearcher-forcing's LD is very small?</p>",
      "rawMarkdown": "![](https://i.ibb.co/VL3CY55/1620612915-1.png).  Thanks you for your excellent code, but there may be a bug in do_valid. Since the predict[i] should start from 0 instead of 1. It means that the tearcher-forcing's LD is very small?",
      "replies": [
        {
          "id": 1299750,
          "postDate": "2021-05-10T02:30:26.637Z",
          "content": "<p>i haven't checked in details and different versions of the code are slightly different.<br>\nyou should check of the predict and ground truth include the &lt;sos&gt; </p>\n<p>i think your reasoning is correct. it is a bug.<br>\nyes, with teach forcing, the LD distance is very small. this can be verified by the log loss, which then gives the accuracy of prediction (almost 99% if the input previous token is correct)</p>",
          "rawMarkdown": "i haven't checked in details and different versions of the code are slightly different.\nyou should check of the predict and ground truth include the <sos\\> \n\ni think your reasoning is correct. it is a bug.\nyes, with teach forcing, the LD distance is very small. this can be verified by the log loss, which then gives the accuracy of prediction (almost 99% if the input previous token is correct)",
          "votes": 1
        },
        {
          "id": 1299774,
          "postDate": "2021-05-10T03:13:51Z",
          "content": "<p>this also exposes the trick:</p>\n<ul>\n<li>there are very few mistakes (if you consider single token wise)</li>\n<li>then why is the leaderboard/cv LD large? It is because of the seq-to-seq modeling effect.  </li>\n</ul>\n<p>there should be a better way to model this problem</p>",
          "rawMarkdown": "this also exposes the trick:\n- there are very few mistakes (if you consider single token wise)\n- then why is the leaderboard/cv LD large? It is because of the seq-to-seq modeling effect.  \n\nthere should be a better way to model this problem",
          "votes": 2
        },
        {
          "id": 1305325,
          "postDate": "2021-05-13T08:00:48.397Z",
          "content": "<p>As there may be seq-to-seq modeling effect, I modified your code to support Self-critical Sequence Training<a href=\"https://arxiv.org/abs/1612.00563\" target=\"_blank\"></a> as every image caption model does. <br>\nI start from your \"/checkpoint/00922000_model.pth\"  CV(1.41)(40000validation)<br>\nAfter trained with RL, \"rl/00936100_model.pth\", CV(1.38)(40000validation)<br>\nIt's too slow…… The sample_reward - baseline_reward at every iteration:<br>\n<img src=\"https://i.ibb.co/gFkRY4X/1620892655-1.png\" alt=\"\"></p>",
          "rawMarkdown": "As there may be seq-to-seq modeling effect, I modified your code to support Self-critical Sequence Training[](https://arxiv.org/abs/1612.00563) as every image caption model does. \nI start from your \"/checkpoint/00922000_model.pth\"  CV(1.41)(40000validation)\nAfter trained with RL, \"rl/00936100_model.pth\", CV(1.38)(40000validation)\nIt's too slow...... The sample_reward - baseline_reward at every iteration:\n![](https://i.ibb.co/gFkRY4X/1620892655-1.png)",
          "votes": 1
        },
        {
          "id": 1305332,
          "postDate": "2021-05-13T08:03:42.237Z",
          "content": "<p>At first, I thought it is the secret for the top teams. But now, it's too slow, and it may be little unstable. </p>",
          "rawMarkdown": "At first, I thought it is the secret for the top teams. But now, it's too slow, and it may be little unstable. "
        },
        {
          "id": 1305341,
          "postDate": "2021-05-13T08:08:34.507Z",
          "content": "<p>\" the secret for the top teams\" … maybe not the model itself …</p>\n<p>data and post/pre processing are the key (i think)</p>\n<hr>\n<p>\" But now, it's too slow\"<br>\nyou don't have to apply it to every samples. </p>",
          "rawMarkdown": "\" the secret for the top teams\" ... maybe not the model itself ...\n\ndata and post/pre processing are the key (i think)\n\n---\n\n\" But now, it's too slow\"\nyou don't have to apply it to every samples. \n",
          "votes": 2
        },
        {
          "id": 1305359,
          "postDate": "2021-05-13T08:17:24.603Z",
          "content": "<p>The baseline reward used in those papers is  a biased estimator. </p>\n<p>Some probabilistic approaches may help ^^</p>",
          "rawMarkdown": "The baseline reward used in those papers is  a biased estimator. \n\nSome probabilistic approaches may help ^^",
          "votes": 1
        },
        {
          "id": 1305687,
          "postDate": "2021-05-13T12:30:57.750Z",
          "content": "<p>i would recommend google paper: <a href=\"https://arxiv.org/abs/2003.10580\" target=\"_blank\">https://arxiv.org/abs/2003.10580</a></p>",
          "rawMarkdown": "i would recommend google paper: https://arxiv.org/abs/2003.10580"
        }
      ]
    },
    {
      "id": 1297571,
      "postDate": "2021-05-08T06:16:18.433Z",
      "content": "<p>Trying follow [2021-apr-24] by coding to the submit function and reproduce the CV score 1.35 (without teacher forcing) for 40000 valid datapoint  ( using the shared checkpoint 00922000_model.pth )</p>\n<p>However, I am only able to get CV 1.99.  Must be making some silly mistakes somewhere. </p>\n<p>In <code>fairseq_model.py</code>, the <code>forward_argmax_decode</code> was commented out. I just assume that this part is complete and can be used without any modification.  Am I correct in assuming this ?  </p>\n<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  Can you confirm that [2021-apr-24]/ <code>forward_argmax_decode</code> can be used without modification and isn't part of the exercise ? </p>\n<p>Anyone else have similar issue with reproducing the CV for  [2021-apr-24]  ?  Cares to share any Gotcha ? </p>",
      "rawMarkdown": "Trying follow [2021-apr-24] by coding to the submit function and reproduce the CV score 1.35 (without teacher forcing) for 40000 valid datapoint  ( using the shared checkpoint 00922000_model.pth )\n\nHowever, I am only able to get CV 1.99.  Must be making some silly mistakes somewhere. \n\nIn `fairseq_model.py`, the `forward_argmax_decode` was commented out. I just assume that this part is complete and can be used without any modification.  Am I correct in assuming this ?  \n\n@hengck23  Can you confirm that [2021-apr-24]/ `forward_argmax_decode` can be used without modification and isn't part of the exercise ? \n\nAnyone else have similar issue with reproducing the CV for  [2021-apr-24]  ?  Cares to share any Gotcha ? \n\n",
      "replies": [
        {
          "id": 1298531,
          "postDate": "2021-05-09T01:35:53.947Z",
          "content": "<p>I haven't looked at the 24/04 update in any detail, but I think you take into consideration the 0.8 scaling used in the example provided. </p>",
          "rawMarkdown": "I haven't looked at the 24/04 update in any detail, but I think you take into consideration the 0.8 scaling used in the example provided. "
        },
        {
          "id": 1298597,
          "postDate": "2021-05-09T03:52:05.193Z",
          "content": "<p>it thought i mention that \"there is no submission code or jit inference. This is left as an exercise for you. It is easy to modify the code\"</p>\n<p>you can look at \"2021-apr-25\"</p>\n<hr>\n<p>you should compare results with image and patch input (for same input image size).<br>\nThe results should be smiliar.</p>\n<p>but note that the test image has salt and pepper noise. this gives unwanted patch which should be removed</p>",
          "rawMarkdown": "it thought i mention that \"there is no submission code or jit inference. This is left as an exercise for you. It is easy to modify the code\"\n\nyou can look at \"2021-apr-25\"\n\n---\n\nyou should compare results with image and patch input (for same input image size).\nThe results should be smiliar.\n\nbut note that the test image has salt and pepper noise. this gives unwanted patch which should be removed"
        }
      ]
    },
    {
      "id": 1295819,
      "postDate": "2021-05-06T18:00:16.397Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thank you for your great work!<br>\nCan I ask you a question, please.<br>\nI'm using transformer decoder and my CV score is ~4.0. However, LB score is terrible - ~34.0. <br>\nI've investigated that if I put image together with the target (like during training) I get a normal score, and if I'm predicting every token per timestamp (like when I don't know true InChI), the final score becomes much worser. <br>\nI thought it could be because I'm not using teacher forcing during training. Did you have such a problem? <br>\nOr maybe it is just my mistake in inference implementation.</p>",
      "rawMarkdown": "@hengck23 Thank you for your great work!\nCan I ask you a question, please.\nI'm using transformer decoder and my CV score is ~4.0. However, LB score is terrible - ~34.0. \nI've investigated that if I put image together with the target (like during training) I get a normal score, and if I'm predicting every token per timestamp (like when I don't know true InChI), the final score becomes much worser. \nI thought it could be because I'm not using teacher forcing during training. Did you have such a problem? \nOr maybe it is just my mistake in inference implementation.\n",
      "replies": [
        {
          "id": 1295825,
          "postDate": "2021-05-06T18:11:22.953Z",
          "content": "<p>Do you rotate test images? There are images that are taller than how wide they are. You should rotate them back to be horizontal.</p>",
          "rawMarkdown": "Do you rotate test images? There are images that are taller than how wide they are. You should rotate them back to be horizontal."
        },
        {
          "id": 1295845,
          "postDate": "2021-05-06T18:33:44.397Z",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> thank you for your reply!<br>\nYes, i do rotate images<br>\nBut the thing is that if I take some InChI from train and predict it using both approaches, I'll get different results, and the approach that is similar to inference gives worser results.<br>\nSo it seems that is not a rotation problem here…<br>\nHave you used transformer decoder and you don't have such problem? Maybe there is really something wrong in my inference implementation… The thing is I have implemented transformer from scratch, maybe the inference process is something more complicated than I think it is…</p>",
          "rawMarkdown": "@nofreewill thank you for your reply!\nYes, i do rotate images\nBut the thing is that if I take some InChI from train and predict it using both approaches, I'll get different results, and the approach that is similar to inference gives worser results.\nSo it seems that is not a rotation problem here...\nHave you used transformer decoder and you don't have such problem? Maybe there is really something wrong in my inference implementation... The thing is I have implemented transformer from scratch, maybe the inference process is something more complicated than I think it is..."
        },
        {
          "id": 1295866,
          "postDate": "2021-05-06T18:58:55.713Z",
          "content": "<p>But you have ~4.0 CV, right? You predicted that step-by-step and used the same code for predicting test images, right? Or am I missing something?<br>\nI don't really understand what you mean by \"putting image together with target\". Do you mean that you get your 4.0 score if you use teacher forcing in your validation?</p>",
          "rawMarkdown": "But you have ~4.0 CV, right? You predicted that step-by-step and used the same code for predicting test images, right? Or am I missing something?\nI don't really understand what you mean by \"putting image together with target\". Do you mean that you get your 4.0 score if you use teacher forcing in your validation?"
        },
        {
          "id": 1295910,
          "postDate": "2021-05-06T19:55:35.320Z",
          "content": "<p>During training I use teacher forcing permanently. <br>\nThen, on validation data, because I know ground truth, I can compare two strategies:</p>\n<ol>\n<li>Predict InChI using teacher forcing (like in train);</li>\n<li>Predict InChI step-by-step, like i would do it when predicting test images.<br>\nAnd the 2 approach gives much worser results (~30.0) compared to first (~4.0) .<br>\nSo, in fact, yes, as you said, I get 4.0 score if I use teacher forcing in my validation.<br>\nSo I thought that the problem is that I don't use step-by-step strategy during training. But in official paper that strategy isn't used and it still gives good results, so it maybe just my wrong implementation of step-by-step prediction. <br>\nThat's why I wanted to ask if it is right to use teacher forcing during training and ignore step-by-step approach, or the second one should be implemented. But maybe it's just something wrong with my step-by-step implementation.<br>\nI understand that I wasn't very good at explaining my problem, but I hope it is more clear now.</li>\n</ol>",
          "rawMarkdown": "During training I use teacher forcing permanently. \nThen, on validation data, because I know ground truth, I can compare two strategies:\n1. Predict InChI using teacher forcing (like in train);\n2. Predict InChI step-by-step, like i would do it when predicting test images.\nAnd the 2 approach gives much worser results (~30.0) compared to first (~4.0) .\nSo, in fact, yes, as you said, I get 4.0 score if I use teacher forcing in my validation.\nSo I thought that the problem is that I don't use step-by-step strategy during training. But in official paper that strategy isn't used and it still gives good results, so it maybe just my wrong implementation of step-by-step prediction. \nThat's why I wanted to ask if it is right to use teacher forcing during training and ignore step-by-step approach, or the second one should be implemented. But maybe it's just something wrong with my step-by-step implementation.\nI understand that I wasn't very good at explaining my problem, but I hope it is more clear now."
        },
        {
          "id": 1295924,
          "postDate": "2021-05-06T20:06:26.093Z",
          "content": "<p>Using the parallel teacher forcing is totally fine and far faster for transformer anyway. So as long as you don't have insane amounts of computing power you should always go for teacher forcing. It will lead to far better performance per compute invested.</p>\n<p>Your sequential inference performance might be worse because of some issue with your positional embeddings. During teacher forcing positional embeddings tend to hurt performance so it won't be obvious. But during sequential generation you need them. So check your positional embeddings.<br>\nOther reasons are of course bugs in your code, especially if you use some kind of memory for the last inputs to speed up the inference (not computing past sequence hidden embeddings twice).</p>\n<p>TLDR: Teacher forcing is a good choice. Your inference issue might be due to wrongly applied positional embeddings or bugs in inference code (inside the transformer itself).</p>",
          "rawMarkdown": "Using the parallel teacher forcing is totally fine and far faster for transformer anyway. So as long as you don't have insane amounts of computing power you should always go for teacher forcing. It will lead to far better performance per compute invested.\n\nYour sequential inference performance might be worse because of some issue with your positional embeddings. During teacher forcing positional embeddings tend to hurt performance so it won't be obvious. But during sequential generation you need them. So check your positional embeddings.\nOther reasons are of course bugs in your code, especially if you use some kind of memory for the last inputs to speed up the inference (not computing past sequence hidden embeddings twice).\n\nTLDR: Teacher forcing is a good choice. Your inference issue might be due to wrongly applied positional embeddings or bugs in inference code (inside the transformer itself).",
          "votes": 1
        },
        {
          "id": 1295943,
          "postDate": "2021-05-06T20:23:30.607Z",
          "content": "<p>I think <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> had a great experience with hengs code. He gives answers to all of them in this discussion.</p>",
          "rawMarkdown": "I think @nofreewill had a great experience with hengs code. He gives answers to all of them in this discussion."
        },
        {
          "id": 1296162,
          "postDate": "2021-05-07T03:19:50.467Z",
          "content": "<p>I haven't tried to compute LD with teacher forcing, but 4.0 sounds like a lot.<br>\nRemember that by teacher forcing you always give your model the right question:</p>\n<ul>\n<li>based on the image and that this is what I've written down already, what's the next thing to write down? Where you always give it the correct thing what already was written down.</li>\n</ul>\n<p>But when you do it step by step, every mistake you make will remain there in all your future questions regarding that image.</p>\n<p>So if your model makes 4.0 worth of LD mistakes even if it always gets the right question, it is no surprise - in my opinion - that it makes 34.0 LD with a lot of wrong questions.</p>",
          "rawMarkdown": "I haven't tried to compute LD with teacher forcing, but 4.0 sounds like a lot.\nRemember that by teacher forcing you always give your model the right question:\n- based on the image and that this is what I've written down already, what's the next thing to write down? Where you always give it the correct thing what already was written down.\n\nBut when you do it step by step, every mistake you make will remain there in all your future questions regarding that image.\n\nSo if your model makes 4.0 worth of LD mistakes even if it always gets the right question, it is no surprise - in my opinion - that it makes 34.0 LD with a lot of wrong questions.",
          "votes": 2
        },
        {
          "id": 1296319,
          "postDate": "2021-05-07T07:24:15.233Z",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> I have my very own code from the start</p>",
          "rawMarkdown": "@morizin I have my very own code from the start",
          "votes": 1
        },
        {
          "id": 1296349,
          "postDate": "2021-05-07T07:58:35.990Z",
          "content": "<p>the first step is to make sure your code has no bug.</p>\n<ul>\n<li>during training forward, you input true token to predict the next token.</li>\n<li>you should select some train samples that has 100% accuracy</li>\n<li>during inference forward, you are doing autoregressive prediction. you use the predicted token as input to estimate the next token. </li>\n<li>but since this selected train sample has 100% accuracy, the inference forward and training forward should give extract the same results. use this to check if your code has bug or not</li>\n</ul>",
          "rawMarkdown": "the first step is to make sure your code has no bug.\n- during training forward, you input true token to predict the next token.\n- you should select some train samples that has 100% accuracy\n- during inference forward, you are doing autoregressive prediction. you use the predicted token as input to estimate the next token. \n- but since this selected train sample has 100% accuracy, the inference forward and training forward should give extract the same results. use this to check if your code has bug or not",
          "votes": 3
        },
        {
          "id": 1296362,
          "postDate": "2021-05-07T08:06:45.433Z",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> Yes, I agree with you, it really makes sense.<br>\nAnd after some investigation on my step-by-step implementation I really think that the problem is not in it.<br>\nI didn't think that 4.0 mistake with teacher forcing will result in such a gap between step-by-step prediction, but it seems that the problem is here. </p>\n<p>Thank you all every much for your help. Will try to increase capacity of my model…</p>",
          "rawMarkdown": "@nofreewill Yes, I agree with you, it really makes sense.\nAnd after some investigation on my step-by-step implementation I really think that the problem is not in it.\nI didn't think that 4.0 mistake with teacher forcing will result in such a gap between step-by-step prediction, but it seems that the problem is here. \n\nThank you all every much for your help. Will try to increase capacity of my model...",
          "votes": 1
        },
        {
          "id": 1296376,
          "postDate": "2021-05-07T08:25:39.340Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <br>\nThank you for your reply!</p>\n<p>Yes, I did this steps that you told about and the results for 100% accuracy are really the same. <br>\nSo it really should be the problem <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> was talking about.</p>\n<p>I saw in your discussion that line: (CV-teacher forcing: 1.27, CV-without teacher forcing : about 1.35 for 40000 validation set).<br>\nSo you tried both approaches too, as far as I can understand. Your results are far better then mine.<br>\nIt seems that the gap in 1.27 and 4.0 in teacher forcing gives a really huge impact on step-by-step prediction score. But maybe there is also something in how models adapt.</p>",
          "rawMarkdown": "@hengck23 \nThank you for your reply!\n\nYes, I did this steps that you told about and the results for 100% accuracy are really the same. \nSo it really should be the problem @nofreewill was talking about.\n\nI saw in your discussion that line: (CV-teacher forcing: 1.27, CV-without teacher forcing : about 1.35 for 40000 validation set).\nSo you tried both approaches too, as far as I can understand. Your results are far better then mine.\nIt seems that the gap in 1.27 and 4.0 in teacher forcing gives a really huge impact on step-by-step prediction score. But maybe there is also something in how models adapt.",
          "votes": 1
        },
        {
          "id": 1296455,
          "postDate": "2021-05-07T09:10:38.533Z",
          "content": "<p>CV-teacher forcing : the result with train forward<br>\nCV-without teacher forcing  :the result with inference forward</p>\n<p>you should break down the score according to seq length to debug the root of the problem.</p>",
          "rawMarkdown": "CV-teacher forcing : the result with train forward\nCV-without teacher forcing  :the result with inference forward\n\nyou should break down the score according to seq length to debug the root of the problem."
        },
        {
          "id": 1296464,
          "postDate": "2021-05-07T09:21:15.400Z",
          "content": "<p>Which CV result should we trust <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>",
          "rawMarkdown": "Which CV result should we trust @hengck23 "
        },
        {
          "id": 1296594,
          "postDate": "2021-05-07T11:08:37.503Z",
          "content": "<p>CV-without teacher forcing :the result with inference forward</p>",
          "rawMarkdown": "CV-without teacher forcing :the result with inference forward"
        },
        {
          "id": 1297747,
          "postDate": "2021-05-08T09:23:42.893Z",
          "content": "<p>You have to trust the CV without teacher forcing, since this is the procedure which will be used to evaluate the test set. </p>",
          "rawMarkdown": "You have to trust the CV without teacher forcing, since this is the procedure which will be used to evaluate the test set. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1295470,
      "postDate": "2021-05-06T13:37:09.597Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> have used batch size of 64 for inference part. Only 1.4 GB of GPU is being used in that case. Can I increase the batch size to 128, or will it have any effect on my predictions ? Is it like we have to use the same batch size for training and inference part ? </p>",
      "rawMarkdown": "@hengck23 have used batch size of 64 for inference part. Only 1.4 GB of GPU is being used in that case. Can I increase the batch size to 128, or will it have any effect on my predictions ? Is it like we have to use the same batch size for training and inference part ? ",
      "replies": [
        {
          "id": 1295632,
          "postDate": "2021-05-06T15:42:06.230Z",
          "content": "<p>Imagine if batch size had any effect on predictions (apart from speed),  that would mean your model is  still learning during inference ^^</p>",
          "rawMarkdown": "Imagine if batch size had any effect on predictions (apart from speed),  that would mean your model is  still learning during inference ^^",
          "votes": 2
        },
        {
          "id": 1295647,
          "postDate": "2021-05-06T15:47:46.803Z",
          "content": "<p>Batch size does not effect inference outcome in any way -&gt; you can increase.<br>\nIt should not. If it does, than the code is wrong.</p>",
          "rawMarkdown": "Batch size does not effect inference outcome in any way -> you can increase.\nIt should not. If it does, than the code is wrong.",
          "votes": 1
        },
        {
          "id": 1295720,
          "postDate": "2021-05-06T16:50:53.237Z",
          "content": "<p><code>will it have any effect on my predictions?</code><br>\nBy having effect I meant the time of inference.</p>",
          "rawMarkdown": " `will it have any effect on my predictions?`\nBy having effect I meant the time of inference."
        },
        {
          "id": 1295732,
          "postDate": "2021-05-06T17:01:29.083Z",
          "content": "<p>Oh, that's another story.<br>\nType <code>watch nvidia-smi</code> in the terminal while you run inference and if it says that your GPU-Util is much below 100% then yes, increasing the batch size should make your inference faster.</p>",
          "rawMarkdown": "Oh, that's another story.\nType `watch nvidia-smi` in the terminal while you run inference and if it says that your GPU-Util is much below 100% then yes, increasing the batch size should make your inference faster.",
          "votes": 1
        },
        {
          "id": 1296236,
          "postDate": "2021-05-07T05:17:40.160Z",
          "content": "<p>Thanks it did increase my inference speed. One more query, how can I train the model with 320 image size. I checked in <code>tnt.py</code> and changed the input size to (3, 320, 320). Do I have to change anything else or is this is only the thing to be done ? Because at many places I have seen <code>image_size = 224</code> for ex: in <code>fairseq_transformer.py</code> . </p>",
          "rawMarkdown": "Thanks it did increase my inference speed. One more query, how can I train the model with 320 image size. I checked in `tnt.py` and changed the input size to (3, 320, 320). Do I have to change anything else or is this is only the thing to be done ? Because at many places I have seen `image_size = 224` for ex: in `fairseq_transformer.py` . ",
          "votes": 1
        },
        {
          "id": 1296242,
          "postDate": "2021-05-07T05:26:23.557Z",
          "content": "<p>Sadly, I have zero clue about that.<br>\n<a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, can you help out?</p>",
          "rawMarkdown": "Sadly, I have zero clue about that.\n@hengck23, can you help out?"
        },
        {
          "id": 1296357,
          "postDate": "2021-05-07T08:03:46.353Z",
          "content": "<p>\":changed the input size to (3, 320, 320)\"</p>\n<p>in theory, you need to interpolate the position embedding weights. but i just re-initialize with random values when in change from 224 to 320 and let it relearn the values.</p>\n<p>apart from this, i think there not other changes.</p>\n<p>tip:</p>\n<p>just try to load the 224 trained model state dict to the modified 320 modified net.<br>\nif there is any error reported due to the wrong size, you can delete the dict key (i.e. retrain from scratch) or copy/repeat/interpolate the values from old state dict to new state dict</p>",
          "rawMarkdown": "\":changed the input size to (3, 320, 320)\"\n\nin theory, you need to interpolate the position embedding weights. but i just re-initialize with random values when in change from 224 to 320 and let it relearn the values.\n\napart from this, i think there not other changes.\n\ntip:\n\njust try to load the 224 trained model state dict to the modified 320 modified net.\nif there is any error reported due to the wrong size, you can delete the dict key (i.e. retrain from scratch) or copy/repeat/interpolate the values from old state dict to new state dict",
          "votes": 2
        },
        {
          "id": 1296433,
          "postDate": "2021-05-07T09:00:09.853Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> I was looking at the run_submit.py file. And I am slightly confused about the following</p>\n<pre><code>       if 'remote' in mode: #1_616_107\n            df_valid = make_fold('test')\n            if gpu_no==0 :  df_valid = df_valid[        : 400_000]\n            if gpu_no==1 :  df_valid = df_valid[ 400_000: 800_000]\n            if gpu_no==2 :  df_valid = df_valid[ 800_000:1200_000]\n            if gpu_no==3 :  df_valid = df_valid[1200_000]\n</code></pre>\n<p>Why do you run inference like this? Is it because you have multiple GPUs and you pass a test df block to each? Also, the code does not cover all the test rows, or does it? </p>\n<p>Sorry if this is a silly question, but I just commented out the subset lines and ran test inference. </p>",
          "rawMarkdown": "@hengck23 I was looking at the run_submit.py file. And I am slightly confused about the following\n\n```\n       if 'remote' in mode: #1_616_107\n            df_valid = make_fold('test')\n            if gpu_no==0 :  df_valid = df_valid[        : 400_000]\n            if gpu_no==1 :  df_valid = df_valid[ 400_000: 800_000]\n            if gpu_no==2 :  df_valid = df_valid[ 800_000:1200_000]\n            if gpu_no==3 :  df_valid = df_valid[1200_000]\n```\n\nWhy do you run inference like this? Is it because you have multiple GPUs and you pass a test df block to each? Also, the code does not cover all the test rows, or does it? \n\nSorry if this is a silly question, but I just commented out the subset lines and ran test inference. "
        },
        {
          "id": 1296440,
          "postDate": "2021-05-07T09:05:14.913Z",
          "content": "<p>Yes he is doing it because he want to run inference in multiple GPUs<br>\nit is too slow to run inference (might take 7 to 8 hours) on a single GPU.<br>\nthis way we can run separate scripts for each GPU hence it speeds up 4x.</p>",
          "rawMarkdown": "Yes he is doing it because he want to run inference in multiple GPUs\nit is too slow to run inference (might take 7 to 8 hours) on a single GPU.\nthis way we can run separate scripts for each GPU hence it speeds up 4x."
        },
        {
          "id": 1296446,
          "postDate": "2021-05-07T09:07:28.043Z",
          "content": "<blockquote>\n  <p>code does not cover all the test rows</p>\n</blockquote>\n<p><code>if gpu_no==3 :  df_valid = df_valid[1200_000]</code><br>\nlacks a colon, right?<br>\nShould be<br>\n<code>if gpu_no==3 :  df_valid = df_valid[1200_000</code><strong>:</strong><code>]</code><br>\ninstead. ?</p>",
          "rawMarkdown": "> code does not cover all the test rows\n\n`if gpu_no==3 :  df_valid = df_valid[1200_000]`\nlacks a colon, right?\nShould be\n`if gpu_no==3 :  df_valid = df_valid[1200_000`**:**`]`\ninstead. ?",
          "votes": 2
        },
        {
          "id": 1296453,
          "postDate": "2021-05-07T09:10:24.637Z",
          "content": "<p>Yes it does lacks a colon. I have 3 GPUs so I have just divided all the predictions in 3 parts.</p>",
          "rawMarkdown": "Yes it does lacks a colon. I have 3 GPUs so I have just divided all the predictions in 3 parts."
        },
        {
          "id": 1296475,
          "postDate": "2021-05-07T09:29:34.710Z",
          "content": "<p>Thanks, this helps! There is so much I don't yet understand in Heng's code. </p>\n<p>Hope to sit and delve over the weekend. </p>",
          "rawMarkdown": "Thanks, this helps! There is so much I don't yet understand in Heng's code. \n\nHope to sit and delve over the weekend. "
        },
        {
          "id": 1307878,
          "postDate": "2021-05-14T18:11:05.920Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1309061,
          "postDate": "2021-05-15T16:25:25.753Z",
          "content": "<p><a href=\"https://www.kaggle.com/atharvaingle\" target=\"_blank\">@atharvaingle</a> did you need any other changes apart from that one line to try on larger images? I would think that you also need to change the img_size arg in TNT and then handle the mismatch in pos some way</p>",
          "rawMarkdown": "@atharvaingle did you need any other changes apart from that one line to try on larger images? I would think that you also need to change the img_size arg in TNT and then handle the mismatch in pos some way"
        }
      ]
    },
    {
      "id": 1293694,
      "postDate": "2021-05-05T05:21:45.680Z",
      "content": "<p>Figured out.   Just need to be more patient and train for more epochs</p>",
      "rawMarkdown": "Figured out.   Just need to be more patient and train for more epochs\n\n",
      "replies": [
        {
          "id": 1293723,
          "postDate": "2021-05-05T05:55:04.307Z",
          "content": "<p>I think your learning rate is much too high. Try to start with 1e-4, reducing on plateau. </p>",
          "rawMarkdown": "I think your learning rate is much too high. Try to start with 1e-4, reducing on plateau. ",
          "votes": 1
        },
        {
          "id": 1293748,
          "postDate": "2021-05-05T06:14:38.960Z",
          "content": "<p>Sorry. that was a typo.</p>\n<p>I did start with lr = 1e-4 / 2 from 00266000_model.pth</p>\n<p>Were you able to replicate the 224 -&gt; 320 improvement starting from <code>00266000_model.pth</code> ? If so, do you mind sharing how many iteration did you train for ? </p>",
          "rawMarkdown": "Sorry. that was a typo.\n\nI did start with lr = 1e-4 / 2 from 00266000_model.pth\n\nWere you able to replicate the 224 -> 320 improvement starting from `00266000_model.pth` ? If so, do you mind sharing how many iteration did you train for ? ",
          "votes": 1
        },
        {
          "id": 1293822,
          "postDate": "2021-05-05T07:27:54.190Z",
          "content": "<p>Try 224 with lower batch size, too. If that makes a huge difference then that might be the boogie.</p>",
          "rawMarkdown": "Try 224 with lower batch size, too. If that makes a huge difference then that might be the boogie.",
          "votes": 1
        },
        {
          "id": 1294019,
          "postDate": "2021-05-05T11:08:20.013Z",
          "content": "<p>Your cv score doesn't look correlated with your loss, have you checked your validation function? There was a bug iirc…</p>",
          "rawMarkdown": "Your cv score doesn't look correlated with your loss, have you checked your validation function? There was a bug iirc...",
          "votes": 1
        },
        {
          "id": 1294058,
          "postDate": "2021-05-05T11:40:13.343Z",
          "content": "<p><a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> do you mean the do_valid function is boogie after heng updated the validation code is boogie out in <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231190#1270712\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "@datafan07 do you mean the do_valid function is boogie after heng updated the validation code is boogie out in [here](https://www.kaggle.com/c/bms-molecular-translation/discussion/231190#1270712)"
        },
        {
          "id": 1294067,
          "postDate": "2021-05-05T11:45:04.170Z",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> Are you mocking me using the word \"boogie\"? :DD</p>",
          "rawMarkdown": "@morizin Are you mocking me using the word \"boogie\"? :DD",
          "votes": 3
        },
        {
          "id": 1294071,
          "postDate": "2021-05-05T11:49:25.467Z",
          "content": "<p>Nope :DD just found that word useful</p>",
          "rawMarkdown": "Nope :DD just found that word useful",
          "votes": 1
        },
        {
          "id": 1294491,
          "postDate": "2021-05-05T17:30:57.810Z",
          "content": "<p>After swapping out the buggy <code>do_valid</code> code, my CV lb metric ( the one displayed during training ) was cut in half.  But this code only affects the VALID LB score displayed during training, correct ?   </p>\n<p>If the CV socre produced by <code>run_submit.py</code> is above 5.0, that means my model was not trained correctly as result of other things not related to the buggy <code>do_valid</code>.</p>",
          "rawMarkdown": "After swapping out the buggy `do_valid` code, my CV lb metric ( the one displayed during training ) was cut in half.  But this code only affects the VALID LB score displayed during training, correct ?   \n\nIf the CV socre produced by `run_submit.py` is above 5.0, that means my model was not trained correctly as result of other things not related to the buggy `do_valid`."
        },
        {
          "id": 1294674,
          "postDate": "2021-05-05T20:08:48.420Z",
          "content": "<p>Whats the bug? Sorry but the link provided above just takes me back to the top of this page.</p>",
          "rawMarkdown": "Whats the bug? Sorry but the link provided above just takes me back to the top of this page."
        },
        {
          "id": 1295764,
          "postDate": "2021-05-06T17:26:40.677Z",
          "content": "<p><a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a><br>\nThat link should redirect you to a comment under this discussion.<br>\nSearch for <code>there is an obvious bug in all validation code during training</code> with ctrl+f to find it.</p>",
          "rawMarkdown": "@pheadrus\nThat link should redirect you to a comment under this discussion.\nSearch for `there is an obvious bug in all validation code during training` with ctrl+f to find it.",
          "votes": 2
        },
        {
          "id": 1296342,
          "postDate": "2021-05-07T07:50:02.113Z",
          "content": "<p>Yep, I got it! Thanks. </p>",
          "rawMarkdown": "Yep, I got it! Thanks. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1292300,
      "postDate": "2021-05-03T20:13:46.383Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> How do you use suggest to use 320X320 size, the pretrained transformers is for 224X224. Or do you write another model def with needed size and use those from scratch? I am asking since the code doesn't implement the 320 size version. Sorry if this was already answered.</p>",
      "rawMarkdown": "@hengck23 How do you use suggest to use 320X320 size, the pretrained transformers is for 224X224. Or do you write another model def with needed size and use those from scratch? I am asking since the code doesn't implement the 320 size version. Sorry if this was already answered.",
      "replies": [
        {
          "id": 1292312,
          "postDate": "2021-05-03T20:21:55.103Z",
          "content": "<p>you can just change the image size argument of the TNT module in the code<br>\nand use the previous train(weights trained for image)weights as initial weights at this point<br>\nthere would be a mismatch between the model and the state dict , since we are taking weight which is trained on images which have unequal patches. to able to solve the issue of mismatch, <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a></p>\n<pre><code>    if initial_checkpoint is not None:\n        f = torch.load(initial_checkpoint, map_location=lambda storage, loc: storage)\n        start_iteration = f['iteration']\n        start_epoch     = f['epoch']\n        state_dict = f['state_dict']\n        del state_dict['cnn.e.patch_pos']\n        del state_dict['text_pos.pos']\n        net.load_state_dict(state_dict, strict=False)  # True\n        #net.load_state_dict(state_dict, strict=True)  # True\n</code></pre>\n<p><br>\nin this <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231190#1276290\" target=\"_blank\">discussion</a></p>",
          "rawMarkdown": "you can just change the image size argument of the TNT module in the code\nand use the previous train(weights trained for image)weights as initial weights at this point\nthere would be a mismatch between the model and the state dict , since we are taking weight which is trained on images which have unequal patches. to able to solve the issue of mismatch, @hengck23\n\n```\n    if initial_checkpoint is not None:\n        f = torch.load(initial_checkpoint, map_location=lambda storage, loc: storage)\n        start_iteration = f['iteration']\n        start_epoch     = f['epoch']\n        state_dict = f['state_dict']\n        del state_dict['cnn.e.patch_pos']\n        del state_dict['text_pos.pos']\n        net.load_state_dict(state_dict, strict=False)  # True\n        #net.load_state_dict(state_dict, strict=True)  # True\n``` \nin this [discussion](https://www.kaggle.com/c/bms-molecular-translation/discussion/231190#1276290)",
          "votes": 1
        },
        {
          "id": 1292349,
          "postDate": "2021-05-03T21:08:05.470Z",
          "content": "<p>But you can't use imagenet TNT weights since they're for 224, right?</p>",
          "rawMarkdown": "But you can't use imagenet TNT weights since they're for 224, right?"
        },
        {
          "id": 1294053,
          "postDate": "2021-05-05T11:38:06.980Z",
          "content": "<p>maybe I don't know. because I think we can still add imagenet weights by removing <br>\ndel state_dict['cnn.e.patch_pos']<br>\nfrom the state dict of TNT imagenet weights<br>\ndon't know because I didn't tried</p>",
          "rawMarkdown": "maybe I don't know. because I think we can still add imagenet weights by removing \ndel state_dict['cnn.e.patch_pos']\nfrom the state dict of TNT imagenet weights\ndon't know because I didn't tried"
        }
      ]
    },
    {
      "id": 1291481,
      "postDate": "2021-05-03T04:55:32.687Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<p>I tried to run [2021-apr-07b] without success.   Any suggestion on how to resolve these errors ? </p>\n<p>1 )  I tried running <code>run_check_fairseq_model.py</code>   But  <code>/checkpoint/00235000_model.pth</code> can not be found in your GDrive. I tried to substitute <code>00266000_model.pth</code> you had in  [2021-apr-07b] folder, but I got this error</p>\n<pre><code>RuntimeError: CUDA error: CUBLAS_STATUS_EXECUTION_FAILED when calling `cublasSgemm( handle, opa, opb, m, n, k, &amp;alpha, a, lda, b, ldb, &amp;beta, c, ldc)`\n</code></pre>\n<p>Not sure if this has anything to do with the wrong checkpoint</p>\n<p>2)  i tried running <code>run_train.py</code>.  Same problem with missing checkpoint. I just set initial_checkpoint = None <br>\nBut I encoutered this error </p>\n<pre><code> untimeError: CUDA error: CUBLAS_STATUS_EXECUTION_FAILED when calling `cublasGemmEx( handle, opa, opb, m, n, k, &amp;falpha, a, CUDA_R_16F, lda, b, CUDA_R_16F, ldb, &amp;fbeta, c, CUDA_R_16F, ldc, CUDA_R_32F, CUBLAS_GEMM_DFALT_TENSOR_OP)`\n</code></pre>",
      "rawMarkdown": "@hengck23 \n\nI tried to run [2021-apr-07b] without success.   Any suggestion on how to resolve these errors ? \n\n\n1 )  I tried running `run_check_fairseq_model.py`   But  `/checkpoint/00235000_model.pth` can not be found in your GDrive. I tried to substitute `00266000_model.pth` you had in  [2021-apr-07b] folder, but I got this error\n\n```\nRuntimeError: CUDA error: CUBLAS_STATUS_EXECUTION_FAILED when calling `cublasSgemm( handle, opa, opb, m, n, k, &alpha, a, lda, b, ldb, &beta, c, ldc)`\n\n```\n\nNot sure if this has anything to do with the wrong checkpoint\n\n2)  i tried running `run_train.py`.  Same problem with missing checkpoint. I just set initial_checkpoint = None \nBut I encoutered this error \n\n```\n untimeError: CUDA error: CUBLAS_STATUS_EXECUTION_FAILED when calling `cublasGemmEx( handle, opa, opb, m, n, k, &falpha, a, CUDA_R_16F, lda, b, CUDA_R_16F, ldb, &fbeta, c, CUDA_R_16F, ldc, CUDA_R_32F, CUBLAS_GEMM_DFALT_TENSOR_OP)`\n```\n\n\n\n\n",
      "replies": [
        {
          "id": 1291869,
          "postDate": "2021-05-03T12:01:41.017Z",
          "content": "<p>you can google for \"CUBLAS_STATUS_EXECUTION_FAILED when calling <code>cublasGemmEx( handle, opa, opb, m, n, k, &amp;falpha, a, CUDA_R_16F, lda, b, CUDA_R_16F, ldb, &amp;fbeta, c, CUDA_R_16F, ldc, CUDA_R_32F, CUBLAS_GEMM_DFALT_TENSOR_OP)</code>\"</p>\n<p>i think it is related to pytorch version, GPU card, fp16 and out of memory.<br>\ni would suggest you reduce the batch size to 1 or 2 first</p>",
          "rawMarkdown": "you can google for \"CUBLAS_STATUS_EXECUTION_FAILED when calling `cublasGemmEx( handle, opa, opb, m, n, k, &falpha, a, CUDA_R_16F, lda, b, CUDA_R_16F, ldb, &fbeta, c, CUDA_R_16F, ldc, CUDA_R_32F, CUBLAS_GEMM_DFALT_TENSOR_OP)`\"\n\n\ni think it is related to pytorch version, GPU card, fp16 and out of memory.\ni would suggest you reduce the batch size to 1 or 2 first"
        },
        {
          "id": 1292191,
          "postDate": "2021-05-03T18:14:13.467Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<p>Thanks for the info. I was able to resolve the error and run the training script.</p>\n<p>1 additional question:</p>\n<p>How should I access the initial check_point referenced in '2021-apr-07b/run_train.py' ?</p>\n<p>(IE: 00235000_model.pth  as referenced in <code>2021-apr-07b/log.train.txt</code> )</p>\n<p>I have checked all zip files in the shared GDrive and could not locate this checkpoint.</p>\n<p>Is the checkpoint model result of meaningful pre-training ?  Do you expect starting from scratch significantly degrades the training result ? </p>",
          "rawMarkdown": "@hengck23 \n\nThanks for the info. I was able to resolve the error and run the training script.\n\n1 additional question:\n\nHow should I access the initial check_point referenced in '2021-apr-07b/run_train.py' ?\n\n(IE: 00235000_model.pth  as referenced in `2021-apr-07b/log.train.txt` )\n\nI have checked all zip files in the shared GDrive and could not locate this checkpoint.\n\nIs the checkpoint model result of meaningful pre-training ?  Do you expect starting from scratch significantly degrades the training result ? "
        },
        {
          "id": 1292206,
          "postDate": "2021-05-03T18:35:39.770Z",
          "content": "<p>you have <code>00266000_model.pth</code> weights in the drive you can use them</p>",
          "rawMarkdown": "you have `00266000_model.pth` weights in the drive you can use them"
        },
        {
          "id": 1292345,
          "postDate": "2021-05-03T21:03:10.900Z",
          "content": "<p>Not using the shared code here, but it seems entirely possible, that somebody might get these error:</p>\n<blockquote>\n  <p>Unable to find a valid cuDNN algorithm to run convolution</p>\n</blockquote>\n<p>That most likely means that your GPU ran out of VRAM (memory). Had the error just today. Although googling the error message usually helps (like in my case).<br>\nJust posting this here if somebody has the same error and wonders about the cause.</p>",
          "rawMarkdown": "Not using the shared code here, but it seems entirely possible, that somebody might get these error:\n> Unable to find a valid cuDNN algorithm to run convolution\n\nThat most likely means that your GPU ran out of VRAM (memory). Had the error just today. Although googling the error message usually helps (like in my case).\nJust posting this here if somebody has the same error and wonders about the cause."
        }
      ]
    },
    {
      "id": 1285508,
      "postDate": "2021-04-27T03:17:12.150Z",
      "content": "<p>Can i  ask that where is the file \"df_train.more.csv.pickle\" ? <br>\n🤔</p>",
      "rawMarkdown": "Can i  ask that where is the file \"df_train.more.csv.pickle\" ? \n🤔",
      "replies": [
        {
          "id": 1285632,
          "postDate": "2021-04-27T06:29:28.337Z",
          "content": "<p>\"/2021-apr-06/data\"</p>",
          "rawMarkdown": "\"/2021-apr-06/data\"",
          "votes": 1
        },
        {
          "id": 1285635,
          "postDate": "2021-04-27T06:37:26.110Z",
          "content": "<p>thanks for reply :D</p>",
          "rawMarkdown": "thanks for reply :D"
        },
        {
          "id": 1289595,
          "postDate": "2021-05-01T07:07:00.750Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1278561,
      "postDate": "2021-04-20T03:22:40.577Z",
      "content": "<p>good for clustering:</p>\n<p>SOFT EDIT DISTANCE FOR DIFFERENTIABLE COMPARISON OF<br>\nSYMBOLIC SEQUENCES</p>\n<p>Soft_edit_distance_for_differentiable_comparison_o.pdf</p>",
      "rawMarkdown": "good for clustering:\n\nSOFT EDIT DISTANCE FOR DIFFERENTIABLE COMPARISON OF\nSYMBOLIC SEQUENCES\n\nSoft_edit_distance_for_differentiable_comparison_o.pdf",
      "replies": [
        {
          "id": 1279172,
          "postDate": "2021-04-20T16:57:47.440Z",
          "content": "<p>I think you forgot to link the pdf you mentioned. <a href=\"https://arxiv.org/pdf/1904.12562.pdf\" target=\"_blank\">Soft edit distance paper link.</a></p>",
          "rawMarkdown": "I think you forgot to link the pdf you mentioned. [Soft edit distance paper link.](https://arxiv.org/pdf/1904.12562.pdf)"
        },
        {
          "id": 1323853,
          "postDate": "2021-05-26T13:19:08.877Z",
          "content": "<p>That is an interesting paper, but after experimenting a bit with it, I realized that it can be applied to rather short sequences (20-30 characters or less), as it contains exponentiation to the power of inchi length which can be longer than 200.</p>",
          "rawMarkdown": "That is an interesting paper, but after experimenting a bit with it, I realized that it can be applied to rather short sequences (20-30 characters or less), as it contains exponentiation to the power of inchi length which can be longer than 200."
        },
        {
          "id": 1324228,
          "postDate": "2021-05-26T18:06:29.767Z",
          "content": "<p><a href=\"https://www.kaggle.com/cepheidq\" target=\"_blank\">@cepheidq</a> I start to think (s)he does it intentionally to make others involved…</p>",
          "rawMarkdown": "@cepheidq I start to think (s)he does it intentionally to make others involved..."
        }
      ]
    },
    {
      "id": 1276789,
      "postDate": "2021-04-18T01:29:53.077Z",
      "content": "<p>I have a simple question, forgive me if It is repeated. can you submit your result if you do inference on local machines? if so  how?  thanks!</p>",
      "rawMarkdown": "I have a simple question, forgive me if It is repeated. can you submit your result if you do inference on local machines? if so  how?  thanks!",
      "replies": [
        {
          "id": 1276799,
          "postDate": "2021-04-18T02:07:11.990Z",
          "content": "<p>Go to this page: <a href=\"https://www.kaggle.com/c/bms-molecular-translation/submissions\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/submissions</a> and click on \"submit predictions\". </p>",
          "rawMarkdown": "Go to this page: https://www.kaggle.com/c/bms-molecular-translation/submissions and click on \"submit predictions\". "
        },
        {
          "id": 1277155,
          "postDate": "2021-04-18T13:12:55.667Z",
          "content": "<p>thank you! I did not pay attention.  I thought as usual that we have to submit from the output of committed notebook.</p>",
          "rawMarkdown": "thank you! I did not pay attention.  I thought as usual that we have to submit from the output of committed notebook."
        },
        {
          "id": 1277156,
          "postDate": "2021-04-18T13:12:55.667Z",
          "content": "<p>thank you! I did not pay attention.  I thought as usual that we have to submit from the output of committed notebook.</p>",
          "rawMarkdown": "thank you! I did not pay attention.  I thought as usual that we have to submit from the output of committed notebook."
        }
      ]
    },
    {
      "id": 1273385,
      "postDate": "2021-04-14T09:42:57.910Z",
      "content": "<p>Hi Everyone, <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>,<br>\nanyone on this panel please clear me<br>\nIn file dataset_224.py, <br>\nIs he taking the same files of our dataset or any resized images<br>\nif he uses original sized images, why does he comment the resize function</p>\n<pre><code>def null_augment(r):\n    image = r['image']\n    #image = cv2.resize(image, dsize=(image_size,image_size), interpolation=cv2.INTER_LINEAR)\n    assert image_size==224\n    r['image'] = image\n    return r\n</code></pre>\n<p>But this makes sense in the dataset.py</p>\n<p>Thanks in Advance</p>",
      "rawMarkdown": "Hi Everyone, @hengck23,\nanyone on this panel please clear me\nIn file dataset_224.py, \nIs he taking the same files of our dataset or any resized images\nif he uses original sized images, why does he comment the resize function\n```\ndef null_augment(r):\n    image = r['image']\n    #image = cv2.resize(image, dsize=(image_size,image_size), interpolation=cv2.INTER_LINEAR)\n    assert image_size==224\n    r['image'] = image\n    return r\n```\nBut this makes sense in the dataset.py\n\nThanks in Advance",
      "replies": [
        {
          "id": 1273493,
          "postDate": "2021-04-14T11:59:14.767Z",
          "content": "<p>I guess he used a resized dataset so uncomment the cv2.resize line if you use the original dataset.</p>",
          "rawMarkdown": "I guess he used a resized dataset so uncomment the cv2.resize line if you use the original dataset.",
          "votes": 3
        },
        {
          "id": 1273521,
          "postDate": "2021-04-14T12:41:44.367Z",
          "content": "<p>thank you so much</p>",
          "rawMarkdown": "thank you so much"
        }
      ]
    },
    {
      "id": 1273233,
      "postDate": "2021-04-14T07:19:41.450Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> did you provided the inference script for your model? i cannot find a inference script in your drive</p>",
      "rawMarkdown": "@hengck23 did you provided the inference script for your model? i cannot find a inference script in your drive",
      "replies": [
        {
          "id": 1273494,
          "postDate": "2021-04-14T12:00:35.077Z",
          "content": "<p>Check run_submit.py for inference script. Please don't spam the discussion, the code is very clear, just search by yourself :)</p>",
          "rawMarkdown": "Check run_submit.py for inference script. Please don't spam the discussion, the code is very clear, just search by yourself :)",
          "votes": 1
        },
        {
          "id": 1273522,
          "postDate": "2021-04-14T12:41:49.750Z",
          "content": "<p>Thank you      </p>",
          "rawMarkdown": "Thank you      "
        }
      ]
    },
    {
      "id": 1272646,
      "postDate": "2021-04-13T16:40:05.500Z",
      "content": "<p>Did you think about use GPT to get strong models </p>",
      "rawMarkdown": "Did you think about use GPT to get strong models "
    },
    {
      "id": 1271348,
      "postDate": "2021-04-12T14:06:50.143Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> : Can you tell us how the <code>test_orientation.csv</code> file is generated?</p>",
      "rawMarkdown": "@hengck23 : Can you tell us how the `test_orientation.csv` file is generated?",
      "replies": [
        {
          "id": 1271424,
          "postDate": "2021-04-12T15:40:44.943Z",
          "content": "<p>Quote:</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231190#1267247\" target=\"_blank\">i train a rotation predictor</a></p>\n</blockquote>",
          "rawMarkdown": "Quote:\n> [i train a rotation predictor](https://www.kaggle.com/c/bms-molecular-translation/discussion/231190#1267247)\n"
        }
      ]
    },
    {
      "id": 1270945,
      "postDate": "2021-04-12T07:16:33.157Z",
      "content": "<p>Great work, very helpful!!!!!!!👍</p>",
      "rawMarkdown": "Great work, very helpful!!!!!!!👍"
    },
    {
      "id": 1270547,
      "postDate": "2021-04-11T18:37:52.207Z",
      "content": "<p>list of vision transformer you can try:<br>\n<a href=\"https://github.com/facebookresearch/deit/blob/main/README_cait.md\" target=\"_blank\">https://github.com/facebookresearch/deit/blob/main/README_cait.md</a><br>\n<a href=\"https://github.com/facebookresearch/deit\" target=\"_blank\">https://github.com/facebookresearch/deit</a></p>",
      "rawMarkdown": "list of vision transformer you can try:\nhttps://github.com/facebookresearch/deit/blob/main/README_cait.md\nhttps://github.com/facebookresearch/deit\n"
    },
    {
      "id": 1270406,
      "postDate": "2021-04-11T16:11:02.120Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> : In the file ./2021-apr-07b/code/tnt-s-224-fairseq-v1-1/run_train.py, in the <code>run_train()</code> function from lines 262 to 283, we have:</p>\n<pre><code>if is_mixed_precision:\n    with amp.autocast():\n        #assert(False)\n        logit = net(image, token, length)\n        loss0 = seq_cross_entropy_loss(logit, token, length)\n        #loss0 = seq_anti_focal_cross_entropy_loss(logit, token, length)\n\n    scaler.scale(loss0).backward()\n    #scaler.unscale_(optimizer)\n    #torch.nn.utils.clip_grad_norm_(net.parameters(), 2)\n    scaler.step(optimizer)\n    scaler.update()\n\nelse:\n    assert False\n    # print('fp32')\n    # image_embed = encoder(image)\n    logit, weight = decoder(image_embed, token, length)\n\n    (loss0).backward()\n    optimizer.step()\n</code></pre>\n<p>It's difficult for me to understand the logic of those lines of codes.<br>\nCan you explain us why when <code>is_mixed_precision</code> is False, then we only call the decoder (that will fail due to <code>assert False</code>)?<br>\nCan you also explain us where the <code>encoder</code> and <code>decoder</code> objects are defined?</p>",
      "rawMarkdown": "@hengck23 : In the file ./2021-apr-07b/code/tnt-s-224-fairseq-v1-1/run_train.py, in the `run_train()` function from lines 262 to 283, we have:\n\n```\nif is_mixed_precision:\n\twith amp.autocast():\n\t\t#assert(False)\n\t\tlogit = net(image, token, length)\n\t\tloss0 = seq_cross_entropy_loss(logit, token, length)\n\t\t#loss0 = seq_anti_focal_cross_entropy_loss(logit, token, length)\n\n\tscaler.scale(loss0).backward()\n\t#scaler.unscale_(optimizer)\n\t#torch.nn.utils.clip_grad_norm_(net.parameters(), 2)\n\tscaler.step(optimizer)\n\tscaler.update()\n\nelse:\n\tassert False\n\t# print('fp32')\n\t# image_embed = encoder(image)\n\tlogit, weight = decoder(image_embed, token, length)\n\n\t(loss0).backward()\n\toptimizer.step()\n```\n\nIt's difficult for me to understand the logic of those lines of codes.\nCan you explain us why when `is_mixed_precision` is False, then we only call the decoder (that will fail due to `assert False`)?\nCan you also explain us where the `encoder` and `decoder` objects are defined?",
      "replies": [
        {
          "id": 1270416,
          "postDate": "2021-04-11T16:26:20.233Z",
          "content": "<p>this part of the code is not used (i always train with fp16), The codes (encoder, decoder, etc) are remain from previous project. you can just ignore</p>",
          "rawMarkdown": "this part of the code is not used (i always train with fp16), The codes (encoder, decoder, etc) are remain from previous project. you can just ignore",
          "votes": 1
        }
      ]
    },
    {
      "id": 1269446,
      "postDate": "2021-04-10T14:38:50.333Z",
      "content": "<p>Good pictures!</p>",
      "rawMarkdown": "Good pictures!"
    },
    {
      "id": 1268991,
      "postDate": "2021-04-10T03:05:37.743Z",
      "content": "<p>Thanks for sharing this great code. I see that your transformer didn't mask the padding token when computing self attention. Should I use the <strong>tgt key padding mask</strong> when using PyTorch's Transformer, because i think its more reasonable to mask padding. I'm not familiar with transformer, hope you can solve my confusion, thanks.</p>",
      "rawMarkdown": "Thanks for sharing this great code. I see that your transformer didn't mask the padding token when computing self attention. Should I use the **tgt key padding mask** when using PyTorch's Transformer, because i think its more reasonable to mask padding. I'm not familiar with transformer, hope you can solve my confusion, thanks.",
      "replies": [
        {
          "id": 1269007,
          "postDate": "2021-04-10T03:52:05.150Z",
          "content": "<p>yes you are right, tgt key padding is logically more correct.</p>\n<p>i may want to investigate this later. currently, i am ignoring this because:</p>\n<ul>\n<li>at inference, decode stop when it hit [eos] or [pad]</li>\n<li>all [pad] tokens are later than [eos] and are ignored in the loss function, so these [pad] are never used in training</li>\n<li>lazy to modify my triangle mask at this time …</li>\n</ul>",
          "rawMarkdown": "yes you are right, tgt key padding is logically more correct.\n\ni may want to investigate this later. currently, i am ignoring this because:\n- at inference, decode stop when it hit [eos] or [pad]\n- all [pad] tokens are later than [eos] and are ignored in the loss function, so these [pad] are never used in training\n- lazy to modify my triangle mask at this time ...",
          "votes": 1
        },
        {
          "id": 1269019,
          "postDate": "2021-04-10T04:03:55.890Z",
          "content": "<p>also, the training code can be speedup by considering only max(length) instead of max_length in a single batch training</p>",
          "rawMarkdown": "also, the training code can be speedup by considering only max(length) instead of max\\_length in a single batch training"
        }
      ]
    },
    {
      "id": 1268806,
      "postDate": "2021-04-09T19:28:10.643Z",
      "content": "<p>Thanks for the great repo! When I do inference, batch_size 32 would fail for a 8G GPU for OOM, when I change to smaller batch size like 4, it can run up to 140 or so for OOM. Do you know what might be wrong? Thanks!</p>",
      "rawMarkdown": "Thanks for the great repo! When I do inference, batch_size 32 would fail for a 8G GPU for OOM, when I change to smaller batch size like 4, it can run up to 140 or so for OOM. Do you know what might be wrong? Thanks!",
      "replies": [
        {
          "id": 1269055,
          "postDate": "2021-04-10T05:20:29.747Z",
          "content": "<p>i suggest you use fairseq or other api as they are more memory efficient</p>",
          "rawMarkdown": "i suggest you use fairseq or other api as they are more memory efficient",
          "votes": 1
        },
        {
          "id": 1269162,
          "postDate": "2021-04-10T08:26:41.773Z",
          "content": "<p>samples are sorted by length. as the validation proceeds, the longer sequence uses more memory</p>",
          "rawMarkdown": "samples are sorted by length. as the validation proceeds, the longer sequence uses more memory",
          "votes": 1
        },
        {
          "id": 1272758,
          "postDate": "2021-04-13T18:21:47.260Z",
          "content": "<p>May I ask for 320 and other image sizes, do you still use pretrained model or train from scratch?</p>",
          "rawMarkdown": "May I ask for 320 and other image sizes, do you still use pretrained model or train from scratch?"
        }
      ]
    },
    {
      "id": 1267170,
      "postDate": "2021-04-08T11:07:52.683Z",
      "content": "<p>Great work, it will help a lot! <br>\nAs far as I understood, you rotate images according to test_orientation.csv. Could you please tell me how do you determine the orientation?</p>",
      "rawMarkdown": "Great work, it will help a lot! \nAs far as I understood, you rotate images according to test_orientation.csv. Could you please tell me how do you determine the orientation?",
      "replies": [
        {
          "id": 1267247,
          "postDate": "2021-04-08T12:12:40.160Z",
          "content": "<p>i train a rotation predictor</p>",
          "rawMarkdown": "i train a rotation predictor",
          "votes": 1
        }
      ]
    },
    {
      "id": 1266156,
      "postDate": "2021-04-07T14:27:57.310Z",
      "content": "<p>Thanks for your work <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> ! Very helpful </p>",
      "rawMarkdown": "Thanks for your work @hengck23 ! Very helpful "
    },
    {
      "id": 1316121,
      "postDate": "2021-05-20T10:06:36.377Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1307475,
      "postDate": "2021-05-14T13:13:37.353Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1304509,
      "postDate": "2021-05-12T16:57:23.137Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1303939,
      "postDate": "2021-05-12T10:49:02.340Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1293613,
      "postDate": "2021-05-05T04:17:51.017Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1267624,
      "postDate": "2021-04-08T16:49:26.223Z",
      "content": "<p>This is gold… <br>\nThanks for sharing…</p>",
      "rawMarkdown": "This is gold... \nThanks for sharing..."
    }
  ],
  "comments": [
    {
      "id": 1280094,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-21T14:54:52.107000",
      "content": "<p><br>\nalready implemented (version 2021-apr-24)</p>\n<p>forget about image size!<br>\nwe use non-empty patch as input (i.e. variable input length). empty image patch are discarded</p>\n<p><img src=\"https://i.ibb.co/nQqfV0S/Selection-104.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/NjVv60V/Selection-106.png\" alt=\"\"></p>",
      "votes": 14,
      "replies": [
        {
          "id": 1280103,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-21T15:06:25.563000",
          "content": "<p>maybe next,next version?<br>\n<img src=\"https://i.ibb.co/g7d5nMh/Selection-107.png\" alt=\"\"></p>\n<p>anyone can recommend bi-translation paper? (e.g. a model that can translate FR to EN and also EN to FR)</p>\n<p><a href=\"https://arxiv.org/pdf/1805.11213.pdf\" target=\"_blank\">https://arxiv.org/pdf/1805.11213.pdf</a><br>\n\"Our technique trains a single model for both directions of a language pair, allowing us<br>\nto back-translate source or target monolingual data without requiring an auxiliary<br>\nmodel. \"</p>\n<hr>\n<p>back translation is a data augmentation method: <a href=\"https://amitness.com/back-translation/\" target=\"_blank\">https://amitness.com/back-translation/</a><br>\nback translation can also be used as self-supervised or semi-supervsied learning</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1280118,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-04-21T15:25:59.227000",
          "content": "<p>Can you then please post your CV/LB performance for this technique? Would be interested if it can keep up. And your iteration speed improvements, as that was your main goal as you have stated in a past comment.</p>\n<p>Interesting fact: The length distribution over the image patches is quite similar to the length distribution of my InChI tokenization. Only the long tail is quite longer and the mean is off by 5.<br>\nSo we have direct, decent correlation between effective molecule size and InChI length. Makes sense.<br>\nAs you have the data you could plot this correlation, if your are interested.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1280382,
          "author_name": "Kuts Alexey",
          "author_url": "",
          "post_date": "2021-04-21T21:18:39.973000",
          "content": "<p>very interesting idea <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, thanks!<br>\nI am interested in what order are you going to feed patches into encoder?<br>\nor this does not much matter?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1280415,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-21T23:32:22.590000",
          "content": "<p>it doesn't matter for transformer as it uses positional encoding.<br>\nin fact, it isn't feed in one by one. it is feed in all at once</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1283267,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-24T18:46:49.060000",
          "content": "<p>updated!<br>\n[2021-apr-24]<br>\nimplement patch-based input for transformer encoder. Please refer to the PPT for details.</p>\n<p>please refer to the comments in the main top window.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1283327,
          "author_name": "Thomas SELECK",
          "author_url": "",
          "post_date": "2021-04-24T20:19:10.793000",
          "content": "<p>Can you tell us what is the LB score for the [2021-apr-24] patch based model ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1283606,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-25T04:59:24.790000",
          "content": "<p>i update the results. LB = 2.01</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1272273,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-13T11:25:54.260000",
      "content": "<p>experiment update:<br>\n<img src=\"https://i.ibb.co/9qcKLQy/Selection-029.png\" alt=\"\"></p>\n<p>updated with new results (LB=2.06 for latest patch+coord input, version 2021-apr-24)</p>\n<p><img src=\"https://i.ibb.co/4dZmVVj/Selection-129.png\" alt=\"\"></p>",
      "votes": 10,
      "replies": [
        {
          "id": 1272484,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-04-13T14:11:02.707000",
          "content": "<p>Can I ask <br>\nchanging img_size out in the below code to 320 will make it train with image size from 224 to 320.</p>\n<pre><code>class TNT(nn.Module):\n    \"\"\" Transformer in Transformer - https://arxiv.org/abs/2103.00112\n    \"\"\"\n\n    def __init__(self, img_size=224, patch_size=16, in_chans=3, num_classes=1000, embed_dim=768, in_dim=48, depth=12,\n                 num_heads=12, in_num_head=4, mlp_ratio=4., qkv_bias=False, drop_rate=0., attn_drop_rate=0.,\n                 drop_path_rate=0., norm_layer=nn.LayerNorm, first_stride=4):\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1273217,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-14T07:06:59.550000",
          "content": "<p><img src=\"https://i.ibb.co/vzLgnpk/Selection-028.png\" alt=\"\"></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1274122,
          "author_name": "FP",
          "author_url": "",
          "post_date": "2021-04-15T02:37:00.163000",
          "content": "<p>That's why \"a picture is worth a thousand words\"……thank you Heng!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1278732,
          "author_name": "donutking",
          "author_url": "",
          "post_date": "2021-04-20T08:30:01.263000",
          "content": "<p>Great work. Could you please share how many epochs averagely need to be run to reach those scores?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1279999,
          "author_name": "IMRules",
          "author_url": "",
          "post_date": "2021-04-21T13:07:35.157000",
          "content": "<p>deleted, :D</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1306975,
          "author_name": "Trushant Kalyanpur",
          "author_url": "",
          "post_date": "2021-05-14T07:18:45.823000",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  was that the only line needed to get diff image sizes to work? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1307464,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-05-14T13:03:15.563000",
          "content": "<p>i think yes <a href=\"https://www.kaggle.com/trushk\" target=\"_blank\">@trushk</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1308673,
          "author_name": "tik_boa",
          "author_url": "",
          "post_date": "2021-05-15T11:25:42.990000",
          "content": "<p>what's the df_fold.fine.csv? thank Heng!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1268242,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-09T08:28:44.973000",
      "content": "<p>What do you want to know about transformer application in this competition?</p>\n<p>I see some posts that newbies want to know more about transformers.<br>\nif you can specifically post your question here, maybe i can organize a talk on it (e.g. zoom video) or include some of the problems and solution in my starter kit</p>\n<p>please share your thoughts here!</p>",
      "votes": 10,
      "replies": [
        {
          "id": 1285213,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2021-04-26T17:31:59.080000",
          "content": "<p>I would definitely be interested in attending a talk from you. For me though, I'm okay with transformers. I'm more interested in your overall knowledge of toolkits and optimisation. This is the first time I've seen torch jit and the first time I've seen fairseq. If you want a more concrete request, I'd be happy to watch a talk solely on jit. What is it in general? How does it apply to torch? How do you use it?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1285260,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2021-04-26T18:08:23.747000",
          "content": "<p>Another concrete thing I'd like to learn about is how you organise your code. With many newer Kagglers used to using a notebook, your code base can look quite daunting. I'm not really a \"new\" Kaggler and I'm bewildered by it. In my head it's because \"that's how the real pros do it\".</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1285523,
          "author_name": "Nitin Datta",
          "author_url": "",
          "post_date": "2021-04-27T03:38:55.357000",
          "content": "<p><a href=\"https://www.kaggle.com/alexandersoare\" target=\"_blank\">@alexandersoare</a>,<br>\nTotally agree with you. If more people are interested we can conduct a zoom call/ Google meet and learn a lot of techniques from him. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1284482,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-26T01:26:21.667000",
      "content": "<p>a few good open source to generate new train data by using random perturbation</p>\n<p>ICLR-2020 poster paper \"Augmenting Genetic Algorithms with Deep Neural Networks for exploring the chemical space\"<br>\n<a href=\"https://www.youtube.com/watch?v=9VilhlEXm9w&amp;t=16s\" target=\"_blank\">https://www.youtube.com/watch?v=9VilhlEXm9w&amp;t=16s</a></p>\n<p>EvoMol: a flexible and interpretable evolutionary algorithm for unbiased de novo molecular generation. J Cheminform 12, 55 (2020)<br>\n<a href=\"https://jcheminf.biomedcentral.com/articles/10.1186/s13321-020-00458-z\" target=\"_blank\">https://jcheminf.biomedcentral.com/articles/10.1186/s13321-020-00458-z</a><br>\n<a href=\"https://github.com/jules-leguy/EvoMol\" target=\"_blank\">https://github.com/jules-leguy/EvoMol</a></p>\n<p><img src=\"https://github.com/jules-leguy/EvoMol/raw/master/examples/figures/detailed_expl_tree.png\" alt=\"\"></p>\n<p>using GAN:<br>\n<a href=\"https://github.com/ardigen/mol-cycle-gan\" target=\"_blank\">https://github.com/ardigen/mol-cycle-gan</a></p>",
      "votes": 8,
      "replies": [
        {
          "id": 1291396,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-05-03T02:34:54.820000",
          "content": "<p>for the adventurous<br>\nOptimization of Molecules via Deep Reinforcement Learning<br>\n<a href=\"https://www.nature.com/articles/s41598-019-47148-x\" target=\"_blank\">https://www.nature.com/articles/s41598-019-47148-x</a></p>\n<p><img src=\"https://www.researchgate.net/publication/334663607/figure/fig2/AS:784360231948290@1564017460710/Sample-molecules-in-the-property-optimization-task-a-Optimization-of-penalized-logP.png\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1291397,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-05-03T02:36:31.560000",
          "content": "<p>the simple competition may end up me learning transformer, image caption, attention-lstm, … and even RL</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1295067,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2021-05-06T07:15:11.090000",
          "content": "<p>Wait, so is this allowed? Are we not meant to only stick to the provided InChIs?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1295275,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-05-06T10:52:35.910000",
          "content": "<p>i am using provided InChIs to create more InchIs</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1295365,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2021-05-06T11:57:27.910000",
          "content": "<p>Yes, that's what I'm concerned about. Do you know for sure that is allowed?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1295407,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2021-05-06T12:41:12.127000",
          "content": "<p>Here, check this out <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231318#1267279\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/231318#1267279</a></p>\n<p>BTW not trying to police you. I just want to do this too and checking to see if you've got a clear green light. If not, I suppose I can ask in that thread.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1270712,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-11T23:10:13.607000",
      "content": "<p>there is an obvious bug in all validation code during training:</p>\n<pre><code>correct version\ndef do_valid(net, tokenizer, valid_loader):\n\n    if 1:\n        score = []\n        for i, (p, t) in enumerate(zip(predict, truth)):\n            t = truth[i][1:length[i]-1]     # in the buggy version, i have used 1 instead of i\n            p = predict[i][1:length[i]-1]\n            t = tokenizer.one_predict_to_inchi(t)\n            p = tokenizer.one_predict_to_inchi(p)\n            s = Levenshtein.distance(p, t)\n            score.append(s)\n        lb_score = np.mean(score)\n</code></pre>",
      "votes": 7,
      "replies": [
        {
          "id": 1278624,
          "author_name": "atfujita",
          "author_url": "",
          "post_date": "2021-04-20T05:53:28.510000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>\n<p>This script is correct?<br>\nI have a score gap between this script and compute_lb_score().</p>\n<p>e.g.<br>\nLB: 1.4(this script) vs LB: 1.9(compute_lb_score)<br>\nBoth valid_df are same fold number and using all data.</p>\n<p>Now, I investigate this reason.<br>\nthanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1278676,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-20T06:55:35.233000",
          "content": "<p>i haven't check in details, but should be correct. just note that this validation is under teaching forcing (using forward function).</p>\n<p>the submit one is using autoregressive, forward with argmax function (i.e. decoder, without teacher forcing)</p>\n<hr>\n<p>also check the sampler and df file of the data loader. it could be different</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1278886,
          "author_name": "atfujita",
          "author_url": "",
          "post_date": "2021-04-20T11:23:35.093000",
          "content": "<p>Thanks!</p>\n<blockquote>\n  <p>the submit one is using autoregressive, forward with argmax function (i.e. decoder, without teacher forcing)</p>\n</blockquote>\n<p>I think that this is the cause of that.<br>\nAnd I understood using autoregressive takes long time.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1302551,
          "author_name": "ultra_zhl",
          "author_url": "",
          "post_date": "2021-05-11T15:36:53.587000",
          "content": "<p>I think it shoule be p = predict[i][0:length[i]-2]</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1269047,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-10T05:03:00.400000",
      "content": "<p>here is the latest [2021-apr-07b] version:</p>\n<p><img src=\"https://i.ibb.co/rb0GNHB/Selection-082.png\" alt=\"\"></p>\n<p>if you want to help in this development, you can:</p>\n<ul>\n<li>develop and share code on how to use fairseq batch k beam search</li>\n<li>experiment with more transformer decoder layers, or larger input image size, or use a larger image transformer encoder.  </li>\n<li>share results for training with augmentation</li>\n</ul>\n<p>Note that you do not need to train from scratch as you can also use your previous trained model (e.g. 224 image, 3 layers decoder) as initial checkpoint for your new experiment (e.g. 320 image, 6 layers decoder)</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1270460,
          "author_name": "Thomas SELECK",
          "author_url": "",
          "post_date": "2021-04-11T17:16:14.053000",
          "content": "<p>Why do you need 3-layers-decoder for 224x224 images and 6-layers-decoder for 320x320 images?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1270468,
          "author_name": "Matthieu Planté",
          "author_url": "",
          "post_date": "2021-04-11T17:22:51.130000",
          "content": "<p><a href=\"https://www.kaggle.com/thomasseleck\" target=\"_blank\">@thomasseleck</a> it's not a must it's just some configs to test. You can use 320x320 images with a 3 layers decoder ;) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1273124,
          "author_name": "Robert Kim",
          "author_url": "",
          "post_date": "2021-04-14T05:34:50.497000",
          "content": "<p>Are you also inputing the embedding of the CLS token into the decoder? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1273432,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-14T10:46:34.710000",
          "content": "<p>\"Are you also inputing the embedding of the CLS token into the decoder?\"</p>\n<p>yes, but i don't think it is important</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1276234,
          "author_name": "Charles",
          "author_url": "",
          "post_date": "2021-04-17T09:35:25.697000",
          "content": "<p>\"Note that you do not need to train from scratch as you can also use your previous trained model (e.g. 224 image, 3 layers decoder) as initial checkpoint for your new experiment (e.g. 320 image, 6 layers decoder)\"</p>\n<p>In this scenario, when you perform the state_dict load, how do you address the size mismatch between patch_pos? Because the number of patches changes, do you increase the patch size, or approach it in some other way? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1276290,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-17T10:58:35.043000",
          "content": "<p>delete this key in the state dict and let this part be learned from scratch.</p>\n<pre><code>    if initial_checkpoint is not None:\n        f = torch.load(initial_checkpoint, map_location=lambda storage, loc: storage)\n        start_iteration = f['iteration']\n        start_epoch     = f['epoch']\n        state_dict = f['state_dict']\n        #del state_dict['cnn.e.patch_pos']\n        #del state_dict['text_pos.pos']\n        #net.load_state_dict(state_dict, strict=False)  # True\n        net.load_state_dict(state_dict, strict=True)  # True\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1276726,
          "author_name": "Charles",
          "author_url": "",
          "post_date": "2021-04-17T22:16:46.910000",
          "content": "<p>Thank you, I will see what I can contribute in the next weeks. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1307451,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-14T12:55:35.700000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1266823,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-08T05:18:15.290000",
      "content": "<p>i have decided to use pytorch/fairseq for my next development</p>\n<ul>\n<li>it is actively developed and has lots of pretain model (including attention-LSTM image caption)</li>\n<li>It is part of pytorch ecosystem</li>\n<li>it can export to onnx (tensorRT) and TPU</li>\n<li>it has pre-norm transformer decoder and encoder (i will try to copy my trained weights to its modules)</li>\n<li>it supports caching of keys and values, incremental decoding</li>\n<li>it supports ensemble of seq model</li>\n</ul>\n<p><a href=\"https://github.com/pytorch/fairseq\" target=\"_blank\">https://github.com/pytorch/fairseq</a></p>\n<p>stay tuned on how to use fairseq for this challenge!</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1268248,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-09T08:32:21.760000",
          "content": "<p>just a quick update. I am preparing the next software version, which should be up in 24hrs.<br>\nsome very good news:</p>\n<ol>\n<li><p>my transformer from scratch is almost similar to fairseq version (after all, all are based on the \"what you need is attention\" paper). Hence i can convert my model to fairseq easily. there are little numercial differences, but can be get rid of in fine tunning some iterations</p></li>\n<li><p>fairseq key,value caching speedup inference. now i can run resnet101d at 13 to 14 min at 40_000 test images, reduction of 4 min from the [2021-apr-06a] version</p></li>\n<li><p>there are several transformer based image encoder, e.g. Vit. But I find TNT is magic (transformer in transformer). It can get local CV in the range of 2.0. Unlike resnet, efficient, etc, there is not image striding, hence transformer-based classifier can with with small image size like 224</p></li>\n<li><p>number of transformer decoder layers can improve results (i think) … but more experiments needs to confirm this </p></li>\n</ol>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1323657,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-05-26T11:31:39.413000",
      "content": "<ul>\n<li>2021-apr-25 software version  (i,e,  sparse patch method: keep only patches with black pixels, empty patches are discarded) + modifications</li>\n<li>add augmentation(e.g. noise, rotate, scale, resize artifacts, …)  to mimick test image </li>\n<li>remove scale and rotation predictor in pre-processing </li>\n<li>in testing i revert to use w&gt;h condition to unrotate the image</li>\n</ul>\n<p>the important point to note is that local CV and LB can be very smiliar. 40_000 validations images seem good enough</p>\n<p>training score is about 0.65 (without rdkit inchi validation check). This means that the top kaggler has reached the limit and modeled train=test very well ?</p>\n<p><img src=\"https://i.ibb.co/y4Gchs8/Selection-083.png\" alt=\"\"></p>",
      "votes": 5,
      "replies": [
        {
          "id": 1323717,
          "author_name": "Thomas SELECK",
          "author_url": "",
          "post_date": "2021-05-26T12:08:53.863000",
          "content": "<p>Thanks for your awesome work !</p>\n<p>Is it possible to get the 01495000_model.pth file to try to reproduce the results ?</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1324018,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-05-26T15:00:54.273000",
          "content": "<p>there is no further plan to release code or model. if you want to get the similiar results, your training log should be close to the below:</p>\n<pre><code>   batch_size = 128\n   experiment = ['vit-s0.8-p16-06b4r6', 'run_train.py']\n                      |----- VALID ---|---- TRAIN/BATCH --------------\nrate     iter   epoch | loss  lb(lev) | loss0  loss1  | time          \n----------------------------------------------------------------------\n0.00000  155.0000* 53.51  | 0.005   0.18  | 0.000  0.000  0.000  |  0 hr 00 min\n0.00003  155.1000  53.54  | 0.005   0.18  | 0.005  0.001  0.000  |  0 hr 19 min\n0.00003  155.1806  53.56  | 0.005   0.18  | 0.003  0.000  0.000  |  0 hr 35 min\n0.00003  155.2000  53.56  | 0.005   0.17  | 0.004  0.000  0.000  |  0 hr 39 min\n0.00003  155.3000  53.59  | 0.005   0.18  | 0.005  0.000  0.000  |  0 hr 58 min\n</code></pre>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1327398,
          "author_name": "starsnew",
          "author_url": "",
          "post_date": "2021-05-29T08:30:07.587000",
          "content": "<p>doesn't this patch-based method require image resize?</p>\n<p>I'm trying to figure out how to solve this without resizing the images</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1327405,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-05-29T08:35:13.030000",
          "content": "<p>i build one model with image resize (scale = 0.8) and another one without (i.e. scale=1.0 or original size).</p>\n<p>if you resize image, you have less patches and the model will train faster and uses less memory.</p>\n<p>i haven't validate and test the performance for model without image resize (training is still in progress)</p>\n<p>but the both model have smiliar training loss.<br>\ni don't expect huge improvement for a single model.</p>\n<p>maybe there will be some gain from ensembling, test time augmentation, post-processing etc</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1339725,
          "author_name": "starsnew",
          "author_url": "",
          "post_date": "2021-06-07T12:29:05.423000",
          "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> .This gave me a lot of intuition and ideas.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1324800,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2021-05-27T08:23:52.037000",
      "content": "<p>The day <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> uses github (or any other git hosting solution)… 😄</p>\n<p>Joke aside, thanks for sharing this starter kit!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1304058,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-05-12T11:54:52.813000",
      "content": "<p>some symmetry mages that may cause havoc to your algorithm</p>\n<p><img src=\"https://i.ibb.co/Jp9HsVt/Selection-062.png\" alt=\"\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 1304871,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-05-13T00:34:57.167000",
          "content": "<p>it is easy to detect such symmetric images : just rotate and do a comparison e.g. l2 loss, see if black pixels align, etc</p>\n<p>you can rotate such image as augmentation in training (to create different noise) </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1300309,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-05-10T11:39:29.253000",
      "content": "<p>transformer set decoding, aka object detection</p>\n<p><img src=\"https://i.ibb.co/dQVWGKX/Selection-031.png\" alt=\"\"></p>\n<p><img src=\"https://i.ibb.co/qmhkk1Y/Selection-037.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/xDZFJSm/Selection-038.png\" alt=\"\"></p>\n<p>i had this idea in the past, but stuck at the bond assignment. e.g. a bond has attribute X-Y where X,Y are the numbering of the atoms. I was wondering how to set the ground truth as the numbering could be unknown. Then I realize that it is not a problem if you use dynamic set assignment of ground truth as it is used in object detection</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1300819,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-05-10T18:34:05.297000",
          "content": "<p>instead of having a seq decoder transformer, i am thinking of changing the strategy. there is only one seq in our case (unlike the language model, where there can be many possible valid sequences)</p>\n<p>since there is one and only one ground truth seq (due to only one canonical atom numbering), we can directly predict the atom at the atomic number location</p>\n<p>the query object will have position encoding = atomic numbering. this ensures we would have most of the atom with atomic numbering correct, hence lowering the LD distance</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1300829,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-05-10T18:48:14.800000",
          "content": "<p>That sounds phenomenal!<br>\nI perhaps would go for something like this if there was enough time left to train.</p>\n<p>You do have great ideas, that's for sure!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1296482,
      "author_name": "Charles",
      "author_url": "",
      "post_date": "2021-05-07T09:35:41.723000",
      "content": "<p>Training curve with manually decreasing LR</p>\n<p><a href=\"https://ibb.co/GTHXWcv\"><img src=\"https://i.ibb.co/k1mp4Qg/train-plot.png\" alt=\"train-plot\"></a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 1296485,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-05-07T09:39:32.553000",
          "content": "<p>Thanks for sharing! Is this after every epoch?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1296496,
          "author_name": "Charles",
          "author_url": "",
          "post_date": "2021-05-07T09:47:23.413000",
          "content": "<p>Nope, the total graph is about one epoch (warm restart from another pretrained model with different size). </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1296510,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-05-07T10:01:34.087000",
          "content": "<p>Oh I see now, I'm dumb ^^'</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1296530,
          "author_name": "Charles",
          "author_url": "",
          "post_date": "2021-05-07T10:12:42.917000",
          "content": "<p>There's no such thing as a dumb question! </p>\n<p>Also, I wish I had the resources to train for so many epochs! 😂</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1296556,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-05-07T10:36:23.540000",
          "content": "<p>Now I wish I was a question. :(</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1296588,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-05-07T11:05:51.673000",
          "content": "<p>you always restart training, e.g cyclic training rate, to see if you can get better result.</p>\n<p>you can compare with my train log if you are using the same fold and net parameters</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1296984,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-05-07T16:17:05.283000",
          "content": "<blockquote>\n  <p>Now I wish I was a question. :(</p>\n</blockquote>\n<p>Don't worry, questions can't think - this way they also can't be dumb. So you have a vast advantage there ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1297705,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-05-08T08:31:01.107000",
          "content": "<p>never mind :D</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1298816,
          "author_name": "Phaedrus",
          "author_url": "",
          "post_date": "2021-05-09T08:29:10.523000",
          "content": "<p>When I do manual LR restarts quite a few starting iterations have a higher cv than the checkpoint model. Any ideas why this might be happening or what I can do to fix? - Possible reason is that I deleted the trained pos embeddings even though I was using same image size. I should not do that. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1300324,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-05-10T11:52:45.840000",
          "content": "<p>if you restart training, you should load your previous trained model, including your trained pos encoding. you should modify the load previous state dict code</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1300420,
          "author_name": "Phaedrus",
          "author_url": "",
          "post_date": "2021-05-10T12:52:53.720000",
          "content": "<p>Yep, seeing the desired behavior now. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1307044,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2021-05-14T08:06:45.750000",
          "content": "<p>This is a grossly underrated string of internet bits among the 45 zettabytes in existence</p>\n<blockquote>\n  <p>Now I wish I was a question. :(</p>\n</blockquote>\n<p>cc <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1307186,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-05-14T09:58:43.877000",
          "content": "<p><a href=\"https://www.kaggle.com/alexandersoare\" target=\"_blank\">@alexandersoare</a> thanks for the appreciation! I'm not gonna lie, I had a pretty good laugh when I wrote that. :D</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1288190,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-29T18:33:16.290000",
      "content": "<p><img src=\"https://www.memecreator.org/static/images/memes/4307297.jpg\" alt=\"\"><br>\nAttention is all you need</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1273297,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-14T08:21:51.083000",
      "content": "<p>list of tricks to try:</p>\n<ol>\n<li><p><a href=\"https://arxiv.org/pdf/2006.12000.pdf\" target=\"_blank\">https://arxiv.org/pdf/2006.12000.pdf</a><br>\nSelf-Knowledge Distillation: A Simple Way for Better Generalization</p></li>\n<li><p>Sharpness-Aware Minimization for Efficiently Improving Generalization</p></li>\n<li><p><a href=\"https://cs.nju.edu.cn/wujx/paper/AAAI2021_Tricks.pdf\" target=\"_blank\">https://cs.nju.edu.cn/wujx/paper/AAAI2021_Tricks.pdf</a><br>\nBag of Tricks for Long-Tailed Visual Recognition with Deep Convolutional Neural Networks</p></li>\n</ol>\n<hr>\n<p><a href=\"https://neptune.ai/blog/text-classification-tips-and-tricks-kaggle-competitions\" target=\"_blank\">https://neptune.ai/blog/text-classification-tips-and-tricks-kaggle-competitions</a></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1327626,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2021-05-29T13:23:36.343000",
      "content": "<p>One minor typo here: pre-norm activated <strong>trnasformer</strong> are used. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1314190,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-05-19T02:51:12.303000",
      "content": "<p>'advertisement' …<br>\n<a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240233\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/240233</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1301657,
      "author_name": "Phaedrus",
      "author_url": "",
      "post_date": "2021-05-11T07:20:15.203000",
      "content": "<p>I read somewhere that we can do data aug by reversing the Inchi strings. Has anyone tried this as a data aug or is it a good idea even? Looking for thoughts. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1296091,
      "author_name": "Polapob",
      "author_url": "",
      "post_date": "2021-05-07T01:12:45.840000",
      "content": "<p>I want to ask such a question. How to use 224 image dimension pretrained weight to train higher dimension image.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1297290,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-05-07T22:08:53.193000",
          "content": "<p>It has been found that CNN can actually be <a href=\"https://arxiv.org/pdf/1906.06423.pdf\" target=\"_blank\">finetuned for higher resolutions</a>. So it shouldn't be a problem to just load the pretrained weights and use the models as is <br>\n(The images look totally different anyway, all that pretraining might be useful for are general inductive biases about images. I have never tried a pretrained model so far, but I doubt that it makes more than a really tiny difference after 10+ epochs).</p>\n<p>Although according to the competition rules the pretrained weights should allow for commercial usage. Something few people here seem to take into consideration.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1297386,
          "author_name": "Polapob",
          "author_url": "",
          "post_date": "2021-05-08T01:17:44.160000",
          "content": "<p>Thank you for your answer.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1280273,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2021-04-21T19:15:29.767000",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> How do you train the rotation detector please?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1280416,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-21T23:33:02.343000",
          "content": "<p>rotate the train images. the rotation used is your ground truth label.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1272986,
      "author_name": "Martin Beyer Deicas",
      "author_url": "",
      "post_date": "2021-04-14T01:32:54.670000",
      "content": "<p>Hi! Just dowloaded the last version [2021-apr-07b], on <code>run_train.py</code> I seem to be missing some imports:</p>\n<pre><code>from common import *\nfrom bms import *\n\nfrom lib.net.lookahead import *\nfrom lib.net.radam import *\n</code></pre>\n<p>Where could I find those? Thanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1272999,
          "author_name": "Nitin Datta",
          "author_url": "",
          "post_date": "2021-04-14T02:02:50.707000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/martinbeyerdeicas\" target=\"_blank\">@martinbeyerdeicas</a> <br>\nPlease check  [2021-apr-06] to find those files</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1270199,
      "author_name": "Thomas SELECK",
      "author_url": "",
      "post_date": "2021-04-11T12:16:37.390000",
      "content": "<p>Wow! This is pure gold!</p>\n<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>: I have a few questions for you: How many GPUs did you use to train the \"2021-apr-07b\" version? How much time did it take ?</p>\n<p>Have you done some preprocessing on images before training like noise removal or something else ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1270205,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-11T12:25:41.393000",
          "content": "<p>\"Have you done some preprocessing on images before training like noise removal or something else ?\"<br>\nnot at this moment. But in the future, some preprocessing (e.g. size normalization) may improve results</p>\n<p>\"How many GPUs\"<br>\nI have a good GPU card. I will talk about that later.</p>\n<p>For a normal user, if you start with 224-TNT-S vision transformer, you just need to train with batch=64 for about 10 epoch. see log files for details. decrease your learning rate from 0.001 to 0.00005</p>\n<p>\"This is pure gold!\"<br>\nwe are still in the early part of the competition. At the end of the challenge, gold should be below 1.0,<br>\nsilver probably below 2 ~2.3. I haven't done anything special except train with long iterations with bigger models and better GPU. These results can be easily replicated and hence it is easy for others to catch up over time, using their different model </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1270197,
      "author_name": "Nitin Datta",
      "author_url": "",
      "post_date": "2021-04-11T12:07:06.803000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , <br>\nI am encountering the following error when I try to run your code on Windows. I tried searching solutions online but could not find one…<br>\n<code>\nNameError: name 'Tensor' is not defined\n</code><br>\nDo you have any idea how to solve this? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1270198,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-11T12:12:53.617000",
          "content": "<p><a href=\"https://pytorch.org/docs/stable/jit.html\" target=\"_blank\">https://pytorch.org/docs/stable/jit.html</a></p>\n<p>try torch.Tensor</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1266756,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-08T04:06:10.427000",
      "content": "<p>you should implement this:</p>\n<p><a href=\"https://wandb.ai/pommedeterresautee/speed_training/reports/Train-HuggingFace-Models-Twice-As-Fast--VmlldzoxMDgzOTI\" target=\"_blank\">https://wandb.ai/pommedeterresautee/speed_training/reports/Train-HuggingFace-Models-Twice-As-Fast--VmlldzoxMDgzOTI</a><br>\n<a href=\"https://towardsdatascience.com/divide-hugging-face-transformers-training-time-by-2-or-more-21bf7129db9q-21bf7129db9e\" target=\"_blank\">https://towardsdatascience.com/divide-hugging-face-transformers-training-time-by-2-or-more-21bf7129db9q-21bf7129db9e</a></p>\n<p>Divide Hugging Face Transformers training time by 2 or more with dynamic padding and uniform length batching</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1267248,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-08T12:14:53.733000",
          "content": "<p>here is sample code for key,value caching:<br>\n<a href=\"https://github.com/tunz/transformer-pytorch/blob/master/model/fast_transformer.py\" target=\"_blank\">https://github.com/tunz/transformer-pytorch/blob/master/model/fast_transformer.py</a> <br>\n<a href=\"https://tunz.kr/post/4\" target=\"_blank\">https://tunz.kr/post/4</a></p>\n<p>(i am not implementing this since i am migrating to fairseq)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1314510,
      "author_name": "Drzhuzhe",
      "author_url": "",
      "post_date": "2021-05-19T07:49:54.657000",
      "content": "<p>hi  , I create two notebook base on your version apri 24</p>\n<ol>\n<li>preprocess:  <a href=\"https://www.kaggle.com/drzhuzhe/bms-preprocess-data-parallel\" target=\"_blank\">https://www.kaggle.com/drzhuzhe/bms-preprocess-data-parallel</a></li>\n<li>training: <a href=\"https://www.kaggle.com/drzhuzhe/training-on-gpu-bms/\" target=\"_blank\">https://www.kaggle.com/drzhuzhe/training-on-gpu-bms/</a></li>\n</ol>\n<p>During training I notice CPU memory keeping increase as iterate, result in out of memory crash </p>\n<p>I checkout all network modules <br>\ndelete all cv image and leak variable <br>\nbut memory still keeping increase, </p>\n<p>hardly locate what make this problem happen</p>\n<p>may someone have some tips?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1327845,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2021-05-29T17:03:46.413000",
          "content": "<p>your datasets are private.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1328248,
          "author_name": "Drzhuzhe",
          "author_url": "",
          "post_date": "2021-05-30T04:34:00.507000",
          "content": "<p>I made it public, <br>\nwatch out the BMS-train-full dataset has too much files, will be very slow for kaggle notebook to load it <br>\nIf you download it as a zip file , it will be much faster </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1329265,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2021-05-31T02:50:57.727000",
          "content": "<p>yes, It crashed.  probably torch issue.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1329315,
          "author_name": "Drzhuzhe",
          "author_url": "",
          "post_date": "2021-05-31T03:49:37.257000",
          "content": "<p>no no no ,</p>\n<ol>\n<li><p>to the out of memory problem</p>\n<p>this code is function well before Out of memory after 20k epoch<br>\nthat may due to image dataset load by both cpu and gpu<br>\nthis is definitely not torch issue, maybe hengk is trainning with larger RAM  </p></li>\n<li><p>to very slow to load input data problem </p>\n<p>it will take half a hour to load full dataset<br>\nif you get input files as a zip , it will be much faster<br>\nI zip all file and upload to kaggle dataset,<br>\nit is kaggle dataset automatic unzip my uploaded file and return to over 4000k small files</p></li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1329408,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2021-05-31T06:07:41.153000",
          "content": "<p>I loaded the pretrained weights, It crashed within 2000 epochs for resuming training.  Of course big enough RAM would not have this problem.  I have experiences that Pytorch consumes more ram.  Usually same algorithm if implemented in Tensorflow,  OOM happens less.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1329421,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-05-31T06:13:41.323000",
          "content": "<p>memory for each batch maybe different.<br>\ni think i allocate for the longest sequence length in the batch.</p>\n<p>you can sort your df_train by decreasing length and use sequential sampler to test the max batch size you system can time.</p>\n<p>once you determine the max batch size (and/or seq length), you can revert to random sampler</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1329852,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2021-05-31T12:26:13.110000",
          "content": "<p>I reduce batchsize, It runs so slowly, No crash however Kaggle 9 hours timeout. I don't remember the previous crashed  iteration number. Mostly Hardware RAM  not big enough.  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1290251,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-05-01T18:39:43.463000",
      "content": "<p>to my horror, many of the train images are rotated! (possibly 1 %)</p>\n<p>just to list a few example<br>\n1766b78e6eca<br>\n1b9063aabec5<br>\n1f60964241fd<br>\n1a56532cfc19<br>\n1fbedd22a0b4<br>\n1af43b2be6d3<br>\n1c9477a191a7<br>\n…</p>\n<p>so both test and train have rotated images. Just that the test has much more</p>\n<p>after preparing the images from some software like rdkit?, the organizer simply rotates the image if w&lt;h.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1283914,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-25T11:53:35.933000",
      "content": "<p>i have been using Lookahead+Radam optimizer. I only save trained model parameters and not the optimizer.</p>\n<p>i notice that when i restart training, the optimizer would need to restart of zero (because i didn't save the previous state)<br>\nAnd there is always a drop and improvement in validation loss for the restart.</p>\n<p>i wonder what is the reason for this?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1283843,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-25T10:20:26.333000",
      "content": "<p>for timm vision transformer models, there is \"def resize_pos_embed(posemb, posemb_new, num_tokens=1)\" function that can resize your positional embedding when you finetune your transformer from small size to a bigger image size</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1277621,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-19T02:11:46.217000",
      "content": "<p>there is JAX implementation for TNT<br>\n<a href=\"https://github.com/NZ99/transformer_in_transformer_flax\" target=\"_blank\">https://github.com/NZ99/transformer_in_transformer_flax</a></p>\n<p>anyone want to take this code to TPU?<br>\n(in my latest experience image size 448 is better than 384 …. i think ultimately, maybe we need to go to 640 for a constant input image size. if not, we need to break the input image into patch tokens ….)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1275234,
      "author_name": "Charles",
      "author_url": "",
      "post_date": "2021-04-16T05:40:59.503000",
      "content": "<p>This is great work, very educational. Thanks for your contribution <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>. </p>\n<p>PS: inference on the test set with 224x224 images takes me approximately 5 hours with a single 1080 Ti. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1272950,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-13T23:04:15.513000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1266193,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-07T14:49:39.163000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1266210,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-07T15:01:52.493000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1266757,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-08T04:07:05.047000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1266762,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-08T04:13:22.227000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2765514,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-21T07:05:38.717000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1329882,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-31T12:49:47.550000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1330069,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-31T15:12:25.087000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1316269,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-20T12:27:01.833000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1299749,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-10T02:26:24.857000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1299748,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-10T02:26:04.310000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1299750,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-10T02:30:26.637000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1299774,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-10T03:13:51",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1305325,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-13T08:00:48.397000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1305332,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-13T08:03:42.237000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1305341,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-13T08:08:34.507000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1305359,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-13T08:17:24.603000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1305687,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-13T12:30:57.750000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1297571,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-08T06:16:18.433000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1298531,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-09T01:35:53.947000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1298597,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-09T03:52:05.193000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1295819,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-06T18:00:16.397000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1295825,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-06T18:11:22.953000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1295845,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-06T18:33:44.397000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1295866,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-06T18:58:55.713000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1295910,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-06T19:55:35.320000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1295924,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-06T20:06:26.093000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1295943,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-06T20:23:30.607000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1296162,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-07T03:19:50.467000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1296319,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-07T07:24:15.233000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1296349,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-07T07:58:35.990000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1296362,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-07T08:06:45.433000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1296376,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-07T08:25:39.340000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1296455,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-07T09:10:38.533000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1296464,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-07T09:21:15.400000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1296594,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-07T11:08:37.503000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1297747,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-08T09:23:42.893000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1295470,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-06T13:37:09.597000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1295632,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-06T15:42:06.230000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1295647,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-06T15:47:46.803000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1295720,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-06T16:50:53.237000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1295732,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-06T17:01:29.083000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1296236,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-07T05:17:40.160000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1296242,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-07T05:26:23.557000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1296357,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-07T08:03:46.353000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1296433,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-07T09:00:09.853000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1296440,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-07T09:05:14.913000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1296446,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-07T09:07:28.043000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1296453,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-07T09:10:24.637000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1296475,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-07T09:29:34.710000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1307878,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-14T18:11:05.920000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1309061,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-15T16:25:25.753000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1293694,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-05T05:21:45.680000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1293723,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-05T05:55:04.307000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1293748,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-05T06:14:38.960000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1293822,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-05T07:27:54.190000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1294019,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-05T11:08:20.013000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1294058,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-05T11:40:13.343000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1294067,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-05T11:45:04.170000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1294071,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-05T11:49:25.467000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1294491,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-05T17:30:57.810000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1294674,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-05T20:08:48.420000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1295764,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-06T17:26:40.677000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1296342,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-07T07:50:02.113000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1292300,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-03T20:13:46.383000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1292312,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-03T20:21:55.103000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1292349,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-03T21:08:05.470000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1294053,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-05T11:38:06.980000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1291481,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-03T04:55:32.687000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1291869,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-03T12:01:41.017000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1292191,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-03T18:14:13.467000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1292206,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-03T18:35:39.770000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1292345,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-03T21:03:10.900000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1285508,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-27T03:17:12.150000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1285632,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-27T06:29:28.337000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1285635,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-27T06:37:26.110000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1289595,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-01T07:07:00.750000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1278561,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-20T03:22:40.577000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1279172,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-20T16:57:47.440000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1323853,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-26T13:19:08.877000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1324228,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-26T18:06:29.767000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1276789,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-18T01:29:53.077000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1276799,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-18T02:07:11.990000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1277155,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-18T13:12:55.667000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1277156,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-18T13:12:55.667000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1273385,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-14T09:42:57.910000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1273493,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-14T11:59:14.767000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1273521,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-14T12:41:44.367000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1273233,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-14T07:19:41.450000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1273494,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-14T12:00:35.077000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1273522,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-14T12:41:49.750000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1272646,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-13T16:40:05.500000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1271348,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-12T14:06:50.143000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1271424,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-12T15:40:44.943000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1270945,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-12T07:16:33.157000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1270547,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-11T18:37:52.207000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1270406,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-11T16:11:02.120000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1270416,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-11T16:26:20.233000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1269446,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-10T14:38:50.333000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1268991,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-10T03:05:37.743000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1269007,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-10T03:52:05.150000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1269019,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-10T04:03:55.890000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1268806,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-09T19:28:10.643000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1269055,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-10T05:20:29.747000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1269162,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-10T08:26:41.773000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1272758,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-13T18:21:47.260000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1267170,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-08T11:07:52.683000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1267247,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-08T12:12:40.160000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1266156,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-07T14:27:57.310000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1316121,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-20T10:06:36.377000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1307475,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-14T13:13:37.353000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1304509,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-12T16:57:23.137000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1303939,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-12T10:49:02.340000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1293613,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-05T04:17:51.017000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1267624,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-08T16:49:26.223000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1265992": "~~this is my terribly slow implementation, i couldn't make a submission yet. Anyway, it shows how to use transformer for a beginner ...~~\n\n\n~~i manage to speed up.  resnet101d+transformer  achieves LB of 3.92, taking 3h to predict 1.6 million test images with 4x Ti1080. The local CV is 3.19 (i train more iterations than the intermediate models in google drive)~~\n\nspeed now is  \"using fairseq+jit\", you can do inference at 2 min for 10_000 images with single GPU.\n\n- 224x224 images\n- pre-norm activated trnasformer are used. transformer layer are coded from scrtach.\n- resnet101d+transformer  : performance: 3.3 CV with 40_000 images , in 78 min \n- resnet26d+transformer  :  performance: ~3.7 CV  estimated\n- [2] resnet26d+attention+LSTM and [1]resnet26d+LSTM are also included for your study and experiment\n- train/validation log files, loss curves and intermediate trained models are included\n- i first trained the image encoder using attention+lstm. the pretained model is then used in the transformer to save time. i not sure if this affect performance\n- no augmentation in training.\n- rotation prediction in inference\n\n[1] Show and Tell: A Neural Image Caption Generator\nhttps://arxiv.org/abs/1411.4555\n\n[2] Show, Attend and Tell: Neural Image Caption Generation with Visual Attention\nhttps://arxiv.org/abs/1502.03044\n\n---\nall files are at google drive: https://drive.google.com/drive/folders/1dTfmZxDkDkrnRzz5DOkPOYr9oDIBcrq5?usp=sharing\n\n\n[2021-apr-06]\n ![](https://i.ibb.co/wzjSzWf/Selection-046.png)\n\n- please refer to readme.ppt to run the software and experiments\n\n\n\n[2021-apr-06a]\n- dirty code for torch.jit.script model\n- if i sort the validation samples by length, i can complete 40_000 test samples in 18 min\n\n\n[2021-apr-07a]\n- code to migrate from my transformer to fairseq transformer.\n- using fairseq+jit, you can do interence at 2 min for 10_000 images.\n- fairseq uses cache for key, values.\n\n\n\n[2021-apr-07b]\n- code to to train TNT(transformer in transformer) image encoder+ token transformer decoder\n- 80 min to do inference for 1.6 millions image with 4xTi1080. CV =1.9, LB =3.1\n- use fairseq API\n- you can modify 224 input to 320 to get CV=1.4, LB2.0 !!!\n\n\n\n[2021-apr-24]\n- implement patch-based input for transformer encoder. Please referto th PPT for deails.\n- you learn how to prepare and store image as patches\n- how to modify image transformer to accept variable-length input of patches. How to modify position encoding for patch input.\n- how to set the mask values for transformer encoder and decoder\n- a small trained model (using 0.8 input scale) and training log is provided. the results of patched based image transformer is the similar to image one, but running much faster (CV-teacher forcing: 1.27, CV-without teacher forcing : about 1.35 for 40000 validation set). time: 5.5 min for 40000 validation set, 55 min for 400000 test images (25% of all). LB = 2.01\n- there is no submission code or jit inference. This is left as an exercise for you. It is easy to modify the code\n\n[bug] as i was running the submission code, i note that the some test images have larger num of patch than the train, this affects the positional encoding for the input patch.\n\n\n[2021-apr-25]\n- model py file for the original VIT vision transformer. you can use this to replace the TNT. There is no pretrain model for deep TNT. But there are larger and deeper pretrain model for VIT\n\n---\n\n\nNote:  in some way, it is similar to this keras code:\nhttps://www.kaggle.com/aditya08/imagecaptioning-show-attnd-tell-w-transformer",
    "1280094": "~~next version is coming soon!~~\nalready implemented (version 2021-apr-24)\n\n\nforget about image size!\nwe use non-empty patch as input (i.e. variable input length). empty image patch are discarded\n\n![](https://i.ibb.co/nQqfV0S/Selection-104.png)\n![](https://i.ibb.co/NjVv60V/Selection-106.png)",
    "1272273": "experiment update:\n![](https://i.ibb.co/9qcKLQy/Selection-029.png)\n\nupdated with new results (LB=2.06 for latest patch+coord input, version 2021-apr-24)\n\n![](https://i.ibb.co/4dZmVVj/Selection-129.png)",
    "1268242": "What do you want to know about transformer application in this competition?\n\nI see some posts that newbies want to know more about transformers.\nif you can specifically post your question here, maybe i can organize a talk on it (e.g. zoom video) or include some of the problems and solution in my starter kit\n\nplease share your thoughts here!",
    "1284482": "a few good open source to generate new train data by using random perturbation\n\n ICLR-2020 poster paper \"Augmenting Genetic Algorithms with Deep Neural Networks for exploring the chemical space\"\nhttps://www.youtube.com/watch?v=9VilhlEXm9w&t=16s\n\nEvoMol: a flexible and interpretable evolutionary algorithm for unbiased de novo molecular generation. J Cheminform 12, 55 (2020)\nhttps://jcheminf.biomedcentral.com/articles/10.1186/s13321-020-00458-z\nhttps://github.com/jules-leguy/EvoMol\n\n\n\n![](https://github.com/jules-leguy/EvoMol/raw/master/examples/figures/detailed_expl_tree.png)\n\nusing GAN:\nhttps://github.com/ardigen/mol-cycle-gan",
    "1270712": "there is an obvious bug in all validation code during training:\n\n```\ncorrect version\ndef do_valid(net, tokenizer, valid_loader):\n\n    if 1:\n        score = []\n        for i, (p, t) in enumerate(zip(predict, truth)):\n            t = truth[i][1:length[i]-1]     # in the buggy version, i have used 1 instead of i\n            p = predict[i][1:length[i]-1]\n            t = tokenizer.one_predict_to_inchi(t)\n            p = tokenizer.one_predict_to_inchi(p)\n            s = Levenshtein.distance(p, t)\n            score.append(s)\n        lb_score = np.mean(score)\n\n\n\n```",
    "1269047": "here is the latest [2021-apr-07b] version:\n\n![](https://i.ibb.co/rb0GNHB/Selection-082.png)\n\nif you want to help in this development, you can:\n- develop and share code on how to use fairseq batch k beam search\n- experiment with more transformer decoder layers, or larger input image size, or use a larger image transformer encoder.  \n- share results for training with augmentation\n\nNote that you do not need to train from scratch as you can also use your previous trained model (e.g. 224 image, 3 layers decoder) as initial checkpoint for your new experiment (e.g. 320 image, 6 layers decoder)\n\n\n",
    "1266823": "i have decided to use pytorch/fairseq for my next development\n\n- it is actively developed and has lots of pretain model (including attention-LSTM image caption)\n- It is part of pytorch ecosystem\n- it can export to onnx (tensorRT) and TPU\n- it has pre-norm transformer decoder and encoder (i will try to copy my trained weights to its modules)\n- it supports caching of keys and values, incremental decoding\n- it supports ensemble of seq model\n\nhttps://github.com/pytorch/fairseq\n\nstay tuned on how to use fairseq for this challenge!",
    "1323657": "- 2021-apr-25 software version  (i,e,  sparse patch method: keep only patches with black pixels, empty patches are discarded) + modifications\n- add augmentation(e.g. noise, rotate, scale, resize artifacts, ...)  to mimick test image \n- remove scale and rotation predictor in pre-processing \n- in testing i revert to use w>h condition to unrotate the image\n\nthe important point to note is that local CV and LB can be very smiliar. 40\\_000 validations images seem good enough\n\ntraining score is about 0.65 (without rdkit inchi validation check). This means that the top kaggler has reached the limit and modeled train=test very well ?\n\n![](https://i.ibb.co/y4Gchs8/Selection-083.png)",
    "1324800": "The day @hengck23 uses github (or any other git hosting solution)... 😄\n\nJoke aside, thanks for sharing this starter kit!",
    "1304058": "some symmetry mages that may cause havoc to your algorithm\n\n![](https://i.ibb.co/Jp9HsVt/Selection-062.png)",
    "1300309": "transformer set decoding, aka object detection\n\n![](https://i.ibb.co/dQVWGKX/Selection-031.png)\n\n![](https://i.ibb.co/qmhkk1Y/Selection-037.png)\n![](https://i.ibb.co/xDZFJSm/Selection-038.png)\n\n\ni had this idea in the past, but stuck at the bond assignment. e.g. a bond has attribute X-Y where X,Y are the numbering of the atoms. I was wondering how to set the ground truth as the numbering could be unknown. Then I realize that it is not a problem if you use dynamic set assignment of ground truth as it is used in object detection",
    "1296482": "Training curve with manually decreasing LR\n\n<a href=\"https://ibb.co/GTHXWcv\"><img src=\"https://i.ibb.co/k1mp4Qg/train-plot.png\" alt=\"train-plot\" border=\"0\"></a>",
    "1288190": "![](https://www.memecreator.org/static/images/memes/4307297.jpg)\nAttention is all you need",
    "1273297": "list of tricks to try:\n1. https://arxiv.org/pdf/2006.12000.pdf\nSelf-Knowledge Distillation: A Simple Way for Better Generalization\n\n2. Sharpness-Aware Minimization for Efficiently Improving Generalization\n\n3. https://cs.nju.edu.cn/wujx/paper/AAAI2021_Tricks.pdf\nBag of Tricks for Long-Tailed Visual Recognition with Deep Convolutional Neural Networks\n\n---\nhttps://neptune.ai/blog/text-classification-tips-and-tricks-kaggle-competitions",
    "1327626": "One minor typo here: pre-norm activated **trnasformer** are used. ",
    "1314190": "'advertisement' ...\nhttps://www.kaggle.com/c/siim-covid19-detection/discussion/240233",
    "1301657": "I read somewhere that we can do data aug by reversing the Inchi strings. Has anyone tried this as a data aug or is it a good idea even? Looking for thoughts. ",
    "1296091": "I want to ask such a question. How to use 224 image dimension pretrained weight to train higher dimension image.",
    "1280273": "@hengck23 How do you train the rotation detector please?",
    "1272986": "Hi! Just dowloaded the last version [2021-apr-07b], on `run_train.py` I seem to be missing some imports:\n```\nfrom common import *\nfrom bms import *\n\nfrom lib.net.lookahead import *\nfrom lib.net.radam import *\n```\n\nWhere could I find those? Thanks!",
    "1270199": "Wow! This is pure gold!\n\n@hengck23: I have a few questions for you: How many GPUs did you use to train the \"2021-apr-07b\" version? How much time did it take ?\n\nHave you done some preprocessing on images before training like noise removal or something else ?",
    "1270197": "Hi @hengck23 , \nI am encountering the following error when I try to run your code on Windows. I tried searching solutions online but could not find one...\n`\nNameError: name 'Tensor' is not defined\n`\nDo you have any idea how to solve this? ",
    "1266756": "you should implement this:\n\nhttps://wandb.ai/pommedeterresautee/speed_training/reports/Train-HuggingFace-Models-Twice-As-Fast--VmlldzoxMDgzOTI\nhttps://towardsdatascience.com/divide-hugging-face-transformers-training-time-by-2-or-more-21bf7129db9q-21bf7129db9e\n\nDivide Hugging Face Transformers training time by 2 or more with dynamic padding and uniform length batching",
    "1314510": "hi  , I create two notebook base on your version apri 24\n\n1. preprocess:  https://www.kaggle.com/drzhuzhe/bms-preprocess-data-parallel\n2. training: https://www.kaggle.com/drzhuzhe/training-on-gpu-bms/\n\nDuring training I notice CPU memory keeping increase as iterate, result in out of memory crash \n\nI checkout all network modules \ndelete all cv image and leak variable \nbut memory still keeping increase, \n\nhardly locate what make this problem happen\n\nmay someone have some tips?",
    "1290251": "to my horror, many of the train images are rotated! (possibly 1 %)\n\njust to list a few example\n1766b78e6eca\n1b9063aabec5\n1f60964241fd\n1a56532cfc19\n1fbedd22a0b4\n1af43b2be6d3\n1c9477a191a7\n...\n\nso both test and train have rotated images. Just that the test has much more\n\nafter preparing the images from some software like rdkit?, the organizer simply rotates the image if w<h.",
    "1283914": "i have been using Lookahead+Radam optimizer. I only save trained model parameters and not the optimizer.\n\ni notice that when i restart training, the optimizer would need to restart of zero (because i didn't save the previous state)\nAnd there is always a drop and improvement in validation loss for the restart.\n\ni wonder what is the reason for this?",
    "1283843": "for timm vision transformer models, there is \"def resize_pos_embed(posemb, posemb_new, num_tokens=1)\" function that can resize your positional embedding when you finetune your transformer from small size to a bigger image size",
    "1277621": "there is JAX implementation for TNT\nhttps://github.com/NZ99/transformer_in_transformer_flax\n\nanyone want to take this code to TPU?\n(in my latest experience image size 448 is better than 384 .... i think ultimately, maybe we need to go to 640 for a constant input image size. if not, we need to break the input image into patch tokens ....)\n",
    "1275234": "This is great work, very educational. Thanks for your contribution @hengck23. \n\nPS: inference on the test set with 224x224 images takes me approximately 5 hours with a single 1080 Ti. ",
    "1272950": "i have completed basic implementation and experiments. Now it is time to go to the next stage ... uncovering the magic of the data (i.e. look at the images, errors, etc)\n\na quick study reveals\n1. at first I thought the images are randomly scaled. No, it isn't. there are only 2 scales. There is fixed (and different) padding of the image for the chemical structure for each scale. \n\nyou can double your train images by inter-converting images from one scale to another ...\n\nit seems that there much more to exploit. You can try it yourself  ",
    "1266193": "these files are for my early development and benchmarking performance. i need to change my strategy for future work:\n- use TIMM models and open-source transformer like huggingface, etc\n- the reason is because TIMM models provide support for torch.jit (and onnx for tensorRT)\n- i need to find a transformer accelerator that can run huggingface transformer model, e.g. Nvidia fast transformers, Tecent Turbo Transformers or Bytance LightSeq. Such accelerators provide speedup and also fast beam search etc.\n\n(Note, i will have to check license and copyright issue, but i will take care of these after i get my performance.)\n",
    "2765514": "Thank you for your contributions. I have 1 question. ViT is pre-trained with size 224*224, dimension 768. How to continue training from pretraining effectively (when changing size and dimension)?",
    "1329882": "Haven't tried it yet. For your second question, I would go for mixing the training set and the pseudo-labelled data rather than just the pseudo-labelled. As soon as you do pseudo-labelling you're at risk of overfitting to your pseudo-labels, and to the particular idiosyncrasies of whatever model produced them. To keep that risk lower, it's better to train with more data (and more diverse data).\n\nWould be interested to hear if anyone thinks otherwise.",
    "1316269": "Thank you for your excellent work. I retrained the [2021-apr-24] code, but found that the dev LB (Lev) is 1.29, so I think your shared pre-training model has converged, I added some code and predicted the test patch dataset, but After i submitted the csv, the LB score is above 5(worse than 07a(4.1)). I want to know why there is a huge gap between dev and test, maybe some of my prediction codes are wrong or I missed some details about your code. The only model code I modified is patch_pos, I use 9999 instead of any number exceed max_length in the patch matrix.",
    "1299749": "i note that there is something wrong with the YNakamaTokenizer that i am using.\nthe Tokenizer is learned from train data.\n\nthere are missing dictionary items, e.g\n\n```\nkey: value\n163: 100\n165: 101  #key164 is missing because there is not train data with such key\n166: 102\n\n```\n\nthis exposes another problem. if there is no train data with atom numbering 164, the model cannot make this prediction in test.\n\nI have a feeling that treating this as a image caption problem seems to be wrong. we should decode the image to graph directly (e.g. molfile)",
    "1299748": "![](https://i.ibb.co/VL3CY55/1620612915-1.png).  Thanks you for your excellent code, but there may be a bug in do_valid. Since the predict[i] should start from 0 instead of 1. It means that the tearcher-forcing's LD is very small?",
    "1297571": "Trying follow [2021-apr-24] by coding to the submit function and reproduce the CV score 1.35 (without teacher forcing) for 40000 valid datapoint  ( using the shared checkpoint 00922000_model.pth )\n\nHowever, I am only able to get CV 1.99.  Must be making some silly mistakes somewhere. \n\nIn `fairseq_model.py`, the `forward_argmax_decode` was commented out. I just assume that this part is complete and can be used without any modification.  Am I correct in assuming this ?  \n\n@hengck23  Can you confirm that [2021-apr-24]/ `forward_argmax_decode` can be used without modification and isn't part of the exercise ? \n\nAnyone else have similar issue with reproducing the CV for  [2021-apr-24]  ?  Cares to share any Gotcha ? \n\n",
    "1295819": "@hengck23 Thank you for your great work!\nCan I ask you a question, please.\nI'm using transformer decoder and my CV score is ~4.0. However, LB score is terrible - ~34.0. \nI've investigated that if I put image together with the target (like during training) I get a normal score, and if I'm predicting every token per timestamp (like when I don't know true InChI), the final score becomes much worser. \nI thought it could be because I'm not using teacher forcing during training. Did you have such a problem? \nOr maybe it is just my mistake in inference implementation.\n",
    "1295470": "@hengck23 have used batch size of 64 for inference part. Only 1.4 GB of GPU is being used in that case. Can I increase the batch size to 128, or will it have any effect on my predictions ? Is it like we have to use the same batch size for training and inference part ? ",
    "1293694": "Figured out.   Just need to be more patient and train for more epochs\n\n",
    "1292300": "@hengck23 How do you use suggest to use 320X320 size, the pretrained transformers is for 224X224. Or do you write another model def with needed size and use those from scratch? I am asking since the code doesn't implement the 320 size version. Sorry if this was already answered.",
    "1291481": "@hengck23 \n\nI tried to run [2021-apr-07b] without success.   Any suggestion on how to resolve these errors ? \n\n\n1 )  I tried running `run_check_fairseq_model.py`   But  `/checkpoint/00235000_model.pth` can not be found in your GDrive. I tried to substitute `00266000_model.pth` you had in  [2021-apr-07b] folder, but I got this error\n\n```\nRuntimeError: CUDA error: CUBLAS_STATUS_EXECUTION_FAILED when calling `cublasSgemm( handle, opa, opb, m, n, k, &alpha, a, lda, b, ldb, &beta, c, ldc)`\n\n```\n\nNot sure if this has anything to do with the wrong checkpoint\n\n2)  i tried running `run_train.py`.  Same problem with missing checkpoint. I just set initial_checkpoint = None \nBut I encoutered this error \n\n```\n untimeError: CUDA error: CUBLAS_STATUS_EXECUTION_FAILED when calling `cublasGemmEx( handle, opa, opb, m, n, k, &falpha, a, CUDA_R_16F, lda, b, CUDA_R_16F, ldb, &fbeta, c, CUDA_R_16F, ldc, CUDA_R_32F, CUBLAS_GEMM_DFALT_TENSOR_OP)`\n```\n\n\n\n\n",
    "1285508": "Can i  ask that where is the file \"df_train.more.csv.pickle\" ? \n🤔",
    "1278561": "good for clustering:\n\nSOFT EDIT DISTANCE FOR DIFFERENTIABLE COMPARISON OF\nSYMBOLIC SEQUENCES\n\nSoft_edit_distance_for_differentiable_comparison_o.pdf",
    "1276789": "I have a simple question, forgive me if It is repeated. can you submit your result if you do inference on local machines? if so  how?  thanks!",
    "1273385": "Hi Everyone, @hengck23,\nanyone on this panel please clear me\nIn file dataset_224.py, \nIs he taking the same files of our dataset or any resized images\nif he uses original sized images, why does he comment the resize function\n```\ndef null_augment(r):\n    image = r['image']\n    #image = cv2.resize(image, dsize=(image_size,image_size), interpolation=cv2.INTER_LINEAR)\n    assert image_size==224\n    r['image'] = image\n    return r\n```\nBut this makes sense in the dataset.py\n\nThanks in Advance",
    "1273233": "@hengck23 did you provided the inference script for your model? i cannot find a inference script in your drive",
    "1272646": "Did you think about use GPT to get strong models ",
    "1271348": "@hengck23 : Can you tell us how the `test_orientation.csv` file is generated?",
    "1270945": "Great work, very helpful!!!!!!!👍",
    "1270547": "list of vision transformer you can try:\nhttps://github.com/facebookresearch/deit/blob/main/README_cait.md\nhttps://github.com/facebookresearch/deit\n",
    "1270406": "@hengck23 : In the file ./2021-apr-07b/code/tnt-s-224-fairseq-v1-1/run_train.py, in the `run_train()` function from lines 262 to 283, we have:\n\n```\nif is_mixed_precision:\n\twith amp.autocast():\n\t\t#assert(False)\n\t\tlogit = net(image, token, length)\n\t\tloss0 = seq_cross_entropy_loss(logit, token, length)\n\t\t#loss0 = seq_anti_focal_cross_entropy_loss(logit, token, length)\n\n\tscaler.scale(loss0).backward()\n\t#scaler.unscale_(optimizer)\n\t#torch.nn.utils.clip_grad_norm_(net.parameters(), 2)\n\tscaler.step(optimizer)\n\tscaler.update()\n\nelse:\n\tassert False\n\t# print('fp32')\n\t# image_embed = encoder(image)\n\tlogit, weight = decoder(image_embed, token, length)\n\n\t(loss0).backward()\n\toptimizer.step()\n```\n\nIt's difficult for me to understand the logic of those lines of codes.\nCan you explain us why when `is_mixed_precision` is False, then we only call the decoder (that will fail due to `assert False`)?\nCan you also explain us where the `encoder` and `decoder` objects are defined?",
    "1269446": "Good pictures!",
    "1268991": "Thanks for sharing this great code. I see that your transformer didn't mask the padding token when computing self attention. Should I use the **tgt key padding mask** when using PyTorch's Transformer, because i think its more reasonable to mask padding. I'm not familiar with transformer, hope you can solve my confusion, thanks.",
    "1268806": "Thanks for the great repo! When I do inference, batch_size 32 would fail for a 8G GPU for OOM, when I change to smaller batch size like 4, it can run up to 140 or so for OOM. Do you know what might be wrong? Thanks!",
    "1267170": "Great work, it will help a lot! \nAs far as I understood, you rotate images according to test_orientation.csv. Could you please tell me how do you determine the orientation?",
    "1266156": "Thanks for your work @hengck23 ! Very helpful ",
    "1316121": "",
    "1307475": "",
    "1304509": "",
    "1303939": "",
    "1293613": "",
    "1267624": "This is gold... \nThanks for sharing..."
  }
}