{
  "id": 225856,
  "title": "Resnet + LSTM with attention starter",
  "url": "/competitions/bms-molecular-translation/discussion/225856",
  "author_name": "",
  "post_date": "2021-03-14T10:19:13.427413Z",
  "votes": 141,
  "comment_count": 49,
  "views": 0,
  "content": "<p>I prepared Resnet + LSTM with attention starter code.</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/inchi-preprocess-2\" target=\"_blank\">preprocess</a></li>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter\" target=\"_blank\">training</a></li>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-inference\" target=\"_blank\">inference</a></li>\n</ol>\n<p>The result is CV: 18.7, LB: 20.3</p>\n<p>In inference notebook, I used 2 epochs trained weight. There is much room to improve, for example more epochs, augmentation, larger models, larger size…</p>\n<p>Hope this helps, happy kaggling!</p>",
  "messages": [
    {
      "id": "1237631",
      "postDate": "03/14/2021 10:19:13",
      "content": "<p>I prepared Resnet + LSTM with attention starter code.</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/inchi-preprocess-2\" target=\"_blank\">preprocess</a></li>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter\" target=\"_blank\">training</a></li>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-inference\" target=\"_blank\">inference</a></li>\n</ol>\n<p>The result is CV: 18.7, LB: 20.3</p>\n<p>In inference notebook, I used 2 epochs trained weight. There is much room to improve, for example more epochs, augmentation, larger models, larger size…</p>\n<p>Hope this helps, happy kaggling!</p>",
      "rawMarkdown": "I prepared Resnet + LSTM with attention starter code.\n\n1. [preprocess](https://www.kaggle.com/yasufuminakama/inchi-preprocess-2)\n2. [training](https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter)\n3. [inference](https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-inference)\n\nThe result is CV: 18.7, LB: 20.3\n\nIn inference notebook, I used 2 epochs trained weight. There is much room to improve, for example more epochs, augmentation, larger models, larger size...\n\nHope this helps, happy kaggling!",
      "votes": null
    },
    {
      "id": "1237740",
      "postDate": "03/14/2021 12:15:32",
      "content": "<p>Wouldn't be better to just explain the workflow in the discussion without publishing the actual code?<br>\nI can already see 2 scores exactly at 21.9 like yours at 5th and 6th positions…</p>",
      "rawMarkdown": "Wouldn't be better to just explain the workflow in the discussion without publishing the actual code?\nI can already see 2 scores exactly at 21.9 like yours at 5th and 6th positions...",
      "votes": null
    },
    {
      "id": "1237787",
      "postDate": "03/14/2021 12:54:57",
      "content": "<p>I understand what you say, but it's easier to join the competition for beginners if there is starter code.<br>\nWhen I was beginner, I learned a lot from public notebooks, that's why I publish starter codes for other competitions too. <br>\nYou don't need to be so nervous about current LB position, it's just beginning of the competition. There are a lot of things to improve score.</p>",
      "rawMarkdown": "I understand what you say, but it's easier to join the competition for beginners if there is starter code.\nWhen I was beginner, I learned a lot from public notebooks, that's why I publish starter codes for other competitions too. \nYou don't need to be so nervous about current LB position, it's just beginning of the competition. There are a lot of things to improve score.",
      "votes": null
    },
    {
      "id": "1238267",
      "postDate": "03/14/2021 20:04:36",
      "content": "<p>I agree with you, I learned a lot from public notebooks in my previous competition (MoA) but the problem with beginners is that they tend to deceive themselves with copying and submitting the notebook without trying to understand the code (even at the beginning of the competition) and that's bad but they don't realize it. Also, they are people who try their best to learn and improve their scores and that is really demotivating for people who try their best.</p>\n<p>I personally never copy past code that I don't understand, and my message to <a href=\"https://www.kaggle.com/jacoporepossi\" target=\"_blank\">@jacoporepossi</a> <br>\nis to not give up and like <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> said it's the beginning of the competition and there is room for improvement and thank you for sharing your notebook with us, very helpful.</p>\n<p>By the way, How can we retrain models when we surpass the 9 hour limit of the Kaggle kernel?</p>",
      "rawMarkdown": "I agree with you, I learned a lot from public notebooks in my previous competition (MoA) but the problem with beginners is that they tend to deceive themselves with copying and submitting the notebook without trying to understand the code (even at the beginning of the competition) and that's bad but they don't realize it. Also, they are people who try their best to learn and improve their scores and that is really demotivating for people who try their best.\n\nI personally never copy past code that I don't understand, and my message to @jacoporepossi \nis to not give up and like @yasufuminakama said it's the beginning of the competition and there is room for improvement and thank you for sharing your notebook with us, very helpful.\n\nBy the way, How can we retrain models when we surpass the 9 hour limit of the Kaggle kernel?",
      "votes": null
    },
    {
      "id": "1238286",
      "postDate": "03/14/2021 20:35:03",
      "content": "<p>You can reload saved model states, optimizer states, scheduler states, and restart training.</p>",
      "rawMarkdown": "You can reload saved model states, optimizer states, scheduler states, and restart training.",
      "votes": null
    },
    {
      "id": "1238321",
      "postDate": "03/14/2021 22:01:11",
      "content": "<p>Don't get me wrong, I'm neither angry nor frustrated by the fact that <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> shared his code!! I'm sure it will help a lot of people and that's the beauty of being part of this wonderful community :)<br>\nI'm trying to put myself in someone else's shoes and imagine his frustration at being passed by people copying a notebook without doing anything. <br>\nObviously it's part of the game here in Kaggle, as you said it's just the beginning and there's room for improvement, it's not the end of the world. However if the aim was simply to provide beginners with ideas on how to approach the problem, probably a win-win approach could have been simply to discuss here about the implementation, just as they did <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/224257\" target=\"_blank\">here</a>.</p>\n<p>Btw, there are multiple competitions available for beginners to get started, as well as plenty of open notebooks from past competitions to learn</p>",
      "rawMarkdown": "Don't get me wrong, I'm neither angry nor frustrated by the fact that @yasufuminakama shared his code!! I'm sure it will help a lot of people and that's the beauty of being part of this wonderful community :)\nI'm trying to put myself in someone else's shoes and imagine his frustration at being passed by people copying a notebook without doing anything. \nObviously it's part of the game here in Kaggle, as you said it's just the beginning and there's room for improvement, it's not the end of the world. However if the aim was simply to provide beginners with ideas on how to approach the problem, probably a win-win approach could have been simply to discuss here about the implementation, just as they did [here](https://www.kaggle.com/c/bms-molecular-translation/discussion/224257).\n\nBtw, there are multiple competitions available for beginners to get started, as well as plenty of open notebooks from past competitions to learn",
      "votes": null
    },
    {
      "id": "1238783",
      "postDate": "03/15/2021 09:39:53",
      "content": "<p>I always enjoy your Attention starter discussions! They are very fruitful and helpful for all of us.<br>\nWouldn't be smarter to simply clarify the work process in the conversation without distributing the real code which you do, as many can have plagiarism.<br>\nBeing a starter, I learn from you rather than copying from your, but others don't do so🤷‍♂️ <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> </p>",
      "rawMarkdown": "I always enjoy your Attention starter discussions! They are very fruitful and helpful for all of us.\nWouldn't be smarter to simply clarify the work process in the conversation without distributing the real code which you do, as many can have plagiarism.\nBeing a starter, I learn from you rather than copying from your, but others don't do so🤷‍♂️ @yasufuminakama",
      "votes": null
    },
    {
      "id": "1239430",
      "postDate": "03/15/2021 17:56:10",
      "content": "<p>Enjoying reading u, thanks a lot, your discussion and your notebooks are useful </p>",
      "rawMarkdown": "Enjoying reading u, thanks a lot, your discussion and your notebooks are useful",
      "votes": null
    },
    {
      "id": "1239483",
      "postDate": "03/15/2021 18:52:49",
      "content": "<p>Thank you for sharing. This is a fascinating competition. Cross domain (i.e., image to text) problems get to the heart of what true human-like understanding is. I was aware of the Show, Attend, and Tell paper, but prior to you sharing your work it was too dense for me to work with directly. </p>\n<p>Can I ask, how long did you train the model for to get to CV 18.7?</p>",
      "rawMarkdown": "Thank you for sharing. This is a fascinating competition. Cross domain (i.e., image to text) problems get to the heart of what true human-like understanding is. I was aware of the Show, Attend, and Tell paper, but prior to you sharing your work it was too dense for me to work with directly. \n\nCan I ask, how long did you train the model for to get to CV 18.7?",
      "votes": null
    },
    {
      "id": "1239700",
      "postDate": "03/15/2021 23:53:08",
      "content": "<p>As I wrote, it's 2 epochs, and it took about 7.5h per epoch using kaggle notebooks as you can see training notebook.</p>",
      "rawMarkdown": "As I wrote, it's 2 epochs, and it took about 7.5h per epoch using kaggle notebooks as you can see training notebook.",
      "votes": null
    },
    {
      "id": "1239747",
      "postDate": "03/16/2021 01:20:23",
      "content": "<p>\" being passed by people copying a notebook without doing anything.\"</p>\n<p>that is part of kaggle and part of life.</p>\n<p>one can then speed his work by learning from the note book and improve on it. then you can overtake yourself.</p>\n<p>\"probably a win-win approach\"<br>\nunfortunately, the devils are in the details and not in the methods. there is a huge difference with and without the code.</p>\n<hr>\n<p>my suggestion is:</p>\n<ul>\n<li>you should capitalize on code posted. e.g. run the code with different parameters and observed how accuracy changes, etc. compare the code with yours and see which part is better (or worse)</li>\n<li>how to be better than those who just copy the code and do nothing? understand the code and improve on it.</li>\n</ul>\n<hr>\n<p>one may one to think about, which is more important:</p>\n<ol>\n<li>i got a medal, but I am inside my box (i.e. i am using methods developed from myself)</li>\n<li>i didn't get any medal, but l learned lots of useful techniques from others (methods that are out of my imagination … i have stepped out of my box)</li>\n</ol>",
      "rawMarkdown": "\" being passed by people copying a notebook without doing anything.\"\n\nthat is part of kaggle and part of life.\n\none can then speed his work by learning from the note book and improve on it. then you can overtake yourself.\n\n\"probably a win-win approach\"\nunfortunately, the devils are in the details and not in the methods. there is a huge difference with and without the code.\n\n---\n\nmy suggestion is:\n- you should capitalize on code posted. e.g. run the code with different parameters and observed how accuracy changes, etc. compare the code with yours and see which part is better (or worse)\n- how to be better than those who just copy the code and do nothing? understand the code and improve on it.\n\n---\none may one to think about, which is more important:\n1. i got a medal, but I am inside my box (i.e. i am using methods developed from myself)\n2. i didn't get any medal, but l learned lots of useful techniques from others (methods that are out of my imagination ... i have stepped out of my box)",
      "votes": null
    },
    {
      "id": "1239763",
      "postDate": "03/16/2021 02:01:13",
      "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> </p>\n<p>this should work. you can do it again!</p>\n<p><img src=\"https://i.ibb.co/L1S3cW0/Selection-352.png\" alt=\"\"></p>",
      "rawMarkdown": "yasufuminakama \n\nthis should work. you can do it again!\n\n![](https://i.ibb.co/L1S3cW0/Selection-352.png)",
      "votes": null
    },
    {
      "id": "1239785",
      "postDate": "03/16/2021 02:35:14",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <br>\nGlad to see you in this competition!<br>\nYes, I was thinking this kind of approach, thanks for kindly visualization :)</p>",
      "rawMarkdown": "hengck23 \nGlad to see you in this competition!\nYes, I was thinking this kind of approach, thanks for kindly visualization :)",
      "votes": null
    },
    {
      "id": "1240575",
      "postDate": "03/16/2021 14:04:03",
      "content": "<p>When RANZCR CLiP is over today and the usual suspects arrive here you will see this baseline start to fall on the leaderboard…</p>",
      "rawMarkdown": "When RANZCR CLiP is over today and the usual suspects arrive here you will see this baseline start to fall on the leaderboard...",
      "votes": null
    },
    {
      "id": "1243100",
      "postDate": "03/18/2021 02:45:59",
      "content": "<p>I think <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> has proven that he doesn't need to be worried about people forking his code. That being said, <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> thank you for the clearly explained code. Very helpful in making the \"Show, Attend and Tell\" paper understandable!</p>",
      "rawMarkdown": "I think @yasufuminakama has proven that he doesn't need to be worried about people forking his code. That being said, @yasufuminakama thank you for the clearly explained code. Very helpful in making the \"Show, Attend and Tell\" paper understandable!",
      "votes": null
    },
    {
      "id": "1244269",
      "postDate": "03/18/2021 21:06:30",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thanks for the suggestion.</p>\n<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> Can I assume from your jump to the top of the leaderboard, that the teacher-student approach worked very well?!</p>",
      "rawMarkdown": "hengck23 Thanks for the suggestion.\n\n@yasufuminakama Can I assume from your jump to the top of the leaderboard, that the teacher-student approach worked very well?!",
      "votes": null
    },
    {
      "id": "1244697",
      "postDate": "03/19/2021 06:49:37",
      "content": "<p>I haven't tried it yet, I want to know the result if someone tried it :)</p>",
      "rawMarkdown": "I haven't tried it yet, I want to know the result if someone tried it :)",
      "votes": null
    },
    {
      "id": "1245187",
      "postDate": "03/19/2021 14:50:19",
      "content": "<p>How might one go about making the training code parallel across multiple GPUs? I've tried the following idea in the \"training\" notebook:</p>\n<pre><code>   :\n   :\n    encoder = Encoder(CFG.model_name, pretrained=True)\n    if torch.cuda.device_count() &gt; 1:\n        print(\"Let's use\", torch.cuda.device_count(), \"GPUs!\")\n        encoder = nn.DataParallel(encoder)\n    encoder.to(device)\n   :\n   :\n</code></pre>\n<p>and the same with the <code>decoder</code>.  When I run training I can see the memory use go up evenly across all GPUs so something is working. However, I get the following error:</p>\n<pre><code>RuntimeError: Gather got an input of invalid size: got [64, 158, 193], but expected [64, 154, 193] (gather at /opt/conda/conda-bld/pytorch_1579022034529/work/torch/csrc/cuda/comm.cpp:231)\n</code></pre>\n<p>The batch size is an even multiple of the number of GPUs so that's fine.</p>",
      "rawMarkdown": "How might one go about making the training code parallel across multiple GPUs? I've tried the following idea in the \"training\" notebook:\n\n```\n   :\n   :\n    encoder = Encoder(CFG.model_name, pretrained=True)\n    if torch.cuda.device_count() > 1:\n        print(\"Let's use\", torch.cuda.device_count(), \"GPUs!\")\n        encoder = nn.DataParallel(encoder)\n    encoder.to(device)\n   :\n   :\n```\n\nand the same with the `decoder`.  When I run training I can see the memory use go up evenly across all GPUs so something is working. However, I get the following error:\n\n```\nRuntimeError: Gather got an input of invalid size: got [64, 158, 193], but expected [64, 154, 193] (gather at /opt/conda/conda-bld/pytorch_1579022034529/work/torch/csrc/cuda/comm.cpp:231)\n```\n\nThe batch size is an even multiple of the number of GPUs so that's fine.",
      "votes": null
    },
    {
      "id": "1245263",
      "postDate": "03/19/2021 16:18:45",
      "content": "<p>I haven't tried multiple GPUs training so I'm not sure but it seems label sequence length is not same.</p>",
      "rawMarkdown": "I haven't tried multiple GPUs training so I'm not sure but it seems label sequence length is not same.",
      "votes": null
    },
    {
      "id": "1245300",
      "postDate": "03/19/2021 16:54:06",
      "content": "<p>I found some clues <a href=\"https://pytorch.org/docs/stable/notes/faq.html#my-recurrent-network-doesn-t-work-with-data-parallelism\" target=\"_blank\">here</a></p>\n<blockquote>\n  <p>There is a subtlety in using the pack sequence -&gt; recurrent network -&gt; unpack sequence pattern in a Module with DataParallel or data_parallel(). Input to each the forward() on each device will only be part of the entire input. Because the unpack operation torch.nn.utils.rnn.pad_packed_sequence() by default only pads up to the longest input it sees, i.e., the longest on that particular device, size mismatches will happen when results are gathered together. Therefore, you can instead take advantage of the total_length argument of pad_packed_sequence() to make sure that the forward() calls return sequences of same length.</p>\n</blockquote>",
      "rawMarkdown": "I found some clues [here](https://pytorch.org/docs/stable/notes/faq.html#my-recurrent-network-doesn-t-work-with-data-parallelism)\n\n> There is a subtlety in using the pack sequence -> recurrent network -> unpack sequence pattern in a Module with DataParallel or data_parallel(). Input to each the forward() on each device will only be part of the entire input. Because the unpack operation torch.nn.utils.rnn.pad_packed_sequence() by default only pads up to the longest input it sees, i.e., the longest on that particular device, size mismatches will happen when results are gathered together. Therefore, you can instead take advantage of the total_length argument of pad_packed_sequence() to make sure that the forward() calls return sequences of same length.",
      "votes": null
    },
    {
      "id": "1246228",
      "postDate": "03/20/2021 14:59:40",
      "content": "<p>How do you think about the different between CV and LB ?<br>\nThe estimation of image orientation (w&gt;h or w&lt;h) is not perfect, is it ?</p>",
      "rawMarkdown": "How do you think about the different between CV and LB ?\nThe estimation of image orientation (w>h or w<h) is not perfect, is it ?",
      "votes": null
    },
    {
      "id": "1246476",
      "postDate": "03/20/2021 19:09:39",
      "content": "<p>the actual problem is with this line <br>\n<code>decode_lengths = (caption_lengths - 1).tolist()</code><br>\nin multi-gpu it splits into 4 lists so they get different lengths and on gathering you receive an error. <br>\nFast and dirty way is to run in parallel encoder, and decoder - not, in other words decoder runs on one gpu.<br>\n<code>the_number_of_gpu = torch.cuda.device_count()\n    if the_number_of_gpu &gt; 1:\n        encoder = nn.DataParallel(encoder)\n        # decoder = nn.DataParallel(decoder)  &lt;--- commented\n</code><br>\nin this case gpus won't be used optimally, unfortunately</p>",
      "rawMarkdown": "the actual problem is with this line \n`decode_lengths = (caption_lengths - 1).tolist()`\nin multi-gpu it splits into 4 lists so they get different lengths and on gathering you receive an error. \nFast and dirty way is to run in parallel encoder, and decoder - not, in other words decoder runs on one gpu.\n`the_number_of_gpu = torch.cuda.device_count()\n    if the_number_of_gpu > 1:\n        encoder = nn.DataParallel(encoder)\n        # decoder = nn.DataParallel(decoder)  <--- commented\n`\nin this case gpus won't be used optimally, unfortunately",
      "votes": null
    },
    {
      "id": "1246540",
      "postDate": "03/20/2021 20:32:04",
      "content": "<p>I haven't used <code>nn.DataParallel</code>, but <code>nn.parallel.DistributedDataParallel</code> doesn't throw any error. </p>",
      "rawMarkdown": "I haven't used `nn.DataParallel`, but `nn.parallel.DistributedDataParallel` doesn't throw any error.",
      "votes": null
    },
    {
      "id": "1246541",
      "postDate": "03/20/2021 20:33:31",
      "content": "<p>Not the author=) But I can add few cents… They gap between <code>CV</code> and <code>LB</code> is greatly reduced if you add <code>90</code> degree rotations augmentations during training. (or perhaps any rotation) </p>",
      "rawMarkdown": "Not the author=) But I can add few cents... They gap between `CV` and `LB` is greatly reduced if you add `90` degree rotations augmentations during training. (or perhaps any rotation)",
      "votes": null
    },
    {
      "id": "1246566",
      "postDate": "03/20/2021 21:10:45",
      "content": "<p>IMHO, I don't mind gap if there is correlation between CV and LB.</p>",
      "rawMarkdown": "IMHO, I don't mind gap if there is correlation between CV and LB.",
      "votes": null
    },
    {
      "id": "1246835",
      "postDate": "03/21/2021 06:48:57",
      "content": "<p>Yes, torch states that DistributedDataParallel (DDP) works faster and better than DataParallel like it is stated <a href=\"https://pytorch.org/tutorials/intermediate/ddp_tutorial.html\" target=\"_blank\">here</a>. <br>\n<a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a>  if you would create a kernel with DDP based on <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> code, lots of kagglers would be happy as DDP requires some additional implementations.  :)</p>",
      "rawMarkdown": "Yes, torch states that DistributedDataParallel (DDP) works faster and better than DataParallel like it is stated [here](https://pytorch.org/tutorials/intermediate/ddp_tutorial.html). \n@drhabib  if you would create a kernel with DDP based on @yasufuminakama code, lots of kagglers would be happy as DDP requires some additional implementations.  :)",
      "votes": null
    },
    {
      "id": "1248262",
      "postDate": "03/22/2021 13:30:24",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> for the link to <code>nn.parallel.DistributedDataParallel</code>. That looks very nice. Will try… </p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> if you would create a kernel with DDP based on <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> code, lots of kagglers would be happy as DDP requires some additional implementations. :)</p>\n</blockquote>\n<p>I would hit the upvote so hard that I think it would break my mouse… :-)</p>",
      "rawMarkdown": "Thanks @drhabib for the link to `nn.parallel.DistributedDataParallel`. That looks very nice. Will try... \n\n> @drhabib if you would create a kernel with DDP based on @yasufuminakama code, lots of kagglers would be happy as DDP requires some additional implementations. :)\n\nI would hit the upvote so hard that I think it would break my mouse... :-)",
      "votes": null
    },
    {
      "id": "1249711",
      "postDate": "03/23/2021 13:18:44",
      "content": "<p>I am struggling to get intuition about the shape of the feature tensor coming out of the encoder. This is resnet34 and we have removed the global pooling and fully connected layer and just pass down the last layer. Inspected as</p>\n<pre><code>from torchsummary import summary\nsummary(encoder, (3, 224, 224)\n\n----------------------------------------------------------------\n        Layer (type)               Output Shape         Param #\n================================================================\n            Conv2d-1         [-1, 64, 112, 112]           9,408\n       BatchNorm2d-2         [-1, 64, 112, 112]             128\n   :\n   :\n   :\nConv2d-119            [-1, 512, 7, 7]       2,359,296\n     BatchNorm2d-120            [-1, 512, 7, 7]           1,024\n            ReLU-121            [-1, 512, 7, 7]               0\n      BasicBlock-122            [-1, 512, 7, 7]               0\n        Identity-123            [-1, 512, 7, 7]               0\n        Identity-124            [-1, 512, 7, 7]               0\n          ResNet-125            [-1, 512, 7, 7]               0\n</code></pre>\n<p>So the shape of the feature tensor is <code>[batch size sample index, index of feature map, x, y]</code>.</p>\n<p>The <code>forward</code> method of the Encoder reshapes this feature tensor:</p>\n<pre><code>class Encoder(nn.Module):\n    def __init__(self, model_name='resnet18', pretrained=False):\n        super().__init__()\n        self.cnn = timm.create_model(model_name, pretrained=pretrained)\n        self.n_features = self.cnn.fc.in_features\n        self.cnn.global_pool = nn.Identity()\n        self.cnn.fc = nn.Identity()\n\n    def forward(self, x):\n        bs = x.size(0)\n        features = self.cnn(x)\n        features = features.permute(0, 2, 3, 1)\n        return features\n</code></pre>\n<p><code>features.permute(0, 2, 3, 1)</code> makes the shape of the tensor <code>[batch size sample index, x, y, index of feature map]</code></p>\n<p>I don't understand why we are doing this. I would be grateful if anyone could share some light on this! Thanks!</p>",
      "rawMarkdown": "I am struggling to get intuition about the shape of the feature tensor coming out of the encoder. This is resnet34 and we have removed the global pooling and fully connected layer and just pass down the last layer. Inspected as\n\n```\nfrom torchsummary import summary\nsummary(encoder, (3, 224, 224)\n\n----------------------------------------------------------------\n        Layer (type)               Output Shape         Param #\n================================================================\n            Conv2d-1         [-1, 64, 112, 112]           9,408\n       BatchNorm2d-2         [-1, 64, 112, 112]             128\n   :\n   :\n   :\nConv2d-119            [-1, 512, 7, 7]       2,359,296\n     BatchNorm2d-120            [-1, 512, 7, 7]           1,024\n            ReLU-121            [-1, 512, 7, 7]               0\n      BasicBlock-122            [-1, 512, 7, 7]               0\n        Identity-123            [-1, 512, 7, 7]               0\n        Identity-124            [-1, 512, 7, 7]               0\n          ResNet-125            [-1, 512, 7, 7]               0\n```\n\nSo the shape of the feature tensor is `[batch size sample index, index of feature map, x, y]`.\n\nThe `forward` method of the Encoder reshapes this feature tensor:\n\n```\nclass Encoder(nn.Module):\n    def __init__(self, model_name='resnet18', pretrained=False):\n        super().__init__()\n        self.cnn = timm.create_model(model_name, pretrained=pretrained)\n        self.n_features = self.cnn.fc.in_features\n        self.cnn.global_pool = nn.Identity()\n        self.cnn.fc = nn.Identity()\n\n    def forward(self, x):\n        bs = x.size(0)\n        features = self.cnn(x)\n        features = features.permute(0, 2, 3, 1)\n        return features\n```\n\n`features.permute(0, 2, 3, 1)` makes the shape of the tensor `[batch size sample index, x, y, index of feature map]`\n\nI don't understand why we are doing this. I would be grateful if anyone could share some light on this! Thanks!",
      "votes": null
    },
    {
      "id": "1249735",
      "postDate": "03/23/2021 13:36:42",
      "content": "<p>In general you can debug <code>Pytorch</code> code by just putting print statements in the <code>nn.model class</code></p>\n<pre><code>class Encoder(nn.Module):\n   def __init__(self, model_name='resnet18', pretrained=False):\n       super().__init__()\n       self.cnn = timm.create_model(model_name, pretrained=pretrained)\n       self.n_features = self.cnn.fc.in_features\n       self.cnn.global_pool = nn.Identity()\n       self.cnn.fc = nn.Identity()\n   def forward(self, x):\n       bs = x.size(0)\n       features = self.cnn(x)\n       print (features.shape)\n       print (f'Shape before permutation {features.shape}')\n       features = features.permute(0, 2, 3, 1)\n       print (f'Shape after permutation {features.shape}')\n       return features\n</code></pre>\n<pre><code>#put model in eval mode\nmodel = Encoder().eval()\nbs = 4\nc  = 3\nh  = 224\nw  = 224\nimage = torch.rand(bs, c, h, w)\n#getting features and printing statement \n#no_grad just tells that we don't want to calculate gradient \nwith torch.no_grad():\n   features = model(image)\n</code></pre>\n<p>output should look like:</p>\n<pre><code>Shape before permutation torch.Size([4, 512, 7, 7])\nShape after permutation torch.Size([4, 7, 7, 512])\n</code></pre>\n<p><code>torch.permute</code> just helps to switch features axis… Later we will reshape this features more in decoder to <code>(bs, 7 *7, 512)</code> so we can apply attention and initiate <code>LSTM</code> cell<br>\nNot sure I answered your question =) </p>",
      "rawMarkdown": "In general you can debug `Pytorch` code by just putting print statements in the `nn.model class`\n\n```\nclass Encoder(nn.Module):\n    def __init__(self, model_name='resnet18', pretrained=False):\n        super().__init__()\n        self.cnn = timm.create_model(model_name, pretrained=pretrained)\n        self.n_features = self.cnn.fc.in_features\n        self.cnn.global_pool = nn.Identity()\n        self.cnn.fc = nn.Identity()\n\n    def forward(self, x):\n        bs = x.size(0)\n        features = self.cnn(x)\n        print (features.shape)\n        print (f'Shape before permutation {features.shape}')\n        features = features.permute(0, 2, 3, 1)\n        print (f'Shape after permutation {features.shape}')\n        return features\n```\n\n```\n#put model in eval mode\nmodel = Encoder().eval()\n\nbs = 4\nc  = 3\nh  = 224\nw  = 224\nimage = torch.rand(bs, c, h, w)\n\n#getting features and printing statement \n#no_grad just tells that we don't want to calculate gradient \nwith torch.no_grad():\n    features = model(image)\n```\n\noutput should look like:\n```\nShape before permutation torch.Size([4, 512, 7, 7])\nShape after permutation torch.Size([4, 7, 7, 512])\n```\n\n`torch.permute` just helps to switch features axis... Later we will reshape this features more in decoder to `(bs, 7 *7, 512)` so we can apply attention and initiate `LSTM` cell\n\nNot sure I answered your question =)",
      "votes": null
    },
    {
      "id": "1249751",
      "postDate": "03/23/2021 13:49:27",
      "content": "<p>It's very detailed <a href=\"https://github.com/sgrvinod/a-PyTorch-Tutorial-to-Image-Captioning\" target=\"_blank\">here</a> … from the author</p>",
      "rawMarkdown": "It's very detailed [here](https://github.com/sgrvinod/a-PyTorch-Tutorial-to-Image-Captioning) ... from the author",
      "votes": null
    },
    {
      "id": "1249758",
      "postDate": "03/23/2021 13:51:09",
      "content": "<p>ahhh perfect! thanks for the link =)</p>",
      "rawMarkdown": "ahhh perfect! thanks for the link =)",
      "votes": null
    },
    {
      "id": "1249766",
      "postDate": "03/23/2021 13:59:46",
      "content": "<p>Thanks! I concur on the output shapes. That's what I intended to communicate. I understand what permute is doing technically. What I am asking is <strong>why</strong> we have to do this.</p>\n<p>Perhaps this is the answer</p>\n<blockquote>\n  <p>Later we will reshape this features more in decoder to (bs, 7 *7, 512) so we can apply attention and initiate LSTM cell</p>\n</blockquote>\n<p>Can you point in the code where this is happening? I could understand it if the 7x7 feature map was flattened to a vector, but still not getting this. 😢</p>",
      "rawMarkdown": "Thanks! I concur on the output shapes. That's what I intended to communicate. I understand what permute is doing technically. What I am asking is **why** we have to do this.\n\nPerhaps this is the answer\n\n> Later we will reshape this features more in decoder to (bs, 7 *7, 512) so we can apply attention and initiate LSTM cell\n\nCan you point in the code where this is happening? I could understand it if the 7x7 feature map was flattened to a vector, but still not getting this. :cry:",
      "votes": null
    },
    {
      "id": "1249770",
      "postDate": "03/23/2021 14:02:57",
      "content": "<blockquote>\n  <p>It's very detailed here … from the author</p>\n</blockquote>\n<p>That link just makes things <strong>more</strong> confusing to me, because in his README he is <strong>not</strong> <br>\npermuting the feature tensor. For example, see</p>\n<p><img src=\"https://github.com/sgrvinod/a-PyTorch-Tutorial-to-Image-Captioning/raw/master/img/decoder_att.png\" alt=\"\"></p>\n<p>So here, we maintain the shape of the tensor in the original CNN.</p>\n<p>EDIT:</p>\n<p>Having read more closely the README.md in this repo, he does talk about reshaping the tensor<br>\n<a href=\"https://github.com/sgrvinod/a-PyTorch-Tutorial-to-Image-Captioning#attention-1\" target=\"_blank\">https://github.com/sgrvinod/a-PyTorch-Tutorial-to-Image-Captioning#attention-1</a></p>\n<blockquote>\n  <p>The output of the Encoder is received here and flattened to dimensions N, 14 * 14, 2048. This is just convenient and prevents having to reshape the tensor multiple times.</p>\n</blockquote>",
      "rawMarkdown": "> It's very detailed here … from the author\n\nThat link just makes things **more** confusing to me, because in his README he is **not** \npermuting the feature tensor. For example, see\n\n![](https://github.com/sgrvinod/a-PyTorch-Tutorial-to-Image-Captioning/raw/master/img/decoder_att.png)\n\n\nSo here, we maintain the shape of the tensor in the original CNN.\n\nEDIT:\n\nHaving read more closely the README.md in this repo, he does talk about reshaping the tensor\nhttps://github.com/sgrvinod/a-PyTorch-Tutorial-to-Image-Captioning#attention-1\n\n> The output of the Encoder is received here and flattened to dimensions N, 14 * 14, 2048. This is just convenient and prevents having to reshape the tensor multiple times.",
      "votes": null
    },
    {
      "id": "1249774",
      "postDate": "03/23/2021 14:04:34",
      "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter</a></p>\n<p><code>class DecoderWithAttention(nn.Module):</code></p>\n<pre><code>        encoder_out = encoder_out.view(batch_size, -1, encoder_dim)  # (batch_size, num_pixels, encoder_dim)\n        num_pixels = encoder_out.size(1)\n</code></pre>\n<pre><code># (bs, 7, 7, 512)\n#encoder_dim = 512\nencoder_out.view(batch_size, -1, encoder_dim) #(bs, 7 * 7, encoder_dim) (-1 will flatten axis 1 and 2)\n</code></pre>",
      "rawMarkdown": "https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter\n\n```class DecoderWithAttention(nn.Module):```\n\n```\n        encoder_out = encoder_out.view(batch_size, -1, encoder_dim)  # (batch_size, num_pixels, encoder_dim)\n        num_pixels = encoder_out.size(1)\n```\n\n```\n# (bs, 7, 7, 512)\n#encoder_dim = 512\nencoder_out.view(batch_size, -1, encoder_dim) #(bs, 7 * 7, encoder_dim) (-1 will flatten axis 1 and 2)\n```",
      "votes": null
    },
    {
      "id": "1249782",
      "postDate": "03/23/2021 14:11:13",
      "content": "<blockquote>\n  <p><code>encoder_out.view(batch_size, -1, encoder_dim) #(bs, 7 * 7, encoder_dim) (-1 will flatten axis 1 and 2)</code></p>\n</blockquote>\n<p>That's it!! Thank you!</p>",
      "rawMarkdown": "> `encoder_out.view(batch_size, -1, encoder_dim) #(bs, 7 * 7, encoder_dim) (-1 will flatten axis 1 and 2)`\n\nThat's it!! Thank you!",
      "votes": null
    },
    {
      "id": "1251592",
      "postDate": "03/24/2021 23:36:01",
      "content": "<p>I have tried it. It seems that the simple annotations in the <code>rdkit</code> generated images are not enough for the teacher model to pick up on: stage 1 performance is only marginally better than training with the original images. Perhaps I need to use a stronger encoder to see more significant results. I will share my notebooks this weekend when I have some GPU and hopefully a more experienced Kaggler can find a way of improving it. (I think I just need to generate more annotations via <code>rdkit</code> to make it work i.e. highlight bonds / atoms).  </p>",
      "rawMarkdown": "I have tried it. It seems that the simple annotations in the `rdkit` generated images are not enough for the teacher model to pick up on: stage 1 performance is only marginally better than training with the original images. Perhaps I need to use a stronger encoder to see more significant results. I will share my notebooks this weekend when I have some GPU and hopefully a more experienced Kaggler can find a way of improving it. (I think I just need to generate more annotations via `rdkit` to make it work i.e. highlight bonds / atoms).",
      "votes": null
    },
    {
      "id": "1251606",
      "postDate": "03/25/2021 00:10:13",
      "content": "<p>Thanks for sharing your result!</p>",
      "rawMarkdown": "Thanks for sharing your result!",
      "votes": null
    },
    {
      "id": "1251623",
      "postDate": "03/25/2021 00:39:22",
      "content": "<p>I think this level of annotation would be better:</p>\n<p><img src=\"https://www.rdkit.org/docs/_images/atom_highlights_3.png\" alt=\"\"></p>\n<p>Will try it next and report results. </p>\n<p>Edit: <br>\nSome useful code can be found <a href=\"https://github.com/rdkit/rdkit/blob/master/Docs/Book/data/test_multi_colours.py\" target=\"_blank\">here</a>. We can use <a href=\"https://docs.chemaxon.com/display/docs/smarts.md\" target=\"_blank\">SMART queries</a> to highlight arbitrary substructures of the chemicals. I don't know how to best exploit this, as I know very little about chemistry - I'm a physicist :D</p>\n<p>For example, we can pass SMART queries like: <br>\n<code>[\"B\",\"Br\",\"C\",\"Cl\",\"F\",\"I\",\"N\",\"O\",\"P\",\"S\"]</code><br>\nto draw circles around the atoms and we can pass queries like:<br>\n<code>[\"[r{3-}]\", \"*~*~*\"]</code> <br>\nto highlight rings of radius <code>3</code> or greater and <code>3</code> linear atoms respectively.</p>",
      "rawMarkdown": "I think this level of annotation would be better:\n\n![](https://www.rdkit.org/docs/_images/atom_highlights_3.png)\n\nWill try it next and report results. \n\nEdit: \nSome useful code can be found [here](https://github.com/rdkit/rdkit/blob/master/Docs/Book/data/test_multi_colours.py). We can use [SMART queries](https://docs.chemaxon.com/display/docs/smarts.md) to highlight arbitrary substructures of the chemicals. I don't know how to best exploit this, as I know very little about chemistry - I'm a physicist :D\n\nFor example, we can pass SMART queries like: \n`[\"B\",\"Br\",\"C\",\"Cl\",\"F\",\"I\",\"N\",\"O\",\"P\",\"S\"]`\nto draw circles around the atoms and we can pass queries like:\n`[\"[r{3-}]\", \"*~*~*\"]` \nto highlight rings of radius `3` or greater and `3` linear atoms respectively.",
      "votes": null
    },
    {
      "id": "1252076",
      "postDate": "03/25/2021 12:09:14",
      "content": "<p>Why you  don't care if there is correlation between CV and LB? Can you explain it?</p>",
      "rawMarkdown": "Why you  don't care if there is correlation between CV and LB? Can you explain it?",
      "votes": null
    },
    {
      "id": "1252095",
      "postDate": "03/25/2021 12:22:36",
      "content": "<p>I found <a href=\"https://github.com/arogozhnikov/einops\" target=\"_blank\">einops</a> very helpful for this type of code. <br>\n<code>x = rearrange(x, 'b c h w -&gt; b (h w) c')</code>  is very readable.</p>",
      "rawMarkdown": "I found [einops](https://github.com/arogozhnikov/einops) very helpful for this type of code. \n`x = rearrange(x, 'b c h w -> b (h w) c')`  is very readable.",
      "votes": null
    },
    {
      "id": "1252134",
      "postDate": "03/25/2021 12:53:07",
      "content": "<p>I think there must be gap between Single-Fold CV and LB for most people, but I think if we perform N-fold ensemble, I think CV is much closer to LB. I haven't tried it, Just my intuition :)</p>",
      "rawMarkdown": "I think there must be gap between Single-Fold CV and LB for most people, but I think if we perform N-fold ensemble, I think CV is much closer to LB. I haven't tried it, Just my intuition :)",
      "votes": null
    },
    {
      "id": "1252164",
      "postDate": "03/25/2021 13:21:52",
      "content": "<p><code>einops</code> looks amazing! Very elegant and clear code. Thanks.</p>",
      "rawMarkdown": "`einops` looks amazing! Very elegant and clear code. Thanks.",
      "votes": null
    },
    {
      "id": "1253027",
      "postDate": "03/26/2021 10:03:48",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> for sharing. This is very helpful. </p>",
      "rawMarkdown": "Thank you @yasufuminakama for sharing. This is very helpful.",
      "votes": null
    },
    {
      "id": "1260702",
      "postDate": "04/02/2021 10:58:05",
      "content": "<p>some good explanation here: <a href=\"https://medium.com/analytics-vidhya/image-captioning-with-attention-part-1-e8a5f783f6d3\" target=\"_blank\">https://medium.com/analytics-vidhya/image-captioning-with-attention-part-1-e8a5f783f6d3</a></p>",
      "rawMarkdown": "some good explanation here: https://medium.com/analytics-vidhya/image-captioning-with-attention-part-1-e8a5f783f6d3",
      "votes": null
    },
    {
      "id": "1271143",
      "postDate": "04/12/2021 10:39:01",
      "content": "<p>Wow! This is going to be really helpful for someone like me who has just started to learn about RNNs. Thank you for sharing. 🙏🏼</p>",
      "rawMarkdown": "Wow! This is going to be really helpful for someone like me who has just started to learn about RNNs. Thank you for sharing. 🙏🏼",
      "votes": null
    },
    {
      "id": "1274464",
      "postDate": "04/15/2021 10:15:44",
      "content": "<p>Good work!</p>",
      "rawMarkdown": "Good work!",
      "votes": null
    },
    {
      "id": "1284992",
      "postDate": "04/26/2021 13:29:20",
      "content": "<p>I'm joining this comp a month later and this score is way down the leaderboard. As op points out, this is what normally happens in these comps. The code has been super useful in getting me started on this comp</p>",
      "rawMarkdown": "I'm joining this comp a month later and this score is way down the leaderboard. As op points out, this is what normally happens in these comps. The code has been super useful in getting me started on this comp",
      "votes": null
    },
    {
      "id": "2747540",
      "postDate": "04/12/2024 00:10:19",
      "content": "<p>Super useful!</p>",
      "rawMarkdown": "Super useful!",
      "votes": null
    },
    {
      "id": "3196935",
      "postDate": "05/07/2025 16:28:26",
      "content": "<p>Great Work!</p>",
      "rawMarkdown": "Great Work!",
      "votes": null
    },
    {
      "id": "3306415",
      "postDate": "10/24/2025 13:59:44",
      "content": "<p>Nicely explained </p>",
      "rawMarkdown": "Nicely explained",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1237740,
      "author_name": "jacoporepossi",
      "author_url": "",
      "post_date": "03/14/2021 12:15:32",
      "content": "<p>Wouldn't be better to just explain the workflow in the discussion without publishing the actual code?<br>\nI can already see 2 scores exactly at 21.9 like yours at 5th and 6th positions…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1237787,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "03/14/2021 12:54:57",
          "content": "<p>I understand what you say, but it's easier to join the competition for beginners if there is starter code.<br>\nWhen I was beginner, I learned a lot from public notebooks, that's why I publish starter codes for other competitions too. <br>\nYou don't need to be so nervous about current LB position, it's just beginning of the competition. There are a lot of things to improve score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1238267,
          "author_name": "medali1992",
          "author_url": "",
          "post_date": "03/14/2021 20:04:36",
          "content": "<p>I agree with you, I learned a lot from public notebooks in my previous competition (MoA) but the problem with beginners is that they tend to deceive themselves with copying and submitting the notebook without trying to understand the code (even at the beginning of the competition) and that's bad but they don't realize it. Also, they are people who try their best to learn and improve their scores and that is really demotivating for people who try their best.</p>\n<p>I personally never copy past code that I don't understand, and my message to <a href=\"https://www.kaggle.com/jacoporepossi\" target=\"_blank\">@jacoporepossi</a> <br>\nis to not give up and like <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> said it's the beginning of the competition and there is room for improvement and thank you for sharing your notebook with us, very helpful.</p>\n<p>By the way, How can we retrain models when we surpass the 9 hour limit of the Kaggle kernel?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1238286,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "03/14/2021 20:35:03",
          "content": "<p>You can reload saved model states, optimizer states, scheduler states, and restart training.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1238321,
          "author_name": "jacoporepossi",
          "author_url": "",
          "post_date": "03/14/2021 22:01:11",
          "content": "<p>Don't get me wrong, I'm neither angry nor frustrated by the fact that <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> shared his code!! I'm sure it will help a lot of people and that's the beauty of being part of this wonderful community :)<br>\nI'm trying to put myself in someone else's shoes and imagine his frustration at being passed by people copying a notebook without doing anything. <br>\nObviously it's part of the game here in Kaggle, as you said it's just the beginning and there's room for improvement, it's not the end of the world. However if the aim was simply to provide beginners with ideas on how to approach the problem, probably a win-win approach could have been simply to discuss here about the implementation, just as they did <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/224257\" target=\"_blank\">here</a>.</p>\n<p>Btw, there are multiple competitions available for beginners to get started, as well as plenty of open notebooks from past competitions to learn</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1239747,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/16/2021 01:20:23",
          "content": "<p>\" being passed by people copying a notebook without doing anything.\"</p>\n<p>that is part of kaggle and part of life.</p>\n<p>one can then speed his work by learning from the note book and improve on it. then you can overtake yourself.</p>\n<p>\"probably a win-win approach\"<br>\nunfortunately, the devils are in the details and not in the methods. there is a huge difference with and without the code.</p>\n<hr>\n<p>my suggestion is:</p>\n<ul>\n<li>you should capitalize on code posted. e.g. run the code with different parameters and observed how accuracy changes, etc. compare the code with yours and see which part is better (or worse)</li>\n<li>how to be better than those who just copy the code and do nothing? understand the code and improve on it.</li>\n</ul>\n<hr>\n<p>one may one to think about, which is more important:</p>\n<ol>\n<li>i got a medal, but I am inside my box (i.e. i am using methods developed from myself)</li>\n<li>i didn't get any medal, but l learned lots of useful techniques from others (methods that are out of my imagination … i have stepped out of my box)</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1240575,
          "author_name": "marketneutral",
          "author_url": "",
          "post_date": "03/16/2021 14:04:03",
          "content": "<p>When RANZCR CLiP is over today and the usual suspects arrive here you will see this baseline start to fall on the leaderboard…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1243100,
          "author_name": "jjinho",
          "author_url": "",
          "post_date": "03/18/2021 02:45:59",
          "content": "<p>I think <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> has proven that he doesn't need to be worried about people forking his code. That being said, <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> thank you for the clearly explained code. Very helpful in making the \"Show, Attend and Tell\" paper understandable!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1284992,
          "author_name": "alexandersoare",
          "author_url": "",
          "post_date": "04/26/2021 13:29:20",
          "content": "<p>I'm joining this comp a month later and this score is way down the leaderboard. As op points out, this is what normally happens in these comps. The code has been super useful in getting me started on this comp</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1238783,
      "author_name": "arpit3043",
      "author_url": "",
      "post_date": "03/15/2021 09:39:53",
      "content": "<p>I always enjoy your Attention starter discussions! They are very fruitful and helpful for all of us.<br>\nWouldn't be smarter to simply clarify the work process in the conversation without distributing the real code which you do, as many can have plagiarism.<br>\nBeing a starter, I learn from you rather than copying from your, but others don't do so🤷‍♂️ <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1239430,
      "author_name": "salimkhazem",
      "author_url": "",
      "post_date": "03/15/2021 17:56:10",
      "content": "<p>Enjoying reading u, thanks a lot, your discussion and your notebooks are useful </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1239483,
      "author_name": "marketneutral",
      "author_url": "",
      "post_date": "03/15/2021 18:52:49",
      "content": "<p>Thank you for sharing. This is a fascinating competition. Cross domain (i.e., image to text) problems get to the heart of what true human-like understanding is. I was aware of the Show, Attend, and Tell paper, but prior to you sharing your work it was too dense for me to work with directly. </p>\n<p>Can I ask, how long did you train the model for to get to CV 18.7?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1239700,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "03/15/2021 23:53:08",
          "content": "<p>As I wrote, it's 2 epochs, and it took about 7.5h per epoch using kaggle notebooks as you can see training notebook.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1239763,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/16/2021 02:01:13",
      "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> </p>\n<p>this should work. you can do it again!</p>\n<p><img src=\"https://i.ibb.co/L1S3cW0/Selection-352.png\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1239785,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "03/16/2021 02:35:14",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <br>\nGlad to see you in this competition!<br>\nYes, I was thinking this kind of approach, thanks for kindly visualization :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1244269,
          "author_name": "andypenrose",
          "author_url": "",
          "post_date": "03/18/2021 21:06:30",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thanks for the suggestion.</p>\n<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> Can I assume from your jump to the top of the leaderboard, that the teacher-student approach worked very well?!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1244697,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "03/19/2021 06:49:37",
          "content": "<p>I haven't tried it yet, I want to know the result if someone tried it :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1251592,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "03/24/2021 23:36:01",
          "content": "<p>I have tried it. It seems that the simple annotations in the <code>rdkit</code> generated images are not enough for the teacher model to pick up on: stage 1 performance is only marginally better than training with the original images. Perhaps I need to use a stronger encoder to see more significant results. I will share my notebooks this weekend when I have some GPU and hopefully a more experienced Kaggler can find a way of improving it. (I think I just need to generate more annotations via <code>rdkit</code> to make it work i.e. highlight bonds / atoms).  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1251606,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "03/25/2021 00:10:13",
          "content": "<p>Thanks for sharing your result!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1251623,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "03/25/2021 00:39:22",
          "content": "<p>I think this level of annotation would be better:</p>\n<p><img src=\"https://www.rdkit.org/docs/_images/atom_highlights_3.png\" alt=\"\"></p>\n<p>Will try it next and report results. </p>\n<p>Edit: <br>\nSome useful code can be found <a href=\"https://github.com/rdkit/rdkit/blob/master/Docs/Book/data/test_multi_colours.py\" target=\"_blank\">here</a>. We can use <a href=\"https://docs.chemaxon.com/display/docs/smarts.md\" target=\"_blank\">SMART queries</a> to highlight arbitrary substructures of the chemicals. I don't know how to best exploit this, as I know very little about chemistry - I'm a physicist :D</p>\n<p>For example, we can pass SMART queries like: <br>\n<code>[\"B\",\"Br\",\"C\",\"Cl\",\"F\",\"I\",\"N\",\"O\",\"P\",\"S\"]</code><br>\nto draw circles around the atoms and we can pass queries like:<br>\n<code>[\"[r{3-}]\", \"*~*~*\"]</code> <br>\nto highlight rings of radius <code>3</code> or greater and <code>3</code> linear atoms respectively.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1245187,
      "author_name": "marketneutral",
      "author_url": "",
      "post_date": "03/19/2021 14:50:19",
      "content": "<p>How might one go about making the training code parallel across multiple GPUs? I've tried the following idea in the \"training\" notebook:</p>\n<pre><code>   :\n   :\n    encoder = Encoder(CFG.model_name, pretrained=True)\n    if torch.cuda.device_count() &gt; 1:\n        print(\"Let's use\", torch.cuda.device_count(), \"GPUs!\")\n        encoder = nn.DataParallel(encoder)\n    encoder.to(device)\n   :\n   :\n</code></pre>\n<p>and the same with the <code>decoder</code>.  When I run training I can see the memory use go up evenly across all GPUs so something is working. However, I get the following error:</p>\n<pre><code>RuntimeError: Gather got an input of invalid size: got [64, 158, 193], but expected [64, 154, 193] (gather at /opt/conda/conda-bld/pytorch_1579022034529/work/torch/csrc/cuda/comm.cpp:231)\n</code></pre>\n<p>The batch size is an even multiple of the number of GPUs so that's fine.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1245263,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "03/19/2021 16:18:45",
          "content": "<p>I haven't tried multiple GPUs training so I'm not sure but it seems label sequence length is not same.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1245300,
          "author_name": "marketneutral",
          "author_url": "",
          "post_date": "03/19/2021 16:54:06",
          "content": "<p>I found some clues <a href=\"https://pytorch.org/docs/stable/notes/faq.html#my-recurrent-network-doesn-t-work-with-data-parallelism\" target=\"_blank\">here</a></p>\n<blockquote>\n  <p>There is a subtlety in using the pack sequence -&gt; recurrent network -&gt; unpack sequence pattern in a Module with DataParallel or data_parallel(). Input to each the forward() on each device will only be part of the entire input. Because the unpack operation torch.nn.utils.rnn.pad_packed_sequence() by default only pads up to the longest input it sees, i.e., the longest on that particular device, size mismatches will happen when results are gathered together. Therefore, you can instead take advantage of the total_length argument of pad_packed_sequence() to make sure that the forward() calls return sequences of same length.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1246476,
          "author_name": "ubique",
          "author_url": "",
          "post_date": "03/20/2021 19:09:39",
          "content": "<p>the actual problem is with this line <br>\n<code>decode_lengths = (caption_lengths - 1).tolist()</code><br>\nin multi-gpu it splits into 4 lists so they get different lengths and on gathering you receive an error. <br>\nFast and dirty way is to run in parallel encoder, and decoder - not, in other words decoder runs on one gpu.<br>\n<code>the_number_of_gpu = torch.cuda.device_count()\n    if the_number_of_gpu &gt; 1:\n        encoder = nn.DataParallel(encoder)\n        # decoder = nn.DataParallel(decoder)  &lt;--- commented\n</code><br>\nin this case gpus won't be used optimally, unfortunately</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1246540,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "03/20/2021 20:32:04",
          "content": "<p>I haven't used <code>nn.DataParallel</code>, but <code>nn.parallel.DistributedDataParallel</code> doesn't throw any error. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1246835,
          "author_name": "ubique",
          "author_url": "",
          "post_date": "03/21/2021 06:48:57",
          "content": "<p>Yes, torch states that DistributedDataParallel (DDP) works faster and better than DataParallel like it is stated <a href=\"https://pytorch.org/tutorials/intermediate/ddp_tutorial.html\" target=\"_blank\">here</a>. <br>\n<a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a>  if you would create a kernel with DDP based on <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> code, lots of kagglers would be happy as DDP requires some additional implementations.  :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1248262,
          "author_name": "marketneutral",
          "author_url": "",
          "post_date": "03/22/2021 13:30:24",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> for the link to <code>nn.parallel.DistributedDataParallel</code>. That looks very nice. Will try… </p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> if you would create a kernel with DDP based on <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> code, lots of kagglers would be happy as DDP requires some additional implementations. :)</p>\n</blockquote>\n<p>I would hit the upvote so hard that I think it would break my mouse… :-)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1246228,
      "author_name": "toshik",
      "author_url": "",
      "post_date": "03/20/2021 14:59:40",
      "content": "<p>How do you think about the different between CV and LB ?<br>\nThe estimation of image orientation (w&gt;h or w&lt;h) is not perfect, is it ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1246541,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "03/20/2021 20:33:31",
          "content": "<p>Not the author=) But I can add few cents… They gap between <code>CV</code> and <code>LB</code> is greatly reduced if you add <code>90</code> degree rotations augmentations during training. (or perhaps any rotation) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1246566,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "03/20/2021 21:10:45",
          "content": "<p>IMHO, I don't mind gap if there is correlation between CV and LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1252076,
          "author_name": "fuxungao",
          "author_url": "",
          "post_date": "03/25/2021 12:09:14",
          "content": "<p>Why you  don't care if there is correlation between CV and LB? Can you explain it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1252134,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "03/25/2021 12:53:07",
          "content": "<p>I think there must be gap between Single-Fold CV and LB for most people, but I think if we perform N-fold ensemble, I think CV is much closer to LB. I haven't tried it, Just my intuition :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1249711,
      "author_name": "marketneutral",
      "author_url": "",
      "post_date": "03/23/2021 13:18:44",
      "content": "<p>I am struggling to get intuition about the shape of the feature tensor coming out of the encoder. This is resnet34 and we have removed the global pooling and fully connected layer and just pass down the last layer. Inspected as</p>\n<pre><code>from torchsummary import summary\nsummary(encoder, (3, 224, 224)\n\n----------------------------------------------------------------\n        Layer (type)               Output Shape         Param #\n================================================================\n            Conv2d-1         [-1, 64, 112, 112]           9,408\n       BatchNorm2d-2         [-1, 64, 112, 112]             128\n   :\n   :\n   :\nConv2d-119            [-1, 512, 7, 7]       2,359,296\n     BatchNorm2d-120            [-1, 512, 7, 7]           1,024\n            ReLU-121            [-1, 512, 7, 7]               0\n      BasicBlock-122            [-1, 512, 7, 7]               0\n        Identity-123            [-1, 512, 7, 7]               0\n        Identity-124            [-1, 512, 7, 7]               0\n          ResNet-125            [-1, 512, 7, 7]               0\n</code></pre>\n<p>So the shape of the feature tensor is <code>[batch size sample index, index of feature map, x, y]</code>.</p>\n<p>The <code>forward</code> method of the Encoder reshapes this feature tensor:</p>\n<pre><code>class Encoder(nn.Module):\n    def __init__(self, model_name='resnet18', pretrained=False):\n        super().__init__()\n        self.cnn = timm.create_model(model_name, pretrained=pretrained)\n        self.n_features = self.cnn.fc.in_features\n        self.cnn.global_pool = nn.Identity()\n        self.cnn.fc = nn.Identity()\n\n    def forward(self, x):\n        bs = x.size(0)\n        features = self.cnn(x)\n        features = features.permute(0, 2, 3, 1)\n        return features\n</code></pre>\n<p><code>features.permute(0, 2, 3, 1)</code> makes the shape of the tensor <code>[batch size sample index, x, y, index of feature map]</code></p>\n<p>I don't understand why we are doing this. I would be grateful if anyone could share some light on this! Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1249735,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "03/23/2021 13:36:42",
          "content": "<p>In general you can debug <code>Pytorch</code> code by just putting print statements in the <code>nn.model class</code></p>\n<pre><code>class Encoder(nn.Module):\n   def __init__(self, model_name='resnet18', pretrained=False):\n       super().__init__()\n       self.cnn = timm.create_model(model_name, pretrained=pretrained)\n       self.n_features = self.cnn.fc.in_features\n       self.cnn.global_pool = nn.Identity()\n       self.cnn.fc = nn.Identity()\n   def forward(self, x):\n       bs = x.size(0)\n       features = self.cnn(x)\n       print (features.shape)\n       print (f'Shape before permutation {features.shape}')\n       features = features.permute(0, 2, 3, 1)\n       print (f'Shape after permutation {features.shape}')\n       return features\n</code></pre>\n<pre><code>#put model in eval mode\nmodel = Encoder().eval()\nbs = 4\nc  = 3\nh  = 224\nw  = 224\nimage = torch.rand(bs, c, h, w)\n#getting features and printing statement \n#no_grad just tells that we don't want to calculate gradient \nwith torch.no_grad():\n   features = model(image)\n</code></pre>\n<p>output should look like:</p>\n<pre><code>Shape before permutation torch.Size([4, 512, 7, 7])\nShape after permutation torch.Size([4, 7, 7, 512])\n</code></pre>\n<p><code>torch.permute</code> just helps to switch features axis… Later we will reshape this features more in decoder to <code>(bs, 7 *7, 512)</code> so we can apply attention and initiate <code>LSTM</code> cell<br>\nNot sure I answered your question =) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1249751,
          "author_name": "sapr3s",
          "author_url": "",
          "post_date": "03/23/2021 13:49:27",
          "content": "<p>It's very detailed <a href=\"https://github.com/sgrvinod/a-PyTorch-Tutorial-to-Image-Captioning\" target=\"_blank\">here</a> … from the author</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1249758,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "03/23/2021 13:51:09",
          "content": "<p>ahhh perfect! thanks for the link =)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1249766,
          "author_name": "marketneutral",
          "author_url": "",
          "post_date": "03/23/2021 13:59:46",
          "content": "<p>Thanks! I concur on the output shapes. That's what I intended to communicate. I understand what permute is doing technically. What I am asking is <strong>why</strong> we have to do this.</p>\n<p>Perhaps this is the answer</p>\n<blockquote>\n  <p>Later we will reshape this features more in decoder to (bs, 7 *7, 512) so we can apply attention and initiate LSTM cell</p>\n</blockquote>\n<p>Can you point in the code where this is happening? I could understand it if the 7x7 feature map was flattened to a vector, but still not getting this. 😢</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1249770,
          "author_name": "marketneutral",
          "author_url": "",
          "post_date": "03/23/2021 14:02:57",
          "content": "<blockquote>\n  <p>It's very detailed here … from the author</p>\n</blockquote>\n<p>That link just makes things <strong>more</strong> confusing to me, because in his README he is <strong>not</strong> <br>\npermuting the feature tensor. For example, see</p>\n<p><img src=\"https://github.com/sgrvinod/a-PyTorch-Tutorial-to-Image-Captioning/raw/master/img/decoder_att.png\" alt=\"\"></p>\n<p>So here, we maintain the shape of the tensor in the original CNN.</p>\n<p>EDIT:</p>\n<p>Having read more closely the README.md in this repo, he does talk about reshaping the tensor<br>\n<a href=\"https://github.com/sgrvinod/a-PyTorch-Tutorial-to-Image-Captioning#attention-1\" target=\"_blank\">https://github.com/sgrvinod/a-PyTorch-Tutorial-to-Image-Captioning#attention-1</a></p>\n<blockquote>\n  <p>The output of the Encoder is received here and flattened to dimensions N, 14 * 14, 2048. This is just convenient and prevents having to reshape the tensor multiple times.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1249774,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "03/23/2021 14:04:34",
          "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter</a></p>\n<p><code>class DecoderWithAttention(nn.Module):</code></p>\n<pre><code>        encoder_out = encoder_out.view(batch_size, -1, encoder_dim)  # (batch_size, num_pixels, encoder_dim)\n        num_pixels = encoder_out.size(1)\n</code></pre>\n<pre><code># (bs, 7, 7, 512)\n#encoder_dim = 512\nencoder_out.view(batch_size, -1, encoder_dim) #(bs, 7 * 7, encoder_dim) (-1 will flatten axis 1 and 2)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1249782,
          "author_name": "marketneutral",
          "author_url": "",
          "post_date": "03/23/2021 14:11:13",
          "content": "<blockquote>\n  <p><code>encoder_out.view(batch_size, -1, encoder_dim) #(bs, 7 * 7, encoder_dim) (-1 will flatten axis 1 and 2)</code></p>\n</blockquote>\n<p>That's it!! Thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1252095,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "03/25/2021 12:22:36",
          "content": "<p>I found <a href=\"https://github.com/arogozhnikov/einops\" target=\"_blank\">einops</a> very helpful for this type of code. <br>\n<code>x = rearrange(x, 'b c h w -&gt; b (h w) c')</code>  is very readable.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1252164,
          "author_name": "marketneutral",
          "author_url": "",
          "post_date": "03/25/2021 13:21:52",
          "content": "<p><code>einops</code> looks amazing! Very elegant and clear code. Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1253027,
      "author_name": "olusesiadebisi",
      "author_url": "",
      "post_date": "03/26/2021 10:03:48",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> for sharing. This is very helpful. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1260702,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/02/2021 10:58:05",
      "content": "<p>some good explanation here: <a href=\"https://medium.com/analytics-vidhya/image-captioning-with-attention-part-1-e8a5f783f6d3\" target=\"_blank\">https://medium.com/analytics-vidhya/image-captioning-with-attention-part-1-e8a5f783f6d3</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1271143,
      "author_name": "adityakadiwal",
      "author_url": "",
      "post_date": "04/12/2021 10:39:01",
      "content": "<p>Wow! This is going to be really helpful for someone like me who has just started to learn about RNNs. Thank you for sharing. 🙏🏼</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1274464,
      "author_name": "zavodrobotov",
      "author_url": "",
      "post_date": "04/15/2021 10:15:44",
      "content": "<p>Good work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2747540,
      "author_name": "itskavyagupta",
      "author_url": "",
      "post_date": "04/12/2024 00:10:19",
      "content": "<p>Super useful!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3196935,
      "author_name": "meesamr",
      "author_url": "",
      "post_date": "05/07/2025 16:28:26",
      "content": "<p>Great Work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3306415,
      "author_name": "muhammadehsan02",
      "author_url": "",
      "post_date": "10/24/2025 13:59:44",
      "content": "<p>Nicely explained </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1237631": "I prepared Resnet + LSTM with attention starter code.\n\n1. [preprocess](https://www.kaggle.com/yasufuminakama/inchi-preprocess-2)\n2. [training](https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter)\n3. [inference](https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-inference)\n\nThe result is CV: 18.7, LB: 20.3\n\nIn inference notebook, I used 2 epochs trained weight. There is much room to improve, for example more epochs, augmentation, larger models, larger size...\n\nHope this helps, happy kaggling!",
    "1237740": "Wouldn't be better to just explain the workflow in the discussion without publishing the actual code?\nI can already see 2 scores exactly at 21.9 like yours at 5th and 6th positions...",
    "1237787": "I understand what you say, but it's easier to join the competition for beginners if there is starter code.\nWhen I was beginner, I learned a lot from public notebooks, that's why I publish starter codes for other competitions too. \nYou don't need to be so nervous about current LB position, it's just beginning of the competition. There are a lot of things to improve score.",
    "1238267": "I agree with you, I learned a lot from public notebooks in my previous competition (MoA) but the problem with beginners is that they tend to deceive themselves with copying and submitting the notebook without trying to understand the code (even at the beginning of the competition) and that's bad but they don't realize it. Also, they are people who try their best to learn and improve their scores and that is really demotivating for people who try their best.\n\nI personally never copy past code that I don't understand, and my message to @jacoporepossi \nis to not give up and like @yasufuminakama said it's the beginning of the competition and there is room for improvement and thank you for sharing your notebook with us, very helpful.\n\nBy the way, How can we retrain models when we surpass the 9 hour limit of the Kaggle kernel?",
    "1238286": "You can reload saved model states, optimizer states, scheduler states, and restart training.",
    "1238321": "Don't get me wrong, I'm neither angry nor frustrated by the fact that @yasufuminakama shared his code!! I'm sure it will help a lot of people and that's the beauty of being part of this wonderful community :)\nI'm trying to put myself in someone else's shoes and imagine his frustration at being passed by people copying a notebook without doing anything. \nObviously it's part of the game here in Kaggle, as you said it's just the beginning and there's room for improvement, it's not the end of the world. However if the aim was simply to provide beginners with ideas on how to approach the problem, probably a win-win approach could have been simply to discuss here about the implementation, just as they did [here](https://www.kaggle.com/c/bms-molecular-translation/discussion/224257).\n\nBtw, there are multiple competitions available for beginners to get started, as well as plenty of open notebooks from past competitions to learn",
    "1238783": "I always enjoy your Attention starter discussions! They are very fruitful and helpful for all of us.\nWouldn't be smarter to simply clarify the work process in the conversation without distributing the real code which you do, as many can have plagiarism.\nBeing a starter, I learn from you rather than copying from your, but others don't do so🤷‍♂️ @yasufuminakama",
    "1239430": "Enjoying reading u, thanks a lot, your discussion and your notebooks are useful",
    "1239483": "Thank you for sharing. This is a fascinating competition. Cross domain (i.e., image to text) problems get to the heart of what true human-like understanding is. I was aware of the Show, Attend, and Tell paper, but prior to you sharing your work it was too dense for me to work with directly. \n\nCan I ask, how long did you train the model for to get to CV 18.7?",
    "1239700": "As I wrote, it's 2 epochs, and it took about 7.5h per epoch using kaggle notebooks as you can see training notebook.",
    "1239747": "\" being passed by people copying a notebook without doing anything.\"\n\nthat is part of kaggle and part of life.\n\none can then speed his work by learning from the note book and improve on it. then you can overtake yourself.\n\n\"probably a win-win approach\"\nunfortunately, the devils are in the details and not in the methods. there is a huge difference with and without the code.\n\n---\n\nmy suggestion is:\n- you should capitalize on code posted. e.g. run the code with different parameters and observed how accuracy changes, etc. compare the code with yours and see which part is better (or worse)\n- how to be better than those who just copy the code and do nothing? understand the code and improve on it.\n\n---\none may one to think about, which is more important:\n1. i got a medal, but I am inside my box (i.e. i am using methods developed from myself)\n2. i didn't get any medal, but l learned lots of useful techniques from others (methods that are out of my imagination ... i have stepped out of my box)",
    "1239763": "yasufuminakama \n\nthis should work. you can do it again!\n\n![](https://i.ibb.co/L1S3cW0/Selection-352.png)",
    "1239785": "hengck23 \nGlad to see you in this competition!\nYes, I was thinking this kind of approach, thanks for kindly visualization :)",
    "1240575": "When RANZCR CLiP is over today and the usual suspects arrive here you will see this baseline start to fall on the leaderboard...",
    "1243100": "I think @yasufuminakama has proven that he doesn't need to be worried about people forking his code. That being said, @yasufuminakama thank you for the clearly explained code. Very helpful in making the \"Show, Attend and Tell\" paper understandable!",
    "1244269": "hengck23 Thanks for the suggestion.\n\n@yasufuminakama Can I assume from your jump to the top of the leaderboard, that the teacher-student approach worked very well?!",
    "1244697": "I haven't tried it yet, I want to know the result if someone tried it :)",
    "1245187": "How might one go about making the training code parallel across multiple GPUs? I've tried the following idea in the \"training\" notebook:\n\n```\n   :\n   :\n    encoder = Encoder(CFG.model_name, pretrained=True)\n    if torch.cuda.device_count() > 1:\n        print(\"Let's use\", torch.cuda.device_count(), \"GPUs!\")\n        encoder = nn.DataParallel(encoder)\n    encoder.to(device)\n   :\n   :\n```\n\nand the same with the `decoder`.  When I run training I can see the memory use go up evenly across all GPUs so something is working. However, I get the following error:\n\n```\nRuntimeError: Gather got an input of invalid size: got [64, 158, 193], but expected [64, 154, 193] (gather at /opt/conda/conda-bld/pytorch_1579022034529/work/torch/csrc/cuda/comm.cpp:231)\n```\n\nThe batch size is an even multiple of the number of GPUs so that's fine.",
    "1245263": "I haven't tried multiple GPUs training so I'm not sure but it seems label sequence length is not same.",
    "1245300": "I found some clues [here](https://pytorch.org/docs/stable/notes/faq.html#my-recurrent-network-doesn-t-work-with-data-parallelism)\n\n> There is a subtlety in using the pack sequence -> recurrent network -> unpack sequence pattern in a Module with DataParallel or data_parallel(). Input to each the forward() on each device will only be part of the entire input. Because the unpack operation torch.nn.utils.rnn.pad_packed_sequence() by default only pads up to the longest input it sees, i.e., the longest on that particular device, size mismatches will happen when results are gathered together. Therefore, you can instead take advantage of the total_length argument of pad_packed_sequence() to make sure that the forward() calls return sequences of same length.",
    "1246228": "How do you think about the different between CV and LB ?\nThe estimation of image orientation (w>h or w<h) is not perfect, is it ?",
    "1246476": "the actual problem is with this line \n`decode_lengths = (caption_lengths - 1).tolist()`\nin multi-gpu it splits into 4 lists so they get different lengths and on gathering you receive an error. \nFast and dirty way is to run in parallel encoder, and decoder - not, in other words decoder runs on one gpu.\n`the_number_of_gpu = torch.cuda.device_count()\n    if the_number_of_gpu > 1:\n        encoder = nn.DataParallel(encoder)\n        # decoder = nn.DataParallel(decoder)  <--- commented\n`\nin this case gpus won't be used optimally, unfortunately",
    "1246540": "I haven't used `nn.DataParallel`, but `nn.parallel.DistributedDataParallel` doesn't throw any error.",
    "1246541": "Not the author=) But I can add few cents... They gap between `CV` and `LB` is greatly reduced if you add `90` degree rotations augmentations during training. (or perhaps any rotation)",
    "1246566": "IMHO, I don't mind gap if there is correlation between CV and LB.",
    "1246835": "Yes, torch states that DistributedDataParallel (DDP) works faster and better than DataParallel like it is stated [here](https://pytorch.org/tutorials/intermediate/ddp_tutorial.html). \n@drhabib  if you would create a kernel with DDP based on @yasufuminakama code, lots of kagglers would be happy as DDP requires some additional implementations.  :)",
    "1248262": "Thanks @drhabib for the link to `nn.parallel.DistributedDataParallel`. That looks very nice. Will try... \n\n> @drhabib if you would create a kernel with DDP based on @yasufuminakama code, lots of kagglers would be happy as DDP requires some additional implementations. :)\n\nI would hit the upvote so hard that I think it would break my mouse... :-)",
    "1249711": "I am struggling to get intuition about the shape of the feature tensor coming out of the encoder. This is resnet34 and we have removed the global pooling and fully connected layer and just pass down the last layer. Inspected as\n\n```\nfrom torchsummary import summary\nsummary(encoder, (3, 224, 224)\n\n----------------------------------------------------------------\n        Layer (type)               Output Shape         Param #\n================================================================\n            Conv2d-1         [-1, 64, 112, 112]           9,408\n       BatchNorm2d-2         [-1, 64, 112, 112]             128\n   :\n   :\n   :\nConv2d-119            [-1, 512, 7, 7]       2,359,296\n     BatchNorm2d-120            [-1, 512, 7, 7]           1,024\n            ReLU-121            [-1, 512, 7, 7]               0\n      BasicBlock-122            [-1, 512, 7, 7]               0\n        Identity-123            [-1, 512, 7, 7]               0\n        Identity-124            [-1, 512, 7, 7]               0\n          ResNet-125            [-1, 512, 7, 7]               0\n```\n\nSo the shape of the feature tensor is `[batch size sample index, index of feature map, x, y]`.\n\nThe `forward` method of the Encoder reshapes this feature tensor:\n\n```\nclass Encoder(nn.Module):\n    def __init__(self, model_name='resnet18', pretrained=False):\n        super().__init__()\n        self.cnn = timm.create_model(model_name, pretrained=pretrained)\n        self.n_features = self.cnn.fc.in_features\n        self.cnn.global_pool = nn.Identity()\n        self.cnn.fc = nn.Identity()\n\n    def forward(self, x):\n        bs = x.size(0)\n        features = self.cnn(x)\n        features = features.permute(0, 2, 3, 1)\n        return features\n```\n\n`features.permute(0, 2, 3, 1)` makes the shape of the tensor `[batch size sample index, x, y, index of feature map]`\n\nI don't understand why we are doing this. I would be grateful if anyone could share some light on this! Thanks!",
    "1249735": "In general you can debug `Pytorch` code by just putting print statements in the `nn.model class`\n\n```\nclass Encoder(nn.Module):\n    def __init__(self, model_name='resnet18', pretrained=False):\n        super().__init__()\n        self.cnn = timm.create_model(model_name, pretrained=pretrained)\n        self.n_features = self.cnn.fc.in_features\n        self.cnn.global_pool = nn.Identity()\n        self.cnn.fc = nn.Identity()\n\n    def forward(self, x):\n        bs = x.size(0)\n        features = self.cnn(x)\n        print (features.shape)\n        print (f'Shape before permutation {features.shape}')\n        features = features.permute(0, 2, 3, 1)\n        print (f'Shape after permutation {features.shape}')\n        return features\n```\n\n```\n#put model in eval mode\nmodel = Encoder().eval()\n\nbs = 4\nc  = 3\nh  = 224\nw  = 224\nimage = torch.rand(bs, c, h, w)\n\n#getting features and printing statement \n#no_grad just tells that we don't want to calculate gradient \nwith torch.no_grad():\n    features = model(image)\n```\n\noutput should look like:\n```\nShape before permutation torch.Size([4, 512, 7, 7])\nShape after permutation torch.Size([4, 7, 7, 512])\n```\n\n`torch.permute` just helps to switch features axis... Later we will reshape this features more in decoder to `(bs, 7 *7, 512)` so we can apply attention and initiate `LSTM` cell\n\nNot sure I answered your question =)",
    "1249751": "It's very detailed [here](https://github.com/sgrvinod/a-PyTorch-Tutorial-to-Image-Captioning) ... from the author",
    "1249758": "ahhh perfect! thanks for the link =)",
    "1249766": "Thanks! I concur on the output shapes. That's what I intended to communicate. I understand what permute is doing technically. What I am asking is **why** we have to do this.\n\nPerhaps this is the answer\n\n> Later we will reshape this features more in decoder to (bs, 7 *7, 512) so we can apply attention and initiate LSTM cell\n\nCan you point in the code where this is happening? I could understand it if the 7x7 feature map was flattened to a vector, but still not getting this. :cry:",
    "1249770": "> It's very detailed here … from the author\n\nThat link just makes things **more** confusing to me, because in his README he is **not** \npermuting the feature tensor. For example, see\n\n![](https://github.com/sgrvinod/a-PyTorch-Tutorial-to-Image-Captioning/raw/master/img/decoder_att.png)\n\n\nSo here, we maintain the shape of the tensor in the original CNN.\n\nEDIT:\n\nHaving read more closely the README.md in this repo, he does talk about reshaping the tensor\nhttps://github.com/sgrvinod/a-PyTorch-Tutorial-to-Image-Captioning#attention-1\n\n> The output of the Encoder is received here and flattened to dimensions N, 14 * 14, 2048. This is just convenient and prevents having to reshape the tensor multiple times.",
    "1249774": "https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter\n\n```class DecoderWithAttention(nn.Module):```\n\n```\n        encoder_out = encoder_out.view(batch_size, -1, encoder_dim)  # (batch_size, num_pixels, encoder_dim)\n        num_pixels = encoder_out.size(1)\n```\n\n```\n# (bs, 7, 7, 512)\n#encoder_dim = 512\nencoder_out.view(batch_size, -1, encoder_dim) #(bs, 7 * 7, encoder_dim) (-1 will flatten axis 1 and 2)\n```",
    "1249782": "> `encoder_out.view(batch_size, -1, encoder_dim) #(bs, 7 * 7, encoder_dim) (-1 will flatten axis 1 and 2)`\n\nThat's it!! Thank you!",
    "1251592": "I have tried it. It seems that the simple annotations in the `rdkit` generated images are not enough for the teacher model to pick up on: stage 1 performance is only marginally better than training with the original images. Perhaps I need to use a stronger encoder to see more significant results. I will share my notebooks this weekend when I have some GPU and hopefully a more experienced Kaggler can find a way of improving it. (I think I just need to generate more annotations via `rdkit` to make it work i.e. highlight bonds / atoms).",
    "1251606": "Thanks for sharing your result!",
    "1251623": "I think this level of annotation would be better:\n\n![](https://www.rdkit.org/docs/_images/atom_highlights_3.png)\n\nWill try it next and report results. \n\nEdit: \nSome useful code can be found [here](https://github.com/rdkit/rdkit/blob/master/Docs/Book/data/test_multi_colours.py). We can use [SMART queries](https://docs.chemaxon.com/display/docs/smarts.md) to highlight arbitrary substructures of the chemicals. I don't know how to best exploit this, as I know very little about chemistry - I'm a physicist :D\n\nFor example, we can pass SMART queries like: \n`[\"B\",\"Br\",\"C\",\"Cl\",\"F\",\"I\",\"N\",\"O\",\"P\",\"S\"]`\nto draw circles around the atoms and we can pass queries like:\n`[\"[r{3-}]\", \"*~*~*\"]` \nto highlight rings of radius `3` or greater and `3` linear atoms respectively.",
    "1252076": "Why you  don't care if there is correlation between CV and LB? Can you explain it?",
    "1252095": "I found [einops](https://github.com/arogozhnikov/einops) very helpful for this type of code. \n`x = rearrange(x, 'b c h w -> b (h w) c')`  is very readable.",
    "1252134": "I think there must be gap between Single-Fold CV and LB for most people, but I think if we perform N-fold ensemble, I think CV is much closer to LB. I haven't tried it, Just my intuition :)",
    "1252164": "`einops` looks amazing! Very elegant and clear code. Thanks.",
    "1253027": "Thank you @yasufuminakama for sharing. This is very helpful.",
    "1260702": "some good explanation here: https://medium.com/analytics-vidhya/image-captioning-with-attention-part-1-e8a5f783f6d3",
    "1271143": "Wow! This is going to be really helpful for someone like me who has just started to learn about RNNs. Thank you for sharing. 🙏🏼",
    "1274464": "Good work!",
    "1284992": "I'm joining this comp a month later and this score is way down the leaderboard. As op points out, this is what normally happens in these comps. The code has been super useful in getting me started on this comp",
    "2747540": "Super useful!",
    "3196935": "Great Work!",
    "3306415": "Nicely explained"
  },
  "source": "meta"
}