{
  "id": 244266,
  "title": "18th place solution",
  "url": "/competitions/bms-molecular-translation/writeups/nofreewill-18th-place-solution",
  "author_name": "",
  "post_date": "2021-06-06T15:54:38.633Z",
  "votes": 19,
  "comment_count": 9,
  "views": 0,
  "content": "<p><strong>Output:</strong><br>\nI used 4096 byte pair encodings created from plain InChI text with <a href=\"https://github.com/google/sentencepiece\" target=\"_blank\">sentencepiece</a>.</p>\n<p><strong>Model:</strong><br>\nCNN -&gt; Encoder -&gt; Decoder<br>\nCNN is a custom model, built up from 5 layers of these:</p>\n<pre><code>def forward(self, x_in):\n    x = self.bnl1(x_in)\n    x = self.bnl2(x) + self.bnl4(self.bnl3(x))\n    x = torch.cat((x, self.pool(x_in)), dim=1) if self.r else x\n    return x\n</code></pre>\n<p>where bnl[1234] is convolution, followed by batch norm and relu.<br>\nEncoder and Decoder are Transformer-based.</p>\n<p><strong>Sampling molecules:</strong><br>\nI removed samples that had no carbon or had no /b layer after /i (due to too few examples).<br>\nI calculated a weight for each molecule based on it's:</p>\n<ul>\n<li>complexity: <code>x.count('-') + x.count('-') * x.count('(') + x.count('(')</code></li>\n<li>atom count: number of all the atoms it's made up of</li>\n<li>atom rarity: I calculate how many times (M) an atom is in all samples (N). For each molecule I sum up one of {1/M, 1/(N-M)} values for each atom, depending on if it is in that molecule or not.</li>\n<li>layer rarity: the same as for atoms</li>\n</ul>\n<p>I take these values to a power and add them together with some reasonable coefficient -&gt; I have a final sampling weight for each molecule.</p>\n<p><strong>Training:</strong><br>\nCross entropy loss; Adam; One cycle LR schedule; Mixed Precision + FP32 fine-tuning.<br>\nScheduled attention dropout and drophead. I also used <a href=\"https://github.com/google/sentencepiece\" target=\"_blank\">sentencepiece</a>'s BPE-dropout.<br>\nI've trained 5 models on different aspect ratios. Then I also trained each of them on yet a different aspect ratio for a few more epochs.<br>\nTheir individual LD on my validation are as follows:<br>\n<code>img_sizes = [(160,224), (192,192),</code><br>\n<code>(256,544), (224,640), (320,480),</code><br>\n<code>(288,512), (320,480), (480,320),</code><br>\n<code>(288,512), (256,576), (320,448),</code><br>\n<code>(384,384), (352,416), (192,768),]</code><br>\n<code>raw_lds = [1.5233, 1.5275,</code><br>\n<code>1.0955, 1.1850, 1.0699,</code><br>\n<code>1.1706, 1.0629, 1.3790,</code><br>\n<code>1.0664, 1.1234, 1.0832,</code><br>\n<code>1.4073, 1.2095, 2.1064,]</code><br>\n<code>norm_lds = [1.3661, 1.3879,</code><br>\n<code>0.9438, 1.0244, 0.9248,</code><br>\n<code>1.0205, 0.9280, 1.2254,</code><br>\n<code>0.9415, 0.9923, 0.9400,</code><br>\n<code>1.2318, 1.0646, 1.9285,]</code></p>\n<p><strong>Prediction:</strong><br>\nI've predicted all molecules with all models separately and normalized all of the predictions.<br>\nThen I searched those that have the same normalized form.<br>\nI've predicted the rest of them (~390k) with ensembled beam search of size 16 and searched those that are not valid predictions. These (~16K) I predicted with beam size of 64.<br>\nTaking all these predictions I only had ~5k molecules that I had no valid prediction for.<br>\nI've predicted the 390k molecules with a bunch of other settings (with different beam size and ensemble weights), and tried to sample the valid ones with different strategies, but I just couldn't break the 0.70 barrier.</p>\n<p>Very interestingly, I had a lot of submissions that had quite an LD between them but all of them were 0.70 on LB. I even had two submissions with 0.11 LD between them, and they're still 0.70 and 0.71 on LB.<br>\nI couldn't figure out if there could be any potential in this \"phenomenon\" with which I could make better predictions.</p>\n<p><strong>Didn't work:</strong></p>\n<ul>\n<li>Label smoothing</li>\n<li>Additional targets on encoder output</li>\n<li>Encoder level mixup (decoder works on the two targets individually from the input to the output)</li>\n<li>Sorting molecules by InChI BPE lenght for decoder speedup. I forgot to randomize, lost a whole week trying to figure out what went wrong. It felt incredible, as this happened right after I bought a 3090 … :DDD</li>\n</ul>\n<p><strong>Thank you all:</strong><br>\nThis was a very interesting problem to work on, and was a huge pleasure to compete with all of you!<br>\nEven though I didn't end up in the gold zone, I don't regret spending that much money on a 3090, as so many worked so hard on this competition, and I'm more than happy that I could reach 18th place among so many so talented people!</p>\n<p><strong>Special thanks to:</strong><br>\n<a href=\"https://www.kaggle.com/stassl\" target=\"_blank\">@stassl</a> and <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> for the idea of normalization, and<br>\n<a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for the importance of validation</p>\n<p>Thank you!</p>",
  "messages": [
    {
      "id": "1337721",
      "postDate": "06/05/2021 20:27:55",
      "content": "<p><strong>Output:</strong><br>\nI used 4096 byte pair encodings created from plain InChI text with <a href=\"https://github.com/google/sentencepiece\" target=\"_blank\">sentencepiece</a>.</p>\n<p><strong>Model:</strong><br>\nCNN -&gt; Encoder -&gt; Decoder<br>\nCNN is a custom model, built up from 5 layers of these:</p>\n<pre><code>def forward(self, x_in):\n    x = self.bnl1(x_in)\n    x = self.bnl2(x) + self.bnl4(self.bnl3(x))\n    x = torch.cat((x, self.pool(x_in)), dim=1) if self.r else x\n    return x\n</code></pre>\n<p>where bnl[1234] is convolution, followed by batch norm and relu.<br>\nEncoder and Decoder are Transformer-based.</p>\n<p><strong>Sampling molecules:</strong><br>\nI removed samples that had no carbon or had no /b layer after /i (due to too few examples).<br>\nI calculated a weight for each molecule based on it's:</p>\n<ul>\n<li>complexity: <code>x.count('-') + x.count('-') * x.count('(') + x.count('(')</code></li>\n<li>atom count: number of all the atoms it's made up of</li>\n<li>atom rarity: I calculate how many times (M) an atom is in all samples (N). For each molecule I sum up one of {1/M, 1/(N-M)} values for each atom, depending on if it is in that molecule or not.</li>\n<li>layer rarity: the same as for atoms</li>\n</ul>\n<p>I take these values to a power and add them together with some reasonable coefficient -&gt; I have a final sampling weight for each molecule.</p>\n<p><strong>Training:</strong><br>\nCross entropy loss; Adam; One cycle LR schedule; Mixed Precision + FP32 fine-tuning.<br>\nScheduled attention dropout and drophead. I also used <a href=\"https://github.com/google/sentencepiece\" target=\"_blank\">sentencepiece</a>'s BPE-dropout.<br>\nI've trained 5 models on different aspect ratios. Then I also trained each of them on yet a different aspect ratio for a few more epochs.<br>\nTheir individual LD on my validation are as follows:<br>\n<code>img_sizes = [(160,224), (192,192),</code><br>\n<code>(256,544), (224,640), (320,480),</code><br>\n<code>(288,512), (320,480), (480,320),</code><br>\n<code>(288,512), (256,576), (320,448),</code><br>\n<code>(384,384), (352,416), (192,768),]</code><br>\n<code>raw_lds = [1.5233, 1.5275,</code><br>\n<code>1.0955, 1.1850, 1.0699,</code><br>\n<code>1.1706, 1.0629, 1.3790,</code><br>\n<code>1.0664, 1.1234, 1.0832,</code><br>\n<code>1.4073, 1.2095, 2.1064,]</code><br>\n<code>norm_lds = [1.3661, 1.3879,</code><br>\n<code>0.9438, 1.0244, 0.9248,</code><br>\n<code>1.0205, 0.9280, 1.2254,</code><br>\n<code>0.9415, 0.9923, 0.9400,</code><br>\n<code>1.2318, 1.0646, 1.9285,]</code></p>\n<p><strong>Prediction:</strong><br>\nI've predicted all molecules with all models separately and normalized all of the predictions.<br>\nThen I searched those that have the same normalized form.<br>\nI've predicted the rest of them (~390k) with ensembled beam search of size 16 and searched those that are not valid predictions. These (~16K) I predicted with beam size of 64.<br>\nTaking all these predictions I only had ~5k molecules that I had no valid prediction for.<br>\nI've predicted the 390k molecules with a bunch of other settings (with different beam size and ensemble weights), and tried to sample the valid ones with different strategies, but I just couldn't break the 0.70 barrier.</p>\n<p>Very interestingly, I had a lot of submissions that had quite an LD between them but all of them were 0.70 on LB. I even had two submissions with 0.11 LD between them, and they're still 0.70 and 0.71 on LB.<br>\nI couldn't figure out if there could be any potential in this \"phenomenon\" with which I could make better predictions.</p>\n<p><strong>Didn't work:</strong></p>\n<ul>\n<li>Label smoothing</li>\n<li>Additional targets on encoder output</li>\n<li>Encoder level mixup (decoder works on the two targets individually from the input to the output)</li>\n<li>Sorting molecules by InChI BPE lenght for decoder speedup. I forgot to randomize, lost a whole week trying to figure out what went wrong. It felt incredible, as this happened right after I bought a 3090 … :DDD</li>\n</ul>\n<p><strong>Thank you all:</strong><br>\nThis was a very interesting problem to work on, and was a huge pleasure to compete with all of you!<br>\nEven though I didn't end up in the gold zone, I don't regret spending that much money on a 3090, as so many worked so hard on this competition, and I'm more than happy that I could reach 18th place among so many so talented people!</p>\n<p><strong>Special thanks to:</strong><br>\n<a href=\"https://www.kaggle.com/stassl\" target=\"_blank\">@stassl</a> and <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> for the idea of normalization, and<br>\n<a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for the importance of validation</p>\n<p>Thank you!</p>",
      "rawMarkdown": "**Output:**\nI used 4096 byte pair encodings created from plain InChI text with [sentencepiece](https://github.com/google/sentencepiece).\n\n**Model:**\nCNN -> Encoder -> Decoder\nCNN is a custom model, built up from 5 layers of these:\n```\ndef forward(self, x_in):\n    x = self.bnl1(x_in)\n    x = self.bnl2(x) + self.bnl4(self.bnl3(x))\n    x = torch.cat((x, self.pool(x_in)), dim=1) if self.r else x\n    return x\n```\nwhere bnl[1234] is convolution, followed by batch norm and relu.\nEncoder and Decoder are Transformer-based.\n\n**Sampling molecules:**\nI removed samples that had no carbon or had no /b layer after /i (due to too few examples).\nI calculated a weight for each molecule based on it's:\n - complexity: `x.count('-') + x.count('-') * x.count('(') + x.count('(')`\n - atom count: number of all the atoms it's made up of\n - atom rarity: I calculate how many times (M) an atom is in all samples (N). For each molecule I sum up one of {1/M, 1/(N-M)} values for each atom, depending on if it is in that molecule or not.\n - layer rarity: the same as for atoms\n\nI take these values to a power and add them together with some reasonable coefficient -> I have a final sampling weight for each molecule.\n\n**Training:**\nCross entropy loss; Adam; One cycle LR schedule; Mixed Precision + FP32 fine-tuning.\nScheduled attention dropout and drophead. I also used [sentencepiece](https://github.com/google/sentencepiece)'s BPE-dropout.\nI've trained 5 models on different aspect ratios. Then I also trained each of them on yet a different aspect ratio for a few more epochs.\nTheir individual LD on my validation are as follows:\n`img_sizes = [(160,224), (192,192),`\n`(256,544), (224,640), (320,480),`\n`(288,512), (320,480), (480,320),`\n`(288,512), (256,576), (320,448),`\n`(384,384), (352,416), (192,768),]`\n`raw_lds = [1.5233, 1.5275,`\n`1.0955, 1.1850, 1.0699,`\n`1.1706, 1.0629, 1.3790,`\n`1.0664, 1.1234, 1.0832,`\n`1.4073, 1.2095, 2.1064,]`\n`norm_lds = [1.3661, 1.3879,`\n`0.9438, 1.0244, 0.9248,`\n`1.0205, 0.9280, 1.2254,`\n`0.9415, 0.9923, 0.9400,`\n`1.2318, 1.0646, 1.9285,]`\n\n**Prediction:**\nI've predicted all molecules with all models separately and normalized all of the predictions.\nThen I searched those that have the same normalized form.\nI've predicted the rest of them (~390k) with ensembled beam search of size 16 and searched those that are not valid predictions. These (~16K) I predicted with beam size of 64.\nTaking all these predictions I only had ~5k molecules that I had no valid prediction for.\nI've predicted the 390k molecules with a bunch of other settings (with different beam size and ensemble weights), and tried to sample the valid ones with different strategies, but I just couldn't break the 0.70 barrier.\n\nVery interestingly, I had a lot of submissions that had quite an LD between them but all of them were 0.70 on LB. I even had two submissions with 0.11 LD between them, and they're still 0.70 and 0.71 on LB.\nI couldn't figure out if there could be any potential in this \"phenomenon\" with which I could make better predictions.\n\n**Didn't work:**\n - Label smoothing\n - Additional targets on encoder output\n - Encoder level mixup (decoder works on the two targets individually from the input to the output)\n - Sorting molecules by InChI BPE lenght for decoder speedup. I forgot to randomize, lost a whole week trying to figure out what went wrong. It felt incredible, as this happened right after I bought a 3090 ... :DDD\n\n**Thank you all:**\nThis was a very interesting problem to work on, and was a huge pleasure to compete with all of you!\nEven though I didn't end up in the gold zone, I don't regret spending that much money on a 3090, as so many worked so hard on this competition, and I'm more than happy that I could reach 18th place among so many so talented people!\n\n**Special thanks to:**\n@stassl and @yasufuminakama for the idea of normalization, and\n@hengck23 for the importance of validation\n\nThank you!",
      "votes": null
    },
    {
      "id": "1337724",
      "postDate": "06/05/2021 20:36:24",
      "content": "<blockquote>\n  <p>Very interestingly, I had a lot of submissions that had quite an LD between them but all of them were 0.70 on LB. I even had two submissions with 0.11 LD between them, and they're still 0.70 and 0.71 on LB.<br>\n  I couldn't figure out if there could be any potential in this \"phenomenon\" with which I could make better predictions.</p>\n</blockquote>\n<p>Could it be due to the fact that only 25% used for LB while 0.11 LD computed for compete submission?</p>",
      "rawMarkdown": "> Very interestingly, I had a lot of submissions that had quite an LD between them but all of them were 0.70 on LB. I even had two submissions with 0.11 LD between them, and they're still 0.70 and 0.71 on LB.\n> I couldn't figure out if there could be any potential in this \"phenomenon\" with which I could make better predictions.\n\nCould it be due to the fact that only 25% used for LB while 0.11 LD computed for compete submission?",
      "votes": null
    },
    {
      "id": "1337728",
      "postDate": "06/05/2021 20:39:11",
      "content": "<p>0.11 LD is on the whole submissions (all 1.6M molecules).<br>\n~0.43 was the LD on the 390k molecules.</p>\n<p>That 0.11 difference means that the 0.71 LB submission is better in 0.05 characters but also is worse in 0.06 other characters than the 0.70 LB submission.</p>",
      "rawMarkdown": "0.11 LD is on the whole submissions (all 1.6M molecules).\n~0.43 was the LD on the 390k molecules.\n\nThat 0.11 difference means that the 0.71 LB submission is better in 0.05 characters but also is worse in 0.06 other characters than the 0.70 LB submission.",
      "votes": null
    },
    {
      "id": "1337729",
      "postDate": "06/05/2021 20:41:16",
      "content": "<p>You joined kaggle 4 years ago, yet this is your very first comment? How? :D</p>",
      "rawMarkdown": "You joined kaggle 4 years ago, yet this is your very first comment? How? :D",
      "votes": null
    },
    {
      "id": "1337738",
      "postDate": "06/05/2021 20:56:52",
      "content": "<p>Well, I was about to join in 2011 but delayed. This particular competition is interesting to me.</p>",
      "rawMarkdown": "Well, I was about to join in 2011 but delayed. This particular competition is interesting to me.",
      "votes": null
    },
    {
      "id": "1337740",
      "postDate": "06/05/2021 21:00:10",
      "content": "<p>Indeed it was enjoyable!! :)</p>",
      "rawMarkdown": "Indeed it was enjoyable!! :)",
      "votes": null
    },
    {
      "id": "1358578",
      "postDate": "06/20/2021 15:31:53",
      "content": "<p>Congrats! Glad you got that boost in the last few days.<br>\nQuestion: why do you do FP32 fine-tuning? I read somewhere that mixed precision (on PyTorch) is not supposed to have any negative impact on the outcome.</p>",
      "rawMarkdown": "Congrats! Glad you got that boost in the last few days.\nQuestion: why do you do FP32 fine-tuning? I read somewhere that mixed precision (on PyTorch) is not supposed to have any negative impact on the outcome.",
      "votes": null
    },
    {
      "id": "1358941",
      "postDate": "06/20/2021 23:09:39",
      "content": "<p>Thank you! :))</p>\n<p>Did you use <code>cascading and agreement</code> even before my mad discussion about the importance of validation?</p>\n<p>I read somewhere in the discussions that FP32 fine-tuning helped for someone. I didn't check it, but would make sense if it helps a little, so I just went with it.</p>",
      "rawMarkdown": "Thank you! :))\n\nDid you use `cascading and agreement` even before my mad discussion about the importance of validation?\n\nI read somewhere in the discussions that FP32 fine-tuning helped for someone. I didn't check it, but would make sense if it helps a little, so I just went with it.",
      "votes": null
    },
    {
      "id": "1364817",
      "postDate": "06/25/2021 08:26:35",
      "content": "<p>Lol yes I did use that before your post.</p>\n<p>Ah okay thanks, worth knowing that data point.</p>",
      "rawMarkdown": "Lol yes I did use that before your post.\n\nAh okay thanks, worth knowing that data point.",
      "votes": null
    },
    {
      "id": "3513844",
      "postDate": "08/17/2026 18:58:20",
      "content": "<p>I get it now, I am also still waiting for a really interesting competition like this ever since.</p>",
      "rawMarkdown": "I get it now, I am also still waiting for a really interesting competition like this ever since.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1337724,
      "author_name": "rednikotin",
      "author_url": "",
      "post_date": "06/05/2021 20:36:24",
      "content": "<blockquote>\n  <p>Very interestingly, I had a lot of submissions that had quite an LD between them but all of them were 0.70 on LB. I even had two submissions with 0.11 LD between them, and they're still 0.70 and 0.71 on LB.<br>\n  I couldn't figure out if there could be any potential in this \"phenomenon\" with which I could make better predictions.</p>\n</blockquote>\n<p>Could it be due to the fact that only 25% used for LB while 0.11 LD computed for compete submission?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1337728,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "06/05/2021 20:39:11",
          "content": "<p>0.11 LD is on the whole submissions (all 1.6M molecules).<br>\n~0.43 was the LD on the 390k molecules.</p>\n<p>That 0.11 difference means that the 0.71 LB submission is better in 0.05 characters but also is worse in 0.06 other characters than the 0.70 LB submission.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1337729,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "06/05/2021 20:41:16",
          "content": "<p>You joined kaggle 4 years ago, yet this is your very first comment? How? :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1337738,
          "author_name": "rednikotin",
          "author_url": "",
          "post_date": "06/05/2021 20:56:52",
          "content": "<p>Well, I was about to join in 2011 but delayed. This particular competition is interesting to me.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3513844,
              "author_name": "nofreewill",
              "author_url": "",
              "post_date": "08/17/2026 18:58:20",
              "content": "<p>I get it now, I am also still waiting for a really interesting competition like this ever since.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1337740,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "06/05/2021 21:00:10",
          "content": "<p>Indeed it was enjoyable!! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1358578,
      "author_name": "alexandersoare",
      "author_url": "",
      "post_date": "06/20/2021 15:31:53",
      "content": "<p>Congrats! Glad you got that boost in the last few days.<br>\nQuestion: why do you do FP32 fine-tuning? I read somewhere that mixed precision (on PyTorch) is not supposed to have any negative impact on the outcome.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1358941,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "06/20/2021 23:09:39",
          "content": "<p>Thank you! :))</p>\n<p>Did you use <code>cascading and agreement</code> even before my mad discussion about the importance of validation?</p>\n<p>I read somewhere in the discussions that FP32 fine-tuning helped for someone. I didn't check it, but would make sense if it helps a little, so I just went with it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1364817,
          "author_name": "alexandersoare",
          "author_url": "",
          "post_date": "06/25/2021 08:26:35",
          "content": "<p>Lol yes I did use that before your post.</p>\n<p>Ah okay thanks, worth knowing that data point.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1337721": "**Output:**\nI used 4096 byte pair encodings created from plain InChI text with [sentencepiece](https://github.com/google/sentencepiece).\n\n**Model:**\nCNN -> Encoder -> Decoder\nCNN is a custom model, built up from 5 layers of these:\n```\ndef forward(self, x_in):\n    x = self.bnl1(x_in)\n    x = self.bnl2(x) + self.bnl4(self.bnl3(x))\n    x = torch.cat((x, self.pool(x_in)), dim=1) if self.r else x\n    return x\n```\nwhere bnl[1234] is convolution, followed by batch norm and relu.\nEncoder and Decoder are Transformer-based.\n\n**Sampling molecules:**\nI removed samples that had no carbon or had no /b layer after /i (due to too few examples).\nI calculated a weight for each molecule based on it's:\n - complexity: `x.count('-') + x.count('-') * x.count('(') + x.count('(')`\n - atom count: number of all the atoms it's made up of\n - atom rarity: I calculate how many times (M) an atom is in all samples (N). For each molecule I sum up one of {1/M, 1/(N-M)} values for each atom, depending on if it is in that molecule or not.\n - layer rarity: the same as for atoms\n\nI take these values to a power and add them together with some reasonable coefficient -> I have a final sampling weight for each molecule.\n\n**Training:**\nCross entropy loss; Adam; One cycle LR schedule; Mixed Precision + FP32 fine-tuning.\nScheduled attention dropout and drophead. I also used [sentencepiece](https://github.com/google/sentencepiece)'s BPE-dropout.\nI've trained 5 models on different aspect ratios. Then I also trained each of them on yet a different aspect ratio for a few more epochs.\nTheir individual LD on my validation are as follows:\n`img_sizes = [(160,224), (192,192),`\n`(256,544), (224,640), (320,480),`\n`(288,512), (320,480), (480,320),`\n`(288,512), (256,576), (320,448),`\n`(384,384), (352,416), (192,768),]`\n`raw_lds = [1.5233, 1.5275,`\n`1.0955, 1.1850, 1.0699,`\n`1.1706, 1.0629, 1.3790,`\n`1.0664, 1.1234, 1.0832,`\n`1.4073, 1.2095, 2.1064,]`\n`norm_lds = [1.3661, 1.3879,`\n`0.9438, 1.0244, 0.9248,`\n`1.0205, 0.9280, 1.2254,`\n`0.9415, 0.9923, 0.9400,`\n`1.2318, 1.0646, 1.9285,]`\n\n**Prediction:**\nI've predicted all molecules with all models separately and normalized all of the predictions.\nThen I searched those that have the same normalized form.\nI've predicted the rest of them (~390k) with ensembled beam search of size 16 and searched those that are not valid predictions. These (~16K) I predicted with beam size of 64.\nTaking all these predictions I only had ~5k molecules that I had no valid prediction for.\nI've predicted the 390k molecules with a bunch of other settings (with different beam size and ensemble weights), and tried to sample the valid ones with different strategies, but I just couldn't break the 0.70 barrier.\n\nVery interestingly, I had a lot of submissions that had quite an LD between them but all of them were 0.70 on LB. I even had two submissions with 0.11 LD between them, and they're still 0.70 and 0.71 on LB.\nI couldn't figure out if there could be any potential in this \"phenomenon\" with which I could make better predictions.\n\n**Didn't work:**\n - Label smoothing\n - Additional targets on encoder output\n - Encoder level mixup (decoder works on the two targets individually from the input to the output)\n - Sorting molecules by InChI BPE lenght for decoder speedup. I forgot to randomize, lost a whole week trying to figure out what went wrong. It felt incredible, as this happened right after I bought a 3090 ... :DDD\n\n**Thank you all:**\nThis was a very interesting problem to work on, and was a huge pleasure to compete with all of you!\nEven though I didn't end up in the gold zone, I don't regret spending that much money on a 3090, as so many worked so hard on this competition, and I'm more than happy that I could reach 18th place among so many so talented people!\n\n**Special thanks to:**\n@stassl and @yasufuminakama for the idea of normalization, and\n@hengck23 for the importance of validation\n\nThank you!",
    "1337724": "> Very interestingly, I had a lot of submissions that had quite an LD between them but all of them were 0.70 on LB. I even had two submissions with 0.11 LD between them, and they're still 0.70 and 0.71 on LB.\n> I couldn't figure out if there could be any potential in this \"phenomenon\" with which I could make better predictions.\n\nCould it be due to the fact that only 25% used for LB while 0.11 LD computed for compete submission?",
    "1337728": "0.11 LD is on the whole submissions (all 1.6M molecules).\n~0.43 was the LD on the 390k molecules.\n\nThat 0.11 difference means that the 0.71 LB submission is better in 0.05 characters but also is worse in 0.06 other characters than the 0.70 LB submission.",
    "1337729": "You joined kaggle 4 years ago, yet this is your very first comment? How? :D",
    "1337738": "Well, I was about to join in 2011 but delayed. This particular competition is interesting to me.",
    "1337740": "Indeed it was enjoyable!! :)",
    "1358578": "Congrats! Glad you got that boost in the last few days.\nQuestion: why do you do FP32 fine-tuning? I read somewhere that mixed precision (on PyTorch) is not supposed to have any negative impact on the outcome.",
    "1358941": "Thank you! :))\n\nDid you use `cascading and agreement` even before my mad discussion about the importance of validation?\n\nI read somewhere in the discussions that FP32 fine-tuning helped for someone. I didn't check it, but would make sense if it helps a little, so I just went with it.",
    "1364817": "Lol yes I did use that before your post.\n\nAh okay thanks, worth knowing that data point.",
    "3513844": "I get it now, I am also still waiting for a really interesting competition like this ever since."
  },
  "source": "meta"
}