{
  "id": 244166,
  "title": "13th Place Team (a story from a Kaggle newbie's view)",
  "url": "/competitions/bms-molecular-translation/writeups/mfl-eindhoven-13th-place-team-a-story-from-a-kaggl",
  "author_name": "",
  "post_date": "2021-06-07T23:39:04.550Z",
  "votes": 30,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hello all,</p>\n<p>It was an amazing competition. I would like to thank my team members ( <a href=\"https://www.kaggle.com/eakdag\" target=\"_blank\">@eakdag</a>, <a href=\"https://www.kaggle.com/ubique\" target=\"_blank\">@ubique</a>, <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a>, <a href=\"https://www.kaggle.com/proletheus\" target=\"_blank\">@proletheus</a> ) for their great effort. We missed the gold medal at the last moment, but I learned a lot and had nice friends.</p>\n<p>As a Kaggle newbie, I will explain the story of our team from my view. It could be a bit long so I divided it into sections and I will fill them in time.</p>\n<h3><strong>[Introduction]</strong></h3>\n<p>I and Erkut were friends from high school (MFL) at Konya/Turkey. After 12 years we met in Eindhoven/Netherlands by chance. We are both pursuing a PhD at TU/e. I am working on cheminformatics applications and Erkut is working on computer vision. We had discussed trying a Kaggle competition together 2 weeks before the BMS competition started. When this competition popped up, it was an amazing opportunity :D. So we set up \"MFL Eindhoven\" team.</p>\n<h3><strong>[Preparetion]</strong></h3>\n<p>I had no prior experience with computer vision except playing with some toy datasets at the beginning of the competition. I took a quick <a href=\"https://www.kaggle.com/learn/computer-vision\" target=\"_blank\">computer vision course</a> from Kaggle. Then I developed <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/229168#1255349\" target=\"_blank\">my first kernel which detects the stereochemical layer</a> in the image. </p>\n<p>At the same time, I analyzed InChI and listed possible solution ideas. (see the picture of my early research in the comments)</p>\n<h3><strong>[Training First Models]</strong></h3>\n<p>First, we started image captioning with <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>'s famous <a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter\" target=\"_blank\">ResNet+LSTM notebook</a>. However, we had limited resources and training was super slow. We wanted to make many different experiments, so we needed a faster way. (<strong>LB:7.06</strong>)</p>\n<p><a href=\"https://www.kaggle.com/markwijkhuizen/tensorflow-tpu-training-baseline-lb-16-92\" target=\"_blank\">Mark's TPU notebook </a>: <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> shared an amazing end-to-end TPU notebook. We used eff_B1+LSTM for training our models. Each epoch was taking around (30 min to 1 hour) which allowed us to make the following experiments:</p>\n<ul>\n<li>Larger Image size</li>\n<li>Different tokenizer </li>\n<li>Train with only longer sequences<br>\n(see comments for the list of other possible improvements)</li>\n</ul>\n<p>Now, we had several models (<strong>LB: btw 4.0 to 7.0</strong>) and it was time to ensemble them which will boost our score. <br>\n<img src=\"https://i.imgur.com/Hovv3C0.png\" alt=\"\"><br>\nPS: Special thanks to <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> ! These experiments would not have been possible without your sharing.</p>\n<h3><strong>[Boost of Merge Algorithms]</strong></h3>\n<p><strong>Gold Rush</strong><br>\nOur first merge algorithm was \"Gold Rush\". It simply merges two submissions by selecting valid molecules detected by Rdkit. When we apply Gold Rush to our submissions from different models, our score improved to <strong>LB: ~3.50</strong>.<br>\n<img src=\"https://i.imgur.com/UX4r21n.png\" alt=\"\"></p>\n<p><strong>The Collector of Lost Souls</strong><br>\nAfter we merge all submissions, there are still around 100k invalid molecules. In order to find them, we developed an algorithm that takes inference from each epoch for only invalid molecules. Using this method, we collected 20k more valid molecules and our score improved to <strong>LB: ~3.30</strong>.</p>\n<p>During that time <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> shared his <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231190\" target=\"_blank\">transformer and the inference script</a>. We inferred two submissions (<strong>LB:3.10 and 3.11</strong>). Then we also included these submissions into Gold Rush and we got a high improvement to <strong>LB:2.30</strong>. </p>\n<p><strong>Multi Merge</strong><br>\nAfter this point, I started to analyze my validation data. I found that avg error of valid molecules was 0.8. So even if we found all valid molecules, 0.8 was the top score we could reach. Then I started to compare different submissions and found different valid InChIs for the same molecules. It meant that valid molecules also need to be validated :D.  So I developed a new algorithm based on voting and minimum distance. That really worked and our score improved to <strong>LB: ~2.00</strong>.<br>\n<img src=\"https://i.imgur.com/5nHXrdf.png\" alt=\"\"><br>\nNow, we had a great merge algorithm and we were in a good LB position, it was a good time to develop a team-up strategy…</p>\n<h3><strong>[Team Up Strategy]</strong></h3>\n<p>Our algorithm boosted 3.xx models to 2.00 level, so if we find new team members with high-scored models we could have played for top positions. Our algorithm gets more benefit from diverse models, so we always took this into account. We decided to go until 5 members. After each member joined our score would increase and we would have a better chance to merge with higher LBs. So we decided to go round by round and started to invite other top LBs. </p>\n<p><strong>PS:</strong> There were some suspicious accounts at the top positions, we decided not to invite them.</p>\n<h3><strong>[Boost of New Members]</strong></h3>\n<p>In the first round, <a href=\"https://www.kaggle.com/ubique\" target=\"_blank\">@ubique</a> accepted our invitation. He had SMILES model trained with ResNest (<strong>LB:2.03</strong>).  When we added his models in our algorithm we reached (<strong>LB:1.16</strong>).</p>\n<p>Before the second round, <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a>  contacted us for a possible merge. He had just finished another competition and recently joined this one. He had an InChI Transformer model (<strong>LB:3.71</strong>) which predicts each layer separately. We thought that his experience would be very useful for the team and probably will improve his model quickly (it exactly came true). Since his InChI layers ensembled from separate models, valid ones were very reliable. His model boosted us to (<strong>LB:0.99</strong>).</p>\n<p>We reached that level, using relatively weak models. We needed a locomotive single model (preferably InChI). In the last round, <a href=\"https://www.kaggle.com/proletheus\" target=\"_blank\">@proletheus</a> joined our team, and we found our locomotive (<strong>LB:1.42</strong> Transformer InChI). By including his model into multi-merge, we reached (<strong>LB:0.78</strong>). </p>\n<h3><strong>[Developed Models]</strong></h3>\n<p>After that point, every day we worked hard and improved our models. In the end, our merged score reached (<strong>LB:0.62</strong>). The main models we developed are listed with scores (before pseudo-labeling) below.</p>\n<ul>\n<li>InChI / Transformer / <a href=\"https://www.kaggle.com/proletheus\" target=\"_blank\">@proletheus</a> / <strong>LB: 1.37</strong></li>\n<li>InChI (layered) / Transformer / <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a>  / <strong>LB: 2.01</strong></li>\n<li>InChI  / eff_B1+LSTM / <a href=\"https://www.kaggle.com/sorkun\" target=\"_blank\">@sorkun</a> &amp; <a href=\"https://www.kaggle.com/eakdag\" target=\"_blank\">@eakdag</a> / <strong>LB: 3.89</strong></li>\n<li>SMILES / eff_B1+LSTM / <a href=\"https://www.kaggle.com/sorkun\" target=\"_blank\">@sorkun</a> &amp; <a href=\"https://www.kaggle.com/eakdag\" target=\"_blank\">@eakdag</a>/ <strong>LB: 1.89</strong></li>\n<li>SMILES / ResNest+GRU / <a href=\"https://www.kaggle.com/ubique\" target=\"_blank\">@ubique</a> / <strong>LB: 2.03</strong></li>\n<li>SELFIES / Transformer  / <a href=\"https://www.kaggle.com/proletheus\" target=\"_blank\">@proletheus</a> / <strong>LB: 2.52</strong></li>\n<li>SELFIES / eff_B1+LSTM / <a href=\"https://www.kaggle.com/sorkun\" target=\"_blank\">@sorkun</a> &amp; <a href=\"https://www.kaggle.com/eakdag\" target=\"_blank\">@eakdag</a> / <strong>LB: 3.74</strong></li>\n</ul>\n<h3><strong>[Applied Methods]</strong></h3>\n<ul>\n<li>Molecular Reps (InChI, SMILES, SELFIES)</li>\n<li>OHEM</li>\n<li>Pseudo-labeling (improved single models but not improved the merged score)</li>\n</ul>",
  "messages": [
    {
      "id": "1337156",
      "postDate": "06/05/2021 12:57:11",
      "content": "<p>Hello all,</p>\n<p>It was an amazing competition. I would like to thank my team members ( <a href=\"https://www.kaggle.com/eakdag\" target=\"_blank\">@eakdag</a>, <a href=\"https://www.kaggle.com/ubique\" target=\"_blank\">@ubique</a>, <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a>, <a href=\"https://www.kaggle.com/proletheus\" target=\"_blank\">@proletheus</a> ) for their great effort. We missed the gold medal at the last moment, but I learned a lot and had nice friends.</p>\n<p>As a Kaggle newbie, I will explain the story of our team from my view. It could be a bit long so I divided it into sections and I will fill them in time.</p>\n<h3><strong>[Introduction]</strong></h3>\n<p>I and Erkut were friends from high school (MFL) at Konya/Turkey. After 12 years we met in Eindhoven/Netherlands by chance. We are both pursuing a PhD at TU/e. I am working on cheminformatics applications and Erkut is working on computer vision. We had discussed trying a Kaggle competition together 2 weeks before the BMS competition started. When this competition popped up, it was an amazing opportunity :D. So we set up \"MFL Eindhoven\" team.</p>\n<h3><strong>[Preparetion]</strong></h3>\n<p>I had no prior experience with computer vision except playing with some toy datasets at the beginning of the competition. I took a quick <a href=\"https://www.kaggle.com/learn/computer-vision\" target=\"_blank\">computer vision course</a> from Kaggle. Then I developed <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/229168#1255349\" target=\"_blank\">my first kernel which detects the stereochemical layer</a> in the image. </p>\n<p>At the same time, I analyzed InChI and listed possible solution ideas. (see the picture of my early research in the comments)</p>\n<h3><strong>[Training First Models]</strong></h3>\n<p>First, we started image captioning with <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>'s famous <a href=\"https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter\" target=\"_blank\">ResNet+LSTM notebook</a>. However, we had limited resources and training was super slow. We wanted to make many different experiments, so we needed a faster way. (<strong>LB:7.06</strong>)</p>\n<p><a href=\"https://www.kaggle.com/markwijkhuizen/tensorflow-tpu-training-baseline-lb-16-92\" target=\"_blank\">Mark's TPU notebook </a>: <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> shared an amazing end-to-end TPU notebook. We used eff_B1+LSTM for training our models. Each epoch was taking around (30 min to 1 hour) which allowed us to make the following experiments:</p>\n<ul>\n<li>Larger Image size</li>\n<li>Different tokenizer </li>\n<li>Train with only longer sequences<br>\n(see comments for the list of other possible improvements)</li>\n</ul>\n<p>Now, we had several models (<strong>LB: btw 4.0 to 7.0</strong>) and it was time to ensemble them which will boost our score. <br>\n<img src=\"https://i.imgur.com/Hovv3C0.png\" alt=\"\"><br>\nPS: Special thanks to <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> ! These experiments would not have been possible without your sharing.</p>\n<h3><strong>[Boost of Merge Algorithms]</strong></h3>\n<p><strong>Gold Rush</strong><br>\nOur first merge algorithm was \"Gold Rush\". It simply merges two submissions by selecting valid molecules detected by Rdkit. When we apply Gold Rush to our submissions from different models, our score improved to <strong>LB: ~3.50</strong>.<br>\n<img src=\"https://i.imgur.com/UX4r21n.png\" alt=\"\"></p>\n<p><strong>The Collector of Lost Souls</strong><br>\nAfter we merge all submissions, there are still around 100k invalid molecules. In order to find them, we developed an algorithm that takes inference from each epoch for only invalid molecules. Using this method, we collected 20k more valid molecules and our score improved to <strong>LB: ~3.30</strong>.</p>\n<p>During that time <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> shared his <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231190\" target=\"_blank\">transformer and the inference script</a>. We inferred two submissions (<strong>LB:3.10 and 3.11</strong>). Then we also included these submissions into Gold Rush and we got a high improvement to <strong>LB:2.30</strong>. </p>\n<p><strong>Multi Merge</strong><br>\nAfter this point, I started to analyze my validation data. I found that avg error of valid molecules was 0.8. So even if we found all valid molecules, 0.8 was the top score we could reach. Then I started to compare different submissions and found different valid InChIs for the same molecules. It meant that valid molecules also need to be validated :D.  So I developed a new algorithm based on voting and minimum distance. That really worked and our score improved to <strong>LB: ~2.00</strong>.<br>\n<img src=\"https://i.imgur.com/5nHXrdf.png\" alt=\"\"><br>\nNow, we had a great merge algorithm and we were in a good LB position, it was a good time to develop a team-up strategy…</p>\n<h3><strong>[Team Up Strategy]</strong></h3>\n<p>Our algorithm boosted 3.xx models to 2.00 level, so if we find new team members with high-scored models we could have played for top positions. Our algorithm gets more benefit from diverse models, so we always took this into account. We decided to go until 5 members. After each member joined our score would increase and we would have a better chance to merge with higher LBs. So we decided to go round by round and started to invite other top LBs. </p>\n<p><strong>PS:</strong> There were some suspicious accounts at the top positions, we decided not to invite them.</p>\n<h3><strong>[Boost of New Members]</strong></h3>\n<p>In the first round, <a href=\"https://www.kaggle.com/ubique\" target=\"_blank\">@ubique</a> accepted our invitation. He had SMILES model trained with ResNest (<strong>LB:2.03</strong>).  When we added his models in our algorithm we reached (<strong>LB:1.16</strong>).</p>\n<p>Before the second round, <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a>  contacted us for a possible merge. He had just finished another competition and recently joined this one. He had an InChI Transformer model (<strong>LB:3.71</strong>) which predicts each layer separately. We thought that his experience would be very useful for the team and probably will improve his model quickly (it exactly came true). Since his InChI layers ensembled from separate models, valid ones were very reliable. His model boosted us to (<strong>LB:0.99</strong>).</p>\n<p>We reached that level, using relatively weak models. We needed a locomotive single model (preferably InChI). In the last round, <a href=\"https://www.kaggle.com/proletheus\" target=\"_blank\">@proletheus</a> joined our team, and we found our locomotive (<strong>LB:1.42</strong> Transformer InChI). By including his model into multi-merge, we reached (<strong>LB:0.78</strong>). </p>\n<h3><strong>[Developed Models]</strong></h3>\n<p>After that point, every day we worked hard and improved our models. In the end, our merged score reached (<strong>LB:0.62</strong>). The main models we developed are listed with scores (before pseudo-labeling) below.</p>\n<ul>\n<li>InChI / Transformer / <a href=\"https://www.kaggle.com/proletheus\" target=\"_blank\">@proletheus</a> / <strong>LB: 1.37</strong></li>\n<li>InChI (layered) / Transformer / <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a>  / <strong>LB: 2.01</strong></li>\n<li>InChI  / eff_B1+LSTM / <a href=\"https://www.kaggle.com/sorkun\" target=\"_blank\">@sorkun</a> &amp; <a href=\"https://www.kaggle.com/eakdag\" target=\"_blank\">@eakdag</a> / <strong>LB: 3.89</strong></li>\n<li>SMILES / eff_B1+LSTM / <a href=\"https://www.kaggle.com/sorkun\" target=\"_blank\">@sorkun</a> &amp; <a href=\"https://www.kaggle.com/eakdag\" target=\"_blank\">@eakdag</a>/ <strong>LB: 1.89</strong></li>\n<li>SMILES / ResNest+GRU / <a href=\"https://www.kaggle.com/ubique\" target=\"_blank\">@ubique</a> / <strong>LB: 2.03</strong></li>\n<li>SELFIES / Transformer  / <a href=\"https://www.kaggle.com/proletheus\" target=\"_blank\">@proletheus</a> / <strong>LB: 2.52</strong></li>\n<li>SELFIES / eff_B1+LSTM / <a href=\"https://www.kaggle.com/sorkun\" target=\"_blank\">@sorkun</a> &amp; <a href=\"https://www.kaggle.com/eakdag\" target=\"_blank\">@eakdag</a> / <strong>LB: 3.74</strong></li>\n</ul>\n<h3><strong>[Applied Methods]</strong></h3>\n<ul>\n<li>Molecular Reps (InChI, SMILES, SELFIES)</li>\n<li>OHEM</li>\n<li>Pseudo-labeling (improved single models but not improved the merged score)</li>\n</ul>",
      "rawMarkdown": "Hello all,\n\nIt was an amazing competition. I would like to thank my team members ( @eakdag, @ubique, @aerdem4, @proletheus ) for their great effort. We missed the gold medal at the last moment, but I learned a lot and had nice friends.\n\nAs a Kaggle newbie, I will explain the story of our team from my view. It could be a bit long so I divided it into sections and I will fill them in time.\n\n### **[Introduction]**\n\nI and Erkut were friends from high school (MFL) at Konya/Turkey. After 12 years we met in Eindhoven/Netherlands by chance. We are both pursuing a PhD at TU/e. I am working on cheminformatics applications and Erkut is working on computer vision. We had discussed trying a Kaggle competition together 2 weeks before the BMS competition started. When this competition popped up, it was an amazing opportunity :D. So we set up \"MFL Eindhoven\" team.\n\n### **[Preparetion]**\n\nI had no prior experience with computer vision except playing with some toy datasets at the beginning of the competition. I took a quick [computer vision course](https://www.kaggle.com/learn/computer-vision) from Kaggle. Then I developed [my first kernel which detects the stereochemical layer](https://www.kaggle.com/c/bms-molecular-translation/discussion/229168#1255349) in the image. \n\nAt the same time, I analyzed InChI and listed possible solution ideas. (see the picture of my early research in the comments)\n\n### **[Training First Models]**\n\nFirst, we started image captioning with @yasufuminakama's famous [ResNet+LSTM notebook](https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter). However, we had limited resources and training was super slow. We wanted to make many different experiments, so we needed a faster way. (**LB:7.06**)\n\n[Mark's TPU notebook ](https://www.kaggle.com/markwijkhuizen/tensorflow-tpu-training-baseline-lb-16-92): @markwijkhuizen shared an amazing end-to-end TPU notebook. We used eff_B1+LSTM for training our models. Each epoch was taking around (30 min to 1 hour) which allowed us to make the following experiments:\n\n- Larger Image size\n- Different tokenizer \n- Train with only longer sequences\n(see comments for the list of other possible improvements)\n\nNow, we had several models (**LB: btw 4.0 to 7.0**) and it was time to ensemble them which will boost our score. \n![](https://i.imgur.com/Hovv3C0.png)\nPS: Special thanks to @markwijkhuizen ! These experiments would not have been possible without your sharing.\n\n\n### **[Boost of Merge Algorithms]**\n\n**Gold Rush**\nOur first merge algorithm was \"Gold Rush\". It simply merges two submissions by selecting valid molecules detected by Rdkit. When we apply Gold Rush to our submissions from different models, our score improved to **LB: ~3.50**.\n![](https://i.imgur.com/UX4r21n.png)\n\n\n**The Collector of Lost Souls**\nAfter we merge all submissions, there are still around 100k invalid molecules. In order to find them, we developed an algorithm that takes inference from each epoch for only invalid molecules. Using this method, we collected 20k more valid molecules and our score improved to **LB: ~3.30**.\n\nDuring that time @hengck23 shared his [transformer and the inference script](https://www.kaggle.com/c/bms-molecular-translation/discussion/231190). We inferred two submissions (**LB:3.10 and 3.11**). Then we also included these submissions into Gold Rush and we got a high improvement to **LB:2.30**. \n\n\n**Multi Merge**\nAfter this point, I started to analyze my validation data. I found that avg error of valid molecules was 0.8. So even if we found all valid molecules, 0.8 was the top score we could reach. Then I started to compare different submissions and found different valid InChIs for the same molecules. It meant that valid molecules also need to be validated :D.  So I developed a new algorithm based on voting and minimum distance. That really worked and our score improved to **LB: ~2.00**.\n![](https://i.imgur.com/5nHXrdf.png)\nNow, we had a great merge algorithm and we were in a good LB position, it was a good time to develop a team-up strategy...\n\n### **[Team Up Strategy]**\nOur algorithm boosted 3.xx models to 2.00 level, so if we find new team members with high-scored models we could have played for top positions. Our algorithm gets more benefit from diverse models, so we always took this into account. We decided to go until 5 members. After each member joined our score would increase and we would have a better chance to merge with higher LBs. So we decided to go round by round and started to invite other top LBs. \n\n**PS:** There were some suspicious accounts at the top positions, we decided not to invite them.\n\n###  **[Boost of New Members]**\nIn the first round, @ubique accepted our invitation. He had SMILES model trained with ResNest (**LB:2.03**).  When we added his models in our algorithm we reached (**LB:1.16**).\n\nBefore the second round, @aerdem4  contacted us for a possible merge. He had just finished another competition and recently joined this one. He had an InChI Transformer model (**LB:3.71**) which predicts each layer separately. We thought that his experience would be very useful for the team and probably will improve his model quickly (it exactly came true). Since his InChI layers ensembled from separate models, valid ones were very reliable. His model boosted us to (**LB:0.99**).\n\nWe reached that level, using relatively weak models. We needed a locomotive single model (preferably InChI). In the last round, @proletheus joined our team, and we found our locomotive (**LB:1.42** Transformer InChI). By including his model into multi-merge, we reached (**LB:0.78**). \n\n###  **[Developed Models]**\nAfter that point, every day we worked hard and improved our models. In the end, our merged score reached (**LB:0.62**). The main models we developed are listed with scores (before pseudo-labeling) below.\n\n- InChI / Transformer / @proletheus / **LB: 1.37**\n- InChI (layered) / Transformer / @aerdem4  / **LB: 2.01**\n- InChI  / eff_B1+LSTM / @sorkun & @eakdag / **LB: 3.89**\n- SMILES / eff_B1+LSTM / @sorkun & @eakdag/ **LB: 1.89**\n- SMILES / ResNest+GRU / @ubique / **LB: 2.03**\n- SELFIES / Transformer  / @proletheus / **LB: 2.52**\n- SELFIES / eff_B1+LSTM / @sorkun & @eakdag / **LB: 3.74**\n\n###  **[Applied Methods]**\n\n- Molecular Reps (InChI, SMILES, SELFIES)\n- OHEM\n- Pseudo-labeling (improved single models but not improved the merged score)",
      "votes": null
    },
    {
      "id": "1337768",
      "postDate": "06/05/2021 21:53:00",
      "content": "<p><strong>Preparation</strong> section is updated. </p>\n<p>Preliminary research on InChI and some Ideas…<br>\n<img src=\"https://i.imgur.com/MrDswM7.jpg\" alt=\"\"></p>",
      "rawMarkdown": "**Preparation** section is updated. \n\nPreliminary research on InChI and some Ideas...\n![](https://i.imgur.com/MrDswM7.jpg)",
      "votes": null
    },
    {
      "id": "1338374",
      "postDate": "06/06/2021 11:47:34",
      "content": "<p><strong>Training First Models</strong> section is updated.</p>\n<p>More ideas about possible improvements…<br>\n<img src=\"https://i.imgur.com/4naXMsM.png\" alt=\"\"></p>",
      "rawMarkdown": "**Training First Models** section is updated.\n\nMore ideas about possible improvements...\n![](https://i.imgur.com/4naXMsM.png)",
      "votes": null
    },
    {
      "id": "1338461",
      "postDate": "06/06/2021 13:32:53",
      "content": "<p>Congrats for the silver! Look forward to your full write-up (and your story)<br>\nDid u mainly use Kaggle TPU in the beginning stage?</p>",
      "rawMarkdown": "Congrats for the silver! Look forward to your full write-up (and your story)\nDid u mainly use Kaggle TPU in the beginning stage?",
      "votes": null
    },
    {
      "id": "1339013",
      "postDate": "06/07/2021 00:26:38",
      "content": "<p>Thanks! Yes, we used Kaggle and Colab for our TPU experiments. </p>",
      "rawMarkdown": "Thanks! Yes, we used Kaggle and Colab for our TPU experiments.",
      "votes": null
    },
    {
      "id": "1339017",
      "postDate": "06/07/2021 00:32:19",
      "content": "<p>Nice! <br>\nI didnt take part in the competition but I know the dataset is gigantic. Running an experiment  could take so long. Did u do anything special in order to run your experiments more efficiently in Kaggle/ Colab?</p>",
      "rawMarkdown": "Nice! \nI didnt take part in the competition but I know the dataset is gigantic. Running an experiment  could take so long. Did u do anything special in order to run your experiments more efficiently in Kaggle/ Colab?",
      "votes": null
    },
    {
      "id": "1339028",
      "postDate": "06/07/2021 01:04:40",
      "content": "<p>What I see is TPU memory allows you to go very large batch sizes like 128*8. I used <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a>'s reference notebook for TPU experiments. It took around 30 mins for each epoch, so I could try different ideas within a day. </p>",
      "rawMarkdown": "What I see is TPU memory allows you to go very large batch sizes like 128*8. I used @markwijkhuizen's reference notebook for TPU experiments. It took around 30 mins for each epoch, so I could try different ideas within a day.",
      "votes": null
    },
    {
      "id": "1339035",
      "postDate": "06/07/2021 01:23:08",
      "content": "<p><strong>Boost of Merge Algorithms</strong> section is updated.</p>\n<p>Multi Merge algorithm first draft…<br>\n<img src=\"https://i.imgur.com/o4Bq1XP.jpg\" alt=\"\"></p>",
      "rawMarkdown": "**Boost of Merge Algorithms** section is updated.\n\nMulti Merge algorithm first draft…\n![](https://i.imgur.com/o4Bq1XP.jpg)",
      "votes": null
    },
    {
      "id": "1340435",
      "postDate": "06/07/2021 23:32:09",
      "content": "<p>The story has been completed. It was an enjoyable but tiring experience for me.<br>\n I hope you also enjoyed reading!</p>\n<p>Here is the final image, a small part of our merge experiments! <br>\n<img src=\"https://i.imgur.com/fMlfEWn.png\" alt=\"\"></p>",
      "rawMarkdown": "The story has been completed. It was an enjoyable but tiring experience for me.\n I hope you also enjoyed reading!\n\nHere is the final image, a small part of our merge experiments! \n![](https://i.imgur.com/fMlfEWn.png)",
      "votes": null
    },
    {
      "id": "1348887",
      "postDate": "06/14/2021 11:06:21",
      "content": "<p>Congrats :) Also, amazing storytelling!</p>",
      "rawMarkdown": "Congrats :) Also, amazing storytelling!",
      "votes": null
    },
    {
      "id": "1351028",
      "postDate": "06/16/2021 03:00:12",
      "content": "<p>Wonderful explanation! 👍 <br>\nI will also helpful for newbies who are still afraid to enter in competitions! 👊 </p>",
      "rawMarkdown": "Wonderful explanation! 👍 \nI will also helpful for newbies who are still afraid to enter in competitions! 👊",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1337768,
      "author_name": "sorkun",
      "author_url": "",
      "post_date": "06/05/2021 21:53:00",
      "content": "<p><strong>Preparation</strong> section is updated. </p>\n<p>Preliminary research on InChI and some Ideas…<br>\n<img src=\"https://i.imgur.com/MrDswM7.jpg\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1338374,
      "author_name": "sorkun",
      "author_url": "",
      "post_date": "06/06/2021 11:47:34",
      "content": "<p><strong>Training First Models</strong> section is updated.</p>\n<p>More ideas about possible improvements…<br>\n<img src=\"https://i.imgur.com/4naXMsM.png\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1338461,
      "author_name": "alexlwh",
      "author_url": "",
      "post_date": "06/06/2021 13:32:53",
      "content": "<p>Congrats for the silver! Look forward to your full write-up (and your story)<br>\nDid u mainly use Kaggle TPU in the beginning stage?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1339013,
          "author_name": "sorkun",
          "author_url": "",
          "post_date": "06/07/2021 00:26:38",
          "content": "<p>Thanks! Yes, we used Kaggle and Colab for our TPU experiments. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1339017,
          "author_name": "alexlwh",
          "author_url": "",
          "post_date": "06/07/2021 00:32:19",
          "content": "<p>Nice! <br>\nI didnt take part in the competition but I know the dataset is gigantic. Running an experiment  could take so long. Did u do anything special in order to run your experiments more efficiently in Kaggle/ Colab?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1339028,
          "author_name": "sorkun",
          "author_url": "",
          "post_date": "06/07/2021 01:04:40",
          "content": "<p>What I see is TPU memory allows you to go very large batch sizes like 128*8. I used <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a>'s reference notebook for TPU experiments. It took around 30 mins for each epoch, so I could try different ideas within a day. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1339035,
      "author_name": "sorkun",
      "author_url": "",
      "post_date": "06/07/2021 01:23:08",
      "content": "<p><strong>Boost of Merge Algorithms</strong> section is updated.</p>\n<p>Multi Merge algorithm first draft…<br>\n<img src=\"https://i.imgur.com/o4Bq1XP.jpg\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1340435,
      "author_name": "sorkun",
      "author_url": "",
      "post_date": "06/07/2021 23:32:09",
      "content": "<p>The story has been completed. It was an enjoyable but tiring experience for me.<br>\n I hope you also enjoyed reading!</p>\n<p>Here is the final image, a small part of our merge experiments! <br>\n<img src=\"https://i.imgur.com/fMlfEWn.png\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1348887,
      "author_name": "cihanyatbaz",
      "author_url": "",
      "post_date": "06/14/2021 11:06:21",
      "content": "<p>Congrats :) Also, amazing storytelling!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1351028,
      "author_name": "mudassar66",
      "author_url": "",
      "post_date": "06/16/2021 03:00:12",
      "content": "<p>Wonderful explanation! 👍 <br>\nI will also helpful for newbies who are still afraid to enter in competitions! 👊 </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1337156": "Hello all,\n\nIt was an amazing competition. I would like to thank my team members ( @eakdag, @ubique, @aerdem4, @proletheus ) for their great effort. We missed the gold medal at the last moment, but I learned a lot and had nice friends.\n\nAs a Kaggle newbie, I will explain the story of our team from my view. It could be a bit long so I divided it into sections and I will fill them in time.\n\n### **[Introduction]**\n\nI and Erkut were friends from high school (MFL) at Konya/Turkey. After 12 years we met in Eindhoven/Netherlands by chance. We are both pursuing a PhD at TU/e. I am working on cheminformatics applications and Erkut is working on computer vision. We had discussed trying a Kaggle competition together 2 weeks before the BMS competition started. When this competition popped up, it was an amazing opportunity :D. So we set up \"MFL Eindhoven\" team.\n\n### **[Preparetion]**\n\nI had no prior experience with computer vision except playing with some toy datasets at the beginning of the competition. I took a quick [computer vision course](https://www.kaggle.com/learn/computer-vision) from Kaggle. Then I developed [my first kernel which detects the stereochemical layer](https://www.kaggle.com/c/bms-molecular-translation/discussion/229168#1255349) in the image. \n\nAt the same time, I analyzed InChI and listed possible solution ideas. (see the picture of my early research in the comments)\n\n### **[Training First Models]**\n\nFirst, we started image captioning with @yasufuminakama's famous [ResNet+LSTM notebook](https://www.kaggle.com/yasufuminakama/inchi-resnet-lstm-with-attention-starter). However, we had limited resources and training was super slow. We wanted to make many different experiments, so we needed a faster way. (**LB:7.06**)\n\n[Mark's TPU notebook ](https://www.kaggle.com/markwijkhuizen/tensorflow-tpu-training-baseline-lb-16-92): @markwijkhuizen shared an amazing end-to-end TPU notebook. We used eff_B1+LSTM for training our models. Each epoch was taking around (30 min to 1 hour) which allowed us to make the following experiments:\n\n- Larger Image size\n- Different tokenizer \n- Train with only longer sequences\n(see comments for the list of other possible improvements)\n\nNow, we had several models (**LB: btw 4.0 to 7.0**) and it was time to ensemble them which will boost our score. \n![](https://i.imgur.com/Hovv3C0.png)\nPS: Special thanks to @markwijkhuizen ! These experiments would not have been possible without your sharing.\n\n\n### **[Boost of Merge Algorithms]**\n\n**Gold Rush**\nOur first merge algorithm was \"Gold Rush\". It simply merges two submissions by selecting valid molecules detected by Rdkit. When we apply Gold Rush to our submissions from different models, our score improved to **LB: ~3.50**.\n![](https://i.imgur.com/UX4r21n.png)\n\n\n**The Collector of Lost Souls**\nAfter we merge all submissions, there are still around 100k invalid molecules. In order to find them, we developed an algorithm that takes inference from each epoch for only invalid molecules. Using this method, we collected 20k more valid molecules and our score improved to **LB: ~3.30**.\n\nDuring that time @hengck23 shared his [transformer and the inference script](https://www.kaggle.com/c/bms-molecular-translation/discussion/231190). We inferred two submissions (**LB:3.10 and 3.11**). Then we also included these submissions into Gold Rush and we got a high improvement to **LB:2.30**. \n\n\n**Multi Merge**\nAfter this point, I started to analyze my validation data. I found that avg error of valid molecules was 0.8. So even if we found all valid molecules, 0.8 was the top score we could reach. Then I started to compare different submissions and found different valid InChIs for the same molecules. It meant that valid molecules also need to be validated :D.  So I developed a new algorithm based on voting and minimum distance. That really worked and our score improved to **LB: ~2.00**.\n![](https://i.imgur.com/5nHXrdf.png)\nNow, we had a great merge algorithm and we were in a good LB position, it was a good time to develop a team-up strategy...\n\n### **[Team Up Strategy]**\nOur algorithm boosted 3.xx models to 2.00 level, so if we find new team members with high-scored models we could have played for top positions. Our algorithm gets more benefit from diverse models, so we always took this into account. We decided to go until 5 members. After each member joined our score would increase and we would have a better chance to merge with higher LBs. So we decided to go round by round and started to invite other top LBs. \n\n**PS:** There were some suspicious accounts at the top positions, we decided not to invite them.\n\n###  **[Boost of New Members]**\nIn the first round, @ubique accepted our invitation. He had SMILES model trained with ResNest (**LB:2.03**).  When we added his models in our algorithm we reached (**LB:1.16**).\n\nBefore the second round, @aerdem4  contacted us for a possible merge. He had just finished another competition and recently joined this one. He had an InChI Transformer model (**LB:3.71**) which predicts each layer separately. We thought that his experience would be very useful for the team and probably will improve his model quickly (it exactly came true). Since his InChI layers ensembled from separate models, valid ones were very reliable. His model boosted us to (**LB:0.99**).\n\nWe reached that level, using relatively weak models. We needed a locomotive single model (preferably InChI). In the last round, @proletheus joined our team, and we found our locomotive (**LB:1.42** Transformer InChI). By including his model into multi-merge, we reached (**LB:0.78**). \n\n###  **[Developed Models]**\nAfter that point, every day we worked hard and improved our models. In the end, our merged score reached (**LB:0.62**). The main models we developed are listed with scores (before pseudo-labeling) below.\n\n- InChI / Transformer / @proletheus / **LB: 1.37**\n- InChI (layered) / Transformer / @aerdem4  / **LB: 2.01**\n- InChI  / eff_B1+LSTM / @sorkun & @eakdag / **LB: 3.89**\n- SMILES / eff_B1+LSTM / @sorkun & @eakdag/ **LB: 1.89**\n- SMILES / ResNest+GRU / @ubique / **LB: 2.03**\n- SELFIES / Transformer  / @proletheus / **LB: 2.52**\n- SELFIES / eff_B1+LSTM / @sorkun & @eakdag / **LB: 3.74**\n\n###  **[Applied Methods]**\n\n- Molecular Reps (InChI, SMILES, SELFIES)\n- OHEM\n- Pseudo-labeling (improved single models but not improved the merged score)",
    "1337768": "**Preparation** section is updated. \n\nPreliminary research on InChI and some Ideas...\n![](https://i.imgur.com/MrDswM7.jpg)",
    "1338374": "**Training First Models** section is updated.\n\nMore ideas about possible improvements...\n![](https://i.imgur.com/4naXMsM.png)",
    "1338461": "Congrats for the silver! Look forward to your full write-up (and your story)\nDid u mainly use Kaggle TPU in the beginning stage?",
    "1339013": "Thanks! Yes, we used Kaggle and Colab for our TPU experiments.",
    "1339017": "Nice! \nI didnt take part in the competition but I know the dataset is gigantic. Running an experiment  could take so long. Did u do anything special in order to run your experiments more efficiently in Kaggle/ Colab?",
    "1339028": "What I see is TPU memory allows you to go very large batch sizes like 128*8. I used @markwijkhuizen's reference notebook for TPU experiments. It took around 30 mins for each epoch, so I could try different ideas within a day.",
    "1339035": "**Boost of Merge Algorithms** section is updated.\n\nMulti Merge algorithm first draft…\n![](https://i.imgur.com/o4Bq1XP.jpg)",
    "1340435": "The story has been completed. It was an enjoyable but tiring experience for me.\n I hope you also enjoyed reading!\n\nHere is the final image, a small part of our merge experiments! \n![](https://i.imgur.com/fMlfEWn.png)",
    "1348887": "Congrats :) Also, amazing storytelling!",
    "1351028": "Wonderful explanation! 👍 \nI will also helpful for newbies who are still afraid to enter in competitions! 👊"
  },
  "source": "meta"
}