{
  "id": 244601,
  "title": "Ensemble directly on submissions",
  "url": "/competitions/bms-molecular-translation/discussion/244601",
  "author_name": "",
  "post_date": "2021-06-07T14:27:24.533058800Z",
  "votes": 26,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I made my script for ensembling, which works directly on set of submissions. So it's easy to try on your model results.</p>\n<p>Link: <a href=\"https://www.kaggle.com/zfturbo/ensemble-code-based-on-inchi-part-and-lev-distance\" target=\"_blank\">https://www.kaggle.com/zfturbo/ensemble-code-based-on-inchi-part-and-lev-distance</a></p>\n<p>Algorithm:</p>\n<ol>\n<li>Script splits Inchi strings on parts based on '/' and optimize each part independently</li>\n<li>Script finds part with minimum Levenshtein distance for each other part and use it as answer</li>\n<li>Script constructs all possible variation of substrings and use the one with minimal distance.</li>\n<li>If RDKit is enabled (you can do it locally) script uses the first answer which is correct according to rdkit in sorted by distance list of possible answers.</li>\n</ol>\n<p>I run it on public script submissions:<br>\n<strong>3.63</strong>: <a href=\"https://www.kaggle.com/yingpengchen/pl-bms-molecular-translation\" target=\"_blank\">https://www.kaggle.com/yingpengchen/pl-bms-molecular-translation</a><br>\n<strong>3.69</strong>: <a href=\"https://www.kaggle.com/marbury/efficientnet-multi-layer-lstm-inference\" target=\"_blank\">https://www.kaggle.com/marbury/efficientnet-multi-layer-lstm-inference</a><br>\n<strong>3.70</strong>: <a href=\"https://www.kaggle.com/dragonzhang/tensorflow-tpu-training-baseline-predictions\" target=\"_blank\">https://www.kaggle.com/dragonzhang/tensorflow-tpu-training-baseline-predictions</a><br>\nScript without RDKit gives: <strong>2.49</strong> (with RDKit enabled it's only: <strong>2.12</strong>)</p>\n<p>Adding<br>\n<strong>4.20</strong>: <a href=\"https://www.kaggle.com/vikrant06/bms-efficientnetv2-tpu-32-epocs-final-lr-e-5\" target=\"_blank\">https://www.kaggle.com/vikrant06/bms-efficientnetv2-tpu-32-epocs-final-lr-e-5</a><br>\nreduce score to <strong>2.30</strong>.</p>\n<p>With more submissions it must give even better score.</p>",
  "messages": [
    {
      "id": "1339904",
      "postDate": "06/07/2021 14:27:24",
      "content": "<p>I made my script for ensembling, which works directly on set of submissions. So it's easy to try on your model results.</p>\n<p>Link: <a href=\"https://www.kaggle.com/zfturbo/ensemble-code-based-on-inchi-part-and-lev-distance\" target=\"_blank\">https://www.kaggle.com/zfturbo/ensemble-code-based-on-inchi-part-and-lev-distance</a></p>\n<p>Algorithm:</p>\n<ol>\n<li>Script splits Inchi strings on parts based on '/' and optimize each part independently</li>\n<li>Script finds part with minimum Levenshtein distance for each other part and use it as answer</li>\n<li>Script constructs all possible variation of substrings and use the one with minimal distance.</li>\n<li>If RDKit is enabled (you can do it locally) script uses the first answer which is correct according to rdkit in sorted by distance list of possible answers.</li>\n</ol>\n<p>I run it on public script submissions:<br>\n<strong>3.63</strong>: <a href=\"https://www.kaggle.com/yingpengchen/pl-bms-molecular-translation\" target=\"_blank\">https://www.kaggle.com/yingpengchen/pl-bms-molecular-translation</a><br>\n<strong>3.69</strong>: <a href=\"https://www.kaggle.com/marbury/efficientnet-multi-layer-lstm-inference\" target=\"_blank\">https://www.kaggle.com/marbury/efficientnet-multi-layer-lstm-inference</a><br>\n<strong>3.70</strong>: <a href=\"https://www.kaggle.com/dragonzhang/tensorflow-tpu-training-baseline-predictions\" target=\"_blank\">https://www.kaggle.com/dragonzhang/tensorflow-tpu-training-baseline-predictions</a><br>\nScript without RDKit gives: <strong>2.49</strong> (with RDKit enabled it's only: <strong>2.12</strong>)</p>\n<p>Adding<br>\n<strong>4.20</strong>: <a href=\"https://www.kaggle.com/vikrant06/bms-efficientnetv2-tpu-32-epocs-final-lr-e-5\" target=\"_blank\">https://www.kaggle.com/vikrant06/bms-efficientnetv2-tpu-32-epocs-final-lr-e-5</a><br>\nreduce score to <strong>2.30</strong>.</p>\n<p>With more submissions it must give even better score.</p>",
      "rawMarkdown": "I made my script for ensembling, which works directly on set of submissions. So it's easy to try on your model results.\n\nLink: https://www.kaggle.com/zfturbo/ensemble-code-based-on-inchi-part-and-lev-distance\n\nAlgorithm:\n1. Script splits Inchi strings on parts based on '/' and optimize each part independently\n2. Script finds part with minimum Levenshtein distance for each other part and use it as answer\n3. Script constructs all possible variation of substrings and use the one with minimal distance.\n4. If RDKit is enabled (you can do it locally) script uses the first answer which is correct according to rdkit in sorted by distance list of possible answers.\n\nI run it on public script submissions:\n**3.63**: https://www.kaggle.com/yingpengchen/pl-bms-molecular-translation\n**3.69**: https://www.kaggle.com/marbury/efficientnet-multi-layer-lstm-inference\n**3.70**: https://www.kaggle.com/dragonzhang/tensorflow-tpu-training-baseline-predictions\nScript without RDKit gives: **2.49** (with RDKit enabled it's only: **2.12**)\n\nAdding\n**4.20**: https://www.kaggle.com/vikrant06/bms-efficientnetv2-tpu-32-epocs-final-lr-e-5\nreduce score to **2.30**.\n\nWith more submissions it must give even better score.",
      "votes": null
    },
    {
      "id": "1340564",
      "postDate": "06/08/2021 05:25:39",
      "content": "<p>Nice work, thanks for sharing. </p>\n<p>I did consider training models separately for each InChI component (i.e. to \"specialise\"), but it was too time consuming and I had doubts as to its effectiveness. Ultimately your solution was a much better approach to exploit the fact that the InChI are given as components! </p>",
      "rawMarkdown": "Nice work, thanks for sharing. \n\nI did consider training models separately for each InChI component (i.e. to \"specialise\"), but it was too time consuming and I had doubts as to its effectiveness. Ultimately your solution was a much better approach to exploit the fact that the InChI are given as components!",
      "votes": null
    },
    {
      "id": "1345185",
      "postDate": "06/11/2021 11:38:53",
      "content": "<p>This was our main Ensembling method. and to be honest we didn't have a single model scoring below 1.00 but this ensembling pull us to 0.65 in the Public LB</p>",
      "rawMarkdown": "This was our main Ensembling method. and to be honest we didn't have a single model scoring below 1.00 but this ensembling pull us to 0.65 in the Public LB",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1340564,
      "author_name": "talktocharles",
      "author_url": "",
      "post_date": "06/08/2021 05:25:39",
      "content": "<p>Nice work, thanks for sharing. </p>\n<p>I did consider training models separately for each InChI component (i.e. to \"specialise\"), but it was too time consuming and I had doubts as to its effectiveness. Ultimately your solution was a much better approach to exploit the fact that the InChI are given as components! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1345185,
      "author_name": "morizin",
      "author_url": "",
      "post_date": "06/11/2021 11:38:53",
      "content": "<p>This was our main Ensembling method. and to be honest we didn't have a single model scoring below 1.00 but this ensembling pull us to 0.65 in the Public LB</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1339904": "I made my script for ensembling, which works directly on set of submissions. So it's easy to try on your model results.\n\nLink: https://www.kaggle.com/zfturbo/ensemble-code-based-on-inchi-part-and-lev-distance\n\nAlgorithm:\n1. Script splits Inchi strings on parts based on '/' and optimize each part independently\n2. Script finds part with minimum Levenshtein distance for each other part and use it as answer\n3. Script constructs all possible variation of substrings and use the one with minimal distance.\n4. If RDKit is enabled (you can do it locally) script uses the first answer which is correct according to rdkit in sorted by distance list of possible answers.\n\nI run it on public script submissions:\n**3.63**: https://www.kaggle.com/yingpengchen/pl-bms-molecular-translation\n**3.69**: https://www.kaggle.com/marbury/efficientnet-multi-layer-lstm-inference\n**3.70**: https://www.kaggle.com/dragonzhang/tensorflow-tpu-training-baseline-predictions\nScript without RDKit gives: **2.49** (with RDKit enabled it's only: **2.12**)\n\nAdding\n**4.20**: https://www.kaggle.com/vikrant06/bms-efficientnetv2-tpu-32-epocs-final-lr-e-5\nreduce score to **2.30**.\n\nWith more submissions it must give even better score.",
    "1340564": "Nice work, thanks for sharing. \n\nI did consider training models separately for each InChI component (i.e. to \"specialise\"), but it was too time consuming and I had doubts as to its effectiveness. Ultimately your solution was a much better approach to exploit the fact that the InChI are given as components!",
    "1345185": "This was our main Ensembling method. and to be honest we didn't have a single model scoring below 1.00 but this ensembling pull us to 0.65 in the Public LB"
  },
  "source": "meta"
}