{
  "id": 238551,
  "title": "Basic PP that maybe can help",
  "url": "/competitions/bms-molecular-translation/discussion/238551",
  "author_name": "",
  "post_date": "2021-05-12T14:52:16.311027300Z",
  "votes": 12,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Joined today and the training takes time so in the meanwhile I tried creating a PP filter.</p>\n<p>By counting the sum of an unique predicted string in several submission and create a filter you can gain some extra score. I got -0.3 with some of the top public submissions. It also gives you a visually look where to analyze further and customize models, like having 10 different strings in say 10 different models.<br>\nHope it can help.</p>\n<p>Update: Have made the example code public <a href=\"https://www.kaggle.com/kirderf/basic-postp-with-different-submissions\" target=\"_blank\">https://www.kaggle.com/kirderf/basic-postp-with-different-submissions</a></p>",
  "messages": [
    {
      "id": "1304318",
      "postDate": "05/12/2021 14:52:16",
      "content": "<p>Joined today and the training takes time so in the meanwhile I tried creating a PP filter.</p>\n<p>By counting the sum of an unique predicted string in several submission and create a filter you can gain some extra score. I got -0.3 with some of the top public submissions. It also gives you a visually look where to analyze further and customize models, like having 10 different strings in say 10 different models.<br>\nHope it can help.</p>\n<p>Update: Have made the example code public <a href=\"https://www.kaggle.com/kirderf/basic-postp-with-different-submissions\" target=\"_blank\">https://www.kaggle.com/kirderf/basic-postp-with-different-submissions</a></p>",
      "rawMarkdown": "Joined today and the training takes time so in the meanwhile I tried creating a PP filter.\n\nBy counting the sum of an unique predicted string in several submission and create a filter you can gain some extra score. I got -0.3 with some of the top public submissions. It also gives you a visually look where to analyze further and customize models, like having 10 different strings in say 10 different models.\nHope it can help.\n\nUpdate: Have made the example code public https://www.kaggle.com/kirderf/basic-postp-with-different-submissions",
      "votes": null
    },
    {
      "id": "1306253",
      "postDate": "05/13/2021 17:32:43",
      "content": "<p><a href=\"https://www.kaggle.com/kirderf\" target=\"_blank\">@kirderf</a> - What do you mean check sum of uq. string. Can you elaborate a little?</p>",
      "rawMarkdown": "kirderf - What do you mean check sum of uq. string. Can you elaborate a little?",
      "votes": null
    },
    {
      "id": "1306340",
      "postDate": "05/13/2021 17:50:07",
      "content": "<p>Sum every row the no times an unique string are in different submissions, 0 it may not be a great string, but instead all of them, that's a string to keep. And also, is it in 90% of the subs but not in your best final sub, you maybe should switch that row.</p>",
      "rawMarkdown": "Sum every row the no times an unique string are in different submissions, 0 it may not be a great string, but instead all of them, that's a string to keep. And also, is it in 90% of the subs but not in your best final sub, you maybe should switch that row.",
      "votes": null
    },
    {
      "id": "1307332",
      "postDate": "05/14/2021 11:33:40",
      "content": "<p>Have made the example code  public <a href=\"https://www.kaggle.com/kirderf/basic-postp-with-different-submissions\" target=\"_blank\">https://www.kaggle.com/kirderf/basic-postp-with-different-submissions</a></p>",
      "rawMarkdown": "Have made the example code  public https://www.kaggle.com/kirderf/basic-postp-with-different-submissions",
      "votes": null
    },
    {
      "id": "1308058",
      "postDate": "05/14/2021 23:39:31",
      "content": "<p>This is an easy way of combining diverse models. Thanks for sharing. </p>\n<p>(But competitors who choose to use it should be careful that it doesn't turn into LB probing.)</p>\n<p>When I checked the top five public submissions (LB 4.94 - 8.90), I found 9706 InChI which are the same in all five submissions, but different to the InChI predicted by my best single model (LB &lt; 2.0). </p>",
      "rawMarkdown": "This is an easy way of combining diverse models. Thanks for sharing. \n\n(But competitors who choose to use it should be careful that it doesn't turn into LB probing.)\n\nWhen I checked the top five public submissions (LB 4.94 - 8.90), I found 9706 InChI which are the same in all five submissions, but different to the InChI predicted by my best single model (LB < 2.0).",
      "votes": null
    },
    {
      "id": "1308578",
      "postDate": "05/15/2021 10:08:01",
      "content": "<p>Yes good point. All ensembling, voting, bagging etc are at risk of overfitting the LB if one just validates on the LB not local CV and OOF. <br>\nI can imagine like you said that this prediction differs from your much better score, especially here with over 1M prediction rows, but this is just a basic starter code. But what the code does and the purpose of it, is to first combine submissions with scores in the same region and find rows where the second-best submission has more in common with other predictions compared to the best one. When switching those rows, the score gets better, in this case, can be a coincidence but it might give some help too.</p>",
      "rawMarkdown": "Yes good point. All ensembling, voting, bagging etc are at risk of overfitting the LB if one just validates on the LB not local CV and OOF. \nI can imagine like you said that this prediction differs from your much better score, especially here with over 1M prediction rows, but this is just a basic starter code. But what the code does and the purpose of it, is to first combine submissions with scores in the same region and find rows where the second-best submission has more in common with other predictions compared to the best one. When switching those rows, the score gets better, in this case, can be a coincidence but it might give some help too.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1306253,
      "author_name": "pheadrus",
      "author_url": "",
      "post_date": "05/13/2021 17:32:43",
      "content": "<p><a href=\"https://www.kaggle.com/kirderf\" target=\"_blank\">@kirderf</a> - What do you mean check sum of uq. string. Can you elaborate a little?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1306340,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "05/13/2021 17:50:07",
          "content": "<p>Sum every row the no times an unique string are in different submissions, 0 it may not be a great string, but instead all of them, that's a string to keep. And also, is it in 90% of the subs but not in your best final sub, you maybe should switch that row.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1307332,
      "author_name": "kirderf",
      "author_url": "",
      "post_date": "05/14/2021 11:33:40",
      "content": "<p>Have made the example code  public <a href=\"https://www.kaggle.com/kirderf/basic-postp-with-different-submissions\" target=\"_blank\">https://www.kaggle.com/kirderf/basic-postp-with-different-submissions</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1308058,
      "author_name": "talktocharles",
      "author_url": "",
      "post_date": "05/14/2021 23:39:31",
      "content": "<p>This is an easy way of combining diverse models. Thanks for sharing. </p>\n<p>(But competitors who choose to use it should be careful that it doesn't turn into LB probing.)</p>\n<p>When I checked the top five public submissions (LB 4.94 - 8.90), I found 9706 InChI which are the same in all five submissions, but different to the InChI predicted by my best single model (LB &lt; 2.0). </p>",
      "votes": null,
      "replies": [
        {
          "id": 1308578,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "05/15/2021 10:08:01",
          "content": "<p>Yes good point. All ensembling, voting, bagging etc are at risk of overfitting the LB if one just validates on the LB not local CV and OOF. <br>\nI can imagine like you said that this prediction differs from your much better score, especially here with over 1M prediction rows, but this is just a basic starter code. But what the code does and the purpose of it, is to first combine submissions with scores in the same region and find rows where the second-best submission has more in common with other predictions compared to the best one. When switching those rows, the score gets better, in this case, can be a coincidence but it might give some help too.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1304318": "Joined today and the training takes time so in the meanwhile I tried creating a PP filter.\n\nBy counting the sum of an unique predicted string in several submission and create a filter you can gain some extra score. I got -0.3 with some of the top public submissions. It also gives you a visually look where to analyze further and customize models, like having 10 different strings in say 10 different models.\nHope it can help.\n\nUpdate: Have made the example code public https://www.kaggle.com/kirderf/basic-postp-with-different-submissions",
    "1306253": "kirderf - What do you mean check sum of uq. string. Can you elaborate a little?",
    "1306340": "Sum every row the no times an unique string are in different submissions, 0 it may not be a great string, but instead all of them, that's a string to keep. And also, is it in 90% of the subs but not in your best final sub, you maybe should switch that row.",
    "1307332": "Have made the example code  public https://www.kaggle.com/kirderf/basic-postp-with-different-submissions",
    "1308058": "This is an easy way of combining diverse models. Thanks for sharing. \n\n(But competitors who choose to use it should be careful that it doesn't turn into LB probing.)\n\nWhen I checked the top five public submissions (LB 4.94 - 8.90), I found 9706 InChI which are the same in all five submissions, but different to the InChI predicted by my best single model (LB < 2.0).",
    "1308578": "Yes good point. All ensembling, voting, bagging etc are at risk of overfitting the LB if one just validates on the LB not local CV and OOF. \nI can imagine like you said that this prediction differs from your much better score, especially here with over 1M prediction rows, but this is just a basic starter code. But what the code does and the purpose of it, is to first combine submissions with scores in the same region and find rows where the second-best submission has more in common with other predictions compared to the best one. When switching those rows, the score gets better, in this case, can be a coincidence but it might give some help too."
  },
  "source": "meta"
}