{
  "id": 252625,
  "title": "Optimizing the two submissions",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/252625",
  "author_name": "",
  "post_date": "2021-07-13T06:59:56.672993800Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I would like to discuss ideas and/or knowledge from other competitions on how to optimize the two submissions allowed for this competition.</p>\n<p>The first idea that comes to my mind is creating a MOE for one or more of the important features and try to make the two submissions as uncorrelated as possible while maintaining the signals we get from the feature(s).</p>",
  "messages": [
    {
      "id": "1386009",
      "postDate": "07/13/2021 06:59:56",
      "content": "<p>I would like to discuss ideas and/or knowledge from other competitions on how to optimize the two submissions allowed for this competition.</p>\n<p>The first idea that comes to my mind is creating a MOE for one or more of the important features and try to make the two submissions as uncorrelated as possible while maintaining the signals we get from the feature(s).</p>",
      "rawMarkdown": "I would like to discuss ideas and/or knowledge from other competitions on how to optimize the two submissions allowed for this competition.\n\nThe first idea that comes to my mind is creating a MOE for one or more of the important features and try to make the two submissions as uncorrelated as possible while maintaining the signals we get from the feature(s).",
      "votes": null
    },
    {
      "id": "1394589",
      "postDate": "07/20/2021 12:59:37",
      "content": "<p><a href=\"https://www.kaggle.com/kaito510\" target=\"_blank\">@kaito510</a>, I think that will be one of the major factors in this competition. Some comments:</p>\n<ol>\n<li><p>If you think that the score on the public LB is a good reflection of the likely performance in evaluation on the private LB, then you will probably pick at least one submission on the basis of its visible score. If however you think the top scoring public LB codes are overfitted, then you might not worry about the public score.<br>\n<br></p></li>\n<li><p>You may want to pick two submissions that are quite different from one another, to give you two lottery tickets rather than one </p></li>\n<li><p>You might think about choosing at least one submission that is quite different from what you expect other people to choose. That gives you a region of the competition space to yourself and if you do get lucky, then you should get a gold.</p></li>\n<li><p>There is a risk in this competition of some notebooks throwing errors in evaluation. You might mitigate this by choosing at least one safe and simple option.</p></li>\n</ol>",
      "rawMarkdown": "kaito510, I think that will be one of the major factors in this competition. Some comments:\n\n1. If you think that the score on the public LB is a good reflection of the likely performance in evaluation on the private LB, then you will probably pick at least one submission on the basis of its visible score. If however you think the top scoring public LB codes are overfitted, then you might not worry about the public score.\n<br>\n2. You may want to pick two submissions that are quite different from one another, to give you two lottery tickets rather than one \n\n3. You might think about choosing at least one submission that is quite different from what you expect other people to choose. That gives you a region of the competition space to yourself and if you do get lucky, then you should get a gold.\n\n4. There is a risk in this competition of some notebooks throwing errors in evaluation. You might mitigate this by choosing at least one safe and simple option.",
      "votes": null
    },
    {
      "id": "1394608",
      "postDate": "07/20/2021 13:13:12",
      "content": "<p>Absolutely to point 4 - \"There is a risk in this competition of some notebooks throwing errors in evaluation.\"<br>\nYou may want to see what the July 20, 2021 (estimated date) - Training Set Update data looks like and how your models work with that. <br>\nFrom past competitions like these with future data, there is always something that comes up and having at least one submission that is unlikely to go over time limits or exceed memory or not handle unseen categorical codes, etc. means you will still be in the game.<br>\nAnd now that it has changed so if a notebook gets a failed cell it will stop execution and be marked as Fail/Error <a href=\"https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/249900\" target=\"_blank\">see</a>, think it is highly likely people may get failed submissions. </p>",
      "rawMarkdown": "Absolutely to point 4 - \"There is a risk in this competition of some notebooks throwing errors in evaluation.\"\nYou may want to see what the July 20, 2021 (estimated date) - Training Set Update data looks like and how your models work with that. \nFrom past competitions like these with future data, there is always something that comes up and having at least one submission that is unlikely to go over time limits or exceed memory or not handle unseen categorical codes, etc. means you will still be in the game.\nAnd now that it has changed so if a notebook gets a failed cell it will stop execution and be marked as Fail/Error [see](https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/249900), think it is highly likely people may get failed submissions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1394589,
      "author_name": "jbomitchell",
      "author_url": "",
      "post_date": "07/20/2021 12:59:37",
      "content": "<p><a href=\"https://www.kaggle.com/kaito510\" target=\"_blank\">@kaito510</a>, I think that will be one of the major factors in this competition. Some comments:</p>\n<ol>\n<li><p>If you think that the score on the public LB is a good reflection of the likely performance in evaluation on the private LB, then you will probably pick at least one submission on the basis of its visible score. If however you think the top scoring public LB codes are overfitted, then you might not worry about the public score.<br>\n<br></p></li>\n<li><p>You may want to pick two submissions that are quite different from one another, to give you two lottery tickets rather than one </p></li>\n<li><p>You might think about choosing at least one submission that is quite different from what you expect other people to choose. That gives you a region of the competition space to yourself and if you do get lucky, then you should get a gold.</p></li>\n<li><p>There is a risk in this competition of some notebooks throwing errors in evaluation. You might mitigate this by choosing at least one safe and simple option.</p></li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1394608,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "07/20/2021 13:13:12",
          "content": "<p>Absolutely to point 4 - \"There is a risk in this competition of some notebooks throwing errors in evaluation.\"<br>\nYou may want to see what the July 20, 2021 (estimated date) - Training Set Update data looks like and how your models work with that. <br>\nFrom past competitions like these with future data, there is always something that comes up and having at least one submission that is unlikely to go over time limits or exceed memory or not handle unseen categorical codes, etc. means you will still be in the game.<br>\nAnd now that it has changed so if a notebook gets a failed cell it will stop execution and be marked as Fail/Error <a href=\"https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/249900\" target=\"_blank\">see</a>, think it is highly likely people may get failed submissions. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1386009": "I would like to discuss ideas and/or knowledge from other competitions on how to optimize the two submissions allowed for this competition.\n\nThe first idea that comes to my mind is creating a MOE for one or more of the important features and try to make the two submissions as uncorrelated as possible while maintaining the signals we get from the feature(s).",
    "1394589": "kaito510, I think that will be one of the major factors in this competition. Some comments:\n\n1. If you think that the score on the public LB is a good reflection of the likely performance in evaluation on the private LB, then you will probably pick at least one submission on the basis of its visible score. If however you think the top scoring public LB codes are overfitted, then you might not worry about the public score.\n<br>\n2. You may want to pick two submissions that are quite different from one another, to give you two lottery tickets rather than one \n\n3. You might think about choosing at least one submission that is quite different from what you expect other people to choose. That gives you a region of the competition space to yourself and if you do get lucky, then you should get a gold.\n\n4. There is a risk in this competition of some notebooks throwing errors in evaluation. You might mitigate this by choosing at least one safe and simple option.",
    "1394608": "Absolutely to point 4 - \"There is a risk in this competition of some notebooks throwing errors in evaluation.\"\nYou may want to see what the July 20, 2021 (estimated date) - Training Set Update data looks like and how your models work with that. \nFrom past competitions like these with future data, there is always something that comes up and having at least one submission that is unlikely to go over time limits or exceed memory or not handle unseen categorical codes, etc. means you will still be in the game.\nAnd now that it has changed so if a notebook gets a failed cell it will stop execution and be marked as Fail/Error [see](https://www.kaggle.com/c/mlb-player-digital-engagement-forecasting/discussion/249900), think it is highly likely people may get failed submissions."
  },
  "source": "meta"
}