{
  "id": 305986,
  "title": "1% on public LB - is this a record? ",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/305986",
  "author_name": "",
  "post_date": "2022-02-07T18:43:27.453044400Z",
  "votes": 48,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I don't know if it's a mistake or not, but the public leaderboard accounts for only 1% of the total test set. Trusting a CV instead of LB often turned out to be a good lead, but in this competition it will be a necessity. What are your thoughts?</p>",
  "messages": [
    {
      "id": "1680358",
      "postDate": "02/07/2022 18:43:27",
      "content": "<p>I don't know if it's a mistake or not, but the public leaderboard accounts for only 1% of the total test set. Trusting a CV instead of LB often turned out to be a good lead, but in this competition it will be a necessity. What are your thoughts?</p>",
      "rawMarkdown": "I don't know if it's a mistake or not, but the public leaderboard accounts for only 1% of the total test set. Trusting a CV instead of LB often turned out to be a good lead, but in this competition it will be a necessity. What are your thoughts?",
      "votes": null
    },
    {
      "id": "1680368",
      "postDate": "02/07/2022 18:51:35",
      "content": "<p>For this reason, I would like to share you with this amazing meme 😄 <img src=\"https://i.ibb.co/7KVN74g/05xlvswuhfg41.jpg\" alt=\"05xlvswuhfg41\"></p>",
      "rawMarkdown": "For this reason, I would like to share you with this amazing meme 😄 <img src=\"https://i.ibb.co/7KVN74g/05xlvswuhfg41.jpg\" alt=\"05xlvswuhfg41\" border=\"0\">",
      "votes": null
    },
    {
      "id": "1680385",
      "postDate": "02/07/2022 19:04:10",
      "content": "<p>With 1% data, I wonder if the public LB would even have a meaning!</p>",
      "rawMarkdown": "With 1% data, I wonder if the public LB would even have a meaning!",
      "votes": null
    },
    {
      "id": "1680404",
      "postDate": "02/07/2022 19:23:01",
      "content": "<p>The 1% is an artifact of ignoring the rows that don't have sales in the test time period. You can safely assume the the actual Public leaderboard is calculated from 5-15% of the <em>scored</em> rows.</p>",
      "rawMarkdown": "The 1% is an artifact of ignoring the rows that don't have sales in the test time period. You can safely assume the the actual Public leaderboard is calculated from 5-15% of the *scored* rows.",
      "votes": null
    },
    {
      "id": "1680410",
      "postDate": "02/07/2022 19:27:41",
      "content": "<p>Thanks for the information, this is relieving. <br>\n5-15% ? <br>\nIs there a reason for not sharing the actual value?</p>",
      "rawMarkdown": "Thanks for the information, this is relieving. \n5-15% ? \nIs there a reason for not sharing the actual value?",
      "votes": null
    },
    {
      "id": "1680822",
      "postDate": "02/08/2022 03:35:58",
      "content": "<p>Look how 5% public LB tuns out: <a href=\"https://www.kaggle.com/c/jigsaw-toxic-severity-rating/leaderboard?tab=private\" target=\"_blank\">Jigsaw Competition</a></p>\n<p>Even 5% was unreliable. </p>",
      "rawMarkdown": "Look how 5% public LB tuns out: [Jigsaw Competition](https://www.kaggle.com/c/jigsaw-toxic-severity-rating/leaderboard?tab=private)\n\nEven 5% was unreliable.",
      "votes": null
    },
    {
      "id": "1681046",
      "postDate": "02/08/2022 07:35:09",
      "content": "<p>It means that only about 7% ~ 20% of customers will purchase within next 7 days after training period.<br>\nAm I understanding correctly?</p>",
      "rawMarkdown": "It means that only about 7% ~ 20% of customers will purchase within next 7 days after training period.\nAm I understanding correctly?",
      "votes": null
    },
    {
      "id": "1681630",
      "postDate": "02/08/2022 15:43:53",
      "content": "<p>The last <a href=\" https://www.kaggle.com/c/jigsaw-toxic-severity-rating/leaderboard?tab=private\">NLP</a> competition has only 5% test data for the public Leaderboard.<br>\nThen Bammm on Private leaderboard(1874---&gt;1 )position.</p>\n<p><a href=\"https://postimg.cc/Ty4cj4r1\" target=\"_blank\"><img src=\"https://i.postimg.cc/gJYBp9jV/PRIZE.png\" alt=\"PRIZE.png\"></a></p>",
      "rawMarkdown": "The last <a href=\" https://www.kaggle.com/c/jigsaw-toxic-severity-rating/leaderboard?tab=private\">NLP</a> competition has only 5% test data for the public Leaderboard.\nThen Bammm on Private leaderboard(1874--->1 )position.\n\n\n[![PRIZE.png](https://i.postimg.cc/gJYBp9jV/PRIZE.png)](https://postimg.cc/Ty4cj4r1)",
      "votes": null
    },
    {
      "id": "1682605",
      "postDate": "02/09/2022 09:21:14",
      "content": "<p>One can use the public lb as an additional validation not the only one, that is if no leak is in it from training, best if all the validations used correlates. See it as an extra OOF to the CV validation, one cannot have enough validations 😊 Primary is the CV validation from the training, or/and if one can use a larger validation set, as in Jigsaw the whole training set could be used as validation if the model was trained on another toxic data.<br>\nAll in all to have different conditions of the validation data, to not overfit a small part that the model finds perfect.</p>\n<p>Having said that, this all easy in theory, harder in practice :) It's the battle in the competition and choices of final submissions every time one competes :)</p>",
      "rawMarkdown": "One can use the public lb as an additional validation not the only one, that is if no leak is in it from training, best if all the validations used correlates. See it as an extra OOF to the CV validation, one cannot have enough validations 😊 Primary is the CV validation from the training, or/and if one can use a larger validation set, as in Jigsaw the whole training set could be used as validation if the model was trained on another toxic data.\nAll in all to have different conditions of the validation data, to not overfit a small part that the model finds perfect.\n\nHaving said that, this all easy in theory, harder in practice :) It's the battle in the competition and choices of final submissions every time one competes :)",
      "votes": null
    },
    {
      "id": "1684620",
      "postDate": "02/10/2022 15:55:53",
      "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> - I think it's to prevent us from more accurately figuring out <strong>what % of customers had sales in evaluation period</strong>.<br>\nWith the info they've provided, we can estimate <strong>it's between 3.5% and 30%</strong>.</p>\n<p><a href=\"https://www.kaggle.com/ryotak12\" target=\"_blank\">@ryotak12</a> - Your calculations assume the 1% hasn't been rounded. <br>\nIf we don't make that assumption, it can be between 3.5% and 30% (see calculations below)</p>\n<p>(For comparison, ~5% of customers had sales during last 7 days of training period)</p>\n<h1>Calculations</h1>\n<p>let PCS = percent of customers with sales during evaluation period<br>\nlet PSP = percent of customers with sales included in the Public LB set<br>\nlet PCP = percent of customers included in the Public LB set</p>\n<p>So \\(PCP = PCS * PSP\\)</p>\n<p>The formula for percent_of_customers_with_sales is:<br>\n\\(PCB = \\frac{PCP}{PSP}\\)</p>\n<p>We know that:<br>\n\\( 5\\% &lt;= PSP &lt;= 15\\% \\)<br>\n\\( .51\\% &lt;= PCP &lt;= 1.49\\% \\)</p>\n<p>It follows that:<br>\n\\( (\\,\\frac{.51\\%}{15\\%} = 3.5\\% )\\,  &lt;= PCS &lt;=  (\\,\\frac{1.49\\%}{5\\%} = 30\\%)\\, \\)</p>\n<p>If we assume the 1% is exact, and has not been rounded, then:<br>\n\\( (\\,\\frac{1\\%}{15\\%} = 6.5\\% )\\,  &lt;= PCS &lt;=  (\\,\\frac{1\\%}{5\\%} = 20\\%)\\, \\)</p>",
      "rawMarkdown": "ilu000 - I think it's to prevent us from more accurately figuring out **what % of customers had sales in evaluation period**.\nWith the info they've provided, we can estimate **it's between 3.5% and 30%**.\n\n\n@ryotak12 - Your calculations assume the 1% hasn't been rounded. \nIf we don't make that assumption, it can be between 3.5% and 30% (see calculations below)\n\n(For comparison, ~5% of customers had sales during last 7 days of training period)\n\n\n# Calculations\nlet PCS = percent of customers with sales during evaluation period\nlet PSP = percent of customers with sales included in the Public LB set\nlet PCP = percent of customers included in the Public LB set\n\nSo \\\\(PCP = PCS * PSP\\\\)\n\nThe formula for percent_of_customers_with_sales is:\n\\\\(PCB = \\frac{PCP}{PSP}\\\\)\n\nWe know that:\n\\\\( 5\\% <= PSP <= 15\\% \\\\)\n\\\\( .51\\% <= PCP <= 1.49\\% \\\\)\n\nIt follows that:\n\\\\( (\\,\\frac{.51\\%}{15\\%} = 3.5\\% )\\,  <= PCS <=  (\\,\\frac{1.49\\%}{5\\%} = 30\\%)\\, \\\\)\n\nIf we assume the 1% is exact, and has not been rounded, then:\n\\\\( (\\,\\frac{1\\%}{15\\%} = 6.5\\% )\\,  <= PCS <=  (\\,\\frac{1\\%}{5\\%} = 20\\%)\\, \\\\)",
      "votes": null
    },
    {
      "id": "1684960",
      "postDate": "02/10/2022 22:21:09",
      "content": "<p>Basically, public LB and private LB are much different. Your model may not be general enough to converge the training set.</p>",
      "rawMarkdown": "Basically, public LB and private LB are much different. Your model may not be general enough to converge the training set.",
      "votes": null
    },
    {
      "id": "1749358",
      "postDate": "04/08/2022 14:00:11",
      "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> so is it possible that there are customer_ids (e.g. <code>7034c909b5f39e6b184fe36b9edda3bf62b26509222ad6eb2744b3e1f8c02c2e</code>) which make a purchase neither during the train period, nor during the evaluation period?</p>",
      "rawMarkdown": "inversion so is it possible that there are customer_ids (e.g. `7034c909b5f39e6b184fe36b9edda3bf62b26509222ad6eb2744b3e1f8c02c2e`) which make a purchase neither during the train period, nor during the evaluation period?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1680368,
      "author_name": "vad13irt",
      "author_url": "",
      "post_date": "02/07/2022 18:51:35",
      "content": "<p>For this reason, I would like to share you with this amazing meme 😄 <img src=\"https://i.ibb.co/7KVN74g/05xlvswuhfg41.jpg\" alt=\"05xlvswuhfg41\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1680385,
      "author_name": "monideepde",
      "author_url": "",
      "post_date": "02/07/2022 19:04:10",
      "content": "<p>With 1% data, I wonder if the public LB would even have a meaning!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1680404,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "02/07/2022 19:23:01",
      "content": "<p>The 1% is an artifact of ignoring the rows that don't have sales in the test time period. You can safely assume the the actual Public leaderboard is calculated from 5-15% of the <em>scored</em> rows.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1680410,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "02/07/2022 19:27:41",
          "content": "<p>Thanks for the information, this is relieving. <br>\n5-15% ? <br>\nIs there a reason for not sharing the actual value?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1681046,
          "author_name": "ryotak12",
          "author_url": "",
          "post_date": "02/08/2022 07:35:09",
          "content": "<p>It means that only about 7% ~ 20% of customers will purchase within next 7 days after training period.<br>\nAm I understanding correctly?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1684620,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "02/10/2022 15:55:53",
          "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> - I think it's to prevent us from more accurately figuring out <strong>what % of customers had sales in evaluation period</strong>.<br>\nWith the info they've provided, we can estimate <strong>it's between 3.5% and 30%</strong>.</p>\n<p><a href=\"https://www.kaggle.com/ryotak12\" target=\"_blank\">@ryotak12</a> - Your calculations assume the 1% hasn't been rounded. <br>\nIf we don't make that assumption, it can be between 3.5% and 30% (see calculations below)</p>\n<p>(For comparison, ~5% of customers had sales during last 7 days of training period)</p>\n<h1>Calculations</h1>\n<p>let PCS = percent of customers with sales during evaluation period<br>\nlet PSP = percent of customers with sales included in the Public LB set<br>\nlet PCP = percent of customers included in the Public LB set</p>\n<p>So \\(PCP = PCS * PSP\\)</p>\n<p>The formula for percent_of_customers_with_sales is:<br>\n\\(PCB = \\frac{PCP}{PSP}\\)</p>\n<p>We know that:<br>\n\\( 5\\% &lt;= PSP &lt;= 15\\% \\)<br>\n\\( .51\\% &lt;= PCP &lt;= 1.49\\% \\)</p>\n<p>It follows that:<br>\n\\( (\\,\\frac{.51\\%}{15\\%} = 3.5\\% )\\,  &lt;= PCS &lt;=  (\\,\\frac{1.49\\%}{5\\%} = 30\\%)\\, \\)</p>\n<p>If we assume the 1% is exact, and has not been rounded, then:<br>\n\\( (\\,\\frac{1\\%}{15\\%} = 6.5\\% )\\,  &lt;= PCS &lt;=  (\\,\\frac{1\\%}{5\\%} = 20\\%)\\, \\)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1749358,
          "author_name": "marcogorelli",
          "author_url": "",
          "post_date": "04/08/2022 14:00:11",
          "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> so is it possible that there are customer_ids (e.g. <code>7034c909b5f39e6b184fe36b9edda3bf62b26509222ad6eb2744b3e1f8c02c2e</code>) which make a purchase neither during the train period, nor during the evaluation period?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1680822,
      "author_name": "ympaik",
      "author_url": "",
      "post_date": "02/08/2022 03:35:58",
      "content": "<p>Look how 5% public LB tuns out: <a href=\"https://www.kaggle.com/c/jigsaw-toxic-severity-rating/leaderboard?tab=private\" target=\"_blank\">Jigsaw Competition</a></p>\n<p>Even 5% was unreliable. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1682605,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "02/09/2022 09:21:14",
          "content": "<p>One can use the public lb as an additional validation not the only one, that is if no leak is in it from training, best if all the validations used correlates. See it as an extra OOF to the CV validation, one cannot have enough validations 😊 Primary is the CV validation from the training, or/and if one can use a larger validation set, as in Jigsaw the whole training set could be used as validation if the model was trained on another toxic data.<br>\nAll in all to have different conditions of the validation data, to not overfit a small part that the model finds perfect.</p>\n<p>Having said that, this all easy in theory, harder in practice :) It's the battle in the competition and choices of final submissions every time one competes :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1681630,
      "author_name": "gazu468",
      "author_url": "",
      "post_date": "02/08/2022 15:43:53",
      "content": "<p>The last <a href=\" https://www.kaggle.com/c/jigsaw-toxic-severity-rating/leaderboard?tab=private\">NLP</a> competition has only 5% test data for the public Leaderboard.<br>\nThen Bammm on Private leaderboard(1874---&gt;1 )position.</p>\n<p><a href=\"https://postimg.cc/Ty4cj4r1\" target=\"_blank\"><img src=\"https://i.postimg.cc/gJYBp9jV/PRIZE.png\" alt=\"PRIZE.png\"></a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1684960,
      "author_name": "nebipeker",
      "author_url": "",
      "post_date": "02/10/2022 22:21:09",
      "content": "<p>Basically, public LB and private LB are much different. Your model may not be general enough to converge the training set.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1680358": "I don't know if it's a mistake or not, but the public leaderboard accounts for only 1% of the total test set. Trusting a CV instead of LB often turned out to be a good lead, but in this competition it will be a necessity. What are your thoughts?",
    "1680368": "For this reason, I would like to share you with this amazing meme 😄 <img src=\"https://i.ibb.co/7KVN74g/05xlvswuhfg41.jpg\" alt=\"05xlvswuhfg41\" border=\"0\">",
    "1680385": "With 1% data, I wonder if the public LB would even have a meaning!",
    "1680404": "The 1% is an artifact of ignoring the rows that don't have sales in the test time period. You can safely assume the the actual Public leaderboard is calculated from 5-15% of the *scored* rows.",
    "1680410": "Thanks for the information, this is relieving. \n5-15% ? \nIs there a reason for not sharing the actual value?",
    "1680822": "Look how 5% public LB tuns out: [Jigsaw Competition](https://www.kaggle.com/c/jigsaw-toxic-severity-rating/leaderboard?tab=private)\n\nEven 5% was unreliable.",
    "1681046": "It means that only about 7% ~ 20% of customers will purchase within next 7 days after training period.\nAm I understanding correctly?",
    "1681630": "The last <a href=\" https://www.kaggle.com/c/jigsaw-toxic-severity-rating/leaderboard?tab=private\">NLP</a> competition has only 5% test data for the public Leaderboard.\nThen Bammm on Private leaderboard(1874--->1 )position.\n\n\n[![PRIZE.png](https://i.postimg.cc/gJYBp9jV/PRIZE.png)](https://postimg.cc/Ty4cj4r1)",
    "1682605": "One can use the public lb as an additional validation not the only one, that is if no leak is in it from training, best if all the validations used correlates. See it as an extra OOF to the CV validation, one cannot have enough validations 😊 Primary is the CV validation from the training, or/and if one can use a larger validation set, as in Jigsaw the whole training set could be used as validation if the model was trained on another toxic data.\nAll in all to have different conditions of the validation data, to not overfit a small part that the model finds perfect.\n\nHaving said that, this all easy in theory, harder in practice :) It's the battle in the competition and choices of final submissions every time one competes :)",
    "1684620": "ilu000 - I think it's to prevent us from more accurately figuring out **what % of customers had sales in evaluation period**.\nWith the info they've provided, we can estimate **it's between 3.5% and 30%**.\n\n\n@ryotak12 - Your calculations assume the 1% hasn't been rounded. \nIf we don't make that assumption, it can be between 3.5% and 30% (see calculations below)\n\n(For comparison, ~5% of customers had sales during last 7 days of training period)\n\n\n# Calculations\nlet PCS = percent of customers with sales during evaluation period\nlet PSP = percent of customers with sales included in the Public LB set\nlet PCP = percent of customers included in the Public LB set\n\nSo \\\\(PCP = PCS * PSP\\\\)\n\nThe formula for percent_of_customers_with_sales is:\n\\\\(PCB = \\frac{PCP}{PSP}\\\\)\n\nWe know that:\n\\\\( 5\\% <= PSP <= 15\\% \\\\)\n\\\\( .51\\% <= PCP <= 1.49\\% \\\\)\n\nIt follows that:\n\\\\( (\\,\\frac{.51\\%}{15\\%} = 3.5\\% )\\,  <= PCS <=  (\\,\\frac{1.49\\%}{5\\%} = 30\\%)\\, \\\\)\n\nIf we assume the 1% is exact, and has not been rounded, then:\n\\\\( (\\,\\frac{1\\%}{15\\%} = 6.5\\% )\\,  <= PCS <=  (\\,\\frac{1\\%}{5\\%} = 20\\%)\\, \\\\)",
    "1684960": "Basically, public LB and private LB are much different. Your model may not be general enough to converge the training set.",
    "1749358": "inversion so is it possible that there are customer_ids (e.g. `7034c909b5f39e6b184fe36b9edda3bf62b26509222ad6eb2744b3e1f8c02c2e`) which make a purchase neither during the train period, nor during the evaluation period?"
  },
  "source": "meta"
}