{
  "id": 457861,
  "title": "If the LB scores keep improving at this rate, can I use the public blend as target to train the model?",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/457861",
  "author_name": "",
  "post_date": "2023-11-27T07:34:58.324033900Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>As the title suggests… </p>",
  "messages": [
    {
      "id": "2539606",
      "postDate": "11/27/2023 07:34:58",
      "content": "<p>As the title suggests… </p>",
      "rawMarkdown": "As the title suggests...",
      "votes": null
    },
    {
      "id": "2539622",
      "postDate": "11/27/2023 07:57:01",
      "content": "<p>Pseudolabelling makes sense if you could know which part is public. Otherwise your model may predict something similar to the public kernel for the private part, which you may want to avoid.</p>",
      "rawMarkdown": "Pseudolabelling makes sense if you could know which part is public. Otherwise your model may predict something similar to the public kernel for the private part, which you may want to avoid.",
      "votes": null
    },
    {
      "id": "2539673",
      "postDate": "11/27/2023 08:56:43",
      "content": "<p>Exactly the thing I am asking myself ! <br>\nI think - NO (but not sure), because: <br>\nalthough the rate is impressive the relative change is small - \"normal\" publics are about 0.565+, now we have 0.531,<br>\nso the error is still quite big.</p>\n<p>So we can make such a local experiment:<br>\nsplit local train - test:<br>\ntake some labels from the local train, and distort them such that error would be about 0.565 and 0.531<br>\ntrain on such two variatns of labels<br>\nand compare the results on local test<br>\nI guess the difference will be very small</p>\n<p>If I have more time - I will try to do that, but not sure. </p>\n<p>PS<br>\n<a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> <br>\nthe split to public and private is given to us - see the image on the \"overview\" page.<br>\nThese public notebooks seems to work only on public LB and most probably fall down on private,<br>\nalthouth there is small chance that some ideas from them might be useful - have not thought on that much…</p>",
      "rawMarkdown": "Exactly the thing I am asking myself ! \nI think - NO (but not sure), because: \nalthough the rate is impressive the relative change is small - \"normal\" publics are about 0.565+, now we have 0.531,\nso the error is still quite big.\n\nSo we can make such a local experiment:\nsplit local train - test:\ntake some labels from the local train, and distort them such that error would be about 0.565 and 0.531\ntrain on such two variatns of labels\nand compare the results on local test\nI guess the difference will be very small\n\nIf I have more time - I will try to do that, but not sure. \n\nPS\n@aerdem4 \nthe split to public and private is given to us - see the image on the \"overview\" page.\nThese public notebooks seems to work only on public LB and most probably fall down on private,\nalthouth there is small chance that some ideas from them might be useful - have not thought on that much...",
      "votes": null
    },
    {
      "id": "2539684",
      "postDate": "11/27/2023 09:06:28",
      "content": "<h2>The Model 1.  Predict constant_value (600:1)  .</h2>\n<h5>Algorithm 1:</h5>\n<p>input: (1,18116)  <br>\noutput: (1,1)   the \"constant_value\".</p>\n<p>The constant that minimizes the RMSE for SM_name.</p>\n<h6>Y = (C0 - x) * * 2 + …. + (C18116 - x) * * 2</h6>\n<p>The constant_value is the value of x that minimizes the function Y.</p>\n<p>Input data: <br>\n The constant_value train 600 rows:  rmse_LB_sort_TRAIN_whide2_lb.csv <br>\n The constant_value submit 100 rows (remove constant_value submit   77*2 rows - forecast is weak.).</p>\n<p><a href=\"https://www.kaggle.com/code/olegpush/op2-eda-lb?scriptVersionId=152467205\" target=\"_blank\">https://www.kaggle.com/code/olegpush/op2-eda-lb?scriptVersionId=152467205</a><br>\nthe output of version 38. </p>\n<h2>The Model 2.  Predict  targets  (600:50) PyBoost .</h2>\n<p>It is important to ensure that model has not been trained on LB sm_names only.</p>\n<h5>Loss:</h5>\n<p>I calculate the loss for each iteration by:</p>\n<ol>\n<li>Transform (600:50) target to  (600:181116 ) target.</li>\n<li>Calculate loss  for  (1:181116 ) target  CY = (C0 - x) * * 2 + …. + (C18116 - x) * * 2<br>\nfor 600 rows separate.</li>\n<li>Transform (600:181116 ) to (600:50) targets andf return to pyBoost</li>\n</ol>\n<h2>The Bland of Model 1 and Model 2</h2>\n<p>I used Model 1 to \"move up or move down\" Model 2 prediction for sm_name:</p>\n<h5>step1: minimizes the function  Y = (C0 - x) * * 2 + …. + (C18116 - x) * * 2</h5>\n<p>where  C0 -C18116  - prediction of Model2.</p>\n<p>Calculate using Algorithm1 constant_value for  rows 0-255 submition.csv of best public blend.</p>\n<h5>step2:</h5>\n<p>For each row in submition:</p>\n<pre><code> Y_pred = constant_value_Model1  \n               + Y_pred(:)_Model2 \n               - constant_value_Model2 \n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4172517%2Ff206215699c522a7197223c5645ef485%2FIrrrMG_20231127_141535.jpg?generation=1701083967056127&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "## The Model 1.  Predict constant_value (600:1)  .\n\n##### Algorithm 1:\ninput: (1,18116)  \noutput: (1,1)   the \"constant_value\".\n\nThe constant that minimizes the RMSE for SM_name.\n\n###### Y = (C0 - x) * * 2 + .... + (C18116 - x) * * 2\n\nThe constant_value is the value of x that minimizes the function Y.\n\nInput data: \n The constant_value train 600 rows:  rmse_LB_sort_TRAIN_whide2_lb.csv \n The constant_value submit 100 rows (remove constant_value submit   77*2 rows - forecast is weak.).\n\nhttps://www.kaggle.com/code/olegpush/op2-eda-lb?scriptVersionId=152467205\nthe output of version 38. \n\n## The Model 2.  Predict  targets  (600:50) PyBoost .\nIt is important to ensure that model has not been trained on LB sm_names only.\n\n##### Loss:\nI calculate the loss for each iteration by:\n1. Transform (600:50) target to  (600:181116 ) target.\n2. Calculate loss  for  (1:181116 ) target  CY = (C0 - x) * * 2 + …. + (C18116 - x) * * 2\nfor 600 rows separate.\n3. Transform (600:181116 ) to (600:50) targets andf return to pyBoost\n\n## The Bland of Model 1 and Model 2\nI used Model 1 to \"move up or move down\" Model 2 prediction for sm_name:\n\n##### step1: minimizes the function  Y = (C0 - x) * * 2 + …. + (C18116 - x) * * 2 \nwhere  C0 -C18116  - prediction of Model2.\n\nCalculate using Algorithm1 constant_value for  rows 0-255 submition.csv of best public blend.\n\n##### step2:\nFor each row in submition:\n\n```python\n Y_pred = constant_value_Model1  \n               + Y_pred(1:181116)_Model2 \n               - constant_value_Model2 \n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4172517%2Ff206215699c522a7197223c5645ef485%2FIrrrMG_20231127_141535.jpg?generation=1701083967056127&alt=media)",
      "votes": null
    },
    {
      "id": "2539975",
      "postDate": "11/27/2023 12:37:22",
      "content": "<p>I also think that these blends are likely to be useless for the private LB.</p>\n<p>The last two weeks of this competition look like a crowdsourced overfitting project, in places down to the granularity of choosing whichever model gets the lowest score on each item of the public LB data.</p>\n<p>Obviously enough, each of us needs to decide whether to completely ignore these \"blends of blends\", or whether we think there might somehow be a way of using them to improve our private LB score.</p>\n<p>My own opinion is that there was probably a point in the history of this competition up to which the blends had some value, but that we are now well past that and the new developments are just overfitting.</p>\n<p>Other opinions are available.</p>",
      "rawMarkdown": "I also think that these blends are likely to be useless for the private LB.\n\nThe last two weeks of this competition look like a crowdsourced overfitting project, in places down to the granularity of choosing whichever model gets the lowest score on each item of the public LB data.\n\nObviously enough, each of us needs to decide whether to completely ignore these \"blends of blends\", or whether we think there might somehow be a way of using them to improve our private LB score.\n\nMy own opinion is that there was probably a point in the history of this competition up to which the blends had some value, but that we are now well past that and the new developments are just overfitting.\n\nOther opinions are available.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2539622,
      "author_name": "aerdem4",
      "author_url": "",
      "post_date": "11/27/2023 07:57:01",
      "content": "<p>Pseudolabelling makes sense if you could know which part is public. Otherwise your model may predict something similar to the public kernel for the private part, which you may want to avoid.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2539673,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "11/27/2023 08:56:43",
      "content": "<p>Exactly the thing I am asking myself ! <br>\nI think - NO (but not sure), because: <br>\nalthough the rate is impressive the relative change is small - \"normal\" publics are about 0.565+, now we have 0.531,<br>\nso the error is still quite big.</p>\n<p>So we can make such a local experiment:<br>\nsplit local train - test:<br>\ntake some labels from the local train, and distort them such that error would be about 0.565 and 0.531<br>\ntrain on such two variatns of labels<br>\nand compare the results on local test<br>\nI guess the difference will be very small</p>\n<p>If I have more time - I will try to do that, but not sure. </p>\n<p>PS<br>\n<a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> <br>\nthe split to public and private is given to us - see the image on the \"overview\" page.<br>\nThese public notebooks seems to work only on public LB and most probably fall down on private,<br>\nalthouth there is small chance that some ideas from them might be useful - have not thought on that much…</p>",
      "votes": null,
      "replies": [
        {
          "id": 2539975,
          "author_name": "jbomitchell",
          "author_url": "",
          "post_date": "11/27/2023 12:37:22",
          "content": "<p>I also think that these blends are likely to be useless for the private LB.</p>\n<p>The last two weeks of this competition look like a crowdsourced overfitting project, in places down to the granularity of choosing whichever model gets the lowest score on each item of the public LB data.</p>\n<p>Obviously enough, each of us needs to decide whether to completely ignore these \"blends of blends\", or whether we think there might somehow be a way of using them to improve our private LB score.</p>\n<p>My own opinion is that there was probably a point in the history of this competition up to which the blends had some value, but that we are now well past that and the new developments are just overfitting.</p>\n<p>Other opinions are available.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2539684,
      "author_name": "",
      "author_url": "",
      "post_date": "11/27/2023 09:06:28",
      "content": "<h2>The Model 1.  Predict constant_value (600:1)  .</h2>\n<h5>Algorithm 1:</h5>\n<p>input: (1,18116)  <br>\noutput: (1,1)   the \"constant_value\".</p>\n<p>The constant that minimizes the RMSE for SM_name.</p>\n<h6>Y = (C0 - x) * * 2 + …. + (C18116 - x) * * 2</h6>\n<p>The constant_value is the value of x that minimizes the function Y.</p>\n<p>Input data: <br>\n The constant_value train 600 rows:  rmse_LB_sort_TRAIN_whide2_lb.csv <br>\n The constant_value submit 100 rows (remove constant_value submit   77*2 rows - forecast is weak.).</p>\n<p><a href=\"https://www.kaggle.com/code/olegpush/op2-eda-lb?scriptVersionId=152467205\" target=\"_blank\">https://www.kaggle.com/code/olegpush/op2-eda-lb?scriptVersionId=152467205</a><br>\nthe output of version 38. </p>\n<h2>The Model 2.  Predict  targets  (600:50) PyBoost .</h2>\n<p>It is important to ensure that model has not been trained on LB sm_names only.</p>\n<h5>Loss:</h5>\n<p>I calculate the loss for each iteration by:</p>\n<ol>\n<li>Transform (600:50) target to  (600:181116 ) target.</li>\n<li>Calculate loss  for  (1:181116 ) target  CY = (C0 - x) * * 2 + …. + (C18116 - x) * * 2<br>\nfor 600 rows separate.</li>\n<li>Transform (600:181116 ) to (600:50) targets andf return to pyBoost</li>\n</ol>\n<h2>The Bland of Model 1 and Model 2</h2>\n<p>I used Model 1 to \"move up or move down\" Model 2 prediction for sm_name:</p>\n<h5>step1: minimizes the function  Y = (C0 - x) * * 2 + …. + (C18116 - x) * * 2</h5>\n<p>where  C0 -C18116  - prediction of Model2.</p>\n<p>Calculate using Algorithm1 constant_value for  rows 0-255 submition.csv of best public blend.</p>\n<h5>step2:</h5>\n<p>For each row in submition:</p>\n<pre><code> Y_pred = constant_value_Model1  \n               + Y_pred(:)_Model2 \n               - constant_value_Model2 \n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4172517%2Ff206215699c522a7197223c5645ef485%2FIrrrMG_20231127_141535.jpg?generation=1701083967056127&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2539606": "As the title suggests...",
    "2539622": "Pseudolabelling makes sense if you could know which part is public. Otherwise your model may predict something similar to the public kernel for the private part, which you may want to avoid.",
    "2539673": "Exactly the thing I am asking myself ! \nI think - NO (but not sure), because: \nalthough the rate is impressive the relative change is small - \"normal\" publics are about 0.565+, now we have 0.531,\nso the error is still quite big.\n\nSo we can make such a local experiment:\nsplit local train - test:\ntake some labels from the local train, and distort them such that error would be about 0.565 and 0.531\ntrain on such two variatns of labels\nand compare the results on local test\nI guess the difference will be very small\n\nIf I have more time - I will try to do that, but not sure. \n\nPS\n@aerdem4 \nthe split to public and private is given to us - see the image on the \"overview\" page.\nThese public notebooks seems to work only on public LB and most probably fall down on private,\nalthouth there is small chance that some ideas from them might be useful - have not thought on that much...",
    "2539684": "## The Model 1.  Predict constant_value (600:1)  .\n\n##### Algorithm 1:\ninput: (1,18116)  \noutput: (1,1)   the \"constant_value\".\n\nThe constant that minimizes the RMSE for SM_name.\n\n###### Y = (C0 - x) * * 2 + .... + (C18116 - x) * * 2\n\nThe constant_value is the value of x that minimizes the function Y.\n\nInput data: \n The constant_value train 600 rows:  rmse_LB_sort_TRAIN_whide2_lb.csv \n The constant_value submit 100 rows (remove constant_value submit   77*2 rows - forecast is weak.).\n\nhttps://www.kaggle.com/code/olegpush/op2-eda-lb?scriptVersionId=152467205\nthe output of version 38. \n\n## The Model 2.  Predict  targets  (600:50) PyBoost .\nIt is important to ensure that model has not been trained on LB sm_names only.\n\n##### Loss:\nI calculate the loss for each iteration by:\n1. Transform (600:50) target to  (600:181116 ) target.\n2. Calculate loss  for  (1:181116 ) target  CY = (C0 - x) * * 2 + …. + (C18116 - x) * * 2\nfor 600 rows separate.\n3. Transform (600:181116 ) to (600:50) targets andf return to pyBoost\n\n## The Bland of Model 1 and Model 2\nI used Model 1 to \"move up or move down\" Model 2 prediction for sm_name:\n\n##### step1: minimizes the function  Y = (C0 - x) * * 2 + …. + (C18116 - x) * * 2 \nwhere  C0 -C18116  - prediction of Model2.\n\nCalculate using Algorithm1 constant_value for  rows 0-255 submition.csv of best public blend.\n\n##### step2:\nFor each row in submition:\n\n```python\n Y_pred = constant_value_Model1  \n               + Y_pred(1:181116)_Model2 \n               - constant_value_Model2 \n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4172517%2Ff206215699c522a7197223c5645ef485%2FIrrrMG_20231127_141535.jpg?generation=1701083967056127&alt=media)",
    "2539975": "I also think that these blends are likely to be useless for the private LB.\n\nThe last two weeks of this competition look like a crowdsourced overfitting project, in places down to the granularity of choosing whichever model gets the lowest score on each item of the public LB data.\n\nObviously enough, each of us needs to decide whether to completely ignore these \"blends of blends\", or whether we think there might somehow be a way of using them to improve our private LB score.\n\nMy own opinion is that there was probably a point in the history of this competition up to which the blends had some value, but that we are now well past that and the new developments are just overfitting.\n\nOther opinions are available."
  },
  "source": "meta"
}