{
  "id": 456943,
  "title": "Will there be a huge shakeup?",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/456943",
  "author_name": "",
  "post_date": "2023-11-22T11:49:44.979778800Z",
  "votes": 28,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I tried to build a ML model at very early stage but almost no useful features I found(I don't know the situation now).High LB score notebooks only using blend and blend that is extremely risky.There are almost no GM&amp;Master on the top of leaderboard.Good Luck to everyone!</p>",
  "messages": [
    {
      "id": "2534089",
      "postDate": "11/22/2023 11:49:44",
      "content": "<p>I tried to build a ML model at very early stage but almost no useful features I found(I don't know the situation now).High LB score notebooks only using blend and blend that is extremely risky.There are almost no GM&amp;Master on the top of leaderboard.Good Luck to everyone!</p>",
      "rawMarkdown": "I tried to build a ML model at very early stage but almost no useful features I found(I don't know the situation now).High LB score notebooks only using blend and blend that is extremely risky.There are almost no GM&Master on the top of leaderboard.Good Luck to everyone!",
      "votes": null
    },
    {
      "id": "2534097",
      "postDate": "11/22/2023 12:03:08",
      "content": "<p>in my opinion it will be casino !</p>",
      "rawMarkdown": "in my opinion it will be casino !",
      "votes": null
    },
    {
      "id": "2534165",
      "postDate": "11/22/2023 13:13:44",
      "content": "<p>Will be huge shaking going on ..</p>",
      "rawMarkdown": "Will be huge shaking going on ..",
      "votes": null
    },
    {
      "id": "2534401",
      "postDate": "11/22/2023 15:40:14",
      "content": "<p><a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> I was almost sure in great shake-up at the begining,<br>\nbut now I am less sure. <br>\nVery lately but seems we found some ways to correspond CV-LB better than public schemes, still not perfect.<br>\nSome solo-models like pyboost and NN can give 0.570+  on their own.</p>\n<p>The main problem we have only 614 samples, but on the other hand we have 18211 targets most of which are noise near zero.<br>\nNevetheless the lack of samples seems to be compensated by abudance of targets  - so overall we have quite a big amount of data and laws of large numbers quite work to reduce the random variation from sample to sample. </p>\n<p>Of course there will be shake-up, but hopefully not like ICR </p>",
      "rawMarkdown": "senkin13 I was almost sure in great shake-up at the begining,\nbut now I am less sure. \nVery lately but seems we found some ways to correspond CV-LB better than public schemes, still not perfect.\nSome solo-models like pyboost and NN can give 0.570+  on their own.\n\nThe main problem we have only 614 samples, but on the other hand we have 18211 targets most of which are noise near zero.\nNevetheless the lack of samples seems to be compensated by abudance of targets  - so overall we have quite a big amount of data and laws of large numbers quite work to reduce the random variation from sample to sample. \n\nOf course there will be shake-up, but hopefully not like ICR",
      "votes": null
    },
    {
      "id": "2534420",
      "postDate": "11/22/2023 15:53:37",
      "content": "<p>You make me interested what's the solid ways you found, I think I will have a try at last days.</p>",
      "rawMarkdown": "You make me interested what's the solid ways you found, I think I will have a try at last days.",
      "votes": null
    },
    {
      "id": "2534434",
      "postDate": "11/22/2023 16:03:57",
      "content": "<p>My best solo model was 0.551 </p>",
      "rawMarkdown": "My best solo model was 0.551",
      "votes": null
    },
    {
      "id": "2534457",
      "postDate": "11/22/2023 16:40:16",
      "content": "<p>Strangely enough a variation-modification of random folds seems works better than AmbrosM and MT schemes - the ones quite logical and discussed here on forum. We also see some other metrics than mrrmse better correlate with LB, but let me not disclose what metrics for the moment ))) On the the hand these findings seems to me NOT critical… Some models and tricks work, despite at the beginning I was desperate that nothing works well…</p>",
      "rawMarkdown": "Strangely enough a variation-modification of random folds seems works better than AmbrosM and MT schemes - the ones quite logical and discussed here on forum. We also see some other metrics than mrrmse better correlate with LB, but let me not disclose what metrics for the moment ))) On the the hand these findings seems to me NOT critical... Some models and tricks work, despite at the beginning I was desperate that nothing works well...",
      "votes": null
    },
    {
      "id": "2534883",
      "postDate": "11/23/2023 00:58:03",
      "content": "<p>i am surprised that you did not even use ensemble🫡</p>",
      "rawMarkdown": "i am surprised that you did not even use ensemble🫡",
      "votes": null
    },
    {
      "id": "2535312",
      "postDate": "11/23/2023 08:45:21",
      "content": "<p>Thanks for sharing ! <br>\nThat model training time - minutes or hours or  days   ?<br>\nThanks in advance </p>",
      "rawMarkdown": "Thanks for sharing ! \nThat model training time - minutes or hours or  days   ?\nThanks in advance",
      "votes": null
    },
    {
      "id": "2535530",
      "postDate": "11/23/2023 12:08:29",
      "content": "<p>Do you mean solo model with k-fold cross validation?</p>",
      "rawMarkdown": "Do you mean solo model with k-fold cross validation?",
      "votes": null
    },
    {
      "id": "2535531",
      "postDate": "11/23/2023 12:09:45",
      "content": "<p>Definitely, given that the hidden test data is 61%</p>",
      "rawMarkdown": "Definitely, given that the hidden test data is 61%",
      "votes": null
    },
    {
      "id": "2536047",
      "postDate": "11/23/2023 21:33:16",
      "content": "<p>I see it like this, but obviously it's just my opinion. At the top of the LB, roughly the gold zone, there are about 10 teams who clearly have submissions stronger than the public ensembles. After that, the next ~400 teams have scores that could easily be coming from overfitted public ensembles. I have no idea what the real strength of those teams will prove to be in the private LB.</p>\n<p>Unless the public ensembles turn out to be much more robust than many of us expect, there's going to be a massive shake-up in the silver and bronze zones. </p>",
      "rawMarkdown": "I see it like this, but obviously it's just my opinion. At the top of the LB, roughly the gold zone, there are about 10 teams who clearly have submissions stronger than the public ensembles. After that, the next ~400 teams have scores that could easily be coming from overfitted public ensembles. I have no idea what the real strength of those teams will prove to be in the private LB.\n\nUnless the public ensembles turn out to be much more robust than many of us expect, there's going to be a massive shake-up in the silver and bronze zones.",
      "votes": null
    },
    {
      "id": "2536048",
      "postDate": "11/23/2023 21:33:31",
      "content": "<p>Does the number of entries seem too high up in the leaderboard? Can it be an overfit?</p>",
      "rawMarkdown": "Does the number of entries seem too high up in the leaderboard? Can it be an overfit?",
      "votes": null
    },
    {
      "id": "2536091",
      "postDate": "11/23/2023 22:40:19",
      "content": "<p>For sure hehehe, it will be fun xD</p>",
      "rawMarkdown": "For sure hehehe, it will be fun xD",
      "votes": null
    },
    {
      "id": "2536212",
      "postDate": "11/24/2023 03:21:36",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> , thanks for your sharing about your pyboost notebook. I would like to ask if the pyboost result of <a href=\"https://www.kaggle.com/datasets/alexandervc/open-problems-2-submits-collection?select=LB604_PyboostLeaveOneOutTsvd50maxdepth10ntrees1000lr001subsample1colsample02_Alex_nbV19.csv\" target=\"_blank\">0574</a> you open-sourced is from a single model?  Did you use pseudo-labeling for training? Could you share the CV of this model?</p>",
      "rawMarkdown": "Hi @alexandervc , thanks for your sharing about your pyboost notebook. I would like to ask if the pyboost result of [0574](https://www.kaggle.com/datasets/alexandervc/open-problems-2-submits-collection?select=LB604_PyboostLeaveOneOutTsvd50maxdepth10ntrees1000lr001subsample1colsample02_Alex_nbV19.csv) you open-sourced is from a single model?  Did you use pseudo-labeling for training? Could you share the CV of this model?\u0000",
      "votes": null
    },
    {
      "id": "2536327",
      "postDate": "11/24/2023 06:00:21",
      "content": "<p>The training time is contingent on various factors, including batch size, model complexity, available resources, and so forth.</p>",
      "rawMarkdown": "The training time is contingent on various factors, including batch size, model complexity, available resources, and so forth.",
      "votes": null
    },
    {
      "id": "2537739",
      "postDate": "11/25/2023 12:53:26",
      "content": "<p>Are you using pseudo labeling for data augmentation by any chance?</p>",
      "rawMarkdown": "Are you using pseudo labeling for data augmentation by any chance?",
      "votes": null
    },
    {
      "id": "2537863",
      "postDate": "11/25/2023 14:57:51",
      "content": "<p>Hi, thanks for kind words ! <br>\n0.574 - not mine, I do not quite understand it, although spent some time on it.<br>\nFrom what I see: <br>\n1) It is single model , but averaging different train subsets<br>\n2) Not sure how to make correct validation of it (at least on my folds) <br>\n3) it is  NOT using  pseudolabel, but uses some simple, but mysterious augmentation tricks - some  samples multiplexed in train - I am afraid  the exact way was found by overfiting to LB, but cannot be sure.</p>",
      "rawMarkdown": "Hi, thanks for kind words ! \n0.574 - not mine, I do not quite understand it, although spent some time on it.\nFrom what I see: \n1) It is single model , but averaging different train subsets\n2) Not sure how to make correct validation of it (at least on my folds) \n3) it is  NOT using  pseudolabel, but uses some simple, but mysterious augmentation tricks - some  samples multiplexed in train - I am afraid  the exact way was found by overfiting to LB, but cannot be sure.",
      "votes": null
    },
    {
      "id": "2544514",
      "postDate": "11/30/2023 22:03:47",
      "content": "<p>This competition is quite different from the former one, but I should have read your 2nd place solution in first OP. You have already revealed codes and methods.</p>",
      "rawMarkdown": "This competition is quite different from the former one, but I should have read your 2nd place solution in first OP. You have already revealed codes and methods.",
      "votes": null
    },
    {
      "id": "2995920",
      "postDate": "09/22/2024 18:38:56",
      "content": "<p>That model training time is hours?</p>",
      "rawMarkdown": "That model training time is hours?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2534097,
      "author_name": "iraqbot",
      "author_url": "",
      "post_date": "11/22/2023 12:03:08",
      "content": "<p>in my opinion it will be casino !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2534165,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "11/22/2023 13:13:44",
      "content": "<p>Will be huge shaking going on ..</p>",
      "votes": null,
      "replies": [
        {
          "id": 2535531,
          "author_name": "jeannkouagou",
          "author_url": "",
          "post_date": "11/23/2023 12:09:45",
          "content": "<p>Definitely, given that the hidden test data is 61%</p>",
          "votes": null,
          "replies": [
            {
              "id": 2536048,
              "author_name": "mehdiafshari",
              "author_url": "",
              "post_date": "11/23/2023 21:33:31",
              "content": "<p>Does the number of entries seem too high up in the leaderboard? Can it be an overfit?</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2536091,
          "author_name": "rapela",
          "author_url": "",
          "post_date": "11/23/2023 22:40:19",
          "content": "<p>For sure hehehe, it will be fun xD</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2534401,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "11/22/2023 15:40:14",
      "content": "<p><a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> I was almost sure in great shake-up at the begining,<br>\nbut now I am less sure. <br>\nVery lately but seems we found some ways to correspond CV-LB better than public schemes, still not perfect.<br>\nSome solo-models like pyboost and NN can give 0.570+  on their own.</p>\n<p>The main problem we have only 614 samples, but on the other hand we have 18211 targets most of which are noise near zero.<br>\nNevetheless the lack of samples seems to be compensated by abudance of targets  - so overall we have quite a big amount of data and laws of large numbers quite work to reduce the random variation from sample to sample. </p>\n<p>Of course there will be shake-up, but hopefully not like ICR </p>",
      "votes": null,
      "replies": [
        {
          "id": 2534420,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "11/22/2023 15:53:37",
          "content": "<p>You make me interested what's the solid ways you found, I think I will have a try at last days.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2534457,
              "author_name": "alexandervc",
              "author_url": "",
              "post_date": "11/22/2023 16:40:16",
              "content": "<p>Strangely enough a variation-modification of random folds seems works better than AmbrosM and MT schemes - the ones quite logical and discussed here on forum. We also see some other metrics than mrrmse better correlate with LB, but let me not disclose what metrics for the moment ))) On the the hand these findings seems to me NOT critical… Some models and tricks work, despite at the beginning I was desperate that nothing works well…</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2534434,
          "author_name": "eliork",
          "author_url": "",
          "post_date": "11/22/2023 16:03:57",
          "content": "<p>My best solo model was 0.551 </p>",
          "votes": null,
          "replies": [
            {
              "id": 2534883,
              "author_name": "blumenkranz7",
              "author_url": "",
              "post_date": "11/23/2023 00:58:03",
              "content": "<p>i am surprised that you did not even use ensemble🫡</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2535312,
              "author_name": "alexandervc",
              "author_url": "",
              "post_date": "11/23/2023 08:45:21",
              "content": "<p>Thanks for sharing ! <br>\nThat model training time - minutes or hours or  days   ?<br>\nThanks in advance </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2536327,
                  "author_name": "eliork",
                  "author_url": "",
                  "post_date": "11/24/2023 06:00:21",
                  "content": "<p>The training time is contingent on various factors, including batch size, model complexity, available resources, and so forth.</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2995920,
                  "author_name": "",
                  "author_url": "",
                  "post_date": "09/22/2024 18:38:56",
                  "content": "<p>That model training time is hours?</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 2535530,
              "author_name": "jeannkouagou",
              "author_url": "",
              "post_date": "11/23/2023 12:08:29",
              "content": "<p>Do you mean solo model with k-fold cross validation?</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2537739,
              "author_name": "jainam213",
              "author_url": "",
              "post_date": "11/25/2023 12:53:26",
              "content": "<p>Are you using pseudo labeling for data augmentation by any chance?</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2536047,
          "author_name": "jbomitchell",
          "author_url": "",
          "post_date": "11/23/2023 21:33:16",
          "content": "<p>I see it like this, but obviously it's just my opinion. At the top of the LB, roughly the gold zone, there are about 10 teams who clearly have submissions stronger than the public ensembles. After that, the next ~400 teams have scores that could easily be coming from overfitted public ensembles. I have no idea what the real strength of those teams will prove to be in the private LB.</p>\n<p>Unless the public ensembles turn out to be much more robust than many of us expect, there's going to be a massive shake-up in the silver and bronze zones. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2536212,
          "author_name": "phoenixzero77",
          "author_url": "",
          "post_date": "11/24/2023 03:21:36",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> , thanks for your sharing about your pyboost notebook. I would like to ask if the pyboost result of <a href=\"https://www.kaggle.com/datasets/alexandervc/open-problems-2-submits-collection?select=LB604_PyboostLeaveOneOutTsvd50maxdepth10ntrees1000lr001subsample1colsample02_Alex_nbV19.csv\" target=\"_blank\">0574</a> you open-sourced is from a single model?  Did you use pseudo-labeling for training? Could you share the CV of this model?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2537863,
              "author_name": "alexandervc",
              "author_url": "",
              "post_date": "11/25/2023 14:57:51",
              "content": "<p>Hi, thanks for kind words ! <br>\n0.574 - not mine, I do not quite understand it, although spent some time on it.<br>\nFrom what I see: <br>\n1) It is single model , but averaging different train subsets<br>\n2) Not sure how to make correct validation of it (at least on my folds) <br>\n3) it is  NOT using  pseudolabel, but uses some simple, but mysterious augmentation tricks - some  samples multiplexed in train - I am afraid  the exact way was found by overfiting to LB, but cannot be sure.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2544514,
      "author_name": "hdynamics",
      "author_url": "",
      "post_date": "11/30/2023 22:03:47",
      "content": "<p>This competition is quite different from the former one, but I should have read your 2nd place solution in first OP. You have already revealed codes and methods.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2534089": "I tried to build a ML model at very early stage but almost no useful features I found(I don't know the situation now).High LB score notebooks only using blend and blend that is extremely risky.There are almost no GM&Master on the top of leaderboard.Good Luck to everyone!",
    "2534097": "in my opinion it will be casino !",
    "2534165": "Will be huge shaking going on ..",
    "2534401": "senkin13 I was almost sure in great shake-up at the begining,\nbut now I am less sure. \nVery lately but seems we found some ways to correspond CV-LB better than public schemes, still not perfect.\nSome solo-models like pyboost and NN can give 0.570+  on their own.\n\nThe main problem we have only 614 samples, but on the other hand we have 18211 targets most of which are noise near zero.\nNevetheless the lack of samples seems to be compensated by abudance of targets  - so overall we have quite a big amount of data and laws of large numbers quite work to reduce the random variation from sample to sample. \n\nOf course there will be shake-up, but hopefully not like ICR",
    "2534420": "You make me interested what's the solid ways you found, I think I will have a try at last days.",
    "2534434": "My best solo model was 0.551",
    "2534457": "Strangely enough a variation-modification of random folds seems works better than AmbrosM and MT schemes - the ones quite logical and discussed here on forum. We also see some other metrics than mrrmse better correlate with LB, but let me not disclose what metrics for the moment ))) On the the hand these findings seems to me NOT critical... Some models and tricks work, despite at the beginning I was desperate that nothing works well...",
    "2534883": "i am surprised that you did not even use ensemble🫡",
    "2535312": "Thanks for sharing ! \nThat model training time - minutes or hours or  days   ?\nThanks in advance",
    "2535530": "Do you mean solo model with k-fold cross validation?",
    "2535531": "Definitely, given that the hidden test data is 61%",
    "2536047": "I see it like this, but obviously it's just my opinion. At the top of the LB, roughly the gold zone, there are about 10 teams who clearly have submissions stronger than the public ensembles. After that, the next ~400 teams have scores that could easily be coming from overfitted public ensembles. I have no idea what the real strength of those teams will prove to be in the private LB.\n\nUnless the public ensembles turn out to be much more robust than many of us expect, there's going to be a massive shake-up in the silver and bronze zones.",
    "2536048": "Does the number of entries seem too high up in the leaderboard? Can it be an overfit?",
    "2536091": "For sure hehehe, it will be fun xD",
    "2536212": "Hi @alexandervc , thanks for your sharing about your pyboost notebook. I would like to ask if the pyboost result of [0574](https://www.kaggle.com/datasets/alexandervc/open-problems-2-submits-collection?select=LB604_PyboostLeaveOneOutTsvd50maxdepth10ntrees1000lr001subsample1colsample02_Alex_nbV19.csv) you open-sourced is from a single model?  Did you use pseudo-labeling for training? Could you share the CV of this model?\u0000",
    "2536327": "The training time is contingent on various factors, including batch size, model complexity, available resources, and so forth.",
    "2537739": "Are you using pseudo labeling for data augmentation by any chance?",
    "2537863": "Hi, thanks for kind words ! \n0.574 - not mine, I do not quite understand it, although spent some time on it.\nFrom what I see: \n1) It is single model , but averaging different train subsets\n2) Not sure how to make correct validation of it (at least on my folds) \n3) it is  NOT using  pseudolabel, but uses some simple, but mysterious augmentation tricks - some  samples multiplexed in train - I am afraid  the exact way was found by overfiting to LB, but cannot be sure.",
    "2544514": "This competition is quite different from the former one, but I should have read your 2nd place solution in first OP. You have already revealed codes and methods.",
    "2995920": "That model training time is hours?"
  },
  "source": "meta"
}