{
  "id": 452515,
  "title": "\"Terribly\" scored models - are useful in current blends - any idea why ? ",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/452515",
  "author_name": "",
  "post_date": "2023-11-02T12:31:33.539540500Z",
  "votes": 18,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Top public blends use some models which have LB score like 0.72 - which is quite worse than even just zeros-only-prediction (LB0.666). </p>\n<p>It is clearly theoretically possible that such \"terrible\" models can improve blends, but practically it seems to me that is the first time when it is so widely used.</p>\n<p>I wonder  why it works ? and naively we can try to substitute \"terrible\" models by something better and get a better score - any ideas in that direction ?</p>\n<p>PS </p>\n<p>There is similar observation for the last year data, but it was only for some cases with CV, not widely used for LB improvement:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/363230\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/363230</a></p>",
  "messages": [
    {
      "id": "2509530",
      "postDate": "11/02/2023 12:31:33",
      "content": "<p>Top public blends use some models which have LB score like 0.72 - which is quite worse than even just zeros-only-prediction (LB0.666). </p>\n<p>It is clearly theoretically possible that such \"terrible\" models can improve blends, but practically it seems to me that is the first time when it is so widely used.</p>\n<p>I wonder  why it works ? and naively we can try to substitute \"terrible\" models by something better and get a better score - any ideas in that direction ?</p>\n<p>PS </p>\n<p>There is similar observation for the last year data, but it was only for some cases with CV, not widely used for LB improvement:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/363230\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/363230</a></p>",
      "rawMarkdown": "Top public blends use some models which have LB score like 0.72 - which is quite worse than even just zeros-only-prediction (LB0.666). \n\nIt is clearly theoretically possible that such \"terrible\" models can improve blends, but practically it seems to me that is the first time when it is so widely used.\n\nI wonder  why it works ? and naively we can try to substitute \"terrible\" models by something better and get a better score - any ideas in that direction ?\n\nPS \n\nThere is similar observation for the last year data, but it was only for some cases with CV, not widely used for LB improvement:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/363230",
      "votes": null
    },
    {
      "id": "2509982",
      "postDate": "11/02/2023 17:12:34",
      "content": "<p>I had the same doubts for long time. The most logical guess:<br>\nAll predicitons have strong bias. The underfitting 0.72 solution happen to have a bias (towards the public test data) that's in the opposite direction of all other overfitting predictions.</p>\n<p>I would like to think that this is unlikely to generalize to the private test. If we want to substitute this \"terrible\" models by something better, we would need something that is \"terrible\" in the similar direction, but towards the private test. Not impossible, but probably not worth the efforts for the risk. </p>",
      "rawMarkdown": "I had the same doubts for long time. The most logical guess:\nAll predicitons have strong bias. The underfitting 0.72 solution happen to have a bias (towards the public test data) that's in the opposite direction of all other overfitting predictions.\n\nI would like to think that this is unlikely to generalize to the private test. If we want to substitute this \"terrible\" models by something better, we would need something that is \"terrible\" in the similar direction, but towards the private test. Not impossible, but probably not worth the efforts for the risk.",
      "votes": null
    },
    {
      "id": "2510144",
      "postDate": "11/02/2023 18:46:30",
      "content": "<p>improved my LB score by 0.001 by blending in the 0.72 solution after reading this. But I felt so dumb do thing…..probably not a good idea to select that as the final submission.</p>",
      "rawMarkdown": "improved my LB score by 0.001 by blending in the 0.72 solution after reading this. But I felt so dumb do thing.....probably not a good idea to select that as the final submission.",
      "votes": null
    },
    {
      "id": "2510500",
      "postDate": "11/03/2023 03:53:28",
      "content": "<p>No medals in the competition, so a lot of time is spent on hard work this time.</p>\n<p>Right now I am trying to improve my public score. However, if you look at the private LBs from last year's OP, the public and private rankings were quite different.</p>\n<p>Is it common for Public and Private rankings to switch places in other competitions?</p>\n<p>I would like to incorporate a \"terrible\" scoring model in order to raise the Public score as much as possible, but could it also be a factor in lowering the final score if it doesn't work well with the Private?　</p>\n<p>Should I look for a robust model that I can use for private scoring?</p>",
      "rawMarkdown": "No medals in the competition, so a lot of time is spent on hard work this time.\n\nRight now I am trying to improve my public score. However, if you look at the private LBs from last year's OP, the public and private rankings were quite different.\n\nIs it common for Public and Private rankings to switch places in other competitions?\n\nI would like to incorporate a \"terrible\" scoring model in order to raise the Public score as much as possible, but could it also be a factor in lowering the final score if it doesn't work well with the Private?　\n\nShould I look for a robust model that I can use for private scoring?",
      "votes": null
    },
    {
      "id": "2510520",
      "postDate": "11/03/2023 04:21:36",
      "content": "<p>At some point soon someone will post a discussion subject on \"will there be a shake-up' in the final LB.  Been doing this for 5 years and only been surprised a couple of times about the size of the shakeup.  My guess - big shakeup but only because so many folks appear to be doing much the same thing in the silver and bronze range.</p>",
      "rawMarkdown": "At some point soon someone will post a discussion subject on \"will there be a shake-up' in the final LB.  Been doing this for 5 years and only been surprised a couple of times about the size of the shakeup.  My guess - big shakeup but only because so many folks appear to be doing much the same thing in the silver and bronze range.",
      "votes": null
    },
    {
      "id": "2510524",
      "postDate": "11/03/2023 04:25:33",
      "content": "<p>Couple of shared notebooks that are blending have plots showing the prediction distributions.  In line with <a href=\"https://www.kaggle.com/qihuaz\" target=\"_blank\">qihauz</a> comments, the 0.72 solution has a distrubution centered off from the others.  </p>",
      "rawMarkdown": "Couple of shared notebooks that are blending have plots showing the prediction distributions.  In line with [qihauz](https://www.kaggle.com/qihuaz) comments, the 0.72 solution has a distrubution centered off from the others.",
      "votes": null
    },
    {
      "id": "2510611",
      "postDate": "11/03/2023 06:25:09",
      "content": "<p>Check this out: <a href=\"https://www.kaggle.com/competitions/icr-identify-age-related-conditions/leaderboard\" target=\"_blank\">https://www.kaggle.com/competitions/icr-identify-age-related-conditions/leaderboard</a> . It's probably one of the most significant shake-ups I've seen. I believe a substantial shake-up is inevitable due to overfitting and data noise.</p>\n<p>In my opinion, a potentially more effective approach involves predicting the gene expressions for adata test (y) from multiome data (X), rather than directly predicting for Differential Expression (let's refer to it as Z). This  involves predicting  y(gene expression after compound treatment) from X (gene expression at baseline - multiome data) and applying LIMMA (though I'm not well-versed in its specifics) on a fully prepared dataset, , resulting in Z (Differential Expression). I'm currently working on this approach, and logically, I believe it's worth trying which has less chances of overfitting.</p>",
      "rawMarkdown": "Check this out: https://www.kaggle.com/competitions/icr-identify-age-related-conditions/leaderboard . It's probably one of the most significant shake-ups I've seen. I believe a substantial shake-up is inevitable due to overfitting and data noise.\n\nIn my opinion, a potentially more effective approach involves predicting the gene expressions for adata test (y) from multiome data (X), rather than directly predicting for Differential Expression (let's refer to it as Z). This  involves predicting  y(gene expression after compound treatment) from X (gene expression at baseline - multiome data) and applying LIMMA (though I'm not well-versed in its specifics) on a fully prepared dataset, , resulting in Z (Differential Expression). I'm currently working on this approach, and logically, I believe it's worth trying which has less chances of overfitting.",
      "votes": null
    },
    {
      "id": "2510617",
      "postDate": "11/03/2023 06:36:24",
      "content": "<p>Some models like xgboost tend to make predictions with a tight distribution (i.e. they have low variance). Since the public LB's ground truth also has a relatively tight distribution, having a low variance prediction is advantageous.</p>\n<p>Conversely, the 0.72 model's biased distribution may capture some meaningful outliers. But they might be too aggressive on other values, resulting in poor performance.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7863678%2F4c4e5032751bbed8fa5cd44c2bc86f59%2FWechatIMG1145.png?generation=1698993096193473&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Some models like xgboost tend to make predictions with a tight distribution (i.e. they have low variance). Since the public LB's ground truth also has a relatively tight distribution, having a low variance prediction is advantageous.\n\nConversely, the 0.72 model's biased distribution may capture some meaningful outliers. But they might be too aggressive on other values, resulting in poor performance.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7863678%2F4c4e5032751bbed8fa5cd44c2bc86f59%2FWechatIMG1145.png?generation=1698993096193473&alt=media)",
      "votes": null
    },
    {
      "id": "2510814",
      "postDate": "11/03/2023 09:11:50",
      "content": "<p>Hi! So you think of making a dataset with multiome data as an X-train and the adata as an Y-train. But there is no multiome for the X-test. Sombody have already mentioned that we could not actually use this approach, since we does not have single-cell data for test <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/445883#2475551\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/445883#2475551</a></p>",
      "rawMarkdown": "Hi! So you think of making a dataset with multiome data as an X-train and the adata as an Y-train. But there is no multiome for the X-test. Sombody have already mentioned that we could not actually use this approach, since we does not have single-cell data for test https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/445883#2475551",
      "votes": null
    },
    {
      "id": "2510922",
      "postDate": "11/03/2023 11:11:14",
      "content": "<p>Cool ! Thanks for sharing !<br>\nWhat is average correlation between your solution old and the 0.720 one ? <br>\nI guess it should be quite low. I guess only uncorrelated \"terrible\" solutions can uplift . <br>\nPS<br>\nExamples of average correlations are computed here: <a href=\"https://www.kaggle.com/code/alexandervc/ensemble-op-correlation-analysis/notebook?scriptVersionId=149154945\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/ensemble-op-correlation-analysis/notebook?scriptVersionId=149154945</a></p>",
      "rawMarkdown": "Cool ! Thanks for sharing !\nWhat is average correlation between your solution old and the 0.720 one ? \nI guess it should be quite low. I guess only uncorrelated \"terrible\" solutions can uplift . \nPS\nExamples of average correlations are computed here: https://www.kaggle.com/code/alexandervc/ensemble-op-correlation-analysis/notebook?scriptVersionId=149154945",
      "votes": null
    },
    {
      "id": "2511078",
      "postDate": "11/03/2023 13:26:08",
      "content": "<p>Correct me if i am wrong. We have all celltype baseline counts(multiome files) and taking that as train + adding corresponding sm_name to predict test adata meta then running limma. Still figuring out this method, will update soon. </p>",
      "rawMarkdown": "Correct me if i am wrong. We have all celltype baseline counts(multiome files) and taking that as train + adding corresponding sm_name to predict test adata meta then running limma. Still figuring out this method, will update soon.",
      "votes": null
    },
    {
      "id": "2511322",
      "postDate": "11/03/2023 15:29:46",
      "content": "<p>PS<br>\nand what weight coefficient did you put in that blend ? About 0.1, 0.05 ? </p>",
      "rawMarkdown": "PS\nand what weight coefficient did you put in that blend ? About 0.1, 0.05 ?",
      "votes": null
    },
    {
      "id": "2515674",
      "postDate": "11/07/2023 05:31:59",
      "content": "<p>Looking forward, since the pipeline of such an approach is still a big confusion to m me</p>",
      "rawMarkdown": "Looking forward, since the pipeline of such an approach is still a big confusion to m me",
      "votes": null
    },
    {
      "id": "2516656",
      "postDate": "11/07/2023 21:07:22",
      "content": "<p>I tried 0.1 and 0.05 weights for the 0.72 public solution , 0.05 gave a better result. I will have a look at the correlations</p>",
      "rawMarkdown": "I tried 0.1 and 0.05 weights for the 0.72 public solution , 0.05 gave a better result. I will have a look at the correlations",
      "votes": null
    },
    {
      "id": "2517040",
      "postDate": "11/08/2023 07:06:16",
      "content": "<p>First, any score (of model) has to be calculated from a sample, that says <strong>score[k]=score(k th sample of data) have their own distribution</strong>, so they can be hacked, or maybe when the final data is revealed, the leaderboard will change.<br>\nSecond, the errors of models may be uncorrelated, which means the mistake that one model makes may cancel out with another.<br>\n<br>\n(after a closer look it seems that some submissions are not what is said in the titles)<br>\n<strong>recall that gbm models are local means, and that most NN/ linear SVR consist of smooth lines</strong><br>\nit is possible that one is too bumpy and another learns gradient from the point masses (this is just a guess).<br>\nAlso, note that modern algorithms on <strong>small datasets</strong> can hack the cross validation (see Kuhn Chapter 1). This dataset with not a lot of rows may suffer.</p>\n<p><strong>Edit: I think this also means how we cross validate and select model will be a factor</strong> As modern ML and autoML can hack CV, we really need to think of what domain definition say about the model inductive. While it is good to achieve high score, the goal is to find a model that smiles at the end of the game.<br>\nSo this <a href=\"https://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes#Ridge-CV-scoring.-Simple-example\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes#Ridge-CV-scoring.-Simple-example</a><br>\nmight be an interesting topic to look at</p>",
      "rawMarkdown": "First, any score (of model) has to be calculated from a sample, that says **score[k]=score(k th sample of data) have their own distribution**, so they can be hacked, or maybe when the final data is revealed, the leaderboard will change.\nSecond, the errors of models may be uncorrelated, which means the mistake that one model makes may cancel out with another.\n~~Third, look at the different approaches to a good public score now:\nhttps://www.kaggle.com/code/haoyuwangc/1-op2-eda-linearsvr-regressor-nn-ki-4744b4 gives linear svr and NN.\nhttps://www.kaggle.com/code/zulqarnainali/advanced-feature-expansion-lightgbm gives lightgbm~~\n(after a closer look it seems that some submissions are not what is said in the titles)\n**recall that gbm models are local means, and that most NN/ linear SVR consist of smooth lines**\nit is possible that one is too bumpy and another learns gradient from the point masses (this is just a guess).\nAlso, note that modern algorithms on **small datasets** can hack the cross validation (see Kuhn Chapter 1). This dataset with not a lot of rows may suffer.\n\n**Edit: I think this also means how we cross validate and select model will be a factor** As modern ML and autoML can hack CV, we really need to think of what domain definition say about the model inductive. While it is good to achieve high score, the goal is to find a model that smiles at the end of the game.\nSo this https://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes#Ridge-CV-scoring.-Simple-example\nmight be an interesting topic to look at",
      "votes": null
    },
    {
      "id": "2518557",
      "postDate": "11/09/2023 12:16:09",
      "content": "<p><a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> Usually Kaggle gives us a lot of surprises like boost from bad models, probabilities multiplicators, manually created rules and so on. </p>\n<p>The case here could be realistic if the worse model works dramatically different in some area in comparison with the usual blend - try to take a look how the prediction looks like against each other on the scatter plot or maybe adversarial validation can help here. If you find something strange in this graph, it could be interesting to check stacking instead of blending to increase this diversion even more. </p>",
      "rawMarkdown": "alexandervc Usually Kaggle gives us a lot of surprises like boost from bad models, probabilities multiplicators, manually created rules and so on. \n\nThe case here could be realistic if the worse model works dramatically different in some area in comparison with the usual blend - try to take a look how the prediction looks like against each other on the scatter plot or maybe adversarial validation can help here. If you find something strange in this graph, it could be interesting to check stacking instead of blending to increase this diversion even more.",
      "votes": null
    },
    {
      "id": "2520745",
      "postDate": "11/11/2023 06:28:55",
      "content": "<p>Sorry for the delayed response; I got caught up in something. Essentially, even without metadata for the test, we can predict the raw counts for the test using baseline raw counts (multiome) and compounds. Here's an example: Let's say X is the gene expression of an NK cell without treatment, and after reacting with compound X, the gene expression becomes Y. The difference between these, let's call it Z, is the differential expression.</p>\n<p>So, it looks like:</p>\n<p>X - &gt;&gt; baseline gene expression.<br>\nXx - &gt;&gt;  Y (gene expression after treatment).<br>\nX - Y - &gt;&gt;  Z (differential expression).</p>\n<p>Using this approach, instead of predicting Z from merely cell_type and compound, we predict Y and then use it to predict Z (DE). And we don't run Limma in this. Hope this clarifies things a bit!</p>\n<p>P.S. - Differential expression is not just the difference here (competition) but (-log10(p-value) * sign(LFC)), and LFC is the estimated log-fold change in expression between the treatment and control condition after shrinkage, as calculated by Limma. Considering even this, the above approach sounds good. Correct me if there's anything wrong with this approach :)</p>",
      "rawMarkdown": "Sorry for the delayed response; I got caught up in something. Essentially, even without metadata for the test, we can predict the raw counts for the test using baseline raw counts (multiome) and compounds. Here's an example: Let's say X is the gene expression of an NK cell without treatment, and after reacting with compound X, the gene expression becomes Y. The difference between these, let's call it Z, is the differential expression.\n\nSo, it looks like:\n\nX - >> baseline gene expression.\nXx - >>  Y (gene expression after treatment).\nX - Y - >>  Z (differential expression).\n\nUsing this approach, instead of predicting Z from merely cell_type and compound, we predict Y and then use it to predict Z (DE). And we don't run Limma in this. Hope this clarifies things a bit!\n\nP.S. - Differential expression is not just the difference here (competition) but (-log10(p-value) * sign(LFC)), and LFC is the estimated log-fold change in expression between the treatment and control condition after shrinkage, as calculated by Limma. Considering even this, the above approach sounds good. Correct me if there's anything wrong with this approach :)",
      "votes": null
    },
    {
      "id": "2521743",
      "postDate": "11/12/2023 03:19:16",
      "content": "<p>AHH, I think you are right! Thanks for the explanation) I'm afraid too little time left to learn all the required processing like pseudobulking and Lima…</p>",
      "rawMarkdown": "AHH, I think you are right! Thanks for the explanation) I'm afraid too little time left to learn all the required processing like pseudobulking and Lima...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2509982,
      "author_name": "qihuaz",
      "author_url": "",
      "post_date": "11/02/2023 17:12:34",
      "content": "<p>I had the same doubts for long time. The most logical guess:<br>\nAll predicitons have strong bias. The underfitting 0.72 solution happen to have a bias (towards the public test data) that's in the opposite direction of all other overfitting predictions.</p>\n<p>I would like to think that this is unlikely to generalize to the private test. If we want to substitute this \"terrible\" models by something better, we would need something that is \"terrible\" in the similar direction, but towards the private test. Not impossible, but probably not worth the efforts for the risk. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2510144,
      "author_name": "qihuaz",
      "author_url": "",
      "post_date": "11/02/2023 18:46:30",
      "content": "<p>improved my LB score by 0.001 by blending in the 0.72 solution after reading this. But I felt so dumb do thing…..probably not a good idea to select that as the final submission.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2510922,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "11/03/2023 11:11:14",
          "content": "<p>Cool ! Thanks for sharing !<br>\nWhat is average correlation between your solution old and the 0.720 one ? <br>\nI guess it should be quite low. I guess only uncorrelated \"terrible\" solutions can uplift . <br>\nPS<br>\nExamples of average correlations are computed here: <a href=\"https://www.kaggle.com/code/alexandervc/ensemble-op-correlation-analysis/notebook?scriptVersionId=149154945\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/ensemble-op-correlation-analysis/notebook?scriptVersionId=149154945</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2511322,
              "author_name": "alexandervc",
              "author_url": "",
              "post_date": "11/03/2023 15:29:46",
              "content": "<p>PS<br>\nand what weight coefficient did you put in that blend ? About 0.1, 0.05 ? </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2516656,
                  "author_name": "qihuaz",
                  "author_url": "",
                  "post_date": "11/07/2023 21:07:22",
                  "content": "<p>I tried 0.1 and 0.05 weights for the 0.72 public solution , 0.05 gave a better result. I will have a look at the correlations</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2510500,
      "author_name": "yoshifumimiya",
      "author_url": "",
      "post_date": "11/03/2023 03:53:28",
      "content": "<p>No medals in the competition, so a lot of time is spent on hard work this time.</p>\n<p>Right now I am trying to improve my public score. However, if you look at the private LBs from last year's OP, the public and private rankings were quite different.</p>\n<p>Is it common for Public and Private rankings to switch places in other competitions?</p>\n<p>I would like to incorporate a \"terrible\" scoring model in order to raise the Public score as much as possible, but could it also be a factor in lowering the final score if it doesn't work well with the Private?　</p>\n<p>Should I look for a robust model that I can use for private scoring?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2510520,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "11/03/2023 04:21:36",
          "content": "<p>At some point soon someone will post a discussion subject on \"will there be a shake-up' in the final LB.  Been doing this for 5 years and only been surprised a couple of times about the size of the shakeup.  My guess - big shakeup but only because so many folks appear to be doing much the same thing in the silver and bronze range.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2510611,
          "author_name": "kishanvavdara",
          "author_url": "",
          "post_date": "11/03/2023 06:25:09",
          "content": "<p>Check this out: <a href=\"https://www.kaggle.com/competitions/icr-identify-age-related-conditions/leaderboard\" target=\"_blank\">https://www.kaggle.com/competitions/icr-identify-age-related-conditions/leaderboard</a> . It's probably one of the most significant shake-ups I've seen. I believe a substantial shake-up is inevitable due to overfitting and data noise.</p>\n<p>In my opinion, a potentially more effective approach involves predicting the gene expressions for adata test (y) from multiome data (X), rather than directly predicting for Differential Expression (let's refer to it as Z). This  involves predicting  y(gene expression after compound treatment) from X (gene expression at baseline - multiome data) and applying LIMMA (though I'm not well-versed in its specifics) on a fully prepared dataset, , resulting in Z (Differential Expression). I'm currently working on this approach, and logically, I believe it's worth trying which has less chances of overfitting.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2510814,
              "author_name": "antoninadolgorukova",
              "author_url": "",
              "post_date": "11/03/2023 09:11:50",
              "content": "<p>Hi! So you think of making a dataset with multiome data as an X-train and the adata as an Y-train. But there is no multiome for the X-test. Sombody have already mentioned that we could not actually use this approach, since we does not have single-cell data for test <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/445883#2475551\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/445883#2475551</a></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2511078,
                  "author_name": "kishanvavdara",
                  "author_url": "",
                  "post_date": "11/03/2023 13:26:08",
                  "content": "<p>Correct me if i am wrong. We have all celltype baseline counts(multiome files) and taking that as train + adding corresponding sm_name to predict test adata meta then running limma. Still figuring out this method, will update soon. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2515674,
                      "author_name": "antoninadolgorukova",
                      "author_url": "",
                      "post_date": "11/07/2023 05:31:59",
                      "content": "<p>Looking forward, since the pipeline of such an approach is still a big confusion to m me</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2520745,
                          "author_name": "kishanvavdara",
                          "author_url": "",
                          "post_date": "11/11/2023 06:28:55",
                          "content": "<p>Sorry for the delayed response; I got caught up in something. Essentially, even without metadata for the test, we can predict the raw counts for the test using baseline raw counts (multiome) and compounds. Here's an example: Let's say X is the gene expression of an NK cell without treatment, and after reacting with compound X, the gene expression becomes Y. The difference between these, let's call it Z, is the differential expression.</p>\n<p>So, it looks like:</p>\n<p>X - &gt;&gt; baseline gene expression.<br>\nXx - &gt;&gt;  Y (gene expression after treatment).<br>\nX - Y - &gt;&gt;  Z (differential expression).</p>\n<p>Using this approach, instead of predicting Z from merely cell_type and compound, we predict Y and then use it to predict Z (DE). And we don't run Limma in this. Hope this clarifies things a bit!</p>\n<p>P.S. - Differential expression is not just the difference here (competition) but (-log10(p-value) * sign(LFC)), and LFC is the estimated log-fold change in expression between the treatment and control condition after shrinkage, as calculated by Limma. Considering even this, the above approach sounds good. Correct me if there's anything wrong with this approach :)</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2521743,
                              "author_name": "antoninadolgorukova",
                              "author_url": "",
                              "post_date": "11/12/2023 03:19:16",
                              "content": "<p>AHH, I think you are right! Thanks for the explanation) I'm afraid too little time left to learn all the required processing like pseudobulking and Lima…</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2510524,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "11/03/2023 04:25:33",
      "content": "<p>Couple of shared notebooks that are blending have plots showing the prediction distributions.  In line with <a href=\"https://www.kaggle.com/qihuaz\" target=\"_blank\">qihauz</a> comments, the 0.72 solution has a distrubution centered off from the others.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2510617,
      "author_name": "superdanielshao",
      "author_url": "",
      "post_date": "11/03/2023 06:36:24",
      "content": "<p>Some models like xgboost tend to make predictions with a tight distribution (i.e. they have low variance). Since the public LB's ground truth also has a relatively tight distribution, having a low variance prediction is advantageous.</p>\n<p>Conversely, the 0.72 model's biased distribution may capture some meaningful outliers. But they might be too aggressive on other values, resulting in poor performance.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7863678%2F4c4e5032751bbed8fa5cd44c2bc86f59%2FWechatIMG1145.png?generation=1698993096193473&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2517040,
      "author_name": "hli111111",
      "author_url": "",
      "post_date": "11/08/2023 07:06:16",
      "content": "<p>First, any score (of model) has to be calculated from a sample, that says <strong>score[k]=score(k th sample of data) have their own distribution</strong>, so they can be hacked, or maybe when the final data is revealed, the leaderboard will change.<br>\nSecond, the errors of models may be uncorrelated, which means the mistake that one model makes may cancel out with another.<br>\n<br>\n(after a closer look it seems that some submissions are not what is said in the titles)<br>\n<strong>recall that gbm models are local means, and that most NN/ linear SVR consist of smooth lines</strong><br>\nit is possible that one is too bumpy and another learns gradient from the point masses (this is just a guess).<br>\nAlso, note that modern algorithms on <strong>small datasets</strong> can hack the cross validation (see Kuhn Chapter 1). This dataset with not a lot of rows may suffer.</p>\n<p><strong>Edit: I think this also means how we cross validate and select model will be a factor</strong> As modern ML and autoML can hack CV, we really need to think of what domain definition say about the model inductive. While it is good to achieve high score, the goal is to find a model that smiles at the end of the game.<br>\nSo this <a href=\"https://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes#Ridge-CV-scoring.-Simple-example\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes#Ridge-CV-scoring.-Simple-example</a><br>\nmight be an interesting topic to look at</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2518557,
      "author_name": "alexryzhkov",
      "author_url": "",
      "post_date": "11/09/2023 12:16:09",
      "content": "<p><a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> Usually Kaggle gives us a lot of surprises like boost from bad models, probabilities multiplicators, manually created rules and so on. </p>\n<p>The case here could be realistic if the worse model works dramatically different in some area in comparison with the usual blend - try to take a look how the prediction looks like against each other on the scatter plot or maybe adversarial validation can help here. If you find something strange in this graph, it could be interesting to check stacking instead of blending to increase this diversion even more. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2509530": "Top public blends use some models which have LB score like 0.72 - which is quite worse than even just zeros-only-prediction (LB0.666). \n\nIt is clearly theoretically possible that such \"terrible\" models can improve blends, but practically it seems to me that is the first time when it is so widely used.\n\nI wonder  why it works ? and naively we can try to substitute \"terrible\" models by something better and get a better score - any ideas in that direction ?\n\nPS \n\nThere is similar observation for the last year data, but it was only for some cases with CV, not widely used for LB improvement:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/363230",
    "2509982": "I had the same doubts for long time. The most logical guess:\nAll predicitons have strong bias. The underfitting 0.72 solution happen to have a bias (towards the public test data) that's in the opposite direction of all other overfitting predictions.\n\nI would like to think that this is unlikely to generalize to the private test. If we want to substitute this \"terrible\" models by something better, we would need something that is \"terrible\" in the similar direction, but towards the private test. Not impossible, but probably not worth the efforts for the risk.",
    "2510144": "improved my LB score by 0.001 by blending in the 0.72 solution after reading this. But I felt so dumb do thing.....probably not a good idea to select that as the final submission.",
    "2510500": "No medals in the competition, so a lot of time is spent on hard work this time.\n\nRight now I am trying to improve my public score. However, if you look at the private LBs from last year's OP, the public and private rankings were quite different.\n\nIs it common for Public and Private rankings to switch places in other competitions?\n\nI would like to incorporate a \"terrible\" scoring model in order to raise the Public score as much as possible, but could it also be a factor in lowering the final score if it doesn't work well with the Private?　\n\nShould I look for a robust model that I can use for private scoring?",
    "2510520": "At some point soon someone will post a discussion subject on \"will there be a shake-up' in the final LB.  Been doing this for 5 years and only been surprised a couple of times about the size of the shakeup.  My guess - big shakeup but only because so many folks appear to be doing much the same thing in the silver and bronze range.",
    "2510524": "Couple of shared notebooks that are blending have plots showing the prediction distributions.  In line with [qihauz](https://www.kaggle.com/qihuaz) comments, the 0.72 solution has a distrubution centered off from the others.",
    "2510611": "Check this out: https://www.kaggle.com/competitions/icr-identify-age-related-conditions/leaderboard . It's probably one of the most significant shake-ups I've seen. I believe a substantial shake-up is inevitable due to overfitting and data noise.\n\nIn my opinion, a potentially more effective approach involves predicting the gene expressions for adata test (y) from multiome data (X), rather than directly predicting for Differential Expression (let's refer to it as Z). This  involves predicting  y(gene expression after compound treatment) from X (gene expression at baseline - multiome data) and applying LIMMA (though I'm not well-versed in its specifics) on a fully prepared dataset, , resulting in Z (Differential Expression). I'm currently working on this approach, and logically, I believe it's worth trying which has less chances of overfitting.",
    "2510617": "Some models like xgboost tend to make predictions with a tight distribution (i.e. they have low variance). Since the public LB's ground truth also has a relatively tight distribution, having a low variance prediction is advantageous.\n\nConversely, the 0.72 model's biased distribution may capture some meaningful outliers. But they might be too aggressive on other values, resulting in poor performance.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7863678%2F4c4e5032751bbed8fa5cd44c2bc86f59%2FWechatIMG1145.png?generation=1698993096193473&alt=media)",
    "2510814": "Hi! So you think of making a dataset with multiome data as an X-train and the adata as an Y-train. But there is no multiome for the X-test. Sombody have already mentioned that we could not actually use this approach, since we does not have single-cell data for test https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/445883#2475551",
    "2510922": "Cool ! Thanks for sharing !\nWhat is average correlation between your solution old and the 0.720 one ? \nI guess it should be quite low. I guess only uncorrelated \"terrible\" solutions can uplift . \nPS\nExamples of average correlations are computed here: https://www.kaggle.com/code/alexandervc/ensemble-op-correlation-analysis/notebook?scriptVersionId=149154945",
    "2511078": "Correct me if i am wrong. We have all celltype baseline counts(multiome files) and taking that as train + adding corresponding sm_name to predict test adata meta then running limma. Still figuring out this method, will update soon.",
    "2511322": "PS\nand what weight coefficient did you put in that blend ? About 0.1, 0.05 ?",
    "2515674": "Looking forward, since the pipeline of such an approach is still a big confusion to m me",
    "2516656": "I tried 0.1 and 0.05 weights for the 0.72 public solution , 0.05 gave a better result. I will have a look at the correlations",
    "2517040": "First, any score (of model) has to be calculated from a sample, that says **score[k]=score(k th sample of data) have their own distribution**, so they can be hacked, or maybe when the final data is revealed, the leaderboard will change.\nSecond, the errors of models may be uncorrelated, which means the mistake that one model makes may cancel out with another.\n~~Third, look at the different approaches to a good public score now:\nhttps://www.kaggle.com/code/haoyuwangc/1-op2-eda-linearsvr-regressor-nn-ki-4744b4 gives linear svr and NN.\nhttps://www.kaggle.com/code/zulqarnainali/advanced-feature-expansion-lightgbm gives lightgbm~~\n(after a closer look it seems that some submissions are not what is said in the titles)\n**recall that gbm models are local means, and that most NN/ linear SVR consist of smooth lines**\nit is possible that one is too bumpy and another learns gradient from the point masses (this is just a guess).\nAlso, note that modern algorithms on **small datasets** can hack the cross validation (see Kuhn Chapter 1). This dataset with not a lot of rows may suffer.\n\n**Edit: I think this also means how we cross validate and select model will be a factor** As modern ML and autoML can hack CV, we really need to think of what domain definition say about the model inductive. While it is good to achieve high score, the goal is to find a model that smiles at the end of the game.\nSo this https://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes#Ridge-CV-scoring.-Simple-example\nmight be an interesting topic to look at",
    "2518557": "alexandervc Usually Kaggle gives us a lot of surprises like boost from bad models, probabilities multiplicators, manually created rules and so on. \n\nThe case here could be realistic if the worse model works dramatically different in some area in comparison with the usual blend - try to take a look how the prediction looks like against each other on the scatter plot or maybe adversarial validation can help here. If you find something strange in this graph, it could be interesting to check stacking instead of blending to increase this diversion even more.",
    "2520745": "Sorry for the delayed response; I got caught up in something. Essentially, even without metadata for the test, we can predict the raw counts for the test using baseline raw counts (multiome) and compounds. Here's an example: Let's say X is the gene expression of an NK cell without treatment, and after reacting with compound X, the gene expression becomes Y. The difference between these, let's call it Z, is the differential expression.\n\nSo, it looks like:\n\nX - >> baseline gene expression.\nXx - >>  Y (gene expression after treatment).\nX - Y - >>  Z (differential expression).\n\nUsing this approach, instead of predicting Z from merely cell_type and compound, we predict Y and then use it to predict Z (DE). And we don't run Limma in this. Hope this clarifies things a bit!\n\nP.S. - Differential expression is not just the difference here (competition) but (-log10(p-value) * sign(LFC)), and LFC is the estimated log-fold change in expression between the treatment and control condition after shrinkage, as calculated by Limma. Considering even this, the above approach sounds good. Correct me if there's anything wrong with this approach :)",
    "2521743": "AHH, I think you are right! Thanks for the explanation) I'm afraid too little time left to learn all the required processing like pseudobulking and Lima..."
  },
  "source": "meta"
}