{
  "id": 346518,
  "title": " Pearson Correlation Coefficient. Here we go again! ",
  "url": "/competitions/open-problems-multimodal/discussion/346518",
  "author_name": "",
  "post_date": "2022-08-20T01:13:33.025803800Z",
  "votes": 12,
  "comment_count": 16,
  "views": 0,
  "content": "<h1>Kaggle Competitions Evaluated by Pearson Coefficient</h1>\n<p>Ubiquant Market Prediction<br>\n<a href=\"https://www.kaggle.com/competitions/ubiquant-market-prediction/overview\" target=\"_blank\">https://www.kaggle.com/competitions/ubiquant-market-prediction/overview</a></p>\n<p>G-Research Crypto Forecasting<br>\n<a href=\"https://www.kaggle.com/competitions/g-research-crypto-forecasting/overview/evaluation\" target=\"_blank\">https://www.kaggle.com/competitions/g-research-crypto-forecasting/overview/evaluation</a></p>\n<p>U.S. Patent Phrase to Phrase Matching<br>\n<a href=\"https://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/overview/evaluation\" target=\"_blank\">https://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/overview/evaluation</a></p>\n<h1>DISCUSSION TOPICS:</h1>\n<p>Pearson vs Spearman Correlation Coefficient  By Saurabh Thakur<br>\n<a href=\"https://www.kaggle.com/discussions/getting-started/171951\" target=\"_blank\">https://www.kaggle.com/discussions/getting-started/171951</a></p>\n<p>Correlation : Pearson v/s Spearman  By Fatakdawala<br>\n<a href=\"https://www.kaggle.com/discussions/getting-started/186658\" target=\"_blank\">https://www.kaggle.com/discussions/getting-started/186658</a></p>\n<p>Contest Evaluation Criteria - Pearson correlation coefficient - Summary for starters By Kalilur Rahman<br>\n<a href=\"https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/302153\" target=\"_blank\">https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/302153</a></p>\n<p>Best Loss function that optimize Pearson correlation  By Amed<br>\n<a href=\"https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/302181\" target=\"_blank\">https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/302181</a></p>\n<p>Custom Pearson Metric for LGBMRegressor  By RDizzI3<br>\n<a href=\"https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/302480\" target=\"_blank\">https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/302480</a></p>\n<p>Understanding Pearson Correlation Coefficient(PCC) By Poteman<br>\n<a href=\"https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/311260\" target=\"_blank\">https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/311260</a></p>\n<p>A simple way to use keras for the pearson correlation coefficient By Laura Fink<br>\n<a href=\"https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/316589\" target=\"_blank\">https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/316589</a></p>\n<p>Why using pearson correlation as loss is tricky - A brief analysis  By Martin BB<br>\n<a href=\"https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/306322\" target=\"_blank\">https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/306322</a></p>\n<p>Directly maximizing Pearson correlation got poor cv. (With CustomLoss code) By Masaki Aota<br>\n<a href=\"https://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/discussion/327502\" target=\"_blank\">https://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/discussion/327502</a></p>\n<h1>Why Pearson Correlation??</h1>\n<p>By Abhranta Panigrahi<br>\n<a href=\"https://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/discussion/317851\" target=\"_blank\">https://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/discussion/317851</a></p>\n<p>\"I have a genuine doubt. Can anyone please explain why we are using pearson correlation ?\"  Abhranta's words.</p>\n<p>Whenever someone is shy to ask something they use the expression \"genuine doubt\"  That's so funny, it's almost to say \"I'm sorry for asking that question.\"</p>\n<h1>Spoiler alert: if your answer was \"easy to interpret, easy to code and is intuitive.\" You're going to get a triple easy, intuitive downvote : )</h1>\n<p>I simply can help laughing at that. If you want to downvote. Go ahead.</p>",
  "messages": [
    {
      "id": "1906538",
      "postDate": "08/20/2022 01:13:33",
      "content": "<h1>Kaggle Competitions Evaluated by Pearson Coefficient</h1>\n<p>Ubiquant Market Prediction<br>\n<a href=\"https://www.kaggle.com/competitions/ubiquant-market-prediction/overview\" target=\"_blank\">https://www.kaggle.com/competitions/ubiquant-market-prediction/overview</a></p>\n<p>G-Research Crypto Forecasting<br>\n<a href=\"https://www.kaggle.com/competitions/g-research-crypto-forecasting/overview/evaluation\" target=\"_blank\">https://www.kaggle.com/competitions/g-research-crypto-forecasting/overview/evaluation</a></p>\n<p>U.S. Patent Phrase to Phrase Matching<br>\n<a href=\"https://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/overview/evaluation\" target=\"_blank\">https://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/overview/evaluation</a></p>\n<h1>DISCUSSION TOPICS:</h1>\n<p>Pearson vs Spearman Correlation Coefficient  By Saurabh Thakur<br>\n<a href=\"https://www.kaggle.com/discussions/getting-started/171951\" target=\"_blank\">https://www.kaggle.com/discussions/getting-started/171951</a></p>\n<p>Correlation : Pearson v/s Spearman  By Fatakdawala<br>\n<a href=\"https://www.kaggle.com/discussions/getting-started/186658\" target=\"_blank\">https://www.kaggle.com/discussions/getting-started/186658</a></p>\n<p>Contest Evaluation Criteria - Pearson correlation coefficient - Summary for starters By Kalilur Rahman<br>\n<a href=\"https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/302153\" target=\"_blank\">https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/302153</a></p>\n<p>Best Loss function that optimize Pearson correlation  By Amed<br>\n<a href=\"https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/302181\" target=\"_blank\">https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/302181</a></p>\n<p>Custom Pearson Metric for LGBMRegressor  By RDizzI3<br>\n<a href=\"https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/302480\" target=\"_blank\">https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/302480</a></p>\n<p>Understanding Pearson Correlation Coefficient(PCC) By Poteman<br>\n<a href=\"https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/311260\" target=\"_blank\">https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/311260</a></p>\n<p>A simple way to use keras for the pearson correlation coefficient By Laura Fink<br>\n<a href=\"https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/316589\" target=\"_blank\">https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/316589</a></p>\n<p>Why using pearson correlation as loss is tricky - A brief analysis  By Martin BB<br>\n<a href=\"https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/306322\" target=\"_blank\">https://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/306322</a></p>\n<p>Directly maximizing Pearson correlation got poor cv. (With CustomLoss code) By Masaki Aota<br>\n<a href=\"https://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/discussion/327502\" target=\"_blank\">https://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/discussion/327502</a></p>\n<h1>Why Pearson Correlation??</h1>\n<p>By Abhranta Panigrahi<br>\n<a href=\"https://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/discussion/317851\" target=\"_blank\">https://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/discussion/317851</a></p>\n<p>\"I have a genuine doubt. Can anyone please explain why we are using pearson correlation ?\"  Abhranta's words.</p>\n<p>Whenever someone is shy to ask something they use the expression \"genuine doubt\"  That's so funny, it's almost to say \"I'm sorry for asking that question.\"</p>\n<h1>Spoiler alert: if your answer was \"easy to interpret, easy to code and is intuitive.\" You're going to get a triple easy, intuitive downvote : )</h1>\n<p>I simply can help laughing at that. If you want to downvote. Go ahead.</p>",
      "rawMarkdown": "#Kaggle Competitions Evaluated by Pearson Coefficient\n\nUbiquant Market Prediction\nhttps://www.kaggle.com/competitions/ubiquant-market-prediction/overview\n\nG-Research Crypto Forecasting\nhttps://www.kaggle.com/competitions/g-research-crypto-forecasting/overview/evaluation\n\nU.S. Patent Phrase to Phrase Matching\nhttps://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/overview/evaluation\n\n#DISCUSSION TOPICS:\n\nPearson vs Spearman Correlation Coefficient  By Saurabh Thakur\nhttps://www.kaggle.com/discussions/getting-started/171951\n\nCorrelation : Pearson v/s Spearman  By Fatakdawala\nhttps://www.kaggle.com/discussions/getting-started/186658\n\nContest Evaluation Criteria - Pearson correlation coefficient - Summary for starters By Kalilur Rahman\nhttps://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/302153\n\nBest Loss function that optimize Pearson correlation  By Amed\nhttps://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/302181\n\nCustom Pearson Metric for LGBMRegressor  By RDizzI3\nhttps://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/302480\n\nUnderstanding Pearson Correlation Coefficient(PCC) By Poteman\nhttps://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/311260\n\nA simple way to use keras for the pearson correlation coefficient By Laura Fink\nhttps://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/316589\n\nWhy using pearson correlation as loss is tricky - A brief analysis  By Martin BB\nhttps://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/306322\n\nDirectly maximizing Pearson correlation got poor cv. (With CustomLoss code) By Masaki Aota\nhttps://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/discussion/327502\n\n#Why Pearson Correlation?? \n\n By Abhranta Panigrahi\nhttps://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/discussion/317851\n\n\"I have a genuine doubt. Can anyone please explain why we are using pearson correlation ?\"  Abhranta's words.\n\nWhenever someone is shy to ask something they use the expression \"genuine doubt\"  That's so funny, it's almost to say \"I'm sorry for asking that question.\"\n\n#Spoiler alert: if your answer was \"easy to interpret, easy to code and is intuitive.\" You're going to get a triple easy, intuitive downvote : )\n\nI simply can help laughing at that. If you want to downvote. Go ahead.",
      "votes": null
    },
    {
      "id": "1906792",
      "postDate": "08/20/2022 07:42:39",
      "content": "<p>I would think something like R^2 would make more sense for an evaluation metric.</p>",
      "rawMarkdown": "I would think something like R^2 would make more sense for an evaluation metric.",
      "votes": null
    },
    {
      "id": "1907012",
      "postDate": "08/20/2022 11:54:08",
      "content": "<p>Thank you for your contribution and support JD.<br>\nMaybe we can suggest that to the hosts too : )</p>",
      "rawMarkdown": "Thank you for your contribution and support JD.\nMaybe we can suggest that to the hosts too : )",
      "votes": null
    },
    {
      "id": "1907068",
      "postDate": "08/20/2022 12:35:08",
      "content": "<p>tl dr        ? </p>",
      "rawMarkdown": "tl dr        ?",
      "votes": null
    },
    {
      "id": "1907464",
      "postDate": "08/20/2022 19:58:45",
      "content": "<p>Hi Alexander <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> </p>\n<p>Indeed, it's too long.  TIKTOKish short way:</p>\n<p>\"Why Pearson Correlation?\"</p>\n<p>\" easy to interpret, easy to code and is intuitive.\"</p>\n<p>That answer above got 3 downvotes. </p>\n<p>As short as a TikTok mp3. <br>\nEven the funny comments, I read.</p>",
      "rawMarkdown": "Hi Alexander @alexandervc \n\nIndeed, it's too long.  TIKTOKish short way:\n\n\"Why Pearson Correlation?\"\n\n\" easy to interpret, easy to code and is intuitive.\"\n\nThat answer above got 3 downvotes. \n\nAs short as a TikTok mp3. \nEven the funny comments, I read.",
      "votes": null
    },
    {
      "id": "1908949",
      "postDate": "08/22/2022 06:51:07",
      "content": "<p>This is actually something we thought a lot about, and it has to do with the data format and its interpretation. </p>\n<p>Single-cell RNA-seq data has to be normalized to be interpretable. When cells are profiled, they are sequenced at different depths. This means that there is more data for some cells and less for others purely for technical reasons. We correct for this by normalization to make cells comparable (a correction factor is typically computed per cell). We hosted a multi-modal data integration competition at NeurIPS last year where we normalized the data and then evaluated the prediction task via MSE to make it more accessible as one could easily directly optimize the evaluation metric. However, as the normalization had to be done per batch, and our chosen normalization method (the current SOTA method in the field) produced normalized data that had a different scale per run, prediction became very challenging as evaluated by MSE. Pearson correlation would have fixed this issue. You can read more about that in our PMLR paper on the competition results (<a href=\"http://proceedings.mlr.press/v176/lance22a.html)\" target=\"_blank\">http://proceedings.mlr.press/v176/lance22a.html)</a>.</p>\n<p>Secondly, biologically we don't actually care about the exact scale of the results. We know that we do not measure every molecule in a cell when it's profiled. Typically it's also not clear what proportion of molecules we capture. So this data type is probably always going to be relative. Thus, if you predict values that are 100 times higher than the ground truth data, but self-consistent, then that's just as good as any other scale. With Pearson correlation, we are not penalizing for differences in scales between predictions.</p>",
      "rawMarkdown": "This is actually something we thought a lot about, and it has to do with the data format and its interpretation. \n\nSingle-cell RNA-seq data has to be normalized to be interpretable. When cells are profiled, they are sequenced at different depths. This means that there is more data for some cells and less for others purely for technical reasons. We correct for this by normalization to make cells comparable (a correction factor is typically computed per cell). We hosted a multi-modal data integration competition at NeurIPS last year where we normalized the data and then evaluated the prediction task via MSE to make it more accessible as one could easily directly optimize the evaluation metric. However, as the normalization had to be done per batch, and our chosen normalization method (the current SOTA method in the field) produced normalized data that had a different scale per run, prediction became very challenging as evaluated by MSE. Pearson correlation would have fixed this issue. You can read more about that in our PMLR paper on the competition results (http://proceedings.mlr.press/v176/lance22a.html).\n\nSecondly, biologically we don't actually care about the exact scale of the results. We know that we do not measure every molecule in a cell when it's profiled. Typically it's also not clear what proportion of molecules we capture. So this data type is probably always going to be relative. Thus, if you predict values that are 100 times higher than the ground truth data, but self-consistent, then that's just as good as any other scale. With Pearson correlation, we are not penalizing for differences in scales between predictions.",
      "votes": null
    },
    {
      "id": "1909289",
      "postDate": "08/22/2022 13:47:56",
      "content": "<p>R^2 would be a poor evaluation metric because it has no sign. We care that genes measured as highly expressed in the test set are predicted as highly expressed.</p>",
      "rawMarkdown": "R^2 would be a poor evaluation metric because it has no sign. We care that genes measured as highly expressed in the test set are predicted as highly expressed.",
      "votes": null
    },
    {
      "id": "1909349",
      "postDate": "08/22/2022 14:54:49",
      "content": "<p>Thank you Daniel for clarifying JD's suggestion.</p>",
      "rawMarkdown": "Thank you Daniel for clarifying JD's suggestion.",
      "votes": null
    },
    {
      "id": "1909655",
      "postDate": "08/22/2022 20:00:01",
      "content": "<p>Thank you for the embracing, valuable comment by competition host.  <br>\nThanks also for the paper link.</p>",
      "rawMarkdown": "Thank you for the embracing, valuable comment by competition host.  \nThanks also for the paper link.",
      "votes": null
    },
    {
      "id": "1910882",
      "postDate": "08/23/2022 18:50:17",
      "content": "<p>I'm sorry M. Luecken, I understood all your complimentary explanation about Pearson correlation started by D. Burkhardt</p>\n<p>Though  read \"PMLR paper on the competition results\"</p>\n<p>the link is resulting page 404<br>\n\"404<br>\nFile not found</p>\n<p>Though Vol 176 works:  \"Volume 176: NeurIPS 2021 Competitions and Demonstrations Track, 6-14 December 2021, Online\"<br>\n<a href=\"http://proceedings.mlr.press/v176/\" target=\"_blank\">http://proceedings.mlr.press/v176/</a></p>",
      "rawMarkdown": "I'm sorry M. Luecken, I understood all your complimentary explanation about Pearson correlation started by D. Burkhardt\n\nThough  read \"PMLR paper on the competition results\"\n\nthe link is resulting page 404\n\"404\nFile not found\n\nThough Vol 176 works:  \"Volume 176: NeurIPS 2021 Competitions and Demonstrations Track, 6-14 December 2021, Online\"\nhttp://proceedings.mlr.press/v176/",
      "votes": null
    },
    {
      "id": "1912528",
      "postDate": "08/24/2022 19:38:49",
      "content": "<p>Easy to understand, easy to code, easy to give nan after the first minutes of training.</p>",
      "rawMarkdown": "Easy to understand, easy to code, easy to give nan after the first minutes of training.",
      "votes": null
    },
    {
      "id": "1912535",
      "postDate": "08/24/2022 19:46:42",
      "content": "<p>You've cast a vote Lucas.  On the other topic, someone got a triple downvote for saying quite the same. It was hilarious.</p>\n<p>However, he deserved since he is a notorius \"bronze medals collector\" and rarely give support to anyone. </p>",
      "rawMarkdown": "You've cast a vote Lucas.  On the other topic, someone got a triple downvote for saying quite the same. It was hilarious.\n\nHowever, he deserved since he is a notorius \"bronze medals collector\" and rarely give support to anyone.",
      "votes": null
    },
    {
      "id": "1918001",
      "postDate": "08/29/2022 08:35:19",
      "content": "<p>Just switch to https for link to work</p>",
      "rawMarkdown": "Just switch to https for link to work",
      "votes": null
    },
    {
      "id": "1918331",
      "postDate": "08/29/2022 13:58:29",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/malteluecken\" target=\"_blank\">@malteluecken</a> </p>\n<p>Thanks for clarifying the Pearson correlation.</p>\n<p>I have another question regarding how the final score is calculated. In \"Evaluation\" section, it says:</p>\n<p><em>For each observation in the Multiome data set, we compute the correlation between the ground-truth gene expressions and the predicted gene expressions</em></p>\n<p>Here, is each observation referring to a single cell? Can you clarify this a bit more?</p>\n<p>Thanks!<br>\nZhijian</p>",
      "rawMarkdown": "Hi @malteluecken \n\nThanks for clarifying the Pearson correlation.\n\nI have another question regarding how the final score is calculated. In \"Evaluation\" section, it says:\n\n*For each observation in the Multiome data set, we compute the correlation between the ground-truth gene expressions and the predicted gene expressions*\n\nHere, is each observation referring to a single cell? Can you clarify this a bit more?\n\nThanks!\nZhijian",
      "votes": null
    },
    {
      "id": "2000197",
      "postDate": "10/23/2022 06:04:30",
      "content": "<p><a href=\"https://www.kaggle.com/zhijianli\" target=\"_blank\">@zhijianli</a> , did you get any insight whether it is cell-cell correlation or gene-gene correlation ? Thank you</p>",
      "rawMarkdown": "zhijianli , did you get any insight whether it is cell-cell correlation or gene-gene correlation ? Thank you",
      "votes": null
    },
    {
      "id": "2029787",
      "postDate": "11/14/2022 23:06:44",
      "content": "<p>This is exactly what i was looking for.</p>",
      "rawMarkdown": "This is exactly what i was looking for.",
      "votes": null
    },
    {
      "id": "2030288",
      "postDate": "11/15/2022 10:57:53",
      "content": "<p>Thank you Alviswrw. You're doing great on this competition. </p>",
      "rawMarkdown": "Thank you Alviswrw. You're doing great on this competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1906792,
      "author_name": "jamesdao",
      "author_url": "",
      "post_date": "08/20/2022 07:42:39",
      "content": "<p>I would think something like R^2 would make more sense for an evaluation metric.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1907012,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "08/20/2022 11:54:08",
          "content": "<p>Thank you for your contribution and support JD.<br>\nMaybe we can suggest that to the hosts too : )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1909289,
          "author_name": "danielburkhardt",
          "author_url": "",
          "post_date": "08/22/2022 13:47:56",
          "content": "<p>R^2 would be a poor evaluation metric because it has no sign. We care that genes measured as highly expressed in the test set are predicted as highly expressed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1909349,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "08/22/2022 14:54:49",
          "content": "<p>Thank you Daniel for clarifying JD's suggestion.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1907068,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "08/20/2022 12:35:08",
      "content": "<p>tl dr        ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1907464,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "08/20/2022 19:58:45",
          "content": "<p>Hi Alexander <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> </p>\n<p>Indeed, it's too long.  TIKTOKish short way:</p>\n<p>\"Why Pearson Correlation?\"</p>\n<p>\" easy to interpret, easy to code and is intuitive.\"</p>\n<p>That answer above got 3 downvotes. </p>\n<p>As short as a TikTok mp3. <br>\nEven the funny comments, I read.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1908949,
      "author_name": "malteluecken",
      "author_url": "",
      "post_date": "08/22/2022 06:51:07",
      "content": "<p>This is actually something we thought a lot about, and it has to do with the data format and its interpretation. </p>\n<p>Single-cell RNA-seq data has to be normalized to be interpretable. When cells are profiled, they are sequenced at different depths. This means that there is more data for some cells and less for others purely for technical reasons. We correct for this by normalization to make cells comparable (a correction factor is typically computed per cell). We hosted a multi-modal data integration competition at NeurIPS last year where we normalized the data and then evaluated the prediction task via MSE to make it more accessible as one could easily directly optimize the evaluation metric. However, as the normalization had to be done per batch, and our chosen normalization method (the current SOTA method in the field) produced normalized data that had a different scale per run, prediction became very challenging as evaluated by MSE. Pearson correlation would have fixed this issue. You can read more about that in our PMLR paper on the competition results (<a href=\"http://proceedings.mlr.press/v176/lance22a.html)\" target=\"_blank\">http://proceedings.mlr.press/v176/lance22a.html)</a>.</p>\n<p>Secondly, biologically we don't actually care about the exact scale of the results. We know that we do not measure every molecule in a cell when it's profiled. Typically it's also not clear what proportion of molecules we capture. So this data type is probably always going to be relative. Thus, if you predict values that are 100 times higher than the ground truth data, but self-consistent, then that's just as good as any other scale. With Pearson correlation, we are not penalizing for differences in scales between predictions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1909655,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "08/22/2022 20:00:01",
          "content": "<p>Thank you for the embracing, valuable comment by competition host.  <br>\nThanks also for the paper link.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1910882,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "08/23/2022 18:50:17",
          "content": "<p>I'm sorry M. Luecken, I understood all your complimentary explanation about Pearson correlation started by D. Burkhardt</p>\n<p>Though  read \"PMLR paper on the competition results\"</p>\n<p>the link is resulting page 404<br>\n\"404<br>\nFile not found</p>\n<p>Though Vol 176 works:  \"Volume 176: NeurIPS 2021 Competitions and Demonstrations Track, 6-14 December 2021, Online\"<br>\n<a href=\"http://proceedings.mlr.press/v176/\" target=\"_blank\">http://proceedings.mlr.press/v176/</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1918001,
          "author_name": "kseniyapetrova",
          "author_url": "",
          "post_date": "08/29/2022 08:35:19",
          "content": "<p>Just switch to https for link to work</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1918331,
          "author_name": "zhijianli",
          "author_url": "",
          "post_date": "08/29/2022 13:58:29",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/malteluecken\" target=\"_blank\">@malteluecken</a> </p>\n<p>Thanks for clarifying the Pearson correlation.</p>\n<p>I have another question regarding how the final score is calculated. In \"Evaluation\" section, it says:</p>\n<p><em>For each observation in the Multiome data set, we compute the correlation between the ground-truth gene expressions and the predicted gene expressions</em></p>\n<p>Here, is each observation referring to a single cell? Can you clarify this a bit more?</p>\n<p>Thanks!<br>\nZhijian</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2000197,
          "author_name": "chandanpandey",
          "author_url": "",
          "post_date": "10/23/2022 06:04:30",
          "content": "<p><a href=\"https://www.kaggle.com/zhijianli\" target=\"_blank\">@zhijianli</a> , did you get any insight whether it is cell-cell correlation or gene-gene correlation ? Thank you</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1912528,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "08/24/2022 19:38:49",
      "content": "<p>Easy to understand, easy to code, easy to give nan after the first minutes of training.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1912535,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "08/24/2022 19:46:42",
          "content": "<p>You've cast a vote Lucas.  On the other topic, someone got a triple downvote for saying quite the same. It was hilarious.</p>\n<p>However, he deserved since he is a notorius \"bronze medals collector\" and rarely give support to anyone. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2029787,
      "author_name": "",
      "author_url": "",
      "post_date": "11/14/2022 23:06:44",
      "content": "<p>This is exactly what i was looking for.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2030288,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "11/15/2022 10:57:53",
          "content": "<p>Thank you Alviswrw. You're doing great on this competition. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1906538": "#Kaggle Competitions Evaluated by Pearson Coefficient\n\nUbiquant Market Prediction\nhttps://www.kaggle.com/competitions/ubiquant-market-prediction/overview\n\nG-Research Crypto Forecasting\nhttps://www.kaggle.com/competitions/g-research-crypto-forecasting/overview/evaluation\n\nU.S. Patent Phrase to Phrase Matching\nhttps://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/overview/evaluation\n\n#DISCUSSION TOPICS:\n\nPearson vs Spearman Correlation Coefficient  By Saurabh Thakur\nhttps://www.kaggle.com/discussions/getting-started/171951\n\nCorrelation : Pearson v/s Spearman  By Fatakdawala\nhttps://www.kaggle.com/discussions/getting-started/186658\n\nContest Evaluation Criteria - Pearson correlation coefficient - Summary for starters By Kalilur Rahman\nhttps://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/302153\n\nBest Loss function that optimize Pearson correlation  By Amed\nhttps://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/302181\n\nCustom Pearson Metric for LGBMRegressor  By RDizzI3\nhttps://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/302480\n\nUnderstanding Pearson Correlation Coefficient(PCC) By Poteman\nhttps://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/311260\n\nA simple way to use keras for the pearson correlation coefficient By Laura Fink\nhttps://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/316589\n\nWhy using pearson correlation as loss is tricky - A brief analysis  By Martin BB\nhttps://www.kaggle.com/competitions/ubiquant-market-prediction/discussion/306322\n\nDirectly maximizing Pearson correlation got poor cv. (With CustomLoss code) By Masaki Aota\nhttps://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/discussion/327502\n\n#Why Pearson Correlation?? \n\n By Abhranta Panigrahi\nhttps://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching/discussion/317851\n\n\"I have a genuine doubt. Can anyone please explain why we are using pearson correlation ?\"  Abhranta's words.\n\nWhenever someone is shy to ask something they use the expression \"genuine doubt\"  That's so funny, it's almost to say \"I'm sorry for asking that question.\"\n\n#Spoiler alert: if your answer was \"easy to interpret, easy to code and is intuitive.\" You're going to get a triple easy, intuitive downvote : )\n\nI simply can help laughing at that. If you want to downvote. Go ahead.",
    "1906792": "I would think something like R^2 would make more sense for an evaluation metric.",
    "1907012": "Thank you for your contribution and support JD.\nMaybe we can suggest that to the hosts too : )",
    "1907068": "tl dr        ?",
    "1907464": "Hi Alexander @alexandervc \n\nIndeed, it's too long.  TIKTOKish short way:\n\n\"Why Pearson Correlation?\"\n\n\" easy to interpret, easy to code and is intuitive.\"\n\nThat answer above got 3 downvotes. \n\nAs short as a TikTok mp3. \nEven the funny comments, I read.",
    "1908949": "This is actually something we thought a lot about, and it has to do with the data format and its interpretation. \n\nSingle-cell RNA-seq data has to be normalized to be interpretable. When cells are profiled, they are sequenced at different depths. This means that there is more data for some cells and less for others purely for technical reasons. We correct for this by normalization to make cells comparable (a correction factor is typically computed per cell). We hosted a multi-modal data integration competition at NeurIPS last year where we normalized the data and then evaluated the prediction task via MSE to make it more accessible as one could easily directly optimize the evaluation metric. However, as the normalization had to be done per batch, and our chosen normalization method (the current SOTA method in the field) produced normalized data that had a different scale per run, prediction became very challenging as evaluated by MSE. Pearson correlation would have fixed this issue. You can read more about that in our PMLR paper on the competition results (http://proceedings.mlr.press/v176/lance22a.html).\n\nSecondly, biologically we don't actually care about the exact scale of the results. We know that we do not measure every molecule in a cell when it's profiled. Typically it's also not clear what proportion of molecules we capture. So this data type is probably always going to be relative. Thus, if you predict values that are 100 times higher than the ground truth data, but self-consistent, then that's just as good as any other scale. With Pearson correlation, we are not penalizing for differences in scales between predictions.",
    "1909289": "R^2 would be a poor evaluation metric because it has no sign. We care that genes measured as highly expressed in the test set are predicted as highly expressed.",
    "1909349": "Thank you Daniel for clarifying JD's suggestion.",
    "1909655": "Thank you for the embracing, valuable comment by competition host.  \nThanks also for the paper link.",
    "1910882": "I'm sorry M. Luecken, I understood all your complimentary explanation about Pearson correlation started by D. Burkhardt\n\nThough  read \"PMLR paper on the competition results\"\n\nthe link is resulting page 404\n\"404\nFile not found\n\nThough Vol 176 works:  \"Volume 176: NeurIPS 2021 Competitions and Demonstrations Track, 6-14 December 2021, Online\"\nhttp://proceedings.mlr.press/v176/",
    "1912528": "Easy to understand, easy to code, easy to give nan after the first minutes of training.",
    "1912535": "You've cast a vote Lucas.  On the other topic, someone got a triple downvote for saying quite the same. It was hilarious.\n\nHowever, he deserved since he is a notorius \"bronze medals collector\" and rarely give support to anyone.",
    "1918001": "Just switch to https for link to work",
    "1918331": "Hi @malteluecken \n\nThanks for clarifying the Pearson correlation.\n\nI have another question regarding how the final score is calculated. In \"Evaluation\" section, it says:\n\n*For each observation in the Multiome data set, we compute the correlation between the ground-truth gene expressions and the predicted gene expressions*\n\nHere, is each observation referring to a single cell? Can you clarify this a bit more?\n\nThanks!\nZhijian",
    "2000197": "zhijianli , did you get any insight whether it is cell-cell correlation or gene-gene correlation ? Thank you",
    "2029787": "This is exactly what i was looking for.",
    "2030288": "Thank you Alviswrw. You're doing great on this competition."
  },
  "source": "meta"
}