{
  "id": 360253,
  "title": "Tricks and understanding the metric - predict 138, not 140 targets CITE-seq   (because Pearson - by targets, not sampleswisely (as it usually happens))",
  "url": "/competitions/open-problems-multimodal/discussion/360253",
  "author_name": "",
  "post_date": "2022-10-15T19:04:53.385843800Z",
  "votes": 15,
  "comment_count": 3,
  "views": 0,
  "content": "<p>We need to predict 140 targets for CITEseq,<br>\nbut actually there are only 138 predictions can be made, and the rest 2 - calculated from them. <br>\n(Simialr for Multiome - TWO predictions are UNNECESSARY). <br>\nBecause correlation  coefficient actually \"eats\" two degrees of freedom,<br>\nthat means if you predict is Y, then for any floating scalars  a,b :   a*Y+b - will have EXACTLY THE SAME CORRELATION with any vector. </p>\n<p><strong>Question:</strong> What are your ideas -  what are the  ways to exploit that redundancy ?  </p>\n<p>Obviously that can be many ways : for example normalize targets such that mean(Y) = 0 , and Y_1 = 1,<br>\nand predict only Y_3, Y_4 … , after prediction put: Y_2 = -1 - Y_3-Y_4 - ….. <br>\nSo here are Y_1 and Y_2 are extremely distinguished.<br>\nBut why you should distinguish Y_1 and Y_2 , not Y_3 and Y_10 ? <br>\nWhat can be the best choice to choose selected pair ?<br>\nBut probably such way is far from being the optimal one for any choices of distinguished Y_i , Y_j  ? <br>\nThere should be more clever way…</p>\n<p>What is currently done in many public notebooks - before training : <br>\nJust renormalize: Y -&gt; mean(Y) = 0, std(Y) = 1<br>\nAnd that it is all .<br>\nSo somehow it seems to me that way does not fully exploit the opportunities we have.<br>\nWhat do you think ?<br>\n(By the way sometimes it might improve score from 0.803 to 0.805 - for Ridge on CITEseq<br>\nCompare <a href=\"https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes?scriptVersionId=108200703\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes?scriptVersionId=108200703</a><br>\nwith the previous version)</p>\n<p>PS</p>\n<p>All that is due to Pearson correlation coefficient is calculated \"targetwisely\" not \"samplewisely\" as usually.<br>\nSo despite current LB 0.81+ might seems as high correlation it does NOT mean prediction is any how good.<br>\nPredictions are VERY BAD , because you should compare them with prediction by averages which gives 0.71,<br>\nand about 0.74 if cell type is included. So current predictions are only about 5% \"above zero\". </p>\n<p>PSPS<br>\nThat does not mean the Pearson is bad metric  here - just one should be careful with interpretations. </p>\n<p>PSPSPS<br>\nIs it too long to read ?  ))</p>",
  "messages": [
    {
      "id": "1989190",
      "postDate": "10/15/2022 19:04:53",
      "content": "<p>We need to predict 140 targets for CITEseq,<br>\nbut actually there are only 138 predictions can be made, and the rest 2 - calculated from them. <br>\n(Simialr for Multiome - TWO predictions are UNNECESSARY). <br>\nBecause correlation  coefficient actually \"eats\" two degrees of freedom,<br>\nthat means if you predict is Y, then for any floating scalars  a,b :   a*Y+b - will have EXACTLY THE SAME CORRELATION with any vector. </p>\n<p><strong>Question:</strong> What are your ideas -  what are the  ways to exploit that redundancy ?  </p>\n<p>Obviously that can be many ways : for example normalize targets such that mean(Y) = 0 , and Y_1 = 1,<br>\nand predict only Y_3, Y_4 … , after prediction put: Y_2 = -1 - Y_3-Y_4 - ….. <br>\nSo here are Y_1 and Y_2 are extremely distinguished.<br>\nBut why you should distinguish Y_1 and Y_2 , not Y_3 and Y_10 ? <br>\nWhat can be the best choice to choose selected pair ?<br>\nBut probably such way is far from being the optimal one for any choices of distinguished Y_i , Y_j  ? <br>\nThere should be more clever way…</p>\n<p>What is currently done in many public notebooks - before training : <br>\nJust renormalize: Y -&gt; mean(Y) = 0, std(Y) = 1<br>\nAnd that it is all .<br>\nSo somehow it seems to me that way does not fully exploit the opportunities we have.<br>\nWhat do you think ?<br>\n(By the way sometimes it might improve score from 0.803 to 0.805 - for Ridge on CITEseq<br>\nCompare <a href=\"https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes?scriptVersionId=108200703\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes?scriptVersionId=108200703</a><br>\nwith the previous version)</p>\n<p>PS</p>\n<p>All that is due to Pearson correlation coefficient is calculated \"targetwisely\" not \"samplewisely\" as usually.<br>\nSo despite current LB 0.81+ might seems as high correlation it does NOT mean prediction is any how good.<br>\nPredictions are VERY BAD , because you should compare them with prediction by averages which gives 0.71,<br>\nand about 0.74 if cell type is included. So current predictions are only about 5% \"above zero\". </p>\n<p>PSPS<br>\nThat does not mean the Pearson is bad metric  here - just one should be careful with interpretations. </p>\n<p>PSPSPS<br>\nIs it too long to read ?  ))</p>",
      "rawMarkdown": "We need to predict 140 targets for CITEseq,\nbut actually there are only 138 predictions can be made, and the rest 2 - calculated from them. \n(Simialr for Multiome - TWO predictions are UNNECESSARY). \nBecause correlation  coefficient actually \"eats\" two degrees of freedom,\nthat means if you predict is Y, then for any floating scalars  a,b :   a*Y+b - will have EXACTLY THE SAME CORRELATION with any vector. \n\n**Question:** What are your ideas -  what are the  ways to exploit that redundancy ?  \n\nObviously that can be many ways : for example normalize targets such that mean(Y) = 0 , and Y_1 = 1,\nand predict only Y_3, Y_4 ... , after prediction put: Y_2 = -1 - Y_3-Y_4 - ..... \nSo here are Y_1 and Y_2 are extremely distinguished.\nBut why you should distinguish Y_1 and Y_2 , not Y_3 and Y_10 ? \nWhat can be the best choice to choose selected pair ?\nBut probably such way is far from being the optimal one for any choices of distinguished Y_i , Y_j  ? \nThere should be more clever way...\n \nWhat is currently done in many public notebooks - before training : \nJust renormalize: Y -> mean(Y) = 0, std(Y) = 1\nAnd that it is all .\nSo somehow it seems to me that way does not fully exploit the opportunities we have.\nWhat do you think ?\n(By the way sometimes it might improve score from 0.803 to 0.805 - for Ridge on CITEseq\nCompare https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes?scriptVersionId=108200703\nwith the previous version)\n\nPS\n\nAll that is due to Pearson correlation coefficient is calculated \"targetwisely\" not \"samplewisely\" as usually.\nSo despite current LB 0.81+ might seems as high correlation it does NOT mean prediction is any how good.\nPredictions are VERY BAD , because you should compare them with prediction by averages which gives 0.71,\nand about 0.74 if cell type is included. So current predictions are only about 5% \"above zero\". \n\nPSPS\nThat does not mean the Pearson is bad metric  here - just one should be careful with interpretations. \n\nPSPSPS\nIs it too long to read ?  ))",
      "votes": null
    },
    {
      "id": "2004313",
      "postDate": "10/26/2022 07:23:04",
      "content": "<p><a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> , a small doubt about the line \"All that is due to Pearson correlation coefficient is calculated \"targetwisely\" not \"samplewisely\" as usually.\" Targetwisely here implies gene-gene correlation (for Multiome) and protein-protein correlation (for CITE-seq) ? Also, samplewisely implies cell-cell correlation ? <br>\nIf, Targetwisely here implies gene-gene correlation then the correlation_score(y_true, y_pred)correlation_score funtion used in the notebook is doing cell-cell correlation. It would be helpful if you could clarify my doubts. Thank you. </p>",
      "rawMarkdown": "alexandervc , a small doubt about the line \"All that is due to Pearson correlation coefficient is calculated \"targetwisely\" not \"samplewisely\" as usually.\" Targetwisely here implies gene-gene correlation (for Multiome) and protein-protein correlation (for CITE-seq) ? Also, samplewisely implies cell-cell correlation ? \nIf, Targetwisely here implies gene-gene correlation then the correlation_score(y_true, y_pred)correlation_score funtion used in the notebook is doing cell-cell correlation. It would be helpful if you could clarify my doubts. Thank you.",
      "votes": null
    },
    {
      "id": "2004379",
      "postDate": "10/26/2022 08:22:52",
      "content": "<p><a href=\"https://www.kaggle.com/chandanpandey\" target=\"_blank\">@chandanpandey</a> <br>\nSorry If I was not clear. <br>\nI mean simple things may be not put them in right order:</p>\n<p>Imagine you need to predict JUST ONE target (not multitarget task like here): \"Y\"<br>\nSo \"Y\" is vector along \"sample\" dimension.<br>\nAnd one may try to put a metric to be  corrcoef( Y_pred, Y_true). <br>\nThat I mean \"samplewise\".<br>\nAs, we know it is BAD metric for many reasons:<br>\n(See post by <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a> <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346518\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346518</a> )</p>\n<p>Now if you have Multitarget task, then :  \"Y\" is a matrix.<br>\nSo you have two options - how to calcualte the correlation coefficient - </p>\n<p>One is naive generalazation of the \"samplewise\" recipe -<br>\nSo it would be  average_{over Targets} (  corrcoef( Y_pred[for Target \"T\"] ,   Y_true[ for Target \"T\" ] ).<br>\nAgain it is bad idea. It is NOT used here.</p>\n<p>And you can do \"targetwise\" : <br>\nSo it would be  average_{over sample} (  corrcoef( Y_pred[for Samples \"S\"] ,   Y_true[ for Sample \"S\" ] ).<br>\nThat is what, is used in that competition.<br>\nOrganizers say that last year competition experience suggest that this is better metric, because it is more stable,<br>\nand my small games with data confirms that. But who knows… </p>\n<p>So the code is : </p>\n<pre><code>def correlation_score(y_true, y_pred):\n    \"\"\"Scores the predictions according to the competition rules. \n\n    It is assumed that the predictions are not constant.\n\n    Returns the average of each sample's Pearson correlation coefficient\"\"\"\n    if type(y_true) == pd.DataFrame: y_true = y_true.values\n    if type(y_pred) == pd.DataFrame: y_pred = y_pred.values\n    corrsum = 0\n    for i in range(len(y_true)):\n        corrsum += np.corrcoef(y_true[i], y_pred[i])[1, 0]\n    return corrsum / len(y_true)\n</code></pre>",
      "rawMarkdown": "chandanpandey \nSorry If I was not clear. \nI mean simple things may be not put them in right order:\n\nImagine you need to predict JUST ONE target (not multitarget task like here): \"Y\"\nSo \"Y\" is vector along \"sample\" dimension.\nAnd one may try to put a metric to be  corrcoef( Y_pred, Y_true). \nThat I mean \"samplewise\".\nAs, we know it is BAD metric for many reasons:\n(See post by @mpwolke https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346518 )\n\nNow if you have Multitarget task, then :  \"Y\" is a matrix.\nSo you have two options - how to calcualte the correlation coefficient - \n\nOne is naive generalazation of the \"samplewise\" recipe -\nSo it would be  average_{over Targets} (  corrcoef( Y_pred[for Target \"T\"] ,   Y_true[ for Target \"T\" ] ).\nAgain it is bad idea. It is NOT used here.\n\nAnd you can do \"targetwise\" : \nSo it would be  average_{over sample} (  corrcoef( Y_pred[for Samples \"S\"] ,   Y_true[ for Sample \"S\" ] ).\nThat is what, is used in that competition.\nOrganizers say that last year competition experience suggest that this is better metric, because it is more stable,\nand my small games with data confirms that. But who knows... \n\nSo the code is : \n\n```\ndef correlation_score(y_true, y_pred):\n    \"\"\"Scores the predictions according to the competition rules. \n    \n    It is assumed that the predictions are not constant.\n    \n    Returns the average of each sample's Pearson correlation coefficient\"\"\"\n    if type(y_true) == pd.DataFrame: y_true = y_true.values\n    if type(y_pred) == pd.DataFrame: y_pred = y_pred.values\n    corrsum = 0\n    for i in range(len(y_true)):\n        corrsum += np.corrcoef(y_true[i], y_pred[i])[1, 0]\n    return corrsum / len(y_true)\n\n```",
      "votes": null
    },
    {
      "id": "2005891",
      "postDate": "10/27/2022 09:53:39",
      "content": "<p>Thank you very much </p>",
      "rawMarkdown": "Thank you very much",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2004313,
      "author_name": "chandanpandey",
      "author_url": "",
      "post_date": "10/26/2022 07:23:04",
      "content": "<p><a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> , a small doubt about the line \"All that is due to Pearson correlation coefficient is calculated \"targetwisely\" not \"samplewisely\" as usually.\" Targetwisely here implies gene-gene correlation (for Multiome) and protein-protein correlation (for CITE-seq) ? Also, samplewisely implies cell-cell correlation ? <br>\nIf, Targetwisely here implies gene-gene correlation then the correlation_score(y_true, y_pred)correlation_score funtion used in the notebook is doing cell-cell correlation. It would be helpful if you could clarify my doubts. Thank you. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2004379,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "10/26/2022 08:22:52",
          "content": "<p><a href=\"https://www.kaggle.com/chandanpandey\" target=\"_blank\">@chandanpandey</a> <br>\nSorry If I was not clear. <br>\nI mean simple things may be not put them in right order:</p>\n<p>Imagine you need to predict JUST ONE target (not multitarget task like here): \"Y\"<br>\nSo \"Y\" is vector along \"sample\" dimension.<br>\nAnd one may try to put a metric to be  corrcoef( Y_pred, Y_true). <br>\nThat I mean \"samplewise\".<br>\nAs, we know it is BAD metric for many reasons:<br>\n(See post by <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a> <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346518\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346518</a> )</p>\n<p>Now if you have Multitarget task, then :  \"Y\" is a matrix.<br>\nSo you have two options - how to calcualte the correlation coefficient - </p>\n<p>One is naive generalazation of the \"samplewise\" recipe -<br>\nSo it would be  average_{over Targets} (  corrcoef( Y_pred[for Target \"T\"] ,   Y_true[ for Target \"T\" ] ).<br>\nAgain it is bad idea. It is NOT used here.</p>\n<p>And you can do \"targetwise\" : <br>\nSo it would be  average_{over sample} (  corrcoef( Y_pred[for Samples \"S\"] ,   Y_true[ for Sample \"S\" ] ).<br>\nThat is what, is used in that competition.<br>\nOrganizers say that last year competition experience suggest that this is better metric, because it is more stable,<br>\nand my small games with data confirms that. But who knows… </p>\n<p>So the code is : </p>\n<pre><code>def correlation_score(y_true, y_pred):\n    \"\"\"Scores the predictions according to the competition rules. \n\n    It is assumed that the predictions are not constant.\n\n    Returns the average of each sample's Pearson correlation coefficient\"\"\"\n    if type(y_true) == pd.DataFrame: y_true = y_true.values\n    if type(y_pred) == pd.DataFrame: y_pred = y_pred.values\n    corrsum = 0\n    for i in range(len(y_true)):\n        corrsum += np.corrcoef(y_true[i], y_pred[i])[1, 0]\n    return corrsum / len(y_true)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2005891,
          "author_name": "chandanpandey",
          "author_url": "",
          "post_date": "10/27/2022 09:53:39",
          "content": "<p>Thank you very much </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1989190": "We need to predict 140 targets for CITEseq,\nbut actually there are only 138 predictions can be made, and the rest 2 - calculated from them. \n(Simialr for Multiome - TWO predictions are UNNECESSARY). \nBecause correlation  coefficient actually \"eats\" two degrees of freedom,\nthat means if you predict is Y, then for any floating scalars  a,b :   a*Y+b - will have EXACTLY THE SAME CORRELATION with any vector. \n\n**Question:** What are your ideas -  what are the  ways to exploit that redundancy ?  \n\nObviously that can be many ways : for example normalize targets such that mean(Y) = 0 , and Y_1 = 1,\nand predict only Y_3, Y_4 ... , after prediction put: Y_2 = -1 - Y_3-Y_4 - ..... \nSo here are Y_1 and Y_2 are extremely distinguished.\nBut why you should distinguish Y_1 and Y_2 , not Y_3 and Y_10 ? \nWhat can be the best choice to choose selected pair ?\nBut probably such way is far from being the optimal one for any choices of distinguished Y_i , Y_j  ? \nThere should be more clever way...\n \nWhat is currently done in many public notebooks - before training : \nJust renormalize: Y -> mean(Y) = 0, std(Y) = 1\nAnd that it is all .\nSo somehow it seems to me that way does not fully exploit the opportunities we have.\nWhat do you think ?\n(By the way sometimes it might improve score from 0.803 to 0.805 - for Ridge on CITEseq\nCompare https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes?scriptVersionId=108200703\nwith the previous version)\n\nPS\n\nAll that is due to Pearson correlation coefficient is calculated \"targetwisely\" not \"samplewisely\" as usually.\nSo despite current LB 0.81+ might seems as high correlation it does NOT mean prediction is any how good.\nPredictions are VERY BAD , because you should compare them with prediction by averages which gives 0.71,\nand about 0.74 if cell type is included. So current predictions are only about 5% \"above zero\". \n\nPSPS\nThat does not mean the Pearson is bad metric  here - just one should be careful with interpretations. \n\nPSPSPS\nIs it too long to read ?  ))",
    "2004313": "alexandervc , a small doubt about the line \"All that is due to Pearson correlation coefficient is calculated \"targetwisely\" not \"samplewisely\" as usually.\" Targetwisely here implies gene-gene correlation (for Multiome) and protein-protein correlation (for CITE-seq) ? Also, samplewisely implies cell-cell correlation ? \nIf, Targetwisely here implies gene-gene correlation then the correlation_score(y_true, y_pred)correlation_score funtion used in the notebook is doing cell-cell correlation. It would be helpful if you could clarify my doubts. Thank you.",
    "2004379": "chandanpandey \nSorry If I was not clear. \nI mean simple things may be not put them in right order:\n\nImagine you need to predict JUST ONE target (not multitarget task like here): \"Y\"\nSo \"Y\" is vector along \"sample\" dimension.\nAnd one may try to put a metric to be  corrcoef( Y_pred, Y_true). \nThat I mean \"samplewise\".\nAs, we know it is BAD metric for many reasons:\n(See post by @mpwolke https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346518 )\n\nNow if you have Multitarget task, then :  \"Y\" is a matrix.\nSo you have two options - how to calcualte the correlation coefficient - \n\nOne is naive generalazation of the \"samplewise\" recipe -\nSo it would be  average_{over Targets} (  corrcoef( Y_pred[for Target \"T\"] ,   Y_true[ for Target \"T\" ] ).\nAgain it is bad idea. It is NOT used here.\n\nAnd you can do \"targetwise\" : \nSo it would be  average_{over sample} (  corrcoef( Y_pred[for Samples \"S\"] ,   Y_true[ for Sample \"S\" ] ).\nThat is what, is used in that competition.\nOrganizers say that last year competition experience suggest that this is better metric, because it is more stable,\nand my small games with data confirms that. But who knows... \n\nSo the code is : \n\n```\ndef correlation_score(y_true, y_pred):\n    \"\"\"Scores the predictions according to the competition rules. \n    \n    It is assumed that the predictions are not constant.\n    \n    Returns the average of each sample's Pearson correlation coefficient\"\"\"\n    if type(y_true) == pd.DataFrame: y_true = y_true.values\n    if type(y_pred) == pd.DataFrame: y_pred = y_pred.values\n    corrsum = 0\n    for i in range(len(y_true)):\n        corrsum += np.corrcoef(y_true[i], y_pred[i])[1, 0]\n    return corrsum / len(y_true)\n\n```",
    "2005891": "Thank you very much"
  },
  "source": "meta"
}