{
  "id": 360067,
  "title": "Some questions about high score",
  "url": "/competitions/open-problems-multimodal/discussion/360067",
  "author_name": "",
  "post_date": "2022-10-14T20:36:45.962879100Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi, I notice that in this notebook:<br>\n<a href=\"https://www.kaggle.com/code/jsmithperera/citeseq-keras-multiome-5x5\" target=\"_blank\">https://www.kaggle.com/code/jsmithperera/citeseq-keras-multiome-5x5</a></p>\n<p>It seems that the prediction results for ADT is not very realiable, because the scope is extremely large. However, its correlation score is still pretty high. Are there any problems existing in the evaluation process? Thanks.</p>",
  "messages": [
    {
      "id": "1987795",
      "postDate": "10/14/2022 20:36:45",
      "content": "<p>Hi, I notice that in this notebook:<br>\n<a href=\"https://www.kaggle.com/code/jsmithperera/citeseq-keras-multiome-5x5\" target=\"_blank\">https://www.kaggle.com/code/jsmithperera/citeseq-keras-multiome-5x5</a></p>\n<p>It seems that the prediction results for ADT is not very realiable, because the scope is extremely large. However, its correlation score is still pretty high. Are there any problems existing in the evaluation process? Thanks.</p>",
      "rawMarkdown": "Hi, I notice that in this notebook:\nhttps://www.kaggle.com/code/jsmithperera/citeseq-keras-multiome-5x5\n\nIt seems that the prediction results for ADT is not very realiable, because the scope is extremely large. However, its correlation score is still pretty high. Are there any problems existing in the evaluation process? Thanks.",
      "votes": null
    },
    {
      "id": "1987854",
      "postDate": "10/14/2022 21:14:38",
      "content": "<p>When you use NO model at all - but just take averages over of targets - you already get correlation score very high around 0.71<br>\n(and taking cell types into account you will get even more - do not rememeber exactly - something like 0.74).</p>\n<p>That means even if \"r2-score\" is exactly zero (and so no prediction at all - is happening), the correlation score - is high - 0.7+</p>\n<p>So all the best models with 0.81+ improve very little over the NO model averages.<br>\nThat essentially means that our prediction abilities are very little. </p>\n<p>That is not about reliability , but just about the interpretation of the results.<br>\nThe correlation score is high because variations of targets around means are not so high, so high correlation score does not reflect good prediction. </p>",
      "rawMarkdown": "When you use NO model at all - but just take averages over of targets - you already get correlation score very high around 0.71\n(and taking cell types into account you will get even more - do not rememeber exactly - something like 0.74).\n\nThat means even if \"r2-score\" is exactly zero (and so no prediction at all - is happening), the correlation score - is high - 0.7+\n\nSo all the best models with 0.81+ improve very little over the NO model averages.\nThat essentially means that our prediction abilities are very little. \n\nThat is not about reliability , but just about the interpretation of the results.\nThe correlation score is high because variations of targets around means are not so high, so high correlation score does not reflect good prediction.",
      "votes": null
    },
    {
      "id": "1987871",
      "postDate": "10/14/2022 21:33:17",
      "content": "<p>Good observation, I agree with you. Therefore, I hope the organizer of this competition can provide some explanation to us, otherwise the high rank may not be equal to best predictor. </p>",
      "rawMarkdown": "Good observation, I agree with you. Therefore, I hope the organizer of this competition can provide some explanation to us, otherwise the high rank may not be equal to best predictor.",
      "votes": null
    },
    {
      "id": "1991183",
      "postDate": "10/17/2022 03:15:08",
      "content": "<p>Hi there. I observed the same thing which is pretty strange that the prediction with high score gives values like -97. I think this is because it only cares about the angle between vectors in the space and ignore the scale, i.e. (-1, 2) is the pretty much the same as (-9.8, 22). Normalization may be useful if we really want to use the prediction in the real biological research, clearly -97 is not helpful and makes cells crazy. But for this competition I am not going to change the target function and I will use this large strange values, maybe the model can train better in this scale and be more sensitive.</p>",
      "rawMarkdown": "Hi there. I observed the same thing which is pretty strange that the prediction with high score gives values like -97. I think this is because it only cares about the angle between vectors in the space and ignore the scale, i.e. (-1, 2) is the pretty much the same as (-9.8, 22). Normalization may be useful if we really want to use the prediction in the real biological research, clearly -97 is not helpful and makes cells crazy. But for this competition I am not going to change the target function and I will use this large strange values, maybe the model can train better in this scale and be more sensitive.",
      "votes": null
    },
    {
      "id": "1991949",
      "postDate": "10/17/2022 13:13:34",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jianyulin1999\" target=\"_blank\">@jianyulin1999</a> , the score uses correlation. If you apply a factor 1 million to your submission, your results will be absurd but the correlations (and the score) remain the same. </p>",
      "rawMarkdown": "Hi @jianyulin1999 , the score uses correlation. If you apply a factor 1 million to your submission, your results will be absurd but the correlations (and the score) remain the same.",
      "votes": null
    },
    {
      "id": "2000724",
      "postDate": "10/23/2022 14:05:24",
      "content": "<p>Thank you that is what am guessing, now I have more confidence. Thank you for telling me. </p>",
      "rawMarkdown": "Thank you that is what am guessing, now I have more confidence. Thank you for telling me.",
      "votes": null
    },
    {
      "id": "2003104",
      "postDate": "10/25/2022 09:50:02",
      "content": "<p><a href=\"https://www.kaggle.com/jianyulin1999\" target=\"_blank\">@jianyulin1999</a>, however when you ensemble, it is important that you first normalize each model, otherwise the result is unpredictable.</p>",
      "rawMarkdown": "jianyulin1999, however when you ensemble, it is important that you first normalize each model, otherwise the result is unpredictable.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1987854,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "10/14/2022 21:14:38",
      "content": "<p>When you use NO model at all - but just take averages over of targets - you already get correlation score very high around 0.71<br>\n(and taking cell types into account you will get even more - do not rememeber exactly - something like 0.74).</p>\n<p>That means even if \"r2-score\" is exactly zero (and so no prediction at all - is happening), the correlation score - is high - 0.7+</p>\n<p>So all the best models with 0.81+ improve very little over the NO model averages.<br>\nThat essentially means that our prediction abilities are very little. </p>\n<p>That is not about reliability , but just about the interpretation of the results.<br>\nThe correlation score is high because variations of targets around means are not so high, so high correlation score does not reflect good prediction. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1987871,
          "author_name": "llttyy",
          "author_url": "",
          "post_date": "10/14/2022 21:33:17",
          "content": "<p>Good observation, I agree with you. Therefore, I hope the organizer of this competition can provide some explanation to us, otherwise the high rank may not be equal to best predictor. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1991183,
      "author_name": "jianyulin1999",
      "author_url": "",
      "post_date": "10/17/2022 03:15:08",
      "content": "<p>Hi there. I observed the same thing which is pretty strange that the prediction with high score gives values like -97. I think this is because it only cares about the angle between vectors in the space and ignore the scale, i.e. (-1, 2) is the pretty much the same as (-9.8, 22). Normalization may be useful if we really want to use the prediction in the real biological research, clearly -97 is not helpful and makes cells crazy. But for this competition I am not going to change the target function and I will use this large strange values, maybe the model can train better in this scale and be more sensitive.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1991949,
          "author_name": "gehallak",
          "author_url": "",
          "post_date": "10/17/2022 13:13:34",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jianyulin1999\" target=\"_blank\">@jianyulin1999</a> , the score uses correlation. If you apply a factor 1 million to your submission, your results will be absurd but the correlations (and the score) remain the same. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2000724,
          "author_name": "jianyulin1999",
          "author_url": "",
          "post_date": "10/23/2022 14:05:24",
          "content": "<p>Thank you that is what am guessing, now I have more confidence. Thank you for telling me. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2003104,
          "author_name": "gehallak",
          "author_url": "",
          "post_date": "10/25/2022 09:50:02",
          "content": "<p><a href=\"https://www.kaggle.com/jianyulin1999\" target=\"_blank\">@jianyulin1999</a>, however when you ensemble, it is important that you first normalize each model, otherwise the result is unpredictable.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1987795": "Hi, I notice that in this notebook:\nhttps://www.kaggle.com/code/jsmithperera/citeseq-keras-multiome-5x5\n\nIt seems that the prediction results for ADT is not very realiable, because the scope is extremely large. However, its correlation score is still pretty high. Are there any problems existing in the evaluation process? Thanks.",
    "1987854": "When you use NO model at all - but just take averages over of targets - you already get correlation score very high around 0.71\n(and taking cell types into account you will get even more - do not rememeber exactly - something like 0.74).\n\nThat means even if \"r2-score\" is exactly zero (and so no prediction at all - is happening), the correlation score - is high - 0.7+\n\nSo all the best models with 0.81+ improve very little over the NO model averages.\nThat essentially means that our prediction abilities are very little. \n\nThat is not about reliability , but just about the interpretation of the results.\nThe correlation score is high because variations of targets around means are not so high, so high correlation score does not reflect good prediction.",
    "1987871": "Good observation, I agree with you. Therefore, I hope the organizer of this competition can provide some explanation to us, otherwise the high rank may not be equal to best predictor.",
    "1991183": "Hi there. I observed the same thing which is pretty strange that the prediction with high score gives values like -97. I think this is because it only cares about the angle between vectors in the space and ignore the scale, i.e. (-1, 2) is the pretty much the same as (-9.8, 22). Normalization may be useful if we really want to use the prediction in the real biological research, clearly -97 is not helpful and makes cells crazy. But for this competition I am not going to change the target function and I will use this large strange values, maybe the model can train better in this scale and be more sensitive.",
    "1991949": "Hi @jianyulin1999 , the score uses correlation. If you apply a factor 1 million to your submission, your results will be absurd but the correlations (and the score) remain the same.",
    "2000724": "Thank you that is what am guessing, now I have more confidence. Thank you for telling me.",
    "2003104": "jianyulin1999, however when you ensemble, it is important that you first normalize each model, otherwise the result is unpredictable."
  },
  "source": "meta"
}