{
  "id": 475331,
  "title": "Is the row sum 1 when scoring a submission?",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/475331",
  "author_name": "",
  "post_date": "2024-02-08T02:33:49.579219Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Is the row sum 1 when scoring a submission?</p>\n<p>The sum of the rows in submission.csv must be 1 or an error will occur.<br>\nBut when I look at the targets in train.csv, the sum is not 1.<br>\nTherefore, the results will differ depending on whether you use target in train.csv or softmax(target) as objective.</p>\n<p>for example,</p>\n<pre><code>votes = torch.tensor([[,,]]).()  # votes  train\npredict = torch.nn.functional.softmax(votes, =)  # perfect predict\n\n = \npredict = torch.clip(predict, ,  - )\n\n = votes  ### row  NOT  when scoring a submission\n = torch.clip(, ,  - )\nscore =  * (.() - predict.())\n(score.mean()) # \n\n = torch.nn.functional.softmax(votes, =)  ### row  IS  when scoring a submission\n = torch.clip(, ,  - )\nscore =  * (.() - predict.())\n(score.mean()) # \n</code></pre>\n<p>Also, if the sum of the rows when scoring is not 1, is it okay to say that the score will not be 0.0 even if I make a perfect prediction?</p>",
  "messages": [
    {
      "id": "2642203",
      "postDate": "02/08/2024 02:33:49",
      "content": "<p>Is the row sum 1 when scoring a submission?</p>\n<p>The sum of the rows in submission.csv must be 1 or an error will occur.<br>\nBut when I look at the targets in train.csv, the sum is not 1.<br>\nTherefore, the results will differ depending on whether you use target in train.csv or softmax(target) as objective.</p>\n<p>for example,</p>\n<pre><code>votes = torch.tensor([[,,]]).()  # votes  train\npredict = torch.nn.functional.softmax(votes, =)  # perfect predict\n\n = \npredict = torch.clip(predict, ,  - )\n\n = votes  ### row  NOT  when scoring a submission\n = torch.clip(, ,  - )\nscore =  * (.() - predict.())\n(score.mean()) # \n\n = torch.nn.functional.softmax(votes, =)  ### row  IS  when scoring a submission\n = torch.clip(, ,  - )\nscore =  * (.() - predict.())\n(score.mean()) # \n</code></pre>\n<p>Also, if the sum of the rows when scoring is not 1, is it okay to say that the score will not be 0.0 even if I make a perfect prediction?</p>",
      "rawMarkdown": "Is the row sum 1 when scoring a submission?\n\nThe sum of the rows in submission.csv must be 1 or an error will occur.\nBut when I look at the targets in train.csv, the sum is not 1.\nTherefore, the results will differ depending on whether you use target in train.csv or softmax(target) as objective.\n\nfor example,\n\n```\nvotes = torch.tensor([[5,0,2]]).float()  # votes in train\npredict = torch.nn.functional.softmax(votes, dim=1)  # perfect predict\n\nepsilon = 1e-10\npredict = torch.clip(predict, epsilon, 1 - epsilon)\n\ntarget = votes  ### row sum NOT 1 when scoring a submission\ntarget = torch.clip(target, epsilon, 1 - epsilon)\nscore = target * (target.log() - predict.log())\nprint(score.mean()) # 1.0367\n\ntarget = torch.nn.functional.softmax(votes, dim=1)  ### row sum IS 1 when scoring a submission\ntarget = torch.clip(target, epsilon, 1 - epsilon)\nscore = target * (target.log() - predict.log())\nprint(score.mean()) # 0.0\n```\n\nAlso, if the sum of the rows when scoring is not 1, is it okay to say that the score will not be 0.0 even if I make a perfect prediction?",
      "votes": null
    },
    {
      "id": "2642833",
      "postDate": "02/08/2024 12:55:10",
      "content": "<p>Hi. The targets at train.csv are the absolute accumulation of the votes from the expert annotators. While the example submission has the normalized votations of equally voted 6 classes, 0.1666… So I'd say make a final normalization just in case. The results should be practically identhical.</p>",
      "rawMarkdown": "Hi. The targets at train.csv are the absolute accumulation of the votes from the expert annotators. While the example submission has the normalized votations of equally voted 6 classes, 0.1666... So I'd say make a final normalization just in case. The results should be practically identhical.",
      "votes": null
    },
    {
      "id": "2643609",
      "postDate": "02/09/2024 00:28:54",
      "content": "<p>thank you. However, this issue becomes important when checking the correlation between CV scores and leaderboards. That's why I want accurate information. The example below uses normalization instead of softmax, but the results are the same, as the two scores are different.</p>\n<pre><code>votes = torch.tensor([[,,]]).float()  \npredict = votes / votes.()  \n\nepsilon = \npredict = torch.clip(predict, epsilon,  - epsilon)\n\n = votes  \n = torch.clip(, epsilon,  - epsilon)\nscore =  * (.() - predict.())\n(score.()) \n\n = votes / votes.()  \n = torch.clip(, epsilon,  - epsilon)\nscore =  * (.() - predict.())\n(score.()) \n</code></pre>",
      "rawMarkdown": "thank you. However, this issue becomes important when checking the correlation between CV scores and leaderboards. That's why I want accurate information. The example below uses normalization instead of softmax, but the results are the same, as the two scores are different.\n\n```\nvotes = torch.tensor([[5,0,2]]).float()  # votes in train\npredict = votes / votes.sum()  # perfect predict in normalize\n\nepsilon = 1e-10\npredict = torch.clip(predict, epsilon, 1 - epsilon)\n\ntarget = votes  ### row sum NOT 1 when scoring a submission\ntarget = torch.clip(target, epsilon, 1 - epsilon)\nscore = target * (target.log() - predict.log())\nprint(score.mean()) # 0.5297\n\ntarget = votes / votes.sum()  ### row sum IS 1 when scoring a submission\ntarget = torch.clip(target, epsilon, 1 - epsilon)\nscore = target * (target.log() - predict.log())\nprint(score.mean()) # 0.0\n```",
      "votes": null
    },
    {
      "id": "2643629",
      "postDate": "02/09/2024 01:09:38",
      "content": "<p>To emphasize the importance of this issue, we provide the following example.<br>\nIf row sum NOT 1 when scoring a submission, a perfect prediction is not the best score.<br>\nIn some cases, a different distribution may yield a better score.</p>\n<pre><code>votes = torch.tensor([[,,]]).float()  \npredict1 = votes / votes.()  \npredict2 = torch.tensor([[,,]])  \n\nepsilon = \npredict1 = torch.clip(predict1, epsilon,  - epsilon)\npredict2 = torch.clip(predict2, epsilon,  - epsilon)\n\n = votes  \n = torch.clip(, epsilon,  - epsilon)\nscore1 =  * (.() - predict1.()) \n(score1.()) \nscore2 =  * (.() - predict2.()) \n(score2.()) \n\n = votes / votes.()  \n = torch.clip(, epsilon,  - epsilon)\nscore =  * (.() - predict.()) \n(score.()) \n</code></pre>",
      "rawMarkdown": "To emphasize the importance of this issue, we provide the following example.\nIf row sum NOT 1 when scoring a submission, a perfect prediction is not the best score.\nIn some cases, a different distribution may yield a better score.\n\n```\nvotes = torch.tensor([[5,0,2]]).float()  # votes in train\npredict1 = votes / votes.sum()  # perfect predict in normalize\npredict2 = torch.tensor([[0.5,0,0.5]])  # NOT perfect predict but best score\n\nepsilon = 1e-10\npredict1 = torch.clip(predict1, epsilon, 1 - epsilon)\npredict2 = torch.clip(predict2, epsilon, 1 - epsilon)\n\ntarget = votes  ### row sum NOT 1 when scoring a submission\ntarget = torch.clip(target, epsilon, 1 - epsilon)\nscore1 = target * (target.log() - predict1.log()) # score for \"perfect\" predict\nprint(score1.mean()) # 0.5297\nscore2 = target * (target.log() - predict2.log()) # score for \"NOT perfect\" predict\nprint(score2.mean()) # 0.4621\n\ntarget = votes / votes.sum()  ### row sum IS 1 when scoring a submission\ntarget = torch.clip(target, epsilon, 1 - epsilon)\nscore = target * (target.log() - predict.log()) # score in sum IS 1 when scoring\nprint(score.mean()) # 0.0\n```",
      "votes": null
    },
    {
      "id": "2643644",
      "postDate": "02/09/2024 01:32:29",
      "content": "<p>Kaggle said <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468705#2606605\" target=\"_blank\">here</a> that they create the ground truth for the test LB by dividing by count of number of votes. For example when test has <code>[1, 0, 3, 4, 0, 0]</code> votes then the ground truth for the LB is <code>[1/8, 0/8, 3/8, 4/8, 0/8, 0/8]</code>. So they do not use softmax, they do not use logs, they do not use anything fancy. I will search for a link to Kaggle's post and update my comment.</p>",
      "rawMarkdown": "Kaggle said [here][1] that they create the ground truth for the test LB by dividing by count of number of votes. For example when test has `[1, 0, 3, 4, 0, 0]` votes then the ground truth for the LB is `[1/8, 0/8, 3/8, 4/8, 0/8, 0/8]`. So they do not use softmax, they do not use logs, they do not use anything fancy. I will search for a link to Kaggle's post and update my comment.\n\n[1]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468705#2606605",
      "votes": null
    },
    {
      "id": "2643648",
      "postDate": "02/09/2024 01:42:26",
      "content": "<p>thank you very much. This is the information I was looking for.</p>",
      "rawMarkdown": "thank you very much. This is the information I was looking for.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2642833,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "02/08/2024 12:55:10",
      "content": "<p>Hi. The targets at train.csv are the absolute accumulation of the votes from the expert annotators. While the example submission has the normalized votations of equally voted 6 classes, 0.1666… So I'd say make a final normalization just in case. The results should be practically identhical.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2643609,
          "author_name": "tanreinama",
          "author_url": "",
          "post_date": "02/09/2024 00:28:54",
          "content": "<p>thank you. However, this issue becomes important when checking the correlation between CV scores and leaderboards. That's why I want accurate information. The example below uses normalization instead of softmax, but the results are the same, as the two scores are different.</p>\n<pre><code>votes = torch.tensor([[,,]]).float()  \npredict = votes / votes.()  \n\nepsilon = \npredict = torch.clip(predict, epsilon,  - epsilon)\n\n = votes  \n = torch.clip(, epsilon,  - epsilon)\nscore =  * (.() - predict.())\n(score.()) \n\n = votes / votes.()  \n = torch.clip(, epsilon,  - epsilon)\nscore =  * (.() - predict.())\n(score.()) \n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2643629,
      "author_name": "tanreinama",
      "author_url": "",
      "post_date": "02/09/2024 01:09:38",
      "content": "<p>To emphasize the importance of this issue, we provide the following example.<br>\nIf row sum NOT 1 when scoring a submission, a perfect prediction is not the best score.<br>\nIn some cases, a different distribution may yield a better score.</p>\n<pre><code>votes = torch.tensor([[,,]]).float()  \npredict1 = votes / votes.()  \npredict2 = torch.tensor([[,,]])  \n\nepsilon = \npredict1 = torch.clip(predict1, epsilon,  - epsilon)\npredict2 = torch.clip(predict2, epsilon,  - epsilon)\n\n = votes  \n = torch.clip(, epsilon,  - epsilon)\nscore1 =  * (.() - predict1.()) \n(score1.()) \nscore2 =  * (.() - predict2.()) \n(score2.()) \n\n = votes / votes.()  \n = torch.clip(, epsilon,  - epsilon)\nscore =  * (.() - predict.()) \n(score.()) \n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2643644,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/09/2024 01:32:29",
      "content": "<p>Kaggle said <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468705#2606605\" target=\"_blank\">here</a> that they create the ground truth for the test LB by dividing by count of number of votes. For example when test has <code>[1, 0, 3, 4, 0, 0]</code> votes then the ground truth for the LB is <code>[1/8, 0/8, 3/8, 4/8, 0/8, 0/8]</code>. So they do not use softmax, they do not use logs, they do not use anything fancy. I will search for a link to Kaggle's post and update my comment.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2643648,
          "author_name": "tanreinama",
          "author_url": "",
          "post_date": "02/09/2024 01:42:26",
          "content": "<p>thank you very much. This is the information I was looking for.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2642203": "Is the row sum 1 when scoring a submission?\n\nThe sum of the rows in submission.csv must be 1 or an error will occur.\nBut when I look at the targets in train.csv, the sum is not 1.\nTherefore, the results will differ depending on whether you use target in train.csv or softmax(target) as objective.\n\nfor example,\n\n```\nvotes = torch.tensor([[5,0,2]]).float()  # votes in train\npredict = torch.nn.functional.softmax(votes, dim=1)  # perfect predict\n\nepsilon = 1e-10\npredict = torch.clip(predict, epsilon, 1 - epsilon)\n\ntarget = votes  ### row sum NOT 1 when scoring a submission\ntarget = torch.clip(target, epsilon, 1 - epsilon)\nscore = target * (target.log() - predict.log())\nprint(score.mean()) # 1.0367\n\ntarget = torch.nn.functional.softmax(votes, dim=1)  ### row sum IS 1 when scoring a submission\ntarget = torch.clip(target, epsilon, 1 - epsilon)\nscore = target * (target.log() - predict.log())\nprint(score.mean()) # 0.0\n```\n\nAlso, if the sum of the rows when scoring is not 1, is it okay to say that the score will not be 0.0 even if I make a perfect prediction?",
    "2642833": "Hi. The targets at train.csv are the absolute accumulation of the votes from the expert annotators. While the example submission has the normalized votations of equally voted 6 classes, 0.1666... So I'd say make a final normalization just in case. The results should be practically identhical.",
    "2643609": "thank you. However, this issue becomes important when checking the correlation between CV scores and leaderboards. That's why I want accurate information. The example below uses normalization instead of softmax, but the results are the same, as the two scores are different.\n\n```\nvotes = torch.tensor([[5,0,2]]).float()  # votes in train\npredict = votes / votes.sum()  # perfect predict in normalize\n\nepsilon = 1e-10\npredict = torch.clip(predict, epsilon, 1 - epsilon)\n\ntarget = votes  ### row sum NOT 1 when scoring a submission\ntarget = torch.clip(target, epsilon, 1 - epsilon)\nscore = target * (target.log() - predict.log())\nprint(score.mean()) # 0.5297\n\ntarget = votes / votes.sum()  ### row sum IS 1 when scoring a submission\ntarget = torch.clip(target, epsilon, 1 - epsilon)\nscore = target * (target.log() - predict.log())\nprint(score.mean()) # 0.0\n```",
    "2643629": "To emphasize the importance of this issue, we provide the following example.\nIf row sum NOT 1 when scoring a submission, a perfect prediction is not the best score.\nIn some cases, a different distribution may yield a better score.\n\n```\nvotes = torch.tensor([[5,0,2]]).float()  # votes in train\npredict1 = votes / votes.sum()  # perfect predict in normalize\npredict2 = torch.tensor([[0.5,0,0.5]])  # NOT perfect predict but best score\n\nepsilon = 1e-10\npredict1 = torch.clip(predict1, epsilon, 1 - epsilon)\npredict2 = torch.clip(predict2, epsilon, 1 - epsilon)\n\ntarget = votes  ### row sum NOT 1 when scoring a submission\ntarget = torch.clip(target, epsilon, 1 - epsilon)\nscore1 = target * (target.log() - predict1.log()) # score for \"perfect\" predict\nprint(score1.mean()) # 0.5297\nscore2 = target * (target.log() - predict2.log()) # score for \"NOT perfect\" predict\nprint(score2.mean()) # 0.4621\n\ntarget = votes / votes.sum()  ### row sum IS 1 when scoring a submission\ntarget = torch.clip(target, epsilon, 1 - epsilon)\nscore = target * (target.log() - predict.log()) # score in sum IS 1 when scoring\nprint(score.mean()) # 0.0\n```",
    "2643644": "Kaggle said [here][1] that they create the ground truth for the test LB by dividing by count of number of votes. For example when test has `[1, 0, 3, 4, 0, 0]` votes then the ground truth for the LB is `[1/8, 0/8, 3/8, 4/8, 0/8, 0/8]`. So they do not use softmax, they do not use logs, they do not use anything fancy. I will search for a link to Kaggle's post and update my comment.\n\n[1]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468705#2606605",
    "2643648": "thank you very much. This is the information I was looking for."
  },
  "source": "meta"
}