{
  "id": 476019,
  "title": "Clarification about Ground Truth for Metric Computation",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/476019",
  "author_name": "",
  "post_date": "2024-02-10T19:55:44.301903900Z",
  "votes": 17,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi Kagglers!</p>\n<p>After diving into <code>Overview</code> and <code>Data</code> sections, I still have not understood how to create Ground Truth for competition metric</p>\n<p>We have a fantastic <a href=\"https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook\" target=\"_blank\">notebook</a> which explains how to compute metric. At the same time, after some probing, I have understood that we have to normalize our predicted values, so they sum into 1.0 for each raw (Please correct me if I am wrong).</p>\n<p>However, it is still unclear how to create GTs for each type of harmful brain activity.<br>\nI have tried an obvious hypothesis about using one-hot, based on <code>expert_consensus</code>. If we submit all equal probs (1/6) - we will have LB 1.09, according to <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467021\" target=\"_blank\">this discussion</a>, while if my hypothesis is true, this value should be ~1.791. So my hypothesis is <strong>NOT</strong> True</p>\n<p>Another options is to normalize <code>*_vote</code> columns with:</p>\n<ul>\n<li>softmax</li>\n<li>by sum</li>\n</ul>\n<p>But let's have clarity about this important point! I would be really grateful if anyone can help me with this!</p>",
  "messages": [
    {
      "id": "2646332",
      "postDate": "02/10/2024 19:55:44",
      "content": "<p>Hi Kagglers!</p>\n<p>After diving into <code>Overview</code> and <code>Data</code> sections, I still have not understood how to create Ground Truth for competition metric</p>\n<p>We have a fantastic <a href=\"https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook\" target=\"_blank\">notebook</a> which explains how to compute metric. At the same time, after some probing, I have understood that we have to normalize our predicted values, so they sum into 1.0 for each raw (Please correct me if I am wrong).</p>\n<p>However, it is still unclear how to create GTs for each type of harmful brain activity.<br>\nI have tried an obvious hypothesis about using one-hot, based on <code>expert_consensus</code>. If we submit all equal probs (1/6) - we will have LB 1.09, according to <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467021\" target=\"_blank\">this discussion</a>, while if my hypothesis is true, this value should be ~1.791. So my hypothesis is <strong>NOT</strong> True</p>\n<p>Another options is to normalize <code>*_vote</code> columns with:</p>\n<ul>\n<li>softmax</li>\n<li>by sum</li>\n</ul>\n<p>But let's have clarity about this important point! I would be really grateful if anyone can help me with this!</p>",
      "rawMarkdown": "Hi Kagglers!\n\nAfter diving into `Overview` and `Data` sections, I still have not understood how to create Ground Truth for competition metric\n\nWe have a fantastic [notebook](https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook) which explains how to compute metric. At the same time, after some probing, I have understood that we have to normalize our predicted values, so they sum into 1.0 for each raw (Please correct me if I am wrong).\n\nHowever, it is still unclear how to create GTs for each type of harmful brain activity.\nI have tried an obvious hypothesis about using one-hot, based on `expert_consensus`. If we submit all equal probs (1/6) - we will have LB 1.09, according to [this discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467021), while if my hypothesis is true, this value should be ~1.791. So my hypothesis is **NOT** True\n\nAnother options is to normalize `*_vote` columns with:\n- softmax\n- by sum\n\nBut let's have clarity about this important point! I would be really grateful if anyone can help me with this!",
      "votes": null
    },
    {
      "id": "2646355",
      "postDate": "02/10/2024 20:18:23",
      "content": "<p>Kaggle responded <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468705#2606605\" target=\"_blank\">here</a> that the test data ground truths are created by dividing each row by the total number of counts. For example if vote counts are <code>[1,2,0,0,5,0]</code> then the ground truth is <code>[1/8, 2/8, 0, 0, 5/8, 0]</code>. </p>",
      "rawMarkdown": "Kaggle responded [here][1] that the test data ground truths are created by dividing each row by the total number of counts. For example if vote counts are `[1,2,0,0,5,0]` then the ground truth is `[1/8, 2/8, 0, 0, 5/8, 0]`. \n\n[1]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468705#2606605",
      "votes": null
    },
    {
      "id": "2646365",
      "postDate": "02/10/2024 20:31:22",
      "content": "<p>Thanks! It helps a lot</p>",
      "rawMarkdown": "Thanks! It helps a lot",
      "votes": null
    },
    {
      "id": "2646399",
      "postDate": "02/10/2024 21:23:05",
      "content": "<p>That's right, but [1,0,0,0,0] should not  be an equivalent of [10,0,0,0,0]. I wonder if it makes sense to account for the total number of votes per sample and smooth true labels during training.</p>",
      "rawMarkdown": "That's right, but [1,0,0,0,0] should not  be an equivalent of [10,0,0,0,0]. I wonder if it makes sense to account for the total number of votes per sample and smooth true labels during training.",
      "votes": null
    },
    {
      "id": "2648061",
      "postDate": "02/12/2024 04:11:45",
      "content": "<p>I have done it for the 1D models and don't find any difference in oof loss, maybe with 2D models it's different </p>",
      "rawMarkdown": "I have done it for the 1D models and don't find any difference in oof loss, maybe with 2D models it's different",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2646355,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/10/2024 20:18:23",
      "content": "<p>Kaggle responded <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468705#2606605\" target=\"_blank\">here</a> that the test data ground truths are created by dividing each row by the total number of counts. For example if vote counts are <code>[1,2,0,0,5,0]</code> then the ground truth is <code>[1/8, 2/8, 0, 0, 5/8, 0]</code>. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2646365,
          "author_name": "vladimirsydor",
          "author_url": "",
          "post_date": "02/10/2024 20:31:22",
          "content": "<p>Thanks! It helps a lot</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2646399,
          "author_name": "victorshlepov",
          "author_url": "",
          "post_date": "02/10/2024 21:23:05",
          "content": "<p>That's right, but [1,0,0,0,0] should not  be an equivalent of [10,0,0,0,0]. I wonder if it makes sense to account for the total number of votes per sample and smooth true labels during training.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2648061,
              "author_name": "luispintoc",
              "author_url": "",
              "post_date": "02/12/2024 04:11:45",
              "content": "<p>I have done it for the 1D models and don't find any difference in oof loss, maybe with 2D models it's different </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2646332": "Hi Kagglers!\n\nAfter diving into `Overview` and `Data` sections, I still have not understood how to create Ground Truth for competition metric\n\nWe have a fantastic [notebook](https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook) which explains how to compute metric. At the same time, after some probing, I have understood that we have to normalize our predicted values, so they sum into 1.0 for each raw (Please correct me if I am wrong).\n\nHowever, it is still unclear how to create GTs for each type of harmful brain activity.\nI have tried an obvious hypothesis about using one-hot, based on `expert_consensus`. If we submit all equal probs (1/6) - we will have LB 1.09, according to [this discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467021), while if my hypothesis is true, this value should be ~1.791. So my hypothesis is **NOT** True\n\nAnother options is to normalize `*_vote` columns with:\n- softmax\n- by sum\n\nBut let's have clarity about this important point! I would be really grateful if anyone can help me with this!",
    "2646355": "Kaggle responded [here][1] that the test data ground truths are created by dividing each row by the total number of counts. For example if vote counts are `[1,2,0,0,5,0]` then the ground truth is `[1/8, 2/8, 0, 0, 5/8, 0]`. \n\n[1]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468705#2606605",
    "2646365": "Thanks! It helps a lot",
    "2646399": "That's right, but [1,0,0,0,0] should not  be an equivalent of [10,0,0,0,0]. I wonder if it makes sense to account for the total number of votes per sample and smooth true labels during training.",
    "2648061": "I have done it for the 1D models and don't find any difference in oof loss, maybe with 2D models it's different"
  },
  "source": "meta"
}