{
  "id": 218562,
  "title": "Submission format question about probabilities",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/218562",
  "author_name": "CMHM",
  "post_date": "2021-02-11T06:33:49.897000",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Please forgive the newbie question, but it says here we should submit in the format given in the submission_sample.csv file; but it is unclear exactly how. I made an algorithm that creates probabilities, but should I submit </p>\n<p>A: a boolean 0 or 1 for each listed class</p>\n<p>B: the raw probability my algorithm computed in each listed class</p>\n<p>or <br>\nC: something enhanced by the probability in all classes including those not listed i.e. no ETT and no NGT.<br>\nTo elaborate on C with an example, logically, if I have 0.999 probability of no ETT, then having high probability of other possibilities feels very wrong (although the math might make sense), so I might write some code to not have the same probabilities in certain categories, or favor the highest one.  <br>\nCan someone give me an awnser here?</p>",
  "messages": [
    {
      "id": 1196062,
      "postDate": "2021-02-11T08:25:11.323Z",
      "content": "<p>The format of the submission should be exactly as shown in the <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/data\" target=\"_blank\">sample_submission.csv</a> - i.e. all the columns shown there, but only the columns shown there, in the order they are shown etc. So any additional classes and/or predictions you create for some reason should not end up in there.</p>\n<p>The columns for the different targets should contain numbers (not True or False), which typically would be numbers between 0 and 1 (e.g. 0.2, 0.99, 0.0001, 0.5 etc.) - 0 or 1 is of course also okay. I believe technically the competition metric can be calculated just fine if the numbers are not probabilities - i.e. not between 0 and 1 (e.g. -0.69, 3.21, -5.232, 0.0002 etc., so e.g. logits of probabilities before the sigmoid function gets applied). You may very well be able to submit such numbers outside the 0 to 1 range to the leaderboard (not sure though). In the end, for the competition metric only the ordering of the numbers matters, so you might just as well make them lie between 0 and 1 by rescaling them. And with a tiny bit of extra effort, you could  even try to ensure they are somewhat calibrated (which for a real-life solution to the competition problem you'd definitely want, but the competition metric does not care either way).</p>\n<p>And yes, it may make sense to take into account whether multiple categories can occur at the same time and/or whether certain categories are more or less likely to co-occur. The obvious models to use for this competition (i.e. neural network with a multi-label target) will probably reflect the correlation between targets already, but if you identify some logical constraints you could enforce them (but you might be better off doing so in your loss function during training already). However, note that the absolute values of probabilities do not matter in this competition. I.e. if you only ever predict values between 0.99 and 0.99999, but for all the true zeros you predict 0.99 and for all the true ones you predict 0.99999, then you've done perfectly, as far as the competition metric is concerned (see above).</p>",
      "rawMarkdown": "The format of the submission should be exactly as shown in the [sample_submission.csv](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/data) - i.e. all the columns shown there, but only the columns shown there, in the order they are shown etc. So any additional classes and/or predictions you create for some reason should not end up in there.\n\nThe columns for the different targets should contain numbers (not True or False), which typically would be numbers between 0 and 1 (e.g. 0.2, 0.99, 0.0001, 0.5 etc.) - 0 or 1 is of course also okay. I believe technically the competition metric can be calculated just fine if the numbers are not probabilities - i.e. not between 0 and 1 (e.g. -0.69, 3.21, -5.232, 0.0002 etc., so e.g. logits of probabilities before the sigmoid function gets applied). You may very well be able to submit such numbers outside the 0 to 1 range to the leaderboard (not sure though). In the end, for the competition metric only the ordering of the numbers matters, so you might just as well make them lie between 0 and 1 by rescaling them. And with a tiny bit of extra effort, you could  even try to ensure they are somewhat calibrated (which for a real-life solution to the competition problem you'd definitely want, but the competition metric does not care either way).\n\nAnd yes, it may make sense to take into account whether multiple categories can occur at the same time and/or whether certain categories are more or less likely to co-occur. The obvious models to use for this competition (i.e. neural network with a multi-label target) will probably reflect the correlation between targets already, but if you identify some logical constraints you could enforce them (but you might be better off doing so in your loss function during training already). However, note that the absolute values of probabilities do not matter in this competition. I.e. if you only ever predict values between 0.99 and 0.99999, but for all the true zeros you predict 0.99 and for all the true ones you predict 0.99999, then you've done perfectly, as far as the competition metric is concerned (see above).",
      "votes": 1,
      "replies": [
        {
          "id": 1196931,
          "postDate": "2021-02-11T19:25:12.760Z",
          "content": "<p>Yeah I can confirm that submissions work fine without sigmoid. Dont need predictions constrained between range of 0-1</p>",
          "rawMarkdown": "Yeah I can confirm that submissions work fine without sigmoid. Dont need predictions constrained between range of 0-1",
          "votes": 2
        }
      ]
    },
    {
      "id": 1195936,
      "postDate": "2021-02-11T06:33:49.897Z",
      "content": "<p>Please forgive the newbie question, but it says here we should submit in the format given in the submission_sample.csv file; but it is unclear exactly how. I made an algorithm that creates probabilities, but should I submit </p>\n<p>A: a boolean 0 or 1 for each listed class</p>\n<p>B: the raw probability my algorithm computed in each listed class</p>\n<p>or <br>\nC: something enhanced by the probability in all classes including those not listed i.e. no ETT and no NGT.<br>\nTo elaborate on C with an example, logically, if I have 0.999 probability of no ETT, then having high probability of other possibilities feels very wrong (although the math might make sense), so I might write some code to not have the same probabilities in certain categories, or favor the highest one.  <br>\nCan someone give me an awnser here?</p>",
      "rawMarkdown": "Please forgive the newbie question, but it says here we should submit in the format given in the submission_sample.csv file; but it is unclear exactly how. I made an algorithm that creates probabilities, but should I submit \n\nA: a boolean 0 or 1 for each listed class\n\nB: the raw probability my algorithm computed in each listed class\n\nor \nC: something enhanced by the probability in all classes including those not listed i.e. no ETT and no NGT.\nTo elaborate on C with an example, logically, if I have 0.999 probability of no ETT, then having high probability of other possibilities feels very wrong (although the math might make sense), so I might write some code to not have the same probabilities in certain categories, or favor the highest one.  \nCan someone give me an awnser here?\n",
      "votes": 1
    },
    {
      "id": 1196947,
      "postDate": "2021-02-11T19:39:14.623Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1196062,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2021-02-11T08:25:11.323000",
      "content": "<p>The format of the submission should be exactly as shown in the <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/data\" target=\"_blank\">sample_submission.csv</a> - i.e. all the columns shown there, but only the columns shown there, in the order they are shown etc. So any additional classes and/or predictions you create for some reason should not end up in there.</p>\n<p>The columns for the different targets should contain numbers (not True or False), which typically would be numbers between 0 and 1 (e.g. 0.2, 0.99, 0.0001, 0.5 etc.) - 0 or 1 is of course also okay. I believe technically the competition metric can be calculated just fine if the numbers are not probabilities - i.e. not between 0 and 1 (e.g. -0.69, 3.21, -5.232, 0.0002 etc., so e.g. logits of probabilities before the sigmoid function gets applied). You may very well be able to submit such numbers outside the 0 to 1 range to the leaderboard (not sure though). In the end, for the competition metric only the ordering of the numbers matters, so you might just as well make them lie between 0 and 1 by rescaling them. And with a tiny bit of extra effort, you could  even try to ensure they are somewhat calibrated (which for a real-life solution to the competition problem you'd definitely want, but the competition metric does not care either way).</p>\n<p>And yes, it may make sense to take into account whether multiple categories can occur at the same time and/or whether certain categories are more or less likely to co-occur. The obvious models to use for this competition (i.e. neural network with a multi-label target) will probably reflect the correlation between targets already, but if you identify some logical constraints you could enforce them (but you might be better off doing so in your loss function during training already). However, note that the absolute values of probabilities do not matter in this competition. I.e. if you only ever predict values between 0.99 and 0.99999, but for all the true zeros you predict 0.99 and for all the true ones you predict 0.99999, then you've done perfectly, as far as the competition metric is concerned (see above).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1196931,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2021-02-11T19:25:12.760000",
          "content": "<p>Yeah I can confirm that submissions work fine without sigmoid. Dont need predictions constrained between range of 0-1</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1196947,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-11T19:39:14.623000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1196062": "The format of the submission should be exactly as shown in the [sample_submission.csv](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/data) - i.e. all the columns shown there, but only the columns shown there, in the order they are shown etc. So any additional classes and/or predictions you create for some reason should not end up in there.\n\nThe columns for the different targets should contain numbers (not True or False), which typically would be numbers between 0 and 1 (e.g. 0.2, 0.99, 0.0001, 0.5 etc.) - 0 or 1 is of course also okay. I believe technically the competition metric can be calculated just fine if the numbers are not probabilities - i.e. not between 0 and 1 (e.g. -0.69, 3.21, -5.232, 0.0002 etc., so e.g. logits of probabilities before the sigmoid function gets applied). You may very well be able to submit such numbers outside the 0 to 1 range to the leaderboard (not sure though). In the end, for the competition metric only the ordering of the numbers matters, so you might just as well make them lie between 0 and 1 by rescaling them. And with a tiny bit of extra effort, you could  even try to ensure they are somewhat calibrated (which for a real-life solution to the competition problem you'd definitely want, but the competition metric does not care either way).\n\nAnd yes, it may make sense to take into account whether multiple categories can occur at the same time and/or whether certain categories are more or less likely to co-occur. The obvious models to use for this competition (i.e. neural network with a multi-label target) will probably reflect the correlation between targets already, but if you identify some logical constraints you could enforce them (but you might be better off doing so in your loss function during training already). However, note that the absolute values of probabilities do not matter in this competition. I.e. if you only ever predict values between 0.99 and 0.99999, but for all the true zeros you predict 0.99 and for all the true ones you predict 0.99999, then you've done perfectly, as far as the competition metric is concerned (see above).",
    "1195936": "Please forgive the newbie question, but it says here we should submit in the format given in the submission_sample.csv file; but it is unclear exactly how. I made an algorithm that creates probabilities, but should I submit \n\nA: a boolean 0 or 1 for each listed class\n\nB: the raw probability my algorithm computed in each listed class\n\nor \nC: something enhanced by the probability in all classes including those not listed i.e. no ETT and no NGT.\nTo elaborate on C with an example, logically, if I have 0.999 probability of no ETT, then having high probability of other possibilities feels very wrong (although the math might make sense), so I might write some code to not have the same probabilities in certain categories, or favor the highest one.  \nCan someone give me an awnser here?\n",
    "1196947": ""
  }
}