{
  "id": 549868,
  "title": "Help on accuracy vs quadratic weighted kappa",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/549868",
  "author_name": "",
  "post_date": "2024-12-04T08:56:27.675315Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I am getting <br>\nPrediction Accuracy on Training Set: 99.16%<br>\nPrediction Accuracy on Validation Set: 85.35%<br>\nHowever, my submission score is 0. Is accuracy not an effective metric?</p>",
  "messages": [
    {
      "id": "3063206",
      "postDate": "12/04/2024 08:56:27",
      "content": "<p>I am getting <br>\nPrediction Accuracy on Training Set: 99.16%<br>\nPrediction Accuracy on Validation Set: 85.35%<br>\nHowever, my submission score is 0. Is accuracy not an effective metric?</p>",
      "rawMarkdown": "I am getting \nPrediction Accuracy on Training Set: 99.16%\nPrediction Accuracy on Validation Set: 85.35%\nHowever, my submission score is 0. Is accuracy not an effective metric?",
      "votes": null
    },
    {
      "id": "3063757",
      "postDate": "12/04/2024 20:28:21",
      "content": "<p>Well, first of all, this is a highly imbalanced set of classes if you check the distribution of the classes. Most of the sii classes are 0. So accuracy does not seem to be a good metric. You can call everything to be 0 and get a 'pretty good' accuracy. I would prefer something like a precision better.</p>\n<p>In regards to you getting so high values, check if you have some kind of data leakage. This could be due to splitting the dataset to train and validation after running some preprocessing on both of them. This would leak some information to the validation part, much like training on the train and then testing on the train (in this case, something similar to train).</p>\n<p>Also additional suggestion: See make_scorer and look how you can generate a quadratic cohen scorer so that your algorithm already focuses on maximizing that parameter. Or you can write your own tuning function that maximizes the quadratic cohen score, much like how others did in this competition.</p>",
      "rawMarkdown": "Well, first of all, this is a highly imbalanced set of classes if you check the distribution of the classes. Most of the sii classes are 0. So accuracy does not seem to be a good metric. You can call everything to be 0 and get a 'pretty good' accuracy. I would prefer something like a precision better.\n\nIn regards to you getting so high values, check if you have some kind of data leakage. This could be due to splitting the dataset to train and validation after running some preprocessing on both of them. This would leak some information to the validation part, much like training on the train and then testing on the train (in this case, something similar to train).\n\nAlso additional suggestion: See make_scorer and look how you can generate a quadratic cohen scorer so that your algorithm already focuses on maximizing that parameter. Or you can write your own tuning function that maximizes the quadratic cohen score, much like how others did in this competition.",
      "votes": null
    },
    {
      "id": "3064465",
      "postDate": "12/05/2024 16:28:19",
      "content": "<p>Yes, I split the dataset and did preprocessing on both of them, first on the training set and then on the validation set. I don't know about data leakage. What is it?</p>",
      "rawMarkdown": "Yes, I split the dataset and did preprocessing on both of them, first on the training set and then on the validation set. I don't know about data leakage. What is it?",
      "votes": null
    },
    {
      "id": "3064473",
      "postDate": "12/05/2024 16:38:48",
      "content": "<p>Yes, accuracy might not be an effective metric for your problem, depending on the competition's evaluation metric and the nature of your data. You read the evaluation which is in the evaluation page and perfrom the metrics according to it</p>",
      "rawMarkdown": "Yes, accuracy might not be an effective metric for your problem, depending on the competition's evaluation metric and the nature of your data. You read the evaluation which is in the evaluation page and perfrom the metrics according to it",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3063757,
      "author_name": "alperenduru",
      "author_url": "",
      "post_date": "12/04/2024 20:28:21",
      "content": "<p>Well, first of all, this is a highly imbalanced set of classes if you check the distribution of the classes. Most of the sii classes are 0. So accuracy does not seem to be a good metric. You can call everything to be 0 and get a 'pretty good' accuracy. I would prefer something like a precision better.</p>\n<p>In regards to you getting so high values, check if you have some kind of data leakage. This could be due to splitting the dataset to train and validation after running some preprocessing on both of them. This would leak some information to the validation part, much like training on the train and then testing on the train (in this case, something similar to train).</p>\n<p>Also additional suggestion: See make_scorer and look how you can generate a quadratic cohen scorer so that your algorithm already focuses on maximizing that parameter. Or you can write your own tuning function that maximizes the quadratic cohen score, much like how others did in this competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3064465,
          "author_name": "iitm21f1006461",
          "author_url": "",
          "post_date": "12/05/2024 16:28:19",
          "content": "<p>Yes, I split the dataset and did preprocessing on both of them, first on the training set and then on the validation set. I don't know about data leakage. What is it?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3064473,
      "author_name": "sumit08",
      "author_url": "",
      "post_date": "12/05/2024 16:38:48",
      "content": "<p>Yes, accuracy might not be an effective metric for your problem, depending on the competition's evaluation metric and the nature of your data. You read the evaluation which is in the evaluation page and perfrom the metrics according to it</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3063206": "I am getting \nPrediction Accuracy on Training Set: 99.16%\nPrediction Accuracy on Validation Set: 85.35%\nHowever, my submission score is 0. Is accuracy not an effective metric?",
    "3063757": "Well, first of all, this is a highly imbalanced set of classes if you check the distribution of the classes. Most of the sii classes are 0. So accuracy does not seem to be a good metric. You can call everything to be 0 and get a 'pretty good' accuracy. I would prefer something like a precision better.\n\nIn regards to you getting so high values, check if you have some kind of data leakage. This could be due to splitting the dataset to train and validation after running some preprocessing on both of them. This would leak some information to the validation part, much like training on the train and then testing on the train (in this case, something similar to train).\n\nAlso additional suggestion: See make_scorer and look how you can generate a quadratic cohen scorer so that your algorithm already focuses on maximizing that parameter. Or you can write your own tuning function that maximizes the quadratic cohen score, much like how others did in this competition.",
    "3064465": "Yes, I split the dataset and did preprocessing on both of them, first on the training set and then on the validation set. I don't know about data leakage. What is it?",
    "3064473": "Yes, accuracy might not be an effective metric for your problem, depending on the competition's evaluation metric and the nature of your data. You read the evaluation which is in the evaluation page and perfrom the metrics according to it"
  },
  "source": "meta"
}