{
  "id": 317153,
  "title": "what should be the loss function",
  "url": "/competitions/ultra-mnist/discussion/317153",
  "author_name": "",
  "post_date": "2022-04-05T17:14:53.373606800Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>i have been banging my head at this for some time, and i was wondering i was treating it as a classification problem and using cross entropy loss, but can this be treated as a regression problem ?<br>\nAlso if some one can help me how to implement regression using fastai and resnet model.</p>",
  "messages": [
    {
      "id": "1746343",
      "postDate": "04/05/2022 17:14:53",
      "content": "<p>i have been banging my head at this for some time, and i was wondering i was treating it as a classification problem and using cross entropy loss, but can this be treated as a regression problem ?<br>\nAlso if some one can help me how to implement regression using fastai and resnet model.</p>",
      "rawMarkdown": "i have been banging my head at this for some time, and i was wondering i was treating it as a classification problem and using cross entropy loss, but can this be treated as a regression problem ?\nAlso if some one can help me how to implement regression using fastai and resnet model.",
      "votes": null
    },
    {
      "id": "1746779",
      "postDate": "04/06/2022 05:08:29",
      "content": "<p>This can be treated as a regression problem. But, the metric used for this competition is accuracy. When you treat it as a regression problem, you will use a loss function similar to MSE. The model MIGHT converge to a point where it generates the mean of all the sums (13) to reduce the loss (Since the data is a bit complex). So, it is better to treat it as a classification problem.</p>",
      "rawMarkdown": "This can be treated as a regression problem. But, the metric used for this competition is accuracy. When you treat it as a regression problem, you will use a loss function similar to MSE. The model MIGHT converge to a point where it generates the mean of all the sums (13) to reduce the loss (Since the data is a bit complex). So, it is better to treat it as a classification problem.",
      "votes": null
    },
    {
      "id": "1746808",
      "postDate": "04/06/2022 05:51:46",
      "content": "<p>Please correct me, my understanding is that cross entropy loss tries to seperate the classes as much as possible and in doing so two adjacent numbers will also be treated as far apart as any two numbers which may not be the case. For eg. If the ground truth sum is 20 then 21 and 19 should also be highly possible and this hampers the learning of algorithms</p>\n<p>In my own experience i tried fitting resnet18 on 512px data and it started to overfit the training data even when the valid accuracy was 13% and i thought it should be able to meet the innovation baseline considering that is being run on 300px data. Thinking about the reason this is happening is i wondered if the algorithm seperates the images on pixel patterns that may not be meaningful simply because that is what we were forcing it to, by makig the loss function such that 2 cosecutive numbers will have lower loss than 2 farther apart number i think it will better fit the given problem</p>",
      "rawMarkdown": "Please correct me, my understanding is that cross entropy loss tries to seperate the classes as much as possible and in doing so two adjacent numbers will also be treated as far apart as any two numbers which may not be the case. For eg. If the ground truth sum is 20 then 21 and 19 should also be highly possible and this hampers the learning of algorithms\n\nIn my own experience i tried fitting resnet18 on 512px data and it started to overfit the training data even when the valid accuracy was 13% and i thought it should be able to meet the innovation baseline considering that is being run on 300px data. Thinking about the reason this is happening is i wondered if the algorithm seperates the images on pixel patterns that may not be meaningful simply because that is what we were forcing it to, by makig the loss function such that 2 cosecutive numbers will have lower loss than 2 farther apart number i think it will better fit the given problem",
      "votes": null
    },
    {
      "id": "1747113",
      "postDate": "04/06/2022 11:24:08",
      "content": "<p>Yes, you are right, ce loss does separate classes far apart. Since I haven't tried out regression on this problem, I'm not able to comment further on its efficacy rn.</p>",
      "rawMarkdown": "Yes, you are right, ce loss does separate classes far apart. Since I haven't tried out regression on this problem, I'm not able to comment further on its efficacy rn.",
      "votes": null
    },
    {
      "id": "1747569",
      "postDate": "04/06/2022 18:55:27",
      "content": "<p>i tried the regression, and it converged around the average(13.4), so you were right. i guess ill have to try some other things</p>",
      "rawMarkdown": "i tried the regression, and it converged around the average(13.4), so you were right. i guess ill have to try some other things",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1746779,
      "author_name": "shankarmahadevan",
      "author_url": "",
      "post_date": "04/06/2022 05:08:29",
      "content": "<p>This can be treated as a regression problem. But, the metric used for this competition is accuracy. When you treat it as a regression problem, you will use a loss function similar to MSE. The model MIGHT converge to a point where it generates the mean of all the sums (13) to reduce the loss (Since the data is a bit complex). So, it is better to treat it as a classification problem.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1746808,
          "author_name": "abhishekdasani",
          "author_url": "",
          "post_date": "04/06/2022 05:51:46",
          "content": "<p>Please correct me, my understanding is that cross entropy loss tries to seperate the classes as much as possible and in doing so two adjacent numbers will also be treated as far apart as any two numbers which may not be the case. For eg. If the ground truth sum is 20 then 21 and 19 should also be highly possible and this hampers the learning of algorithms</p>\n<p>In my own experience i tried fitting resnet18 on 512px data and it started to overfit the training data even when the valid accuracy was 13% and i thought it should be able to meet the innovation baseline considering that is being run on 300px data. Thinking about the reason this is happening is i wondered if the algorithm seperates the images on pixel patterns that may not be meaningful simply because that is what we were forcing it to, by makig the loss function such that 2 cosecutive numbers will have lower loss than 2 farther apart number i think it will better fit the given problem</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1747113,
          "author_name": "shankarmahadevan",
          "author_url": "",
          "post_date": "04/06/2022 11:24:08",
          "content": "<p>Yes, you are right, ce loss does separate classes far apart. Since I haven't tried out regression on this problem, I'm not able to comment further on its efficacy rn.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1747569,
          "author_name": "abhishekdasani",
          "author_url": "",
          "post_date": "04/06/2022 18:55:27",
          "content": "<p>i tried the regression, and it converged around the average(13.4), so you were right. i guess ill have to try some other things</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1746343": "i have been banging my head at this for some time, and i was wondering i was treating it as a classification problem and using cross entropy loss, but can this be treated as a regression problem ?\nAlso if some one can help me how to implement regression using fastai and resnet model.",
    "1746779": "This can be treated as a regression problem. But, the metric used for this competition is accuracy. When you treat it as a regression problem, you will use a loss function similar to MSE. The model MIGHT converge to a point where it generates the mean of all the sums (13) to reduce the loss (Since the data is a bit complex). So, it is better to treat it as a classification problem.",
    "1746808": "Please correct me, my understanding is that cross entropy loss tries to seperate the classes as much as possible and in doing so two adjacent numbers will also be treated as far apart as any two numbers which may not be the case. For eg. If the ground truth sum is 20 then 21 and 19 should also be highly possible and this hampers the learning of algorithms\n\nIn my own experience i tried fitting resnet18 on 512px data and it started to overfit the training data even when the valid accuracy was 13% and i thought it should be able to meet the innovation baseline considering that is being run on 300px data. Thinking about the reason this is happening is i wondered if the algorithm seperates the images on pixel patterns that may not be meaningful simply because that is what we were forcing it to, by makig the loss function such that 2 cosecutive numbers will have lower loss than 2 farther apart number i think it will better fit the given problem",
    "1747113": "Yes, you are right, ce loss does separate classes far apart. Since I haven't tried out regression on this problem, I'm not able to comment further on its efficacy rn.",
    "1747569": "i tried the regression, and it converged around the average(13.4), so you were right. i guess ill have to try some other things"
  },
  "source": "meta"
}