{
  "id": 118010,
  "title": "investigation into bi-tempered loss",
  "url": "/competitions/understanding_cloud_organization/discussion/118010",
  "author_name": "",
  "post_date": "2019-11-19T07:00:18.762196Z",
  "votes": 23,
  "comment_count": 4,
  "views": 0,
  "content": "<p>reference paper and google blog:\n<a href=\"https://arxiv.org/abs/1906.03361\">https://arxiv.org/abs/1906.03361</a>\n<a href=\"https://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html\">https://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html</a>\n<a href=\"https://google.github.io/bi-tempered-loss/\">https://google.github.io/bi-tempered-loss/</a></p>\n\n<p>google bi-tempered loss is supposed to fight noisy data. the objective of this work is to see how it will perform on this cloud task. Follow this thread if you are interested in the work </p>\n\n<p>code reference:\n<a href=\"https://github.com/fhopfmueller/bi-tempered-loss-pytorch/commits?author=fhopfmueller&amp;since=2019-09-01&amp;until=2019-09-12\">https://github.com/fhopfmueller/bi-tempered-loss-pytorch/commits?author=fhopfmueller&amp;since=2019-09-01&amp;until=2019-09-12</a></p>\n\n<hr>\n\n<p>code: <a href=\"https://drive.google.com/open?id=1uxMIEM1t2rGSLyYWinqliKDtFnjUJ4Ir\">https://drive.google.com/open?id=1uxMIEM1t2rGSLyYWinqliKDtFnjUJ4Ir</a></p>\n\n<ul>\n<li>we investigate the effect of  bi-tempered-loss for image label classification only</li>\n<li>implement nigh resolution net for image classification</li>\n<li>some lb submission for baseline model and bi-tempered model</li>\n</ul>",
  "messages": [
    {
      "id": "676424",
      "postDate": "11/19/2019 07:00:18",
      "content": "<p>reference paper and google blog:\n<a href=\"https://arxiv.org/abs/1906.03361\">https://arxiv.org/abs/1906.03361</a>\n<a href=\"https://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html\">https://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html</a>\n<a href=\"https://google.github.io/bi-tempered-loss/\">https://google.github.io/bi-tempered-loss/</a></p>\n\n<p>google bi-tempered loss is supposed to fight noisy data. the objective of this work is to see how it will perform on this cloud task. Follow this thread if you are interested in the work </p>\n\n<p>code reference:\n<a href=\"https://github.com/fhopfmueller/bi-tempered-loss-pytorch/commits?author=fhopfmueller&amp;since=2019-09-01&amp;until=2019-09-12\">https://github.com/fhopfmueller/bi-tempered-loss-pytorch/commits?author=fhopfmueller&amp;since=2019-09-01&amp;until=2019-09-12</a></p>\n\n<hr>\n\n<p>code: <a href=\"https://drive.google.com/open?id=1uxMIEM1t2rGSLyYWinqliKDtFnjUJ4Ir\">https://drive.google.com/open?id=1uxMIEM1t2rGSLyYWinqliKDtFnjUJ4Ir</a></p>\n\n<ul>\n<li>we investigate the effect of  bi-tempered-loss for image label classification only</li>\n<li>implement nigh resolution net for image classification</li>\n<li>some lb submission for baseline model and bi-tempered model</li>\n</ul>",
      "rawMarkdown": "reference paper and google blog:\nhttps://arxiv.org/abs/1906.03361\nhttps://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html\nhttps://google.github.io/bi-tempered-loss/\n\ngoogle bi-tempered loss is supposed to fight noisy data. the objective of this work is to see how it will perform on this cloud task. Follow this thread if you are interested in the work \n\ncode reference:\nhttps://github.com/fhopfmueller/bi-tempered-loss-pytorch/commits?author=fhopfmueller&amp;since=2019-09-01&amp;until=2019-09-12\n\n---\n\ncode: https://drive.google.com/open?id=1uxMIEM1t2rGSLyYWinqliKDtFnjUJ4Ir\n\n- we investigate the effect of  bi-tempered-loss for image label classification only\n- implement nigh resolution net for image classification\n- some lb submission for baseline model and bi-tempered model",
      "votes": null
    },
    {
      "id": "678222",
      "postDate": "11/21/2019 05:30:49",
      "content": "<p>here is the first  results.</p>\n\n<p>it is hard to say if there is any improvement. But at least, it doesn't perform worse than baseline model.</p>\n\n<p>(note: if it works, than it can be a good loss for pseudo label in which we expect some degree of noise )</p>",
      "rawMarkdown": "here is the first  results.\n\nit is hard to say if there is any improvement. But at least, it doesn't perform worse than baseline model.\n\n(note: if it works, than it can be a good loss for pseudo label in which we expect some degree of noise )",
      "votes": null
    },
    {
      "id": "678296",
      "postDate": "11/21/2019 07:46:36",
      "content": "<p>results for tempering softmax (heaviness of the tail)</p>",
      "rawMarkdown": "results for tempering softmax (heaviness of the tail)",
      "votes": null
    },
    {
      "id": "679047",
      "postDate": "11/22/2019 07:38:19",
      "content": "<p>here is the final results\n- tempering the exp function  in sigmoid (or softmax) has good effects in noisy labels. the overfitting is less, even at low learning rate. it gives better private and public score.</p>\n\n<ul>\n<li>tempering of the log function for log loss is less effective</li>\n</ul>\n\n<p>```\nclassification of cloud images\nbaseline 0.64441-private-lb ( probed classification error) : 0.22892 \nbaseline results : hrnet18 net (probed classification error at threshold 0.50/0.65) : 0.23250/0.24830 \ntempered loss  results : hrnet18 net (probed classification error at threshold 0.50/0.65) : 0.22998/0.22748</p>\n\n<p>```</p>",
      "rawMarkdown": "here is the final results\n- tempering the exp function  in sigmoid (or softmax) has good effects in noisy labels. the overfitting is less, even at low learning rate. it gives better private and public score.\n\n- tempering of the log function for log loss is less effective\n\n```\nclassification of cloud images\nbaseline 0.64441-private-lb ( probed classification error) : 0.22892 \nbaseline results : hrnet18 net (probed classification error at threshold 0.50/0.65) : 0.23250/0.24830 \ntempered loss  results : hrnet18 net (probed classification error at threshold 0.50/0.65) : 0.22998/0.22748\n \n \n\n```",
      "votes": null
    },
    {
      "id": "1164330",
      "postDate": "01/22/2021 09:57:41",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>   not sure why did some one not see your post.<br>\nhow do we interpret t1/t2 values  if increased  or decreased. M finding it hard to understand. <br>\nPaper says t2=4,t1=0.8 gives best results ,but not sure about the noise level.</p>",
      "rawMarkdown": "hengck23   not sure why did some one not see your post.\nhow do we interpret t1/t2 values  if increased  or decreased. M finding it hard to understand. \nPaper says t2=4,t1=0.8 gives best results ,but not sure about the noise level.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1164330,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "01/22/2021 09:57:41",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>   not sure why did some one not see your post.<br>\nhow do we interpret t1/t2 values  if increased  or decreased. M finding it hard to understand. <br>\nPaper says t2=4,t1=0.8 gives best results ,but not sure about the noise level.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 678222,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/21/2019 05:30:49",
      "content": "<p>here is the first  results.</p>\n\n<p>it is hard to say if there is any improvement. But at least, it doesn't perform worse than baseline model.</p>\n\n<p>(note: if it works, than it can be a good loss for pseudo label in which we expect some degree of noise )</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 678296,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/21/2019 07:46:36",
      "content": "<p>results for tempering softmax (heaviness of the tail)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 679047,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/22/2019 07:38:19",
      "content": "<p>here is the final results\n- tempering the exp function  in sigmoid (or softmax) has good effects in noisy labels. the overfitting is less, even at low learning rate. it gives better private and public score.</p>\n\n<ul>\n<li>tempering of the log function for log loss is less effective</li>\n</ul>\n\n<p>```\nclassification of cloud images\nbaseline 0.64441-private-lb ( probed classification error) : 0.22892 \nbaseline results : hrnet18 net (probed classification error at threshold 0.50/0.65) : 0.23250/0.24830 \ntempered loss  results : hrnet18 net (probed classification error at threshold 0.50/0.65) : 0.22998/0.22748</p>\n\n<p>```</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "676424": "reference paper and google blog:\nhttps://arxiv.org/abs/1906.03361\nhttps://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html\nhttps://google.github.io/bi-tempered-loss/\n\ngoogle bi-tempered loss is supposed to fight noisy data. the objective of this work is to see how it will perform on this cloud task. Follow this thread if you are interested in the work \n\ncode reference:\nhttps://github.com/fhopfmueller/bi-tempered-loss-pytorch/commits?author=fhopfmueller&amp;since=2019-09-01&amp;until=2019-09-12\n\n---\n\ncode: https://drive.google.com/open?id=1uxMIEM1t2rGSLyYWinqliKDtFnjUJ4Ir\n\n- we investigate the effect of  bi-tempered-loss for image label classification only\n- implement nigh resolution net for image classification\n- some lb submission for baseline model and bi-tempered model",
    "678222": "here is the first  results.\n\nit is hard to say if there is any improvement. But at least, it doesn't perform worse than baseline model.\n\n(note: if it works, than it can be a good loss for pseudo label in which we expect some degree of noise )",
    "678296": "results for tempering softmax (heaviness of the tail)",
    "679047": "here is the final results\n- tempering the exp function  in sigmoid (or softmax) has good effects in noisy labels. the overfitting is less, even at low learning rate. it gives better private and public score.\n\n- tempering of the log function for log loss is less effective\n\n```\nclassification of cloud images\nbaseline 0.64441-private-lb ( probed classification error) : 0.22892 \nbaseline results : hrnet18 net (probed classification error at threshold 0.50/0.65) : 0.23250/0.24830 \ntempered loss  results : hrnet18 net (probed classification error at threshold 0.50/0.65) : 0.22998/0.22748\n \n \n\n```",
    "1164330": "hengck23   not sure why did some one not see your post.\nhow do we interpret t1/t2 values  if increased  or decreased. M finding it hard to understand. \nPaper says t2=4,t1=0.8 gives best results ,but not sure about the noise level."
  },
  "source": "meta"
}