{
  "id": 76238,
  "title": "Why pos_weight didn't work?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76238",
  "author_name": "",
  "post_date": "2018-12-31T00:39:49.494610300Z",
  "votes": 4,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I used pytorch and set pos_weight like following.\n<code>\nclass_weight = torch.FloatTensor([15.5]).cuda()\nloss_fn = torch.nn.BCEWithLogitsLoss(reduction='sum', pos_weight=class_weight).cuda()\n</code>\nThe number of data that target=1 is 15.5 times larger, so I set 15.5 for pos_weight.\nBut LB f1 score was worse about 0.015.\nI think one reason is threshold_search. We can find best threshold for f1. So it needn't set pos_weight?</p>",
  "messages": [
    {
      "id": "447969",
      "postDate": "12/31/2018 00:39:49",
      "content": "<p>I used pytorch and set pos_weight like following.\n<code>\nclass_weight = torch.FloatTensor([15.5]).cuda()\nloss_fn = torch.nn.BCEWithLogitsLoss(reduction='sum', pos_weight=class_weight).cuda()\n</code>\nThe number of data that target=1 is 15.5 times larger, so I set 15.5 for pos_weight.\nBut LB f1 score was worse about 0.015.\nI think one reason is threshold_search. We can find best threshold for f1. So it needn't set pos_weight?</p>",
      "rawMarkdown": "I used pytorch and set pos_weight like following.\n```\nclass_weight = torch.FloatTensor([15.5]).cuda()\nloss_fn = torch.nn.BCEWithLogitsLoss(reduction='sum', pos_weight=class_weight).cuda()\n```\nThe number of data that target=1 is 15.5 times larger, so I set 15.5 for pos_weight.\nBut LB f1 score was worse about 0.015.\nI think one reason is threshold_search. We can find best threshold for f1. So it needn't set pos_weight?",
      "votes": null
    },
    {
      "id": "447983",
      "postDate": "12/31/2018 01:47:24",
      "content": "<p>I experimented with weighted loss fairly extensively, and found I got the best results for a weight of about 3:1.  However, it seems to have higher variance then just using a default even weighting.</p>",
      "rawMarkdown": "I experimented with weighted loss fairly extensively, and found I got the best results for a weight of about 3:1.  However, it seems to have higher variance then just using a default even weighting.",
      "votes": null
    },
    {
      "id": "447989",
      "postDate": "12/31/2018 02:03:45",
      "content": "<p>This class_weight is too high for positive sample. You can try to set a dynamic weight for loss function, sometimes it helps improve a little. The max weight I have tried is 1.25.</p>",
      "rawMarkdown": "This class_weight is too high for positive sample. You can try to set a dynamic weight for loss function, sometimes it helps improve a little. The max weight I have tried is 1.25.",
      "votes": null
    },
    {
      "id": "448003",
      "postDate": "12/31/2018 03:11:48",
      "content": "<p>Thank you for your sharing! I usually calculate weight with data distribution for GBDT. But it may not work for NN. Anyway I will try it:)</p>",
      "rawMarkdown": "Thank you for your sharing! I usually calculate weight with data distribution for GBDT. But it may not work for NN. Anyway I will try it:)",
      "votes": null
    },
    {
      "id": "448004",
      "postDate": "12/31/2018 03:12:42",
      "content": "<p>Thank you very much! My pos_weight seems to be too high. I will try less pos_weight.</p>",
      "rawMarkdown": "Thank you very much! My pos_weight seems to be too high. I will try less pos_weight.",
      "votes": null
    },
    {
      "id": "448276",
      "postDate": "12/31/2018 16:59:46",
      "content": "<p>RNNs are more prone to a vanishing/exploding gradient compared to a GBDT or most NN architectures. It's not necessarily a problem with the data or your parameter settings.</p>",
      "rawMarkdown": "RNNs are more prone to a vanishing/exploding gradient compared to a GBDT or most NN architectures. It's not necessarily a problem with the data or your parameter settings.",
      "votes": null
    },
    {
      "id": "448387",
      "postDate": "01/01/2019 01:12:44",
      "content": "<p>How to use class_weight in pytorch</p>",
      "rawMarkdown": "How to use class_weight in pytorch",
      "votes": null
    },
    {
      "id": "448494",
      "postDate": "01/01/2019 10:34:33",
      "content": "<p>Threre is a pos_weight parameter in BCEWithLogitsLoss function</p>",
      "rawMarkdown": "Threre is a pos_weight parameter in BCEWithLogitsLoss function",
      "votes": null
    },
    {
      "id": "448500",
      "postDate": "01/01/2019 10:43:36",
      "content": "<p>thank you  very much!</p>",
      "rawMarkdown": "thank you  very much!",
      "votes": null
    },
    {
      "id": "449444",
      "postDate": "01/03/2019 06:37:25",
      "content": "<p>When I myself used pos_weights nothing changed except the best threshold</p>",
      "rawMarkdown": "When I myself used pos_weights nothing changed except the best threshold",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 447983,
      "author_name": "stevedraper",
      "author_url": "",
      "post_date": "12/31/2018 01:47:24",
      "content": "<p>I experimented with weighted loss fairly extensively, and found I got the best results for a weight of about 3:1.  However, it seems to have higher variance then just using a default even weighting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 448004,
          "author_name": "takuok",
          "author_url": "",
          "post_date": "12/31/2018 03:12:42",
          "content": "<p>Thank you very much! My pos_weight seems to be too high. I will try less pos_weight.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 447989,
      "author_name": "kaggleczs",
      "author_url": "",
      "post_date": "12/31/2018 02:03:45",
      "content": "<p>This class_weight is too high for positive sample. You can try to set a dynamic weight for loss function, sometimes it helps improve a little. The max weight I have tried is 1.25.</p>",
      "votes": null,
      "replies": [
        {
          "id": 448003,
          "author_name": "takuok",
          "author_url": "",
          "post_date": "12/31/2018 03:11:48",
          "content": "<p>Thank you for your sharing! I usually calculate weight with data distribution for GBDT. But it may not work for NN. Anyway I will try it:)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 448276,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "12/31/2018 16:59:46",
          "content": "<p>RNNs are more prone to a vanishing/exploding gradient compared to a GBDT or most NN architectures. It's not necessarily a problem with the data or your parameter settings.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 448387,
          "author_name": "chenhao2334",
          "author_url": "",
          "post_date": "01/01/2019 01:12:44",
          "content": "<p>How to use class_weight in pytorch</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 448494,
          "author_name": "kaggleczs",
          "author_url": "",
          "post_date": "01/01/2019 10:34:33",
          "content": "<p>Threre is a pos_weight parameter in BCEWithLogitsLoss function</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 448500,
          "author_name": "chenhao2334",
          "author_url": "",
          "post_date": "01/01/2019 10:43:36",
          "content": "<p>thank you  very much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 449444,
      "author_name": "mlwhiz",
      "author_url": "",
      "post_date": "01/03/2019 06:37:25",
      "content": "<p>When I myself used pos_weights nothing changed except the best threshold</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "447969": "I used pytorch and set pos_weight like following.\n```\nclass_weight = torch.FloatTensor([15.5]).cuda()\nloss_fn = torch.nn.BCEWithLogitsLoss(reduction='sum', pos_weight=class_weight).cuda()\n```\nThe number of data that target=1 is 15.5 times larger, so I set 15.5 for pos_weight.\nBut LB f1 score was worse about 0.015.\nI think one reason is threshold_search. We can find best threshold for f1. So it needn't set pos_weight?",
    "447983": "I experimented with weighted loss fairly extensively, and found I got the best results for a weight of about 3:1.  However, it seems to have higher variance then just using a default even weighting.",
    "447989": "This class_weight is too high for positive sample. You can try to set a dynamic weight for loss function, sometimes it helps improve a little. The max weight I have tried is 1.25.",
    "448003": "Thank you for your sharing! I usually calculate weight with data distribution for GBDT. But it may not work for NN. Anyway I will try it:)",
    "448004": "Thank you very much! My pos_weight seems to be too high. I will try less pos_weight.",
    "448276": "RNNs are more prone to a vanishing/exploding gradient compared to a GBDT or most NN architectures. It's not necessarily a problem with the data or your parameter settings.",
    "448387": "How to use class_weight in pytorch",
    "448494": "Threre is a pos_weight parameter in BCEWithLogitsLoss function",
    "448500": "thank you  very much!",
    "449444": "When I myself used pos_weights nothing changed except the best threshold"
  },
  "source": "meta"
}