{
  "id": 162002,
  "title": "Classification Threshold?",
  "url": "/competitions/birdsong-recognition/discussion/162002",
  "author_name": "",
  "post_date": "2020-06-27T01:12:40.794424500Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Ok so, turns out we're dealing with multi-label classification. Usually when this is the case I might do something like </p>\n\n<ol>\n<li>train a multi-label classifier on the training dateset</li>\n<li>work out a threshold (to choose which labels to set as positive and which to set as negative) that seems to give a good F1 score on the training dataset</li>\n<li>Use that same threshold on the test dataset and hope all goes well</li>\n</ol>\n\n<p>In this case though we can't use the training dataset to estimate the label density of the test dataset, since the training data only has one label per sample.  In fact, since birdcalls-per-second is highly variable thought the day in a given location, let alone multiple sites, the label density of the test dataset could have a huge range of possible values.</p>\n\n<p>So: the most reasonable approach in my eyes is: </p>\n\n<ol>\n<li>Try to get my hands on some multi-label data that it likely to be similar to the test set (a cetacean posted some interesting stuff in the data disclosure thread) and use this for choosing a threshold.</li>\n<li>And sadly I think this is going to be much more important: find a way to effectively probe the private lb to get a better idea of what the distribution of test labels is like.</li>\n</ol>\n\n<p>You could have the best bird-call identification model ever created but if your threshold estimate is based off data that contains one bird call roughly once every 25 seconds and the test set contains one in every 2 then you have no chance unless you've managed to probe that information. </p>\n\n<p>Does that make sense or am I missing something here?</p>",
  "messages": [
    {
      "id": "903594",
      "postDate": "06/27/2020 01:12:40",
      "content": "<p>Ok so, turns out we're dealing with multi-label classification. Usually when this is the case I might do something like </p>\n\n<ol>\n<li>train a multi-label classifier on the training dateset</li>\n<li>work out a threshold (to choose which labels to set as positive and which to set as negative) that seems to give a good F1 score on the training dataset</li>\n<li>Use that same threshold on the test dataset and hope all goes well</li>\n</ol>\n\n<p>In this case though we can't use the training dataset to estimate the label density of the test dataset, since the training data only has one label per sample.  In fact, since birdcalls-per-second is highly variable thought the day in a given location, let alone multiple sites, the label density of the test dataset could have a huge range of possible values.</p>\n\n<p>So: the most reasonable approach in my eyes is: </p>\n\n<ol>\n<li>Try to get my hands on some multi-label data that it likely to be similar to the test set (a cetacean posted some interesting stuff in the data disclosure thread) and use this for choosing a threshold.</li>\n<li>And sadly I think this is going to be much more important: find a way to effectively probe the private lb to get a better idea of what the distribution of test labels is like.</li>\n</ol>\n\n<p>You could have the best bird-call identification model ever created but if your threshold estimate is based off data that contains one bird call roughly once every 25 seconds and the test set contains one in every 2 then you have no chance unless you've managed to probe that information. </p>\n\n<p>Does that make sense or am I missing something here?</p>",
      "rawMarkdown": "Ok so, turns out we're dealing with multi-label classification. Usually when this is the case I might do something like \n\n1. train a multi-label classifier on the training dateset\n2. work out a threshold (to choose which labels to set as positive and which to set as negative) that seems to give a good F1 score on the training dataset\n3. Use that same threshold on the test dataset and hope all goes well\n\nIn this case though we can't use the training dataset to estimate the label density of the test dataset, since the training data only has one label per sample.  In fact, since birdcalls-per-second is highly variable thought the day in a given location, let alone multiple sites, the label density of the test dataset could have a huge range of possible values.\n\nSo: the most reasonable approach in my eyes is: \n\n1. Try to get my hands on some multi-label data that it likely to be similar to the test set (a cetacean posted some interesting stuff in the data disclosure thread) and use this for choosing a threshold.\n2. And sadly I think this is going to be much more important: find a way to effectively probe the private lb to get a better idea of what the distribution of test labels is like.\n\nYou could have the best bird-call identification model ever created but if your threshold estimate is based off data that contains one bird call roughly once every 25 seconds and the test set contains one in every 2 then you have no chance unless you've managed to probe that information. \n\nDoes that make sense or am I missing something here?",
      "votes": null
    },
    {
      "id": "904624",
      "postDate": "06/27/2020 19:12:42",
      "content": "<p>I would advise against your 2nd idea.</p>\n\n<p>My initial thought was to build from the training set a secondary set which more closely resembles the test set, by its description, and use that as a non-lb test set. I haven't had the opportunity to do this, but it seems somewhat straightforward to both create and label such a set.</p>",
      "rawMarkdown": "I would advise against your 2nd idea.\n\nMy initial thought was to build from the training set a secondary set which more closely resembles the test set, by its description, and use that as a non-lb test set. I haven't had the opportunity to do this, but it seems somewhat straightforward to both create and label such a set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 904624,
      "author_name": "sergiumpopa",
      "author_url": "",
      "post_date": "06/27/2020 19:12:42",
      "content": "<p>I would advise against your 2nd idea.</p>\n\n<p>My initial thought was to build from the training set a secondary set which more closely resembles the test set, by its description, and use that as a non-lb test set. I haven't had the opportunity to do this, but it seems somewhat straightforward to both create and label such a set.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "903594": "Ok so, turns out we're dealing with multi-label classification. Usually when this is the case I might do something like \n\n1. train a multi-label classifier on the training dateset\n2. work out a threshold (to choose which labels to set as positive and which to set as negative) that seems to give a good F1 score on the training dataset\n3. Use that same threshold on the test dataset and hope all goes well\n\nIn this case though we can't use the training dataset to estimate the label density of the test dataset, since the training data only has one label per sample.  In fact, since birdcalls-per-second is highly variable thought the day in a given location, let alone multiple sites, the label density of the test dataset could have a huge range of possible values.\n\nSo: the most reasonable approach in my eyes is: \n\n1. Try to get my hands on some multi-label data that it likely to be similar to the test set (a cetacean posted some interesting stuff in the data disclosure thread) and use this for choosing a threshold.\n2. And sadly I think this is going to be much more important: find a way to effectively probe the private lb to get a better idea of what the distribution of test labels is like.\n\nYou could have the best bird-call identification model ever created but if your threshold estimate is based off data that contains one bird call roughly once every 25 seconds and the test set contains one in every 2 then you have no chance unless you've managed to probe that information. \n\nDoes that make sense or am I missing something here?",
    "904624": "I would advise against your 2nd idea.\n\nMy initial thought was to build from the training set a secondary set which more closely resembles the test set, by its description, and use that as a non-lb test set. I haven't had the opportunity to do this, but it seems somewhat straightforward to both create and label such a set."
  },
  "source": "meta"
}