{
  "id": 172618,
  "title": "Aldfly + Nocall = Good Solution?",
  "url": "/competitions/birdsong-recognition/discussion/172618",
  "author_name": "",
  "post_date": "2020-08-05T19:18:41.410417700Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>That's distribution of the best public solution till now (<a href=\"https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast\">https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast</a>, <a href=\"/ttahara\">@ttahara</a>, another great thanks for strong baseline!):</p>\n\n<p>|  <strong>bird name</strong>| <strong>count</strong> |\n| --- | --- |\n| aldfly |  40|\n| nocall |  30|\n| btnwar |  1|\n| scatan |  1|\n| aldfly hamfly  |  1|\n| comyel |  1|\n| casfin |  1|\n| aldfly comyel |  1|</p>\n\n<p>So, we have dominant values for answers \"aldfly\" and \"nocall\". Thats's strange, but we have had great f1-score with that: 0.58. So, roughly speaking, more than half answers are correct. I have some ideas, why that have happened:\n1. In my opinion, there is only tiny probability, that so many answers \"only aldfly\" can be correct. Chances are, this is only one part of the answer, but our model can't give different labels for one sample (only one label exists during training), so, maybe, it catches just the popular one and that's it.\n2. There is the same observation with answers \"nocall\": that's strange to give us a task to find the answer with many correct ones are \"no answer\". Maybe corresponding sounds don't consists birds' voices as strictly as inside the training dataset. Or, maybe, our model despite everything, tries to give different answers, but can't struggles with 0.5 threshold even with one answer.</p>\n\n<p>What are you thinking about this bird's distribution?</p>",
  "messages": [
    {
      "id": "959646",
      "postDate": "08/05/2020 19:18:41",
      "content": "<p>That's distribution of the best public solution till now (<a href=\"https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast\">https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast</a>, <a href=\"/ttahara\">@ttahara</a>, another great thanks for strong baseline!):</p>\n\n<p>|  <strong>bird name</strong>| <strong>count</strong> |\n| --- | --- |\n| aldfly |  40|\n| nocall |  30|\n| btnwar |  1|\n| scatan |  1|\n| aldfly hamfly  |  1|\n| comyel |  1|\n| casfin |  1|\n| aldfly comyel |  1|</p>\n\n<p>So, we have dominant values for answers \"aldfly\" and \"nocall\". Thats's strange, but we have had great f1-score with that: 0.58. So, roughly speaking, more than half answers are correct. I have some ideas, why that have happened:\n1. In my opinion, there is only tiny probability, that so many answers \"only aldfly\" can be correct. Chances are, this is only one part of the answer, but our model can't give different labels for one sample (only one label exists during training), so, maybe, it catches just the popular one and that's it.\n2. There is the same observation with answers \"nocall\": that's strange to give us a task to find the answer with many correct ones are \"no answer\". Maybe corresponding sounds don't consists birds' voices as strictly as inside the training dataset. Or, maybe, our model despite everything, tries to give different answers, but can't struggles with 0.5 threshold even with one answer.</p>\n\n<p>What are you thinking about this bird's distribution?</p>",
      "rawMarkdown": "That's distribution of the best public solution till now (https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast, @ttahara, another great thanks for strong baseline!):\n\n|  **bird name**| **count** |\n| --- | --- |\n| aldfly |  40|\n| nocall |  30|\n| btnwar |  1|\n| scatan |  1|\n| aldfly hamfly  |  1|\n| comyel |  1|\n| casfin |  1|\n| aldfly comyel |  1|\n\nSo, we have dominant values for answers \"aldfly\" and \"nocall\". Thats's strange, but we have had great f1-score with that: 0.58. So, roughly speaking, more than half answers are correct. I have some ideas, why that have happened:\n1. In my opinion, there is only tiny probability, that so many answers \"only aldfly\" can be correct. Chances are, this is only one part of the answer, but our model can't give different labels for one sample (only one label exists during training), so, maybe, it catches just the popular one and that's it.\n2. There is the same observation with answers \"nocall\": that's strange to give us a task to find the answer with many correct ones are \"no answer\". Maybe corresponding sounds don't consists birds' voices as strictly as inside the training dataset. Or, maybe, our model despite everything, tries to give different answers, but can't struggles with 0.5 threshold even with one answer.\n\nWhat are you thinking about this bird's distribution?",
      "votes": null
    },
    {
      "id": "959817",
      "postDate": "08/05/2020 23:18:44",
      "content": "<p>I would like to know more about this. <a href=\"/shonenkov\">@shonenkov</a> Can you explain how you created your custom birdcall check? Is it just the first few rows of the training set? Ty</p>",
      "rawMarkdown": "I would like to know more about this. @shonenkov Can you explain how you created your custom birdcall check? Is it just the first few rows of the training set? Ty",
      "votes": null
    },
    {
      "id": "959830",
      "postDate": "08/06/2020 00:05:51",
      "content": "<p>I believe you are confusing a test submission with an actual submission.  There are 76 rows in Alex's <a href=\"/shonenkov\">@shonenkov</a> test data, but more in the submit for LB data (time them both).  Plus I don't think there are anywhere near that many aldfly's in the submit data.  <a href=\"/returnofsputnik\">@returnofsputnik</a> if you look in the thread for Alex's data he includes a link to the folder used to make it. </p>",
      "rawMarkdown": "I believe you are confusing a test submission with an actual submission.  There are 76 rows in Alex's @shonenkov test data, but more in the submit for LB data (time them both).  Plus I don't think there are anywhere near that many aldfly's in the submit data.  @returnofsputnik if you look in the thread for Alex's data he includes a link to the folder used to make it.",
      "votes": null
    },
    {
      "id": "959872",
      "postDate": "08/06/2020 01:43:21",
      "content": "<p>Oh, you're right, I missed a part in the cometition description with a hidden test set. Thank you for your response! My observations aren't valid.</p>",
      "rawMarkdown": "Oh, you're right, I missed a part in the cometition description with a hidden test set. Thank you for your response! My observations aren't valid.",
      "votes": null
    },
    {
      "id": "959877",
      "postDate": "08/06/2020 01:48:13",
      "content": "<p>Test set, that I am talking about, is just the first 15 row from the train set with manually added sites: <a href=\"https://www.kaggle.com/shonenkov/prepare-check-dataset\">https://www.kaggle.com/shonenkov/prepare-check-dataset</a> (particulary, now it's obvious, why I get almost identity answers for all my models on it)</p>",
      "rawMarkdown": "Test set, that I am talking about, is just the first 15 row from the train set with manually added sites: https://www.kaggle.com/shonenkov/prepare-check-dataset (particulary, now it's obvious, why I get almost identity answers for all my models on it)",
      "votes": null
    },
    {
      "id": "963350",
      "postDate": "08/09/2020 00:12:58",
      "content": "<p>All good, just didn't want you wasting time on a rabbit trail.  This competition is quite challenging already.  I do have a model with 20% correct predictions after eliminating dead air that might be able to do well with a bit more tweaking.  264+1 categories is quite a few choices.</p>",
      "rawMarkdown": "All good, just didn't want you wasting time on a rabbit trail.  This competition is quite challenging already.  I do have a model with 20% correct predictions after eliminating dead air that might be able to do well with a bit more tweaking.  264+1 categories is quite a few choices.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 959817,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "08/05/2020 23:18:44",
      "content": "<p>I would like to know more about this. <a href=\"/shonenkov\">@shonenkov</a> Can you explain how you created your custom birdcall check? Is it just the first few rows of the training set? Ty</p>",
      "votes": null,
      "replies": [
        {
          "id": 959877,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "08/06/2020 01:48:13",
          "content": "<p>Test set, that I am talking about, is just the first 15 row from the train set with manually added sites: <a href=\"https://www.kaggle.com/shonenkov/prepare-check-dataset\">https://www.kaggle.com/shonenkov/prepare-check-dataset</a> (particulary, now it's obvious, why I get almost identity answers for all my models on it)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 959830,
      "author_name": "meckdahl",
      "author_url": "",
      "post_date": "08/06/2020 00:05:51",
      "content": "<p>I believe you are confusing a test submission with an actual submission.  There are 76 rows in Alex's <a href=\"/shonenkov\">@shonenkov</a> test data, but more in the submit for LB data (time them both).  Plus I don't think there are anywhere near that many aldfly's in the submit data.  <a href=\"/returnofsputnik\">@returnofsputnik</a> if you look in the thread for Alex's data he includes a link to the folder used to make it. </p>",
      "votes": null,
      "replies": [
        {
          "id": 959872,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "08/06/2020 01:43:21",
          "content": "<p>Oh, you're right, I missed a part in the cometition description with a hidden test set. Thank you for your response! My observations aren't valid.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 963350,
          "author_name": "meckdahl",
          "author_url": "",
          "post_date": "08/09/2020 00:12:58",
          "content": "<p>All good, just didn't want you wasting time on a rabbit trail.  This competition is quite challenging already.  I do have a model with 20% correct predictions after eliminating dead air that might be able to do well with a bit more tweaking.  264+1 categories is quite a few choices.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "959646": "That's distribution of the best public solution till now (https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast, @ttahara, another great thanks for strong baseline!):\n\n|  **bird name**| **count** |\n| --- | --- |\n| aldfly |  40|\n| nocall |  30|\n| btnwar |  1|\n| scatan |  1|\n| aldfly hamfly  |  1|\n| comyel |  1|\n| casfin |  1|\n| aldfly comyel |  1|\n\nSo, we have dominant values for answers \"aldfly\" and \"nocall\". Thats's strange, but we have had great f1-score with that: 0.58. So, roughly speaking, more than half answers are correct. I have some ideas, why that have happened:\n1. In my opinion, there is only tiny probability, that so many answers \"only aldfly\" can be correct. Chances are, this is only one part of the answer, but our model can't give different labels for one sample (only one label exists during training), so, maybe, it catches just the popular one and that's it.\n2. There is the same observation with answers \"nocall\": that's strange to give us a task to find the answer with many correct ones are \"no answer\". Maybe corresponding sounds don't consists birds' voices as strictly as inside the training dataset. Or, maybe, our model despite everything, tries to give different answers, but can't struggles with 0.5 threshold even with one answer.\n\nWhat are you thinking about this bird's distribution?",
    "959817": "I would like to know more about this. @shonenkov Can you explain how you created your custom birdcall check? Is it just the first few rows of the training set? Ty",
    "959830": "I believe you are confusing a test submission with an actual submission.  There are 76 rows in Alex's @shonenkov test data, but more in the submit for LB data (time them both).  Plus I don't think there are anywhere near that many aldfly's in the submit data.  @returnofsputnik if you look in the thread for Alex's data he includes a link to the folder used to make it.",
    "959872": "Oh, you're right, I missed a part in the cometition description with a hidden test set. Thank you for your response! My observations aren't valid.",
    "959877": "Test set, that I am talking about, is just the first 15 row from the train set with manually added sites: https://www.kaggle.com/shonenkov/prepare-check-dataset (particulary, now it's obvious, why I get almost identity answers for all my models on it)",
    "963350": "All good, just didn't want you wasting time on a rabbit trail.  This competition is quite challenging already.  I do have a model with 20% correct predictions after eliminating dead air that might be able to do well with a bit more tweaking.  264+1 categories is quite a few choices."
  },
  "source": "meta"
}