{
  "id": 181035,
  "title": "Very low score even the model works well",
  "url": "/competitions/birdsong-recognition/discussion/181035",
  "author_name": "",
  "post_date": "2020-09-07T11:00:28.587258100Z",
  "votes": null,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I trained my own model with some configurations on ResNet. I split 20% of the data for validation data and I can get 85% accuracy for validation data according to my training. Everything looks good so far. However, when I submit my code, I always get weird scores that I dont know how to interpret the result. Here are my submissions I tried:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3130748%2F0207f76d28ff89d0be7e62d4bdae72af%2FScreenshot_3.png?generation=1599476027234777&amp;alt=media\" alt=\"\"></p>\n<p>I used <a href=\"https://www.kaggle.com/shonenkov\" target=\"_blank\">@shonenkov</a>'s  <a href=\"https://www.kaggle.com/shonenkov/sample-submission-using-custom-check/data\" target=\"_blank\">submission check</a> data to verify my code whether it is working and there wasn't any problem about my code. Is there anyone faced with the same issue? How can I handle this problem?</p>",
  "messages": [
    {
      "id": "1001470",
      "postDate": "09/07/2020 11:00:28",
      "content": "<p>I trained my own model with some configurations on ResNet. I split 20% of the data for validation data and I can get 85% accuracy for validation data according to my training. Everything looks good so far. However, when I submit my code, I always get weird scores that I dont know how to interpret the result. Here are my submissions I tried:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3130748%2F0207f76d28ff89d0be7e62d4bdae72af%2FScreenshot_3.png?generation=1599476027234777&amp;alt=media\" alt=\"\"></p>\n<p>I used <a href=\"https://www.kaggle.com/shonenkov\" target=\"_blank\">@shonenkov</a>'s  <a href=\"https://www.kaggle.com/shonenkov/sample-submission-using-custom-check/data\" target=\"_blank\">submission check</a> data to verify my code whether it is working and there wasn't any problem about my code. Is there anyone faced with the same issue? How can I handle this problem?</p>",
      "rawMarkdown": "I trained my own model with some configurations on ResNet. I split 20% of the data for validation data and I can get 85% accuracy for validation data according to my training. Everything looks good so far. However, when I submit my code, I always get weird scores that I dont know how to interpret the result. Here are my submissions I tried:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3130748%2F0207f76d28ff89d0be7e62d4bdae72af%2FScreenshot_3.png?generation=1599476027234777&alt=media)\n\n\nI used @shonenkov's  [submission check](https://www.kaggle.com/shonenkov/sample-submission-using-custom-check/data) data to verify my code whether it is working and there wasn't any problem about my code. Is there anyone faced with the same issue? How can I handle this problem?",
      "votes": null
    },
    {
      "id": "1001510",
      "postDate": "09/07/2020 11:24:42",
      "content": "<p>I got LB 0 with my first successful submission, and to be honest I don't know what I did to fix it.  Make sure you resample input files to 32 kHz, even though host says they are sampled at this frequency.  Also make sure you get the right number of 5 seconds clips for each file.  You can look at public notebooks to see what they do.  I coded something different, but would have copied their code if mine did not work.</p>",
      "rawMarkdown": "I got LB 0 with my first successful submission, and to be honest I don't know what I did to fix it.  Make sure you resample input files to 32 kHz, even though host says they are sampled at this frequency.  Also make sure you get the right number of 5 seconds clips for each file.  You can look at public notebooks to see what they do.  I coded something different, but would have copied their code if mine did not work.",
      "votes": null
    },
    {
      "id": "1001541",
      "postDate": "09/07/2020 11:52:38",
      "content": "<p>I used the files with 32kHZ resampling and also number of outputs match with inputs. I don't know how I got 0.544 score but I couldn't get the same again.</p>",
      "rawMarkdown": "I used the files with 32kHZ resampling and also number of outputs match with inputs. I don't know how I got 0.544 score but I couldn't get the same again.",
      "votes": null
    },
    {
      "id": "1001545",
      "postDate": "09/07/2020 11:54:38",
      "content": "<blockquote>\n  <p>number of outputs match with inputs</p>\n</blockquote>\n<p>How could you know given we cannot see test data?</p>",
      "rawMarkdown": "> number of outputs match with inputs\n\nHow could you know given we cannot see test data?",
      "votes": null
    },
    {
      "id": "1001557",
      "postDate": "09/07/2020 12:02:55",
      "content": "<p>I used birdcall check folder from <a href=\"https://www.kaggle.com/shonenkov/sample-submission-using-custom-check/data\" target=\"_blank\">this link</a> which is a composition of site_1, site_2 and site_3 audio data. My model extracted the same amount of outputs for each site. I thought it will work as the same for submission but apparently it is not </p>",
      "rawMarkdown": "I used birdcall check folder from [this link](https://www.kaggle.com/shonenkov/sample-submission-using-custom-check/data) which is a composition of site_1, site_2 and site_3 audio data. My model extracted the same amount of outputs for each site. I thought it will work as the same for submission but apparently it is not",
      "votes": null
    },
    {
      "id": "1001566",
      "postDate": "09/07/2020 12:14:38",
      "content": "<p>To get exactly <code>0.544</code> you have to predict all test-samples as <code>nocall</code> </p>\n<p>It is very likely that your model just predicted every entry as <code>nocall</code> and this got you the result.</p>",
      "rawMarkdown": "To get exactly `0.544` you have to predict all test-samples as `nocall` \n\nIt is very likely that your model just predicted every entry as `nocall` and this got you the result.",
      "votes": null
    },
    {
      "id": "1001573",
      "postDate": "09/07/2020 12:19:20",
      "content": "<p>Some possibilities - </p>\n<p>You could have missed something in your inference data loading, and it is not equal to your training data loading, leading to wildly inaccurate predictions.<br>\nYou could have data leak in your validation data (for example, training on parts of a file with unique noise pattern, and using another parts of the same file with same unique noise pattern for validation).<br>\nIt also could be as simple as some error in id &gt; name translation for your answers.</p>\n<p>0.544 score is all \"nocall\" submission, all others are valid submissions, but <em>hugely</em> inaccurate with almost no correct answers.</p>",
      "rawMarkdown": "Some possibilities - \n\nYou could have missed something in your inference data loading, and it is not equal to your training data loading, leading to wildly inaccurate predictions.\nYou could have data leak in your validation data (for example, training on parts of a file with unique noise pattern, and using another parts of the same file with same unique noise pattern for validation).\nIt also could be as simple as some error in id > name translation for your answers.\n\n0.544 score is all \"nocall\" submission, all others are valid submissions, but _hugely_ inaccurate with almost no correct answers.",
      "votes": null
    },
    {
      "id": "1001619",
      "postDate": "09/07/2020 12:42:39",
      "content": "<blockquote>\n  <p>I used birdcall check folder from this link </p>\n</blockquote>\n<p>this is not the test data.  As said, you cannot see the test data.  It is only available once you submit your notebook.</p>",
      "rawMarkdown": "> I used birdcall check folder from this link \n\nthis is not the test data.  As said, you cannot see the test data.  It is only available once you submit your notebook.",
      "votes": null
    },
    {
      "id": "1002232",
      "postDate": "09/08/2020 00:10:00",
      "content": "<p>Same here… <br>\nIf I cut my training at maximum validation accuracy (astonishing 90% with training accuracy of 99%) I can hardly touch 0.550 with some thresholding. The results are similar if I cut training at minimum validation loss.</p>\n<p>All, I can realize is that my model is overfitting on our training dataset but fails very badly on that test dataset.</p>",
      "rawMarkdown": "Same here... \nIf I cut my training at maximum validation accuracy (astonishing 90% with training accuracy of 99%) I can hardly touch 0.550 with some thresholding. The results are similar if I cut training at minimum validation loss.\n\nAll, I can realize is that my model is overfitting on our training dataset but fails very badly on that test dataset.",
      "votes": null
    },
    {
      "id": "1002234",
      "postDate": "09/08/2020 00:16:36",
      "content": "<p><a href=\"https://www.kaggle.com/serkavak\" target=\"_blank\">@serkavak</a> </p>\n<p>I see that ur score improved to 560… Will be helpful if you could share what worked for you…</p>",
      "rawMarkdown": "serkavak \n\nI see that ur score improved to 560... Will be helpful if you could share what worked for you...",
      "votes": null
    },
    {
      "id": "1005240",
      "postDate": "09/10/2020 10:49:32",
      "content": "<p>I finally successfully submit my code and I surprised with the result though. I split 20% of the data for validation data in training process and I got 75% to 85% of accuracy with various hyperparameter settings. However for LB, I only got 0.239. Apparently the model is completely overfitting </p>",
      "rawMarkdown": "I finally successfully submit my code and I surprised with the result though. I split 20% of the data for validation data in training process and I got 75% to 85% of accuracy with various hyperparameter settings. However for LB, I only got 0.239. Apparently the model is completely overfitting",
      "votes": null
    },
    {
      "id": "1005241",
      "postDate": "09/10/2020 10:50:32",
      "content": "<p>I forked baseline submission it is not the score i got with my own model</p>",
      "rawMarkdown": "I forked baseline submission it is not the score i got with my own model",
      "votes": null
    },
    {
      "id": "1005253",
      "postDate": "09/10/2020 11:13:10",
      "content": "<p>If it can smooth your pain, I got LB 0 with my first successful submission…</p>",
      "rawMarkdown": "If it can smooth your pain, I got LB 0 with my first successful submission...",
      "votes": null
    },
    {
      "id": "1005259",
      "postDate": "09/10/2020 11:20:28",
      "content": "<p>your code could be like mine… try this, increase threshold crazily till 0.9 or even 0.99<br>\nIf ur score is well below 0.52, it means your model is predicting a bird in many no bird cases… <br>\nso increasing threshold should help…<br>\nor<br>\ndo a 265 class model (with a no bird class) or<br>\ndo a bird-nobird model and then 264 bird model or <br>\nsomething else…</p>\n<p>I tried like 5 different models… using keras… ensemble, direct, with/without noise etc., <br>\nall of them crazily overfits on train set (both training and validation sets - ratio - 0.8 and 0.2) but fails badly on leaderboard… my best is 0.549 with my own code</p>",
      "rawMarkdown": "your code could be like mine... try this, increase threshold crazily till 0.9 or even 0.99\nIf ur score is well below 0.52, it means your model is predicting a bird in many no bird cases... \nso increasing threshold should help...\nor\ndo a 265 class model (with a no bird class) or\ndo a bird-nobird model and then 264 bird model or \nsomething else...\n\nI tried like 5 different models... using keras... ensemble, direct, with/without noise etc., \nall of them crazily overfits on train set (both training and validation sets - ratio - 0.8 and 0.2) but fails badly on leaderboard... my best is 0.549 with my own code",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1001510,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "09/07/2020 11:24:42",
      "content": "<p>I got LB 0 with my first successful submission, and to be honest I don't know what I did to fix it.  Make sure you resample input files to 32 kHz, even though host says they are sampled at this frequency.  Also make sure you get the right number of 5 seconds clips for each file.  You can look at public notebooks to see what they do.  I coded something different, but would have copied their code if mine did not work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1001541,
      "author_name": "serkavak",
      "author_url": "",
      "post_date": "09/07/2020 11:52:38",
      "content": "<p>I used the files with 32kHZ resampling and also number of outputs match with inputs. I don't know how I got 0.544 score but I couldn't get the same again.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1001545,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/07/2020 11:54:38",
          "content": "<blockquote>\n  <p>number of outputs match with inputs</p>\n</blockquote>\n<p>How could you know given we cannot see test data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1001557,
          "author_name": "serkavak",
          "author_url": "",
          "post_date": "09/07/2020 12:02:55",
          "content": "<p>I used birdcall check folder from <a href=\"https://www.kaggle.com/shonenkov/sample-submission-using-custom-check/data\" target=\"_blank\">this link</a> which is a composition of site_1, site_2 and site_3 audio data. My model extracted the same amount of outputs for each site. I thought it will work as the same for submission but apparently it is not </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1001566,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "09/07/2020 12:14:38",
          "content": "<p>To get exactly <code>0.544</code> you have to predict all test-samples as <code>nocall</code> </p>\n<p>It is very likely that your model just predicted every entry as <code>nocall</code> and this got you the result.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1001619,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/07/2020 12:42:39",
          "content": "<blockquote>\n  <p>I used birdcall check folder from this link </p>\n</blockquote>\n<p>this is not the test data.  As said, you cannot see the test data.  It is only available once you submit your notebook.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1005240,
          "author_name": "serkavak",
          "author_url": "",
          "post_date": "09/10/2020 10:49:32",
          "content": "<p>I finally successfully submit my code and I surprised with the result though. I split 20% of the data for validation data in training process and I got 75% to 85% of accuracy with various hyperparameter settings. However for LB, I only got 0.239. Apparently the model is completely overfitting </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1005253,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/10/2020 11:13:10",
          "content": "<p>If it can smooth your pain, I got LB 0 with my first successful submission…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1005259,
          "author_name": "vsvrp1995",
          "author_url": "",
          "post_date": "09/10/2020 11:20:28",
          "content": "<p>your code could be like mine… try this, increase threshold crazily till 0.9 or even 0.99<br>\nIf ur score is well below 0.52, it means your model is predicting a bird in many no bird cases… <br>\nso increasing threshold should help…<br>\nor<br>\ndo a 265 class model (with a no bird class) or<br>\ndo a bird-nobird model and then 264 bird model or <br>\nsomething else…</p>\n<p>I tried like 5 different models… using keras… ensemble, direct, with/without noise etc., <br>\nall of them crazily overfits on train set (both training and validation sets - ratio - 0.8 and 0.2) but fails badly on leaderboard… my best is 0.549 with my own code</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1001573,
      "author_name": "fffrrt",
      "author_url": "",
      "post_date": "09/07/2020 12:19:20",
      "content": "<p>Some possibilities - </p>\n<p>You could have missed something in your inference data loading, and it is not equal to your training data loading, leading to wildly inaccurate predictions.<br>\nYou could have data leak in your validation data (for example, training on parts of a file with unique noise pattern, and using another parts of the same file with same unique noise pattern for validation).<br>\nIt also could be as simple as some error in id &gt; name translation for your answers.</p>\n<p>0.544 score is all \"nocall\" submission, all others are valid submissions, but <em>hugely</em> inaccurate with almost no correct answers.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1002232,
      "author_name": "vsvrp1995",
      "author_url": "",
      "post_date": "09/08/2020 00:10:00",
      "content": "<p>Same here… <br>\nIf I cut my training at maximum validation accuracy (astonishing 90% with training accuracy of 99%) I can hardly touch 0.550 with some thresholding. The results are similar if I cut training at minimum validation loss.</p>\n<p>All, I can realize is that my model is overfitting on our training dataset but fails very badly on that test dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1002234,
          "author_name": "vsvrp1995",
          "author_url": "",
          "post_date": "09/08/2020 00:16:36",
          "content": "<p><a href=\"https://www.kaggle.com/serkavak\" target=\"_blank\">@serkavak</a> </p>\n<p>I see that ur score improved to 560… Will be helpful if you could share what worked for you…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1005241,
          "author_name": "serkavak",
          "author_url": "",
          "post_date": "09/10/2020 10:50:32",
          "content": "<p>I forked baseline submission it is not the score i got with my own model</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1001470": "I trained my own model with some configurations on ResNet. I split 20% of the data for validation data and I can get 85% accuracy for validation data according to my training. Everything looks good so far. However, when I submit my code, I always get weird scores that I dont know how to interpret the result. Here are my submissions I tried:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3130748%2F0207f76d28ff89d0be7e62d4bdae72af%2FScreenshot_3.png?generation=1599476027234777&alt=media)\n\n\nI used @shonenkov's  [submission check](https://www.kaggle.com/shonenkov/sample-submission-using-custom-check/data) data to verify my code whether it is working and there wasn't any problem about my code. Is there anyone faced with the same issue? How can I handle this problem?",
    "1001510": "I got LB 0 with my first successful submission, and to be honest I don't know what I did to fix it.  Make sure you resample input files to 32 kHz, even though host says they are sampled at this frequency.  Also make sure you get the right number of 5 seconds clips for each file.  You can look at public notebooks to see what they do.  I coded something different, but would have copied their code if mine did not work.",
    "1001541": "I used the files with 32kHZ resampling and also number of outputs match with inputs. I don't know how I got 0.544 score but I couldn't get the same again.",
    "1001545": "> number of outputs match with inputs\n\nHow could you know given we cannot see test data?",
    "1001557": "I used birdcall check folder from [this link](https://www.kaggle.com/shonenkov/sample-submission-using-custom-check/data) which is a composition of site_1, site_2 and site_3 audio data. My model extracted the same amount of outputs for each site. I thought it will work as the same for submission but apparently it is not",
    "1001566": "To get exactly `0.544` you have to predict all test-samples as `nocall` \n\nIt is very likely that your model just predicted every entry as `nocall` and this got you the result.",
    "1001573": "Some possibilities - \n\nYou could have missed something in your inference data loading, and it is not equal to your training data loading, leading to wildly inaccurate predictions.\nYou could have data leak in your validation data (for example, training on parts of a file with unique noise pattern, and using another parts of the same file with same unique noise pattern for validation).\nIt also could be as simple as some error in id > name translation for your answers.\n\n0.544 score is all \"nocall\" submission, all others are valid submissions, but _hugely_ inaccurate with almost no correct answers.",
    "1001619": "> I used birdcall check folder from this link \n\nthis is not the test data.  As said, you cannot see the test data.  It is only available once you submit your notebook.",
    "1002232": "Same here... \nIf I cut my training at maximum validation accuracy (astonishing 90% with training accuracy of 99%) I can hardly touch 0.550 with some thresholding. The results are similar if I cut training at minimum validation loss.\n\nAll, I can realize is that my model is overfitting on our training dataset but fails very badly on that test dataset.",
    "1002234": "serkavak \n\nI see that ur score improved to 560... Will be helpful if you could share what worked for you...",
    "1005240": "I finally successfully submit my code and I surprised with the result though. I split 20% of the data for validation data in training process and I got 75% to 85% of accuracy with various hyperparameter settings. However for LB, I only got 0.239. Apparently the model is completely overfitting",
    "1005241": "I forked baseline submission it is not the score i got with my own model",
    "1005253": "If it can smooth your pain, I got LB 0 with my first successful submission...",
    "1005259": "your code could be like mine... try this, increase threshold crazily till 0.9 or even 0.99\nIf ur score is well below 0.52, it means your model is predicting a bird in many no bird cases... \nso increasing threshold should help...\nor\ndo a 265 class model (with a no bird class) or\ndo a bird-nobird model and then 264 bird model or \nsomething else...\n\nI tried like 5 different models... using keras... ensemble, direct, with/without noise etc., \nall of them crazily overfits on train set (both training and validation sets - ratio - 0.8 and 0.2) but fails badly on leaderboard... my best is 0.549 with my own code"
  },
  "source": "meta"
}