{
  "id": 163528,
  "title": "[thoughts on troubleshooting submissions] what happens if we don't submit valid predictions? ",
  "url": "/competitions/birdsong-recognition/discussion/163528",
  "author_name": "",
  "post_date": "2020-07-02T12:07:37.759482800Z",
  "votes": 5,
  "comment_count": 12,
  "views": 0,
  "content": "<p>My last four submissions, all being substantially different, all scored 0.54, which is what an all zero submission would score. </p>\n\n<p>I am using <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159993\">the custom check phase</a> but I imagine that this still does not remove all ways in which one can have a bug.</p>\n\n<p>Here is my problem - on the files from the check phase, I can see my models perform detection (<code>submission.csv</code> results vastly vary between the models). And yet I get exactly 0.54 on my submissions.</p>\n\n<p>If I were not submitting results for all the rows in the test set, would I see an error? If I would be outputting classes that do not match those in the training set, would I see an error?</p>\n\n<p>If not, would it be possible for such functionality to be implemented?</p>\n\n<p>Assuming this functionality is not in place already, the risk is that someone might be doing something useful from the perspective of building a model to address the challenge, and we may never learn about that as they might have a bug in their code, their submission might not work correctly and they might not ever learn about this fact.</p>\n\n<p>BTW, my apologies for raising this if this functionality is already implemented. Thank you very much for all your help and your thoughts on this!</p>",
  "messages": [
    {
      "id": "912339",
      "postDate": "07/02/2020 12:07:37",
      "content": "<p>My last four submissions, all being substantially different, all scored 0.54, which is what an all zero submission would score. </p>\n\n<p>I am using <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159993\">the custom check phase</a> but I imagine that this still does not remove all ways in which one can have a bug.</p>\n\n<p>Here is my problem - on the files from the check phase, I can see my models perform detection (<code>submission.csv</code> results vastly vary between the models). And yet I get exactly 0.54 on my submissions.</p>\n\n<p>If I were not submitting results for all the rows in the test set, would I see an error? If I would be outputting classes that do not match those in the training set, would I see an error?</p>\n\n<p>If not, would it be possible for such functionality to be implemented?</p>\n\n<p>Assuming this functionality is not in place already, the risk is that someone might be doing something useful from the perspective of building a model to address the challenge, and we may never learn about that as they might have a bug in their code, their submission might not work correctly and they might not ever learn about this fact.</p>\n\n<p>BTW, my apologies for raising this if this functionality is already implemented. Thank you very much for all your help and your thoughts on this!</p>",
      "rawMarkdown": "My last four submissions, all being substantially different, all scored 0.54, which is what an all zero submission would score. \n\nI am using [the custom check phase](https://www.kaggle.com/c/birdsong-recognition/discussion/159993) but I imagine that this still does not remove all ways in which one can have a bug.\n\nHere is my problem - on the files from the check phase, I can see my models perform detection (`submission.csv` results vastly vary between the models). And yet I get exactly 0.54 on my submissions.\n\nIf I were not submitting results for all the rows in the test set, would I see an error? If I would be outputting classes that do not match those in the training set, would I see an error?\n\nIf not, would it be possible for such functionality to be implemented?\n\nAssuming this functionality is not in place already, the risk is that someone might be doing something useful from the perspective of building a model to address the challenge, and we may never learn about that as they might have a bug in their code, their submission might not work correctly and they might not ever learn about this fact.\n\nBTW, my apologies for raising this if this functionality is already implemented. Thank you very much for all your help and your thoughts on this!",
      "votes": null
    },
    {
      "id": "913274",
      "postDate": "07/03/2020 05:14:04",
      "content": "<p>I believe it would be worth to get at least any system errors and output of exceptions. Though I do not know how to get the latter without any possibility to leak critical info</p>",
      "rawMarkdown": "I believe it would be worth to get at least any system errors and output of exceptions. Though I do not know how to get the latter without any possibility to leak critical info",
      "votes": null
    },
    {
      "id": "913378",
      "postDate": "07/03/2020 07:25:02",
      "content": "<p>The solution to my mind would be having the platform check two things:</p>\n\n<ul>\n<li>does the submission contain predictions for all the rows in <code>test.csv</code></li>\n<li>are all the predictions valid, whether they are not malformed, etc</li>\n</ul>\n\n<p>ML code is probably one of the hardest code to debug and having this minimal feedback, not introducing much (any?) potential for leakage could be very helpful. Just a blank exception: <code>You didn't provide predictions for all the rows in test.csv or some of the rows were malformed</code>.</p>",
      "rawMarkdown": "The solution to my mind would be having the platform check two things:\n\n- does the submission contain predictions for all the rows in `test.csv`\n- are all the predictions valid, whether they are not malformed, etc\n\nML code is probably one of the hardest code to debug and having this minimal feedback, not introducing much (any?) potential for leakage could be very helpful. Just a blank exception: `You didn't provide predictions for all the rows in test.csv or some of the rows were malformed`.",
      "votes": null
    },
    {
      "id": "913392",
      "postDate": "07/03/2020 07:37:52",
      "content": "<p>Tbh, I am afraid your version 2 might fail on the out of memory exception and therefore the whole submission resets to 'nocall' and that is the reason why your different submissions all get 0.54. Why so? Because your code reads all files into memory and keeps them there. Each 5 seconds wav (output of reading an mp3) takes up ca. 1.5 mb in memory. So 150 files each ca. 10 min means 150 * 10 * 24 * 1.5 = ca. 54 gb of memory. Makes sense?</p>",
      "rawMarkdown": "Tbh, I am afraid your version 2 might fail on the out of memory exception and therefore the whole submission resets to 'nocall' and that is the reason why your different submissions all get 0.54. Why so? Because your code reads all files into memory and keeps them there. Each 5 seconds wav (output of reading an mp3) takes up ca. 1.5 mb in memory. So 150 files each ca. 10 min means 150 * 10 * 24 * 1.5 = ca. 54 gb of memory. Makes sense?",
      "votes": null
    },
    {
      "id": "913395",
      "postDate": "07/03/2020 07:40:25",
      "content": "<p>I don't read all the files into memory. I read them one by one and not store them. Also, it seems one gets an exception that is communicated to them on the submission page when they use too many resources and the kernel resets.</p>",
      "rawMarkdown": "I don't read all the files into memory. I read them one by one and not store them. Also, it seems one gets an exception that is communicated to them on the submission page when they use too many resources and the kernel resets.",
      "votes": null
    },
    {
      "id": "913396",
      "postDate": "07/03/2020 07:40:42",
      "content": "<p>Oh no. You actually read file by file. Sorry</p>",
      "rawMarkdown": "Oh no. You actually read file by file. Sorry",
      "votes": null
    },
    {
      "id": "913485",
      "postDate": "07/03/2020 08:35:56",
      "content": "<p>I think you have a mistake here:\n<code>def __getitem__(self, idx):\n        _, rec_fn, start = self.items[idx]\n        x = self.rec[start*SAMPLE_RATE:(start+5)*SAMPLE_RATE]\n        example = self.get_specs(x)\n        example = self.normalize(example)\n        imgs = example.reshape(-1, 3, 80, 212)\n        return imgs.astype(np.float32)</code></p>\n\n<p>This line <code>x = self.rec[start*SAMPLE_RATE:(start+5)*SAMPLE_RATE]</code></p>\n\n<p>Should be <code>x = self.rec[start*(SAMPLE_RATE*5):(start+1)*(SAMPLE_RATE*5)]</code></p>",
      "rawMarkdown": "I think you have a mistake here:\n`    def __getitem__(self, idx):\n        _, rec_fn, start = self.items[idx]\n        x = self.rec[start*SAMPLE_RATE:(start+5)*SAMPLE_RATE]\n        example = self.get_specs(x)\n        example = self.normalize(example)\n        imgs = example.reshape(-1, 3, 80, 212)\n        return imgs.astype(np.float32)`\n\nThis line `x = self.rec[start*SAMPLE_RATE:(start+5)*SAMPLE_RATE]`\n\nShould be `x = self.rec[start*(SAMPLE_RATE*5):(start+1)*(SAMPLE_RATE*5)]`",
      "votes": null
    },
    {
      "id": "913604",
      "postDate": "07/03/2020 09:56:36",
      "content": "<p>Thank you very much <a href=\"/snovik1975\">@snovik1975</a>, really appreciate that you are going through my submission and checking the code 🙏</p>\n\n<p>I think the line you highlighted might be correct. <code>start</code> is the offset in seconds from file start for the current example. Translating into frames, we want to index into the recording starting with frame that is <code>offset_in_seconds</code> * <code>SAMPLE_RATE</code> up to <code>offset_in_seconds + 5 seconds (duration of example)</code> * <code>SAMPLE_RATE</code>.</p>\n\n<p>It is also a little bit suspect the score is exactly <code>0.54</code>. The two most likely explanations seem to me that either the model is not being run on the audio files for some reason (and the submission from the check phase is submitted) or the model is not outputting any predictions - could be if the threshold is too high or the test data is very different from the data the model was trained on).</p>\n\n<p>Would be very helpful though to receive diagnostic information that could make this easier to figure out.</p>",
      "rawMarkdown": "Thank you very much @snovik1975, really appreciate that you are going through my submission and checking the code 🙏\n\nI think the line you highlighted might be correct. `start` is the offset in seconds from file start for the current example. Translating into frames, we want to index into the recording starting with frame that is `offset_in_seconds` * `SAMPLE_RATE` up to `offset_in_seconds + 5 seconds (duration of example)` * `SAMPLE_RATE`.\n\nIt is also a little bit suspect the score is exactly `0.54`. The two most likely explanations seem to me that either the model is not being run on the audio files for some reason (and the submission from the check phase is submitted) or the model is not outputting any predictions - could be if the threshold is too high or the test data is very different from the data the model was trained on).\n\nWould be very helpful though to receive diagnostic information that could make this easier to figure out.",
      "votes": null
    },
    {
      "id": "913707",
      "postDate": "07/03/2020 11:16:02",
      "content": "<p>Right, my mistake again, your start is not the index of the chunk but the starting second of the chunk. </p>\n\n<p>Well, I am using your inference code to integrate into my pipeline. That is why I am going through it in such details. Sorry for bothering with questions :) </p>\n\n<p>I am using a slightly different setup but looking at the problems with the actually missing test set, I am building in multiple quality checks and defaults, e.g. what if the test mp3 is on not properly decoded, or test data has a number of chunks which is different from the size of mp3, etc. I do not know whether it is relevant but that is what I could do in this situation.</p>\n\n<p>Apparently missing predictions default to 'nocall' in the scoring and if all predictions fail, you get 0.54. Failing here means that the row_id from the submission file does not match with rows from the actual test, e.g. getting the path variables wrong would mean 0.54 out of the box. Since we have values on the LB different from 0.54, it means that there is <em>some</em> scoring going on. I wonder how people got 0 on the LB, as that means that all predictions are wrong however all lines in the submission file are present.</p>\n\n<p>I would not play with thresholds as this is a blackbox and whenever you change the model, the optimal threshold changes and without a proper and careful validation, it is dangerous to assume any specific value and keep it constant.</p>",
      "rawMarkdown": "Right, my mistake again, your start is not the index of the chunk but the starting second of the chunk. \n\nWell, I am using your inference code to integrate into my pipeline. That is why I am going through it in such details. Sorry for bothering with questions :) \n\nI am using a slightly different setup but looking at the problems with the actually missing test set, I am building in multiple quality checks and defaults, e.g. what if the test mp3 is on not properly decoded, or test data has a number of chunks which is different from the size of mp3, etc. I do not know whether it is relevant but that is what I could do in this situation.\n\nApparently missing predictions default to 'nocall' in the scoring and if all predictions fail, you get 0.54. Failing here means that the row_id from the submission file does not match with rows from the actual test, e.g. getting the path variables wrong would mean 0.54 out of the box. Since we have values on the LB different from 0.54, it means that there is *some* scoring going on. I wonder how people got 0 on the LB, as that means that all predictions are wrong however all lines in the submission file are present.\n\nI would not play with thresholds as this is a blackbox and whenever you change the model, the optimal threshold changes and without a proper and careful validation, it is dangerous to assume any specific value and keep it constant.",
      "votes": null
    },
    {
      "id": "913725",
      "postDate": "07/03/2020 11:27:16",
      "content": "<p>Thank you for the reply <a href=\"/snovik1975\">@snovik1975</a> and best of luck with your inference pipeline! This is my <a href=\"https://www.kaggle.com/radek1/esp-starter-pack-v3-lme-trainable-front-57-cls?scriptVersionId=38003004\">latest and greatest</a> submission 🙂Still being evaluated, but I am doing some things slightly differently there (predicting on 5 sec segments for site_1 and site_2, feeding 50 seconds per example for site_3).</p>\n\n<p>Thanks also for the information what happens when one submits and has no valid predictions - that is very useful to know. Silent failures such as this can be very challenging, especially to people newer to kaggle kernels (I haven't used kaggle kernels a lot in the past) or to people new to programming in general.</p>",
      "rawMarkdown": "Thank you for the reply @snovik1975 and best of luck with your inference pipeline! This is my [latest and greatest](https://www.kaggle.com/radek1/esp-starter-pack-v3-lme-trainable-front-57-cls?scriptVersionId=38003004) submission 🙂Still being evaluated, but I am doing some things slightly differently there (predicting on 5 sec segments for site_1 and site_2, feeding 50 seconds per example for site_3).\n\nThanks also for the information what happens when one submits and has no valid predictions - that is very useful to know. Silent failures such as this can be very challenging, especially to people newer to kaggle kernels (I haven't used kaggle kernels a lot in the past) or to people new to programming in general.",
      "votes": null
    },
    {
      "id": "913750",
      "postDate": "07/03/2020 11:44:17",
      "content": "<p>Well, if I look at your submissions file, and assuming that it is the true output of your kernel (now, I do not know what other disclaimers to make with this hidden test setup) then it does not look like mp3 files are approximately 10 min long. At best they are 1 min long. Your file is only 72 lines and has only a couple of recordings/files for for each site.</p>\n\n<p>And this is my 1st kernel competition. I have never ever played with kernels before :)</p>",
      "rawMarkdown": "Well, if I look at your submissions file, and assuming that it is the true output of your kernel (now, I do not know what other disclaimers to make with this hidden test setup) then it does not look like mp3 files are approximately 10 min long. At best they are 1 min long. Your file is only 72 lines and has only a couple of recordings/files for for each site.\n\nAnd this is my 1st kernel competition. I have never ever played with kernels before :)",
      "votes": null
    },
    {
      "id": "914445",
      "postDate": "07/03/2020 22:04:25",
      "content": "<p>I figured out the 72 lines thing. Really, the competition design with the hidden test set is very user unfriendly. It cost me almost a day to understand the set up and logic. I really do not understand the organizers in this regard</p>",
      "rawMarkdown": "I figured out the 72 lines thing. Really, the competition design with the hidden test set is very user unfriendly. It cost me almost a day to understand the set up and logic. I really do not understand the organizers in this regard",
      "votes": null
    },
    {
      "id": "955641",
      "postDate": "08/02/2020 19:09:27",
      "content": "<p>I wasted so many submissions and GPU hours on this. The problem was memory error while processing site_3 data. So frustrating!</p>",
      "rawMarkdown": "I wasted so many submissions and GPU hours on this. The problem was memory error while processing site_3 data. So frustrating!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 913274,
      "author_name": "snovik1975",
      "author_url": "",
      "post_date": "07/03/2020 05:14:04",
      "content": "<p>I believe it would be worth to get at least any system errors and output of exceptions. Though I do not know how to get the latter without any possibility to leak critical info</p>",
      "votes": null,
      "replies": [
        {
          "id": 913378,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "07/03/2020 07:25:02",
          "content": "<p>The solution to my mind would be having the platform check two things:</p>\n\n<ul>\n<li>does the submission contain predictions for all the rows in <code>test.csv</code></li>\n<li>are all the predictions valid, whether they are not malformed, etc</li>\n</ul>\n\n<p>ML code is probably one of the hardest code to debug and having this minimal feedback, not introducing much (any?) potential for leakage could be very helpful. Just a blank exception: <code>You didn't provide predictions for all the rows in test.csv or some of the rows were malformed</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 913392,
          "author_name": "snovik1975",
          "author_url": "",
          "post_date": "07/03/2020 07:37:52",
          "content": "<p>Tbh, I am afraid your version 2 might fail on the out of memory exception and therefore the whole submission resets to 'nocall' and that is the reason why your different submissions all get 0.54. Why so? Because your code reads all files into memory and keeps them there. Each 5 seconds wav (output of reading an mp3) takes up ca. 1.5 mb in memory. So 150 files each ca. 10 min means 150 * 10 * 24 * 1.5 = ca. 54 gb of memory. Makes sense?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 913395,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "07/03/2020 07:40:25",
          "content": "<p>I don't read all the files into memory. I read them one by one and not store them. Also, it seems one gets an exception that is communicated to them on the submission page when they use too many resources and the kernel resets.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 913396,
          "author_name": "snovik1975",
          "author_url": "",
          "post_date": "07/03/2020 07:40:42",
          "content": "<p>Oh no. You actually read file by file. Sorry</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 913485,
          "author_name": "snovik1975",
          "author_url": "",
          "post_date": "07/03/2020 08:35:56",
          "content": "<p>I think you have a mistake here:\n<code>def __getitem__(self, idx):\n        _, rec_fn, start = self.items[idx]\n        x = self.rec[start*SAMPLE_RATE:(start+5)*SAMPLE_RATE]\n        example = self.get_specs(x)\n        example = self.normalize(example)\n        imgs = example.reshape(-1, 3, 80, 212)\n        return imgs.astype(np.float32)</code></p>\n\n<p>This line <code>x = self.rec[start*SAMPLE_RATE:(start+5)*SAMPLE_RATE]</code></p>\n\n<p>Should be <code>x = self.rec[start*(SAMPLE_RATE*5):(start+1)*(SAMPLE_RATE*5)]</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 913604,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "07/03/2020 09:56:36",
          "content": "<p>Thank you very much <a href=\"/snovik1975\">@snovik1975</a>, really appreciate that you are going through my submission and checking the code 🙏</p>\n\n<p>I think the line you highlighted might be correct. <code>start</code> is the offset in seconds from file start for the current example. Translating into frames, we want to index into the recording starting with frame that is <code>offset_in_seconds</code> * <code>SAMPLE_RATE</code> up to <code>offset_in_seconds + 5 seconds (duration of example)</code> * <code>SAMPLE_RATE</code>.</p>\n\n<p>It is also a little bit suspect the score is exactly <code>0.54</code>. The two most likely explanations seem to me that either the model is not being run on the audio files for some reason (and the submission from the check phase is submitted) or the model is not outputting any predictions - could be if the threshold is too high or the test data is very different from the data the model was trained on).</p>\n\n<p>Would be very helpful though to receive diagnostic information that could make this easier to figure out.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 913707,
          "author_name": "snovik1975",
          "author_url": "",
          "post_date": "07/03/2020 11:16:02",
          "content": "<p>Right, my mistake again, your start is not the index of the chunk but the starting second of the chunk. </p>\n\n<p>Well, I am using your inference code to integrate into my pipeline. That is why I am going through it in such details. Sorry for bothering with questions :) </p>\n\n<p>I am using a slightly different setup but looking at the problems with the actually missing test set, I am building in multiple quality checks and defaults, e.g. what if the test mp3 is on not properly decoded, or test data has a number of chunks which is different from the size of mp3, etc. I do not know whether it is relevant but that is what I could do in this situation.</p>\n\n<p>Apparently missing predictions default to 'nocall' in the scoring and if all predictions fail, you get 0.54. Failing here means that the row_id from the submission file does not match with rows from the actual test, e.g. getting the path variables wrong would mean 0.54 out of the box. Since we have values on the LB different from 0.54, it means that there is <em>some</em> scoring going on. I wonder how people got 0 on the LB, as that means that all predictions are wrong however all lines in the submission file are present.</p>\n\n<p>I would not play with thresholds as this is a blackbox and whenever you change the model, the optimal threshold changes and without a proper and careful validation, it is dangerous to assume any specific value and keep it constant.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 913725,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "07/03/2020 11:27:16",
          "content": "<p>Thank you for the reply <a href=\"/snovik1975\">@snovik1975</a> and best of luck with your inference pipeline! This is my <a href=\"https://www.kaggle.com/radek1/esp-starter-pack-v3-lme-trainable-front-57-cls?scriptVersionId=38003004\">latest and greatest</a> submission 🙂Still being evaluated, but I am doing some things slightly differently there (predicting on 5 sec segments for site_1 and site_2, feeding 50 seconds per example for site_3).</p>\n\n<p>Thanks also for the information what happens when one submits and has no valid predictions - that is very useful to know. Silent failures such as this can be very challenging, especially to people newer to kaggle kernels (I haven't used kaggle kernels a lot in the past) or to people new to programming in general.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 913750,
          "author_name": "snovik1975",
          "author_url": "",
          "post_date": "07/03/2020 11:44:17",
          "content": "<p>Well, if I look at your submissions file, and assuming that it is the true output of your kernel (now, I do not know what other disclaimers to make with this hidden test setup) then it does not look like mp3 files are approximately 10 min long. At best they are 1 min long. Your file is only 72 lines and has only a couple of recordings/files for for each site.</p>\n\n<p>And this is my 1st kernel competition. I have never ever played with kernels before :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 914445,
          "author_name": "snovik1975",
          "author_url": "",
          "post_date": "07/03/2020 22:04:25",
          "content": "<p>I figured out the 72 lines thing. Really, the competition design with the hidden test set is very user unfriendly. It cost me almost a day to understand the set up and logic. I really do not understand the organizers in this regard</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 955641,
      "author_name": "ramarlina",
      "author_url": "",
      "post_date": "08/02/2020 19:09:27",
      "content": "<p>I wasted so many submissions and GPU hours on this. The problem was memory error while processing site_3 data. So frustrating!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "912339": "My last four submissions, all being substantially different, all scored 0.54, which is what an all zero submission would score. \n\nI am using [the custom check phase](https://www.kaggle.com/c/birdsong-recognition/discussion/159993) but I imagine that this still does not remove all ways in which one can have a bug.\n\nHere is my problem - on the files from the check phase, I can see my models perform detection (`submission.csv` results vastly vary between the models). And yet I get exactly 0.54 on my submissions.\n\nIf I were not submitting results for all the rows in the test set, would I see an error? If I would be outputting classes that do not match those in the training set, would I see an error?\n\nIf not, would it be possible for such functionality to be implemented?\n\nAssuming this functionality is not in place already, the risk is that someone might be doing something useful from the perspective of building a model to address the challenge, and we may never learn about that as they might have a bug in their code, their submission might not work correctly and they might not ever learn about this fact.\n\nBTW, my apologies for raising this if this functionality is already implemented. Thank you very much for all your help and your thoughts on this!",
    "913274": "I believe it would be worth to get at least any system errors and output of exceptions. Though I do not know how to get the latter without any possibility to leak critical info",
    "913378": "The solution to my mind would be having the platform check two things:\n\n- does the submission contain predictions for all the rows in `test.csv`\n- are all the predictions valid, whether they are not malformed, etc\n\nML code is probably one of the hardest code to debug and having this minimal feedback, not introducing much (any?) potential for leakage could be very helpful. Just a blank exception: `You didn't provide predictions for all the rows in test.csv or some of the rows were malformed`.",
    "913392": "Tbh, I am afraid your version 2 might fail on the out of memory exception and therefore the whole submission resets to 'nocall' and that is the reason why your different submissions all get 0.54. Why so? Because your code reads all files into memory and keeps them there. Each 5 seconds wav (output of reading an mp3) takes up ca. 1.5 mb in memory. So 150 files each ca. 10 min means 150 * 10 * 24 * 1.5 = ca. 54 gb of memory. Makes sense?",
    "913395": "I don't read all the files into memory. I read them one by one and not store them. Also, it seems one gets an exception that is communicated to them on the submission page when they use too many resources and the kernel resets.",
    "913396": "Oh no. You actually read file by file. Sorry",
    "913485": "I think you have a mistake here:\n`    def __getitem__(self, idx):\n        _, rec_fn, start = self.items[idx]\n        x = self.rec[start*SAMPLE_RATE:(start+5)*SAMPLE_RATE]\n        example = self.get_specs(x)\n        example = self.normalize(example)\n        imgs = example.reshape(-1, 3, 80, 212)\n        return imgs.astype(np.float32)`\n\nThis line `x = self.rec[start*SAMPLE_RATE:(start+5)*SAMPLE_RATE]`\n\nShould be `x = self.rec[start*(SAMPLE_RATE*5):(start+1)*(SAMPLE_RATE*5)]`",
    "913604": "Thank you very much @snovik1975, really appreciate that you are going through my submission and checking the code 🙏\n\nI think the line you highlighted might be correct. `start` is the offset in seconds from file start for the current example. Translating into frames, we want to index into the recording starting with frame that is `offset_in_seconds` * `SAMPLE_RATE` up to `offset_in_seconds + 5 seconds (duration of example)` * `SAMPLE_RATE`.\n\nIt is also a little bit suspect the score is exactly `0.54`. The two most likely explanations seem to me that either the model is not being run on the audio files for some reason (and the submission from the check phase is submitted) or the model is not outputting any predictions - could be if the threshold is too high or the test data is very different from the data the model was trained on).\n\nWould be very helpful though to receive diagnostic information that could make this easier to figure out.",
    "913707": "Right, my mistake again, your start is not the index of the chunk but the starting second of the chunk. \n\nWell, I am using your inference code to integrate into my pipeline. That is why I am going through it in such details. Sorry for bothering with questions :) \n\nI am using a slightly different setup but looking at the problems with the actually missing test set, I am building in multiple quality checks and defaults, e.g. what if the test mp3 is on not properly decoded, or test data has a number of chunks which is different from the size of mp3, etc. I do not know whether it is relevant but that is what I could do in this situation.\n\nApparently missing predictions default to 'nocall' in the scoring and if all predictions fail, you get 0.54. Failing here means that the row_id from the submission file does not match with rows from the actual test, e.g. getting the path variables wrong would mean 0.54 out of the box. Since we have values on the LB different from 0.54, it means that there is *some* scoring going on. I wonder how people got 0 on the LB, as that means that all predictions are wrong however all lines in the submission file are present.\n\nI would not play with thresholds as this is a blackbox and whenever you change the model, the optimal threshold changes and without a proper and careful validation, it is dangerous to assume any specific value and keep it constant.",
    "913725": "Thank you for the reply @snovik1975 and best of luck with your inference pipeline! This is my [latest and greatest](https://www.kaggle.com/radek1/esp-starter-pack-v3-lme-trainable-front-57-cls?scriptVersionId=38003004) submission 🙂Still being evaluated, but I am doing some things slightly differently there (predicting on 5 sec segments for site_1 and site_2, feeding 50 seconds per example for site_3).\n\nThanks also for the information what happens when one submits and has no valid predictions - that is very useful to know. Silent failures such as this can be very challenging, especially to people newer to kaggle kernels (I haven't used kaggle kernels a lot in the past) or to people new to programming in general.",
    "913750": "Well, if I look at your submissions file, and assuming that it is the true output of your kernel (now, I do not know what other disclaimers to make with this hidden test setup) then it does not look like mp3 files are approximately 10 min long. At best they are 1 min long. Your file is only 72 lines and has only a couple of recordings/files for for each site.\n\nAnd this is my 1st kernel competition. I have never ever played with kernels before :)",
    "914445": "I figured out the 72 lines thing. Really, the competition design with the hidden test set is very user unfriendly. It cost me almost a day to understand the set up and logic. I really do not understand the organizers in this regard",
    "955641": "I wasted so many submissions and GPU hours on this. The problem was memory error while processing site_3 data. So frustrating!"
  },
  "source": "meta"
}