{
  "id": 122841,
  "title": "Can we get a hint on submission errors?",
  "url": "/competitions/deepfake-detection-challenge/discussion/122841",
  "author_name": "",
  "post_date": "2019-12-23T06:49:42.337065200Z",
  "votes": 34,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I've got a fairly robust workflow I've set up with a bunch of safeguards in place and I still seem to be randomly having submission errors. I had one submission where the only thing I altered was a model file. One submission worked and the other did not for some reason. Both should've come in around 4 hours on the public test set of 4k videos. </p>\n\n<p>I'll list some of the pitfalls I've seen some people falling into in hopes of also figuring out what is possibly going wrong with my submissions sometimes. </p>\n\n<ul>\n<li>Submission succeeds on the public validation but not on public test because public test is 10x the amount of data of the public validation. Have to make sure you're able to do 4000 videos in under 9 hours. </li>\n<li>Have to be able to handle variable length and resolutions of videos. Some people have hard-coded 300 frames into their workflow, but there are some cases where it will only be 299 or 298 frames. Others have wrapped the iterating through frames in a while loop. OpenCV very quietly will return <code>None</code> instead of a blank frame or anything like that. You need to make sure you are checking if OpenCV returns True for the frame before you try to do some preprocessing and prediction on it. Some videos will be 1920x1080 and others will be 1080x1920, Have to make sure your method is able to handle both scenarios. </li>\n<li>Make sure your predictions are bounded between 0-1. This is a given if you have a sigmoid activation function, but still good to check your assumptions. </li>\n<li>Wrap each prediction in a try, except loop block. The way I have organized this is instead of appending each prediction to a list I am assigning each prediction to a dictionary with the key being the videos filename. Then when I assign the predictions I just map them to the sample submissions filenames. That way even if by some weird occurrence a prediction fails only one prediction will fail rather than the whole set of predictions not being properly assigned to the dataframe. </li>\n</ul>\n\n<p>Not sure how with these various safeguards I am still getting <code>Submission Scoring Error</code> with my most recent iteration. It would be nice if rather than allowing us to see the precise errors maybe the organizers could at least release a few of the most common error types. Based on the error I have received I assume my submission file was created but something was improperly formed with it. Outside of the 0-1 range is the only thing I can imagine because my setup should default to 0.5 and have all values entered. </p>",
  "messages": [
    {
      "id": "701152",
      "postDate": "12/23/2019 06:49:42",
      "content": "<p>I've got a fairly robust workflow I've set up with a bunch of safeguards in place and I still seem to be randomly having submission errors. I had one submission where the only thing I altered was a model file. One submission worked and the other did not for some reason. Both should've come in around 4 hours on the public test set of 4k videos. </p>\n\n<p>I'll list some of the pitfalls I've seen some people falling into in hopes of also figuring out what is possibly going wrong with my submissions sometimes. </p>\n\n<ul>\n<li>Submission succeeds on the public validation but not on public test because public test is 10x the amount of data of the public validation. Have to make sure you're able to do 4000 videos in under 9 hours. </li>\n<li>Have to be able to handle variable length and resolutions of videos. Some people have hard-coded 300 frames into their workflow, but there are some cases where it will only be 299 or 298 frames. Others have wrapped the iterating through frames in a while loop. OpenCV very quietly will return <code>None</code> instead of a blank frame or anything like that. You need to make sure you are checking if OpenCV returns True for the frame before you try to do some preprocessing and prediction on it. Some videos will be 1920x1080 and others will be 1080x1920, Have to make sure your method is able to handle both scenarios. </li>\n<li>Make sure your predictions are bounded between 0-1. This is a given if you have a sigmoid activation function, but still good to check your assumptions. </li>\n<li>Wrap each prediction in a try, except loop block. The way I have organized this is instead of appending each prediction to a list I am assigning each prediction to a dictionary with the key being the videos filename. Then when I assign the predictions I just map them to the sample submissions filenames. That way even if by some weird occurrence a prediction fails only one prediction will fail rather than the whole set of predictions not being properly assigned to the dataframe. </li>\n</ul>\n\n<p>Not sure how with these various safeguards I am still getting <code>Submission Scoring Error</code> with my most recent iteration. It would be nice if rather than allowing us to see the precise errors maybe the organizers could at least release a few of the most common error types. Based on the error I have received I assume my submission file was created but something was improperly formed with it. Outside of the 0-1 range is the only thing I can imagine because my setup should default to 0.5 and have all values entered. </p>",
      "rawMarkdown": "I've got a fairly robust workflow I've set up with a bunch of safeguards in place and I still seem to be randomly having submission errors. I had one submission where the only thing I altered was a model file. One submission worked and the other did not for some reason. Both should've come in around 4 hours on the public test set of 4k videos. \n\nI'll list some of the pitfalls I've seen some people falling into in hopes of also figuring out what is possibly going wrong with my submissions sometimes. \n\n- Submission succeeds on the public validation but not on public test because public test is 10x the amount of data of the public validation. Have to make sure you're able to do 4000 videos in under 9 hours. \n- Have to be able to handle variable length and resolutions of videos. Some people have hard-coded 300 frames into their workflow, but there are some cases where it will only be 299 or 298 frames. Others have wrapped the iterating through frames in a while loop. OpenCV very quietly will return `None` instead of a blank frame or anything like that. You need to make sure you are checking if OpenCV returns True for the frame before you try to do some preprocessing and prediction on it. Some videos will be 1920x1080 and others will be 1080x1920, Have to make sure your method is able to handle both scenarios. \n- Make sure your predictions are bounded between 0-1. This is a given if you have a sigmoid activation function, but still good to check your assumptions. \n- Wrap each prediction in a try, except loop block. The way I have organized this is instead of appending each prediction to a list I am assigning each prediction to a dictionary with the key being the videos filename. Then when I assign the predictions I just map them to the sample submissions filenames. That way even if by some weird occurrence a prediction fails only one prediction will fail rather than the whole set of predictions not being properly assigned to the dataframe. \n\nNot sure how with these various safeguards I am still getting `Submission Scoring Error` with my most recent iteration. It would be nice if rather than allowing us to see the precise errors maybe the organizers could at least release a few of the most common error types. Based on the error I have received I assume my submission file was created but something was improperly formed with it. Outside of the 0-1 range is the only thing I can imagine because my setup should default to 0.5 and have all values entered.",
      "votes": null
    },
    {
      "id": "701757",
      "postDate": "12/23/2019 21:56:37",
      "content": "<p>+1</p>",
      "rawMarkdown": "1",
      "votes": null
    },
    {
      "id": "701781",
      "postDate": "12/23/2019 22:48:49",
      "content": "<p>Discovered my issue. With how I was mapping it to the sample submission it would overwrite my initial 0.5 with nans so I run a .fillna(.5) after the mapping and the submission worked after that. Would be nice if I had some visibility into why my prediction was possibly failing in the first place instead of needing to just accept blank predictions and fill them. </p>",
      "rawMarkdown": "Discovered my issue. With how I was mapping it to the sample submission it would overwrite my initial 0.5 with nans so I run a .fillna(.5) after the mapping and the submission worked after that. Would be nice if I had some visibility into why my prediction was possibly failing in the first place instead of needing to just accept blank predictions and fill them.",
      "votes": null
    },
    {
      "id": "703037",
      "postDate": "12/25/2019 13:51:27",
      "content": "<p>Good advice. I'm still trying to make a valid submission.\nMy last version's status is 'Too Many Notebook Output Files'. I guess I should clear the working folder before submission.</p>",
      "rawMarkdown": "Good advice. I'm still trying to make a valid submission.\nMy last version's status is 'Too Many Notebook Output Files'. I guess I should clear the working folder before submission.",
      "votes": null
    },
    {
      "id": "703377",
      "postDate": "12/26/2019 04:21:51",
      "content": "<p>Hi, I have a problem: can we use some face detector( trained by dataset A), and need I upload dataset A in 1GB limit? Thanks</p>",
      "rawMarkdown": "Hi, I have a problem: can we use some face detector( trained by dataset A), and need I upload dataset A in 1GB limit? Thanks",
      "votes": null
    },
    {
      "id": "703691",
      "postDate": "12/26/2019 14:02:27",
      "content": "<p>I think the rules need us to use publicly released pretrained models, otherwise, the dataset A in your question is limited to 1GB.</p>",
      "rawMarkdown": "I think the rules need us to use publicly released pretrained models, otherwise, the dataset A in your question is limited to 1GB.",
      "votes": null
    },
    {
      "id": "704144",
      "postDate": "12/27/2019 05:10:20",
      "content": "<p>Do we need to map the label to the existing file names in sample_submission file? Those file names may not exist in public test videos right? Instead can't we build new set of filenames derived directly from public test and save in output? Is this approach correct?\nAlso should the label be rounded off to 1 decimal?\nPlease advice as I am getting submission score errors and no clues about it.</p>",
      "rawMarkdown": "Do we need to map the label to the existing file names in sample_submission file? Those file names may not exist in public test videos right? Instead can't we build new set of filenames derived directly from public test and save in output? Is this approach correct?\nAlso should the label be rounded off to 1 decimal?\nPlease advice as I am getting submission score errors and no clues about it.",
      "votes": null
    },
    {
      "id": "706811",
      "postDate": "12/30/2019 21:20:49",
      "content": "<p>It would be really great if there was a kaggle specific way of returning an error message when submissions fail. Something like this: </p>\n\n<p>try: \n   submit_df.to_csv('submission.csv', index=False)\nexcept:\n   kaggle_error_string = 'Writing submission file failed\"</p>\n\n<p>Then, if they could pass that error (kaggle_error_string) to the results page, it would be Fantastic ! Even if the length of the string was limited to a few characters it would be very helpful. </p>",
      "rawMarkdown": "It would be really great if there was a kaggle specific way of returning an error message when submissions fail. Something like this: \n\ntry: \n   submit_df.to_csv('submission.csv', index=False)\nexcept:\n   kaggle_error_string = 'Writing submission file failed\"\n\nThen, if they could pass that error (kaggle_error_string) to the results page, it would be Fantastic ! Even if the length of the string was limited to a few characters it would be very helpful.",
      "votes": null
    },
    {
      "id": "706813",
      "postDate": "12/30/2019 21:36:25",
      "content": "<p>But allowing kernels to return arbitrary error messages also allows them to return important information about the test set that Kaggle may want to keep private.</p>",
      "rawMarkdown": "But allowing kernels to return arbitrary error messages also allows them to return important information about the test set that Kaggle may want to keep private.",
      "votes": null
    },
    {
      "id": "706814",
      "postDate": "12/30/2019 21:40:26",
      "content": "<p>Every additional bit of information that they pass is a bit of information for LB probing strategies. The public LB in this competition is only 4000 bits, so they can't give away too many additional bits per submission.</p>",
      "rawMarkdown": "Every additional bit of information that they pass is a bit of information for LB probing strategies. The public LB in this competition is only 4000 bits, so they can't give away too many additional bits per submission.",
      "votes": null
    },
    {
      "id": "706817",
      "postDate": "12/30/2019 21:45:09",
      "content": "<p>One way to help researchers with submission errors and not give away any additional information per submission can be displaying aggregate error types across all participants, with some randomized delay. </p>",
      "rawMarkdown": "One way to help researchers with submission errors and not give away any additional information per submission can be displaying aggregate error types across all participants, with some randomized delay.",
      "votes": null
    },
    {
      "id": "706885",
      "postDate": "12/31/2019 00:05:42",
      "content": "<p>Yeah I think that is the way to go. At least be able to tell us it's a divide by zero error from python that is in 10% of kernels rather than just saying submission error. And since it's aggregated and only showing the top 3 maybe then someone cant use it to probe because their error won't be common enough to determine the error. Could have it with randomized delay or also just have very small precision like only down to the 10% instead of 10.37%</p>",
      "rawMarkdown": "Yeah I think that is the way to go. At least be able to tell us it's a divide by zero error from python that is in 10% of kernels rather than just saying submission error. And since it's aggregated and only showing the top 3 maybe then someone cant use it to probe because their error won't be common enough to determine the error. Could have it with randomized delay or also just have very small precision like only down to the 10% instead of 10.37%",
      "votes": null
    },
    {
      "id": "707370",
      "postDate": "12/31/2019 17:41:54",
      "content": "<p><a href=\"/ryches\">@ryches</a> We know that the <code>sample_submission.csv</code> file that we can see (public to us) contains 400 rows because the visible test set has 400 videos. When we submit the kernel for scoring, will the <code>sample_submission.csv</code> file now have 4000 rows as we know that there are approx 4000 hidden test videos or is this file the same as before (with only 400 rows)?\nI want to just map the output probability with the corresponding \"filename\" from the <code>sample_submission.csv</code> file instead of creating a new dataframe.\nI am also getting the <code>Submission Scoring Error</code> message for the past two submissions when I am creating a new dataframe instead of mapping according to the filename. So, wanted to make sure of it.</p>",
      "rawMarkdown": "ryches We know that the `sample_submission.csv` file that we can see (public to us) contains 400 rows because the visible test set has 400 videos. When we submit the kernel for scoring, will the `sample_submission.csv` file now have 4000 rows as we know that there are approx 4000 hidden test videos or is this file the same as before (with only 400 rows)?\nI want to just map the output probability with the corresponding \"filename\" from the `sample_submission.csv` file instead of creating a new dataframe.\nI am also getting the `Submission Scoring Error` message for the past two submissions when I am creating a new dataframe instead of mapping according to the filename. So, wanted to make sure of it.",
      "votes": null
    },
    {
      "id": "711006",
      "postDate": "01/05/2020 14:48:56",
      "content": "<p>Hi <a href=\"/ryches\">@ryches</a> ,\n    I also get a <code>Submission Scoring Error</code>. But it is not in the case you listed above. It looks very strange because I give 0.5 score for each test video and ensure every test sample path&amp;result is saved in <code>submission.csv</code>\nHere is my kernel: <a href=\"https://www.kaggle.com/xiaofengmao/kernel553c006564\">https://www.kaggle.com/xiaofengmao/kernel553c006564</a></p>\n\n<p>Its really hard to debug for me. I just want to extract audios of test videos but failed. Is it the problem of <code>ffmpeg</code>?</p>",
      "rawMarkdown": "Hi @ryches ,\n    I also get a `Submission Scoring Error`. But it is not in the case you listed above. It looks very strange because I give 0.5 score for each test video and ensure every test sample path&amp;result is saved in `submission.csv`\nHere is my kernel: https://www.kaggle.com/xiaofengmao/kernel553c006564\n\nIts really hard to debug for me. I just want to extract audios of test videos but failed. Is it the problem of `ffmpeg`?",
      "votes": null
    },
    {
      "id": "711390",
      "postDate": "01/06/2020 02:47:44",
      "content": "<p>Sorry to bother. It is caused by too much files in output folder. <code>Submission Scoring Error</code> happens when exceeding the limitation on number of output files.</p>",
      "rawMarkdown": "Sorry to bother. It is caused by too much files in output folder. `Submission Scoring Error` happens when exceeding the limitation on number of output files.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 701757,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "12/23/2019 21:56:37",
      "content": "<p>+1</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 701781,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "12/23/2019 22:48:49",
      "content": "<p>Discovered my issue. With how I was mapping it to the sample submission it would overwrite my initial 0.5 with nans so I run a .fillna(.5) after the mapping and the submission worked after that. Would be nice if I had some visibility into why my prediction was possibly failing in the first place instead of needing to just accept blank predictions and fill them. </p>",
      "votes": null,
      "replies": [
        {
          "id": 703377,
          "author_name": "bluebluedd",
          "author_url": "",
          "post_date": "12/26/2019 04:21:51",
          "content": "<p>Hi, I have a problem: can we use some face detector( trained by dataset A), and need I upload dataset A in 1GB limit? Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 703691,
          "author_name": "fionalxd",
          "author_url": "",
          "post_date": "12/26/2019 14:02:27",
          "content": "<p>I think the rules need us to use publicly released pretrained models, otherwise, the dataset A in your question is limited to 1GB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707370,
          "author_name": "manideep2510",
          "author_url": "",
          "post_date": "12/31/2019 17:41:54",
          "content": "<p><a href=\"/ryches\">@ryches</a> We know that the <code>sample_submission.csv</code> file that we can see (public to us) contains 400 rows because the visible test set has 400 videos. When we submit the kernel for scoring, will the <code>sample_submission.csv</code> file now have 4000 rows as we know that there are approx 4000 hidden test videos or is this file the same as before (with only 400 rows)?\nI want to just map the output probability with the corresponding \"filename\" from the <code>sample_submission.csv</code> file instead of creating a new dataframe.\nI am also getting the <code>Submission Scoring Error</code> message for the past two submissions when I am creating a new dataframe instead of mapping according to the filename. So, wanted to make sure of it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 703037,
      "author_name": "feifeizaici",
      "author_url": "",
      "post_date": "12/25/2019 13:51:27",
      "content": "<p>Good advice. I'm still trying to make a valid submission.\nMy last version's status is 'Too Many Notebook Output Files'. I guess I should clear the working folder before submission.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 704144,
      "author_name": "rajasuman09",
      "author_url": "",
      "post_date": "12/27/2019 05:10:20",
      "content": "<p>Do we need to map the label to the existing file names in sample_submission file? Those file names may not exist in public test videos right? Instead can't we build new set of filenames derived directly from public test and save in output? Is this approach correct?\nAlso should the label be rounded off to 1 decimal?\nPlease advice as I am getting submission score errors and no clues about it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 706811,
      "author_name": "jonmacpherson",
      "author_url": "",
      "post_date": "12/30/2019 21:20:49",
      "content": "<p>It would be really great if there was a kaggle specific way of returning an error message when submissions fail. Something like this: </p>\n\n<p>try: \n   submit_df.to_csv('submission.csv', index=False)\nexcept:\n   kaggle_error_string = 'Writing submission file failed\"</p>\n\n<p>Then, if they could pass that error (kaggle_error_string) to the results page, it would be Fantastic ! Even if the length of the string was limited to a few characters it would be very helpful. </p>",
      "votes": null,
      "replies": [
        {
          "id": 706813,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "12/30/2019 21:36:25",
          "content": "<p>But allowing kernels to return arbitrary error messages also allows them to return important information about the test set that Kaggle may want to keep private.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 706814,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "12/30/2019 21:40:26",
          "content": "<p>Every additional bit of information that they pass is a bit of information for LB probing strategies. The public LB in this competition is only 4000 bits, so they can't give away too many additional bits per submission.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 706817,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "12/30/2019 21:45:09",
      "content": "<p>One way to help researchers with submission errors and not give away any additional information per submission can be displaying aggregate error types across all participants, with some randomized delay. </p>",
      "votes": null,
      "replies": [
        {
          "id": 706885,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "12/31/2019 00:05:42",
          "content": "<p>Yeah I think that is the way to go. At least be able to tell us it's a divide by zero error from python that is in 10% of kernels rather than just saying submission error. And since it's aggregated and only showing the top 3 maybe then someone cant use it to probe because their error won't be common enough to determine the error. Could have it with randomized delay or also just have very small precision like only down to the 10% instead of 10.37%</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 711006,
      "author_name": "xiaofengmao",
      "author_url": "",
      "post_date": "01/05/2020 14:48:56",
      "content": "<p>Hi <a href=\"/ryches\">@ryches</a> ,\n    I also get a <code>Submission Scoring Error</code>. But it is not in the case you listed above. It looks very strange because I give 0.5 score for each test video and ensure every test sample path&amp;result is saved in <code>submission.csv</code>\nHere is my kernel: <a href=\"https://www.kaggle.com/xiaofengmao/kernel553c006564\">https://www.kaggle.com/xiaofengmao/kernel553c006564</a></p>\n\n<p>Its really hard to debug for me. I just want to extract audios of test videos but failed. Is it the problem of <code>ffmpeg</code>?</p>",
      "votes": null,
      "replies": [
        {
          "id": 711390,
          "author_name": "xiaofengmao",
          "author_url": "",
          "post_date": "01/06/2020 02:47:44",
          "content": "<p>Sorry to bother. It is caused by too much files in output folder. <code>Submission Scoring Error</code> happens when exceeding the limitation on number of output files.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "701152": "I've got a fairly robust workflow I've set up with a bunch of safeguards in place and I still seem to be randomly having submission errors. I had one submission where the only thing I altered was a model file. One submission worked and the other did not for some reason. Both should've come in around 4 hours on the public test set of 4k videos. \n\nI'll list some of the pitfalls I've seen some people falling into in hopes of also figuring out what is possibly going wrong with my submissions sometimes. \n\n- Submission succeeds on the public validation but not on public test because public test is 10x the amount of data of the public validation. Have to make sure you're able to do 4000 videos in under 9 hours. \n- Have to be able to handle variable length and resolutions of videos. Some people have hard-coded 300 frames into their workflow, but there are some cases where it will only be 299 or 298 frames. Others have wrapped the iterating through frames in a while loop. OpenCV very quietly will return `None` instead of a blank frame or anything like that. You need to make sure you are checking if OpenCV returns True for the frame before you try to do some preprocessing and prediction on it. Some videos will be 1920x1080 and others will be 1080x1920, Have to make sure your method is able to handle both scenarios. \n- Make sure your predictions are bounded between 0-1. This is a given if you have a sigmoid activation function, but still good to check your assumptions. \n- Wrap each prediction in a try, except loop block. The way I have organized this is instead of appending each prediction to a list I am assigning each prediction to a dictionary with the key being the videos filename. Then when I assign the predictions I just map them to the sample submissions filenames. That way even if by some weird occurrence a prediction fails only one prediction will fail rather than the whole set of predictions not being properly assigned to the dataframe. \n\nNot sure how with these various safeguards I am still getting `Submission Scoring Error` with my most recent iteration. It would be nice if rather than allowing us to see the precise errors maybe the organizers could at least release a few of the most common error types. Based on the error I have received I assume my submission file was created but something was improperly formed with it. Outside of the 0-1 range is the only thing I can imagine because my setup should default to 0.5 and have all values entered.",
    "701757": "1",
    "701781": "Discovered my issue. With how I was mapping it to the sample submission it would overwrite my initial 0.5 with nans so I run a .fillna(.5) after the mapping and the submission worked after that. Would be nice if I had some visibility into why my prediction was possibly failing in the first place instead of needing to just accept blank predictions and fill them.",
    "703037": "Good advice. I'm still trying to make a valid submission.\nMy last version's status is 'Too Many Notebook Output Files'. I guess I should clear the working folder before submission.",
    "703377": "Hi, I have a problem: can we use some face detector( trained by dataset A), and need I upload dataset A in 1GB limit? Thanks",
    "703691": "I think the rules need us to use publicly released pretrained models, otherwise, the dataset A in your question is limited to 1GB.",
    "704144": "Do we need to map the label to the existing file names in sample_submission file? Those file names may not exist in public test videos right? Instead can't we build new set of filenames derived directly from public test and save in output? Is this approach correct?\nAlso should the label be rounded off to 1 decimal?\nPlease advice as I am getting submission score errors and no clues about it.",
    "706811": "It would be really great if there was a kaggle specific way of returning an error message when submissions fail. Something like this: \n\ntry: \n   submit_df.to_csv('submission.csv', index=False)\nexcept:\n   kaggle_error_string = 'Writing submission file failed\"\n\nThen, if they could pass that error (kaggle_error_string) to the results page, it would be Fantastic ! Even if the length of the string was limited to a few characters it would be very helpful.",
    "706813": "But allowing kernels to return arbitrary error messages also allows them to return important information about the test set that Kaggle may want to keep private.",
    "706814": "Every additional bit of information that they pass is a bit of information for LB probing strategies. The public LB in this competition is only 4000 bits, so they can't give away too many additional bits per submission.",
    "706817": "One way to help researchers with submission errors and not give away any additional information per submission can be displaying aggregate error types across all participants, with some randomized delay.",
    "706885": "Yeah I think that is the way to go. At least be able to tell us it's a divide by zero error from python that is in 10% of kernels rather than just saying submission error. And since it's aggregated and only showing the top 3 maybe then someone cant use it to probe because their error won't be common enough to determine the error. Could have it with randomized delay or also just have very small precision like only down to the 10% instead of 10.37%",
    "707370": "ryches We know that the `sample_submission.csv` file that we can see (public to us) contains 400 rows because the visible test set has 400 videos. When we submit the kernel for scoring, will the `sample_submission.csv` file now have 4000 rows as we know that there are approx 4000 hidden test videos or is this file the same as before (with only 400 rows)?\nI want to just map the output probability with the corresponding \"filename\" from the `sample_submission.csv` file instead of creating a new dataframe.\nI am also getting the `Submission Scoring Error` message for the past two submissions when I am creating a new dataframe instead of mapping according to the filename. So, wanted to make sure of it.",
    "711006": "Hi @ryches ,\n    I also get a `Submission Scoring Error`. But it is not in the case you listed above. It looks very strange because I give 0.5 score for each test video and ensure every test sample path&amp;result is saved in `submission.csv`\nHere is my kernel: https://www.kaggle.com/xiaofengmao/kernel553c006564\n\nIts really hard to debug for me. I just want to extract audios of test videos but failed. Is it the problem of `ffmpeg`?",
    "711390": "Sorry to bother. It is caused by too much files in output folder. `Submission Scoring Error` happens when exceeding the limitation on number of output files."
  },
  "source": "meta"
}