{
  "id": 129173,
  "title": "Can you please increase the number of submissions per day before the end?",
  "url": "/competitions/deepfake-detection-challenge/discussion/129173",
  "author_name": "",
  "post_date": "2020-02-05T23:56:59.890300700Z",
  "votes": 3,
  "comment_count": 17,
  "views": 0,
  "content": "<p>@juliaelliott I understand that the low number of submissions was to prevent LB probing. However, as we will near the end of the competition this will become less relevant.\nWould the hosts consider increasing the limit, even to 3 in the last month? Also consider that after team merger deadline there will be much less teams and bigger teams might need more then 2 submissions to check possible options.\nI find that sometimes by making a really stupid mistake (like today, using p instead of 1-p) I am delayed by a full day.\nThanks!</p>",
  "messages": [
    {
      "id": "737933",
      "postDate": "02/05/2020 23:56:59",
      "content": "<p>@juliaelliott I understand that the low number of submissions was to prevent LB probing. However, as we will near the end of the competition this will become less relevant.\nWould the hosts consider increasing the limit, even to 3 in the last month? Also consider that after team merger deadline there will be much less teams and bigger teams might need more then 2 submissions to check possible options.\nI find that sometimes by making a really stupid mistake (like today, using p instead of 1-p) I am delayed by a full day.\nThanks!</p>",
      "rawMarkdown": "juliaelliott I understand that the low number of submissions was to prevent LB probing. However, as we will near the end of the competition this will become less relevant.\nWould the hosts consider increasing the limit, even to 3 in the last month? Also consider that after team merger deadline there will be much less teams and bigger teams might need more then 2 submissions to check possible options.\nI find that sometimes by making a really stupid mistake (like today, using p instead of 1-p) I am delayed by a full day.\nThanks!",
      "votes": null
    },
    {
      "id": "737936",
      "postDate": "02/06/2020 00:04:33",
      "content": "<p>I agree with you!</p>",
      "rawMarkdown": "I agree with you!",
      "votes": null
    },
    {
      "id": "737945",
      "postDate": "02/06/2020 00:41:36",
      "content": "<p>Two submissions a day is just enough, I think, no real reason to increase. The problem that you described should be solved with better testing before submitting. Add metrics to be printed out on commits inside the kernel (on 400 public validation videos), I find it very useful. We have the labels for those videos, it is a subset of the full train dataset. Add <a href=\"https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc\">this dataset</a> to your kernel, get the labels from it, and print out your favorite metrics, that way you know you submit something meaningful. </p>\n\n<p>The models are heavy, taking up to 9h to run, much more to train, one shouldn't really produce more than 2 meaningful iterations per day. </p>\n\n<p>The real issue is that we can't use the submissions that we already have effectively. Just today I got an error \"Notebook Exceeded Allowed Compute\", for the kernel that run successfully before with very small change at the end, 100% not related to running time, memory or disk usage. And corrupted files are soooo annoying. Anyone who participates competitively will try to squeeze maximum from it, trying to probe that may be the first half of the videos are OK and can be used for predictions and so on and so forth. That process alone will burn through submissions like Australian fires.</p>\n\n<p>We need more clarity with the errors, and we need clean public and private test datasets with no unexpected video formats. Before that happens, Kaggle keeps stealing expensive time from hundreds of the participants! </p>",
      "rawMarkdown": "Two submissions a day is just enough, I think, no real reason to increase. The problem that you described should be solved with better testing before submitting. Add metrics to be printed out on commits inside the kernel (on 400 public validation videos), I find it very useful. We have the labels for those videos, it is a subset of the full train dataset. Add [this dataset](https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc) to your kernel, get the labels from it, and print out your favorite metrics, that way you know you submit something meaningful. \n\nThe models are heavy, taking up to 9h to run, much more to train, one shouldn't really produce more than 2 meaningful iterations per day. \n\nThe real issue is that we can't use the submissions that we already have effectively. Just today I got an error \"Notebook Exceeded Allowed Compute\", for the kernel that run successfully before with very small change at the end, 100% not related to running time, memory or disk usage. And corrupted files are soooo annoying. Anyone who participates competitively will try to squeeze maximum from it, trying to probe that may be the first half of the videos are OK and can be used for predictions and so on and so forth. That process alone will burn through submissions like Australian fires.\n\nWe need more clarity with the errors, and we need clean public and private test datasets with no unexpected video formats. Before that happens, Kaggle keeps stealing expensive time from hundreds of the participants!",
      "votes": null
    },
    {
      "id": "737963",
      "postDate": "02/06/2020 01:30:48",
      "content": "<p>I check my submissions fanatically and I have been lucky so far - made one error in submission a while back (but have several indecipherable error messages). But the impossibility of debugging with the \"real\" videos makes this task difficult far beyond the stated purpose.\n  I second the OP's request for more entries, if only to compensate for the debugging difficulty. For one of my error messages, I took a few days to figure it out.</p>",
      "rawMarkdown": "I check my submissions fanatically and I have been lucky so far - made one error in submission a while back (but have several indecipherable error messages). But the impossibility of debugging with the \"real\" videos makes this task difficult far beyond the stated purpose.\n  I second the OP's request for more entries, if only to compensate for the debugging difficulty. For one of my error messages, I took a few days to figure it out.",
      "votes": null
    },
    {
      "id": "738040",
      "postDate": "02/06/2020 04:50:43",
      "content": "<p>I disagree that two is enough. Human errors happen even with best checking. And checking with the public validation set is... Debatable because of cross actors/training \"leak\". And yes, error handling /debugging is VERY difficult in this competition.\nThe submissions don't have to be that lengthy for checking on model improvement. You can check on 1000 entries and fill the rest with 0.5. This will run in 2h max. However, it doesn't help much if you have only 2 submissions! </p>",
      "rawMarkdown": "I disagree that two is enough. Human errors happen even with best checking. And checking with the public validation set is... Debatable because of cross actors/training \"leak\". And yes, error handling /debugging is VERY difficult in this competition.\nThe submissions don't have to be that lengthy for checking on model improvement. You can check on 1000 entries and fill the rest with 0.5. This will run in 2h max. However, it doesn't help much if you have only 2 submissions!",
      "votes": null
    },
    {
      "id": "738205",
      "postDate": "02/06/2020 09:05:51",
      "content": "<p>I propose to at least allow some \"non-countable\" error submissions. Like you can make 1-2 errors per day which won't count as real submissions.</p>",
      "rawMarkdown": "I propose to at least allow some \"non-countable\" error submissions. Like you can make 1-2 errors per day which won't count as real submissions.",
      "votes": null
    },
    {
      "id": "738213",
      "postDate": "02/06/2020 09:15:12",
      "content": "<p>Two submissions a day is enough IMO. This is a code competition, and one of the key aspects of this is for competitors to generate robust and scalable models and pipelines. That means capability not just in terms of accuracy, but also in terms of unforeseen situations &amp; error handling, and dealing with much larger unseen data. </p>\n\n<p>For me, the limitation on computation resource, the fact that you don't EXACTLY know what is wrong during submission process, and that the long scoring time is part and parcel of this competition, and it is part of the fun. Personally, I kind of take this as an engineering challenge in addition to an ML challenge. For instance, what is the best way to squeeze in as much compute with the given allowance, and making sure it doesn't spiral out of control in the long scoring process. </p>\n\n<p>After all, this is the exact nature of some of human's most challenging engineering tasks - spare a though for people who design Mars rovers -  years of work, billions of budget, months of journey, and success/failure will only unfold at the hours/minutes at touching down. I am quite grateful for the fact that I get a small taste of it in the safe and free environment of an online data competition :)</p>",
      "rawMarkdown": "Two submissions a day is enough IMO. This is a code competition, and one of the key aspects of this is for competitors to generate robust and scalable models and pipelines. That means capability not just in terms of accuracy, but also in terms of unforeseen situations &amp; error handling, and dealing with much larger unseen data. \n\nFor me, the limitation on computation resource, the fact that you don't EXACTLY know what is wrong during submission process, and that the long scoring time is part and parcel of this competition, and it is part of the fun. Personally, I kind of take this as an engineering challenge in addition to an ML challenge. For instance, what is the best way to squeeze in as much compute with the given allowance, and making sure it doesn't spiral out of control in the long scoring process. \n\nAfter all, this is the exact nature of some of human's most challenging engineering tasks - spare a though for people who design Mars rovers -  years of work, billions of budget, months of journey, and success/failure will only unfold at the hours/minutes at touching down. I am quite grateful for the fact that I get a small taste of it in the safe and free environment of an online data competition :)",
      "votes": null
    },
    {
      "id": "738337",
      "postDate": "02/06/2020 12:03:30",
      "content": "<p>I don't think that the low number of submissions is to prevent LB probing. I think it is to reduce the high demand for GPUs in Kaggle Notebooks.</p>",
      "rawMarkdown": "I don't think that the low number of submissions is to prevent LB probing. I think it is to reduce the high demand for GPUs in Kaggle Notebooks.",
      "votes": null
    },
    {
      "id": "738616",
      "postDate": "02/06/2020 19:25:50",
      "content": "<p>Even so, increasing it in the last month won't hurt too much. </p>",
      "rawMarkdown": "Even so, increasing it in the last month won't hurt too much.",
      "votes": null
    },
    {
      "id": "738655",
      "postDate": "02/06/2020 20:42:48",
      "content": "<p>I agree with you. BTW I have wrote a workaround script \\in python. I download the generated submission file, pass it to script then it prints loss, plots histogram and checks for range and length of submission file. It helps a lot. Many time error is catched when I see the loss. </p>",
      "rawMarkdown": "I agree with you. BTW I have wrote a workaround script \\in python. I download the generated submission file, pass it to script then it prints loss, plots histogram and checks for range and length of submission file. It helps a lot. Many time error is catched when I see the loss.",
      "votes": null
    },
    {
      "id": "738657",
      "postDate": "02/06/2020 20:43:53",
      "content": "<p><a href=\"/moshel\">@moshel</a> I agree with you.</p>",
      "rawMarkdown": "moshel I agree with you.",
      "votes": null
    },
    {
      "id": "738658",
      "postDate": "02/06/2020 20:44:43",
      "content": "<p>It works because the 400 test videos are from our training data. So we can have their correct labels. </p>",
      "rawMarkdown": "It works because the 400 test videos are from our training data. So we can have their correct labels.",
      "votes": null
    },
    {
      "id": "738688",
      "postDate": "02/06/2020 21:38:27",
      "content": "<p>unfortunately, unless you specifically removed them from the training, you will get a MUCH better result on things that were in the training set as the model \"memorize\" them. This is overfitting, basically. another problem that I have noticed is that my model actually learn to detect the actors real faces, so even unseen videos by same actors will get better results.</p>",
      "rawMarkdown": "unfortunately, unless you specifically removed them from the training, you will get a MUCH better result on things that were in the training set as the model \"memorize\" them. This is overfitting, basically. another problem that I have noticed is that my model actually learn to detect the actors real faces, so even unseen videos by same actors will get better results.",
      "votes": null
    },
    {
      "id": "738690",
      "postDate": "02/06/2020 21:44:06",
      "content": "<p>regarding your script, it is actually neater and more reliable to do the checks in the kernel. just add a line after getting the list of files from the test_videos directory\n<code>\nif len(file_list) == 400:\n    DEBUG_MODE=True\nelse:\n    DEBUG_MODE=False\n</code></p>\n\n<p>and at the end add a cell with test specific checks:</p>\n\n<p><code>\nif DEBUG_MODE:\n         do checks\n</code></p>",
      "rawMarkdown": "regarding your script, it is actually neater and more reliable to do the checks in the kernel. just add a line after getting the list of files from the test_videos directory\n```\nif len(file_list) == 400:\n    DEBUG_MODE=True\nelse:\n    DEBUG_MODE=False\n```\n\nand at the end add a cell with test specific checks:\n\n```\nif DEBUG_MODE:\n         do checks\n```",
      "votes": null
    },
    {
      "id": "738705",
      "postDate": "02/06/2020 22:07:56",
      "content": "<p>Good to know this way of debugging. But still, I prefer separate files for testing. Just personal preference. For the overfitting case, my tests are not to check if the model is good, it is for checking if submission is correct. In your case, it would've caught p and 1-p mistake.</p>",
      "rawMarkdown": "Good to know this way of debugging. But still, I prefer separate files for testing. Just personal preference. For the overfitting case, my tests are not to check if the model is good, it is for checking if submission is correct. In your case, it would've caught p and 1-p mistake.",
      "votes": null
    },
    {
      "id": "738709",
      "postDate": "02/06/2020 22:18:19",
      "content": "<p>Generally, true. However, in this case I also switched the labels (converting the fake/real to 0/1) so no, it didn't.... facepalm....</p>\n\n<p>What I was trying to say was... S%^t happens. 2 submissions limit is generally ok, but I still think we (and the hosts) will benefit from increasing it, at least in the last month.</p>",
      "rawMarkdown": "Generally, true. However, in this case I also switched the labels (converting the fake/real to 0/1) so no, it didn't.... facepalm....\n\nWhat I was trying to say was... S%^t happens. 2 submissions limit is generally ok, but I still think we (and the hosts) will benefit from increasing it, at least in the last month.",
      "votes": null
    },
    {
      "id": "738713",
      "postDate": "02/06/2020 22:26:12",
      "content": "<p>Yeah, I agree with you.</p>",
      "rawMarkdown": "Yeah, I agree with you.",
      "votes": null
    },
    {
      "id": "738755",
      "postDate": "02/07/2020 00:56:12",
      "content": "<p>Maybe Kaggle team could tell us why 2 submissions. I also would like to have 3...</p>",
      "rawMarkdown": "Maybe Kaggle team could tell us why 2 submissions. I also would like to have 3...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 737936,
      "author_name": "phunghieu",
      "author_url": "",
      "post_date": "02/06/2020 00:04:33",
      "content": "<p>I agree with you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 737945,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "02/06/2020 00:41:36",
      "content": "<p>Two submissions a day is just enough, I think, no real reason to increase. The problem that you described should be solved with better testing before submitting. Add metrics to be printed out on commits inside the kernel (on 400 public validation videos), I find it very useful. We have the labels for those videos, it is a subset of the full train dataset. Add <a href=\"https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc\">this dataset</a> to your kernel, get the labels from it, and print out your favorite metrics, that way you know you submit something meaningful. </p>\n\n<p>The models are heavy, taking up to 9h to run, much more to train, one shouldn't really produce more than 2 meaningful iterations per day. </p>\n\n<p>The real issue is that we can't use the submissions that we already have effectively. Just today I got an error \"Notebook Exceeded Allowed Compute\", for the kernel that run successfully before with very small change at the end, 100% not related to running time, memory or disk usage. And corrupted files are soooo annoying. Anyone who participates competitively will try to squeeze maximum from it, trying to probe that may be the first half of the videos are OK and can be used for predictions and so on and so forth. That process alone will burn through submissions like Australian fires.</p>\n\n<p>We need more clarity with the errors, and we need clean public and private test datasets with no unexpected video formats. Before that happens, Kaggle keeps stealing expensive time from hundreds of the participants! </p>",
      "votes": null,
      "replies": [
        {
          "id": 738040,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "02/06/2020 04:50:43",
          "content": "<p>I disagree that two is enough. Human errors happen even with best checking. And checking with the public validation set is... Debatable because of cross actors/training \"leak\". And yes, error handling /debugging is VERY difficult in this competition.\nThe submissions don't have to be that lengthy for checking on model improvement. You can check on 1000 entries and fill the rest with 0.5. This will run in 2h max. However, it doesn't help much if you have only 2 submissions! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 738657,
          "author_name": "ankitsainiankit",
          "author_url": "",
          "post_date": "02/06/2020 20:43:53",
          "content": "<p><a href=\"/moshel\">@moshel</a> I agree with you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 737963,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "02/06/2020 01:30:48",
      "content": "<p>I check my submissions fanatically and I have been lucky so far - made one error in submission a while back (but have several indecipherable error messages). But the impossibility of debugging with the \"real\" videos makes this task difficult far beyond the stated purpose.\n  I second the OP's request for more entries, if only to compensate for the debugging difficulty. For one of my error messages, I took a few days to figure it out.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 738205,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "02/06/2020 09:05:51",
      "content": "<p>I propose to at least allow some \"non-countable\" error submissions. Like you can make 1-2 errors per day which won't count as real submissions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 738213,
      "author_name": "yifanxie",
      "author_url": "",
      "post_date": "02/06/2020 09:15:12",
      "content": "<p>Two submissions a day is enough IMO. This is a code competition, and one of the key aspects of this is for competitors to generate robust and scalable models and pipelines. That means capability not just in terms of accuracy, but also in terms of unforeseen situations &amp; error handling, and dealing with much larger unseen data. </p>\n\n<p>For me, the limitation on computation resource, the fact that you don't EXACTLY know what is wrong during submission process, and that the long scoring time is part and parcel of this competition, and it is part of the fun. Personally, I kind of take this as an engineering challenge in addition to an ML challenge. For instance, what is the best way to squeeze in as much compute with the given allowance, and making sure it doesn't spiral out of control in the long scoring process. </p>\n\n<p>After all, this is the exact nature of some of human's most challenging engineering tasks - spare a though for people who design Mars rovers -  years of work, billions of budget, months of journey, and success/failure will only unfold at the hours/minutes at touching down. I am quite grateful for the fact that I get a small taste of it in the safe and free environment of an online data competition :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 738337,
      "author_name": "felipefonte99",
      "author_url": "",
      "post_date": "02/06/2020 12:03:30",
      "content": "<p>I don't think that the low number of submissions is to prevent LB probing. I think it is to reduce the high demand for GPUs in Kaggle Notebooks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 738616,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "02/06/2020 19:25:50",
          "content": "<p>Even so, increasing it in the last month won't hurt too much. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 738755,
          "author_name": "felipefonte99",
          "author_url": "",
          "post_date": "02/07/2020 00:56:12",
          "content": "<p>Maybe Kaggle team could tell us why 2 submissions. I also would like to have 3...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 738655,
      "author_name": "ankitsainiankit",
      "author_url": "",
      "post_date": "02/06/2020 20:42:48",
      "content": "<p>I agree with you. BTW I have wrote a workaround script \\in python. I download the generated submission file, pass it to script then it prints loss, plots histogram and checks for range and length of submission file. It helps a lot. Many time error is catched when I see the loss. </p>",
      "votes": null,
      "replies": [
        {
          "id": 738658,
          "author_name": "ankitsainiankit",
          "author_url": "",
          "post_date": "02/06/2020 20:44:43",
          "content": "<p>It works because the 400 test videos are from our training data. So we can have their correct labels. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 738688,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "02/06/2020 21:38:27",
          "content": "<p>unfortunately, unless you specifically removed them from the training, you will get a MUCH better result on things that were in the training set as the model \"memorize\" them. This is overfitting, basically. another problem that I have noticed is that my model actually learn to detect the actors real faces, so even unseen videos by same actors will get better results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 738690,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "02/06/2020 21:44:06",
          "content": "<p>regarding your script, it is actually neater and more reliable to do the checks in the kernel. just add a line after getting the list of files from the test_videos directory\n<code>\nif len(file_list) == 400:\n    DEBUG_MODE=True\nelse:\n    DEBUG_MODE=False\n</code></p>\n\n<p>and at the end add a cell with test specific checks:</p>\n\n<p><code>\nif DEBUG_MODE:\n         do checks\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 738705,
          "author_name": "ankitsainiankit",
          "author_url": "",
          "post_date": "02/06/2020 22:07:56",
          "content": "<p>Good to know this way of debugging. But still, I prefer separate files for testing. Just personal preference. For the overfitting case, my tests are not to check if the model is good, it is for checking if submission is correct. In your case, it would've caught p and 1-p mistake.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 738709,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "02/06/2020 22:18:19",
          "content": "<p>Generally, true. However, in this case I also switched the labels (converting the fake/real to 0/1) so no, it didn't.... facepalm....</p>\n\n<p>What I was trying to say was... S%^t happens. 2 submissions limit is generally ok, but I still think we (and the hosts) will benefit from increasing it, at least in the last month.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 738713,
          "author_name": "ankitsainiankit",
          "author_url": "",
          "post_date": "02/06/2020 22:26:12",
          "content": "<p>Yeah, I agree with you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "737933": "juliaelliott I understand that the low number of submissions was to prevent LB probing. However, as we will near the end of the competition this will become less relevant.\nWould the hosts consider increasing the limit, even to 3 in the last month? Also consider that after team merger deadline there will be much less teams and bigger teams might need more then 2 submissions to check possible options.\nI find that sometimes by making a really stupid mistake (like today, using p instead of 1-p) I am delayed by a full day.\nThanks!",
    "737936": "I agree with you!",
    "737945": "Two submissions a day is just enough, I think, no real reason to increase. The problem that you described should be solved with better testing before submitting. Add metrics to be printed out on commits inside the kernel (on 400 public validation videos), I find it very useful. We have the labels for those videos, it is a subset of the full train dataset. Add [this dataset](https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc) to your kernel, get the labels from it, and print out your favorite metrics, that way you know you submit something meaningful. \n\nThe models are heavy, taking up to 9h to run, much more to train, one shouldn't really produce more than 2 meaningful iterations per day. \n\nThe real issue is that we can't use the submissions that we already have effectively. Just today I got an error \"Notebook Exceeded Allowed Compute\", for the kernel that run successfully before with very small change at the end, 100% not related to running time, memory or disk usage. And corrupted files are soooo annoying. Anyone who participates competitively will try to squeeze maximum from it, trying to probe that may be the first half of the videos are OK and can be used for predictions and so on and so forth. That process alone will burn through submissions like Australian fires.\n\nWe need more clarity with the errors, and we need clean public and private test datasets with no unexpected video formats. Before that happens, Kaggle keeps stealing expensive time from hundreds of the participants!",
    "737963": "I check my submissions fanatically and I have been lucky so far - made one error in submission a while back (but have several indecipherable error messages). But the impossibility of debugging with the \"real\" videos makes this task difficult far beyond the stated purpose.\n  I second the OP's request for more entries, if only to compensate for the debugging difficulty. For one of my error messages, I took a few days to figure it out.",
    "738040": "I disagree that two is enough. Human errors happen even with best checking. And checking with the public validation set is... Debatable because of cross actors/training \"leak\". And yes, error handling /debugging is VERY difficult in this competition.\nThe submissions don't have to be that lengthy for checking on model improvement. You can check on 1000 entries and fill the rest with 0.5. This will run in 2h max. However, it doesn't help much if you have only 2 submissions!",
    "738205": "I propose to at least allow some \"non-countable\" error submissions. Like you can make 1-2 errors per day which won't count as real submissions.",
    "738213": "Two submissions a day is enough IMO. This is a code competition, and one of the key aspects of this is for competitors to generate robust and scalable models and pipelines. That means capability not just in terms of accuracy, but also in terms of unforeseen situations &amp; error handling, and dealing with much larger unseen data. \n\nFor me, the limitation on computation resource, the fact that you don't EXACTLY know what is wrong during submission process, and that the long scoring time is part and parcel of this competition, and it is part of the fun. Personally, I kind of take this as an engineering challenge in addition to an ML challenge. For instance, what is the best way to squeeze in as much compute with the given allowance, and making sure it doesn't spiral out of control in the long scoring process. \n\nAfter all, this is the exact nature of some of human's most challenging engineering tasks - spare a though for people who design Mars rovers -  years of work, billions of budget, months of journey, and success/failure will only unfold at the hours/minutes at touching down. I am quite grateful for the fact that I get a small taste of it in the safe and free environment of an online data competition :)",
    "738337": "I don't think that the low number of submissions is to prevent LB probing. I think it is to reduce the high demand for GPUs in Kaggle Notebooks.",
    "738616": "Even so, increasing it in the last month won't hurt too much.",
    "738655": "I agree with you. BTW I have wrote a workaround script \\in python. I download the generated submission file, pass it to script then it prints loss, plots histogram and checks for range and length of submission file. It helps a lot. Many time error is catched when I see the loss.",
    "738657": "moshel I agree with you.",
    "738658": "It works because the 400 test videos are from our training data. So we can have their correct labels.",
    "738688": "unfortunately, unless you specifically removed them from the training, you will get a MUCH better result on things that were in the training set as the model \"memorize\" them. This is overfitting, basically. another problem that I have noticed is that my model actually learn to detect the actors real faces, so even unseen videos by same actors will get better results.",
    "738690": "regarding your script, it is actually neater and more reliable to do the checks in the kernel. just add a line after getting the list of files from the test_videos directory\n```\nif len(file_list) == 400:\n    DEBUG_MODE=True\nelse:\n    DEBUG_MODE=False\n```\n\nand at the end add a cell with test specific checks:\n\n```\nif DEBUG_MODE:\n         do checks\n```",
    "738705": "Good to know this way of debugging. But still, I prefer separate files for testing. Just personal preference. For the overfitting case, my tests are not to check if the model is good, it is for checking if submission is correct. In your case, it would've caught p and 1-p mistake.",
    "738709": "Generally, true. However, in this case I also switched the labels (converting the fake/real to 0/1) so no, it didn't.... facepalm....\n\nWhat I was trying to say was... S%^t happens. 2 submissions limit is generally ok, but I still think we (and the hosts) will benefit from increasing it, at least in the last month.",
    "738713": "Yeah, I agree with you.",
    "738755": "Maybe Kaggle team could tell us why 2 submissions. I also would like to have 3..."
  },
  "source": "meta"
}