{
  "id": 131216,
  "title": "BAD Training Data!",
  "url": "/competitions/deepfake-detection-challenge/discussion/131216",
  "author_name": "",
  "post_date": "2020-02-18T21:43:51.386279Z",
  "votes": 1,
  "comment_count": 16,
  "views": 0,
  "content": "<p>We are frustrated, and I'm sure in very good company.</p>\n\n<p>I feel like the sponsors of this contest fell well short of providing data to train on that accurately fit the test data.</p>\n\n<p>We have looked into cutting edge, innovative ways of analyzing the data and have had impressive success, only to see the test data for Leader Board was vastly inconsistent with the provided data.</p>\n\n<p>After putting in well over 100 hours we have models that will do 95% accuracy on training data get results worse than just putting 0.5 for all.  I have developed an analysis technique of the structure of the underlying file that is even more accurate and the LB data file structures are entirely different, (not the leak).</p>\n\n<p>Am I wrong in feeling the training data should be indicative of the data tested against?</p>\n\n<p>I feel we have put time into a contest that was very poorly designed, and wasted much time due to this.</p>\n\n<p>Mark</p>",
  "messages": [
    {
      "id": "749716",
      "postDate": "02/18/2020 21:43:51",
      "content": "<p>We are frustrated, and I'm sure in very good company.</p>\n\n<p>I feel like the sponsors of this contest fell well short of providing data to train on that accurately fit the test data.</p>\n\n<p>We have looked into cutting edge, innovative ways of analyzing the data and have had impressive success, only to see the test data for Leader Board was vastly inconsistent with the provided data.</p>\n\n<p>After putting in well over 100 hours we have models that will do 95% accuracy on training data get results worse than just putting 0.5 for all.  I have developed an analysis technique of the structure of the underlying file that is even more accurate and the LB data file structures are entirely different, (not the leak).</p>\n\n<p>Am I wrong in feeling the training data should be indicative of the data tested against?</p>\n\n<p>I feel we have put time into a contest that was very poorly designed, and wasted much time due to this.</p>\n\n<p>Mark</p>",
      "rawMarkdown": "We are frustrated, and I'm sure in very good company.\n\nI feel like the sponsors of this contest fell well short of providing data to train on that accurately fit the test data.\n\nWe have looked into cutting edge, innovative ways of analyzing the data and have had impressive success, only to see the test data for Leader Board was vastly inconsistent with the provided data.\n\nAfter putting in well over 100 hours we have models that will do 95% accuracy on training data get results worse than just putting 0.5 for all.  I have developed an analysis technique of the structure of the underlying file that is even more accurate and the LB data file structures are entirely different, (not the leak).\n\nAm I wrong in feeling the training data should be indicative of the data tested against?\n\nI feel we have put time into a contest that was very poorly designed, and wasted much time due to this.\n\nMark",
      "votes": null
    },
    {
      "id": "749728",
      "postDate": "02/18/2020 21:56:41",
      "content": "<p>I don't think the training data is necessarily bad, I think the main difficulty is that there's no separate validation set with unique actors etc., so it is incredibly difficult to judge overfitting without posting to the leaderboard.</p>",
      "rawMarkdown": "I don't think the training data is necessarily bad, I think the main difficulty is that there's no separate validation set with unique actors etc., so it is incredibly difficult to judge overfitting without posting to the leaderboard.",
      "votes": null
    },
    {
      "id": "749763",
      "postDate": "02/18/2020 22:26:39",
      "content": "<p>Yep, i agree. If we could isolate a good cv and make more educated sampling based on the actors, it would be better. If the host just provided the actors code in the training data and made sure there was no overlap, it would have saved us countless hours of grief. So, yes, training data is bad, but not for the reasons you mentioned. We all have 95% accuracy models that performed less than uniform 0.5 in our junk folder... </p>",
      "rawMarkdown": "Yep, i agree. If we could isolate a good cv and make more educated sampling based on the actors, it would be better. If the host just provided the actors code in the training data and made sure there was no overlap, it would have saved us countless hours of grief. So, yes, training data is bad, but not for the reasons you mentioned. We all have 95% accuracy models that performed less than uniform 0.5 in our junk folder...",
      "votes": null
    },
    {
      "id": "749792",
      "postDate": "02/18/2020 22:45:33",
      "content": "<blockquote>\n  <p>and the LB data file structures are entirely different</p>\n</blockquote>\n\n<p>Different codecs or something?</p>\n\n<p>This is disturbing, because I sometimes come across MP4 formats (not in this training data) where even VLC throws errors (decoding failures, some hardware-software-codec combination that was supposed to work, but doesn't). Decoding can be problematic, especially where you won't have a chance to troubleshoot it.</p>\n\n<p><a href=\"/cristiancanton\">@cristiancanton</a>  could you comment on this?</p>\n\n<blockquote>\n  <p>Leader Board was vastly inconsistent with the provided data</p>\n</blockquote>\n\n<p>And the private test data will not be consistent with the public test data (or that's how I understood it).</p>",
      "rawMarkdown": "&gt; and the LB data file structures are entirely different\n\nDifferent codecs or something?\n\nThis is disturbing, because I sometimes come across MP4 formats (not in this training data) where even VLC throws errors (decoding failures, some hardware-software-codec combination that was supposed to work, but doesn't). Decoding can be problematic, especially where you won't have a chance to troubleshoot it.\n\n@cristiancanton  could you comment on this?\n\n&gt; Leader Board was vastly inconsistent with the provided data\n\nAnd the private test data will not be consistent with the public test data (or that's how I understood it).",
      "votes": null
    },
    {
      "id": "749800",
      "postDate": "02/18/2020 22:55:56",
      "content": "<p>In my opinion, there are really few mp4 that are not able decode. It is only a really small portion of it. It sure costed us some time but it is not that bad to say \"poorly designed\". And if you look at those top scorers, they actually have a pretty good score! \nI recommend you to find possible glitch in your inference kernel that might have problem when there's a bad mp4 file; or to improve CV(I am working on that).</p>",
      "rawMarkdown": "In my opinion, there are really few mp4 that are not able decode. It is only a really small portion of it. It sure costed us some time but it is not that bad to say \"poorly designed\". And if you look at those top scorers, they actually have a pretty good score! \nI recommend you to find possible glitch in your inference kernel that might have problem when there's a bad mp4 file; or to improve CV(I am working on that).",
      "votes": null
    },
    {
      "id": "749865",
      "postDate": "02/19/2020 00:22:44",
      "content": "<p>I feel the same, I got training loss of 0.07 and cross validation loss of 0.14 but my inference model got 0.7 loss on LB :/  </p>",
      "rawMarkdown": "I feel the same, I got training loss of 0.07 and cross validation loss of 0.14 but my inference model got 0.7 loss on LB :/",
      "votes": null
    },
    {
      "id": "749880",
      "postDate": "02/19/2020 00:57:18",
      "content": "<pre><code>Thanks for a very concise description of the problem. I've never experienced this in a competition and it is very unsettling. Normally, one can do many simple, useful experiments on the home machine before submitting anything. It is particularly difficult when the submission process is so tedious and the submissions restricted.\n</code></pre>\n\n<p>Having said that, there are some positives. For example, I now give a fair amount of thought to what I want to try (and especially submit) instead of my usual shotgun approach.</p>",
      "rawMarkdown": "Thanks for a very concise description of the problem. I've never experienced this in a competition and it is very unsettling. Normally, one can do many simple, useful experiments on the home machine before submitting anything. It is particularly difficult when the submission process is so tedious and the submissions restricted.\n   Having said that, there are some positives. For example, I now give a fair amount of thought to what I want to try (and especially submit) instead of my usual shotgun approach.",
      "votes": null
    },
    {
      "id": "750415",
      "postDate": "02/19/2020 10:57:57",
      "content": "<p>Also note that the metric is logloss. Accuracy and logloss are pretty much unrelated. You can have 100% accuracy and score worse on logloss than a model that has 50% accuracy. </p>\n\n<p>(First case: predict 0.49 for all real videos and 0.5 for all fake videos = 100% accurate but logloss of ~0.68. Second case: predict 0.5 for all real videos and 1.0 for all fake videos = 50% accurate but logloss of ~0.35.)</p>",
      "rawMarkdown": "Also note that the metric is logloss. Accuracy and logloss are pretty much unrelated. You can have 100% accuracy and score worse on logloss than a model that has 50% accuracy. \n\n(First case: predict 0.49 for all real videos and 0.5 for all fake videos = 100% accurate but logloss of ~0.68. Second case: predict 0.5 for all real videos and 1.0 for all fake videos = 50% accurate but logloss of ~0.35.)",
      "votes": null
    },
    {
      "id": "750859",
      "postDate": "02/19/2020 18:49:29",
      "content": "<p>My best guess is the software that created the training data was entirely different and configured differently than the test data, from my testing comparisons.</p>",
      "rawMarkdown": "My best guess is the software that created the training data was entirely different and configured differently than the test data, from my testing comparisons.",
      "votes": null
    },
    {
      "id": "750863",
      "postDate": "02/19/2020 18:53:59",
      "content": "<p>Of note, after a failed submission that I should have caught, and wasting a valuable submission.  I wrote a function to test the submission for easy errors before submitting.  I made it public yesterday if anyone wants a simple tool to prevent/limit failed submits.  It test the DF before writing to file.</p>\n\n<p><a href=\"https://www.kaggle.com/meckdahl/checksubmittable/\">https://www.kaggle.com/meckdahl/checksubmittable/</a></p>",
      "rawMarkdown": "Of note, after a failed submission that I should have caught, and wasting a valuable submission.  I wrote a function to test the submission for easy errors before submitting.  I made it public yesterday if anyone wants a simple tool to prevent/limit failed submits.  It test the DF before writing to file.\n\nhttps://www.kaggle.com/meckdahl/checksubmittable/",
      "votes": null
    },
    {
      "id": "750868",
      "postDate": "02/19/2020 18:56:12",
      "content": "<p>What might be valuable is some suggested secondary training datasets that better match the test dataset, if anyone has come across some.  That's on my to-do list.</p>",
      "rawMarkdown": "What might be valuable is some suggested secondary training datasets that better match the test dataset, if anyone has come across some.  That's on my to-do list.",
      "votes": null
    },
    {
      "id": "752209",
      "postDate": "02/20/2020 19:42:26",
      "content": "<p>I agree with you, that the data is vry frustrating! I have spent days on improving a model only to see it fail  (or get a bad score) when I submit it. But that's representative of most <em>real world</em> datasets. And deepfakes are a social problem. These models should work in the wild. And as <a href=\"/unkownhihi\">@unkownhihi</a> points out the leaderboard has kagglers who have cracked it (to some degree!) </p>\n\n<p>After all, the sponsors won;t pay $1 mn for an easy problem! </p>",
      "rawMarkdown": "I agree with you, that the data is vry frustrating! I have spent days on improving a model only to see it fail  (or get a bad score) when I submit it. But that's representative of most *real world* datasets. And deepfakes are a social problem. These models should work in the wild. And as @unkownhihi points out the leaderboard has kagglers who have cracked it (to some degree!) \n\nAfter all, the sponsors won;t pay $1 mn for an easy problem!",
      "votes": null
    },
    {
      "id": "752324",
      "postDate": "02/20/2020 21:54:27",
      "content": "<p>The problem is difficult enough without having so much trouble getting proper cv. The selection of actors was the hosts decision and they could cluster it easily. Oh well, bygones</p>",
      "rawMarkdown": "The problem is difficult enough without having so much trouble getting proper cv. The selection of actors was the hosts decision and they could cluster it easily. Oh well, bygones",
      "votes": null
    },
    {
      "id": "752687",
      "postDate": "02/21/2020 09:58:40",
      "content": "<p>All I can say is you are not alone</p>",
      "rawMarkdown": "All I can say is you are not alone",
      "votes": null
    },
    {
      "id": "753017",
      "postDate": "02/21/2020 16:06:35",
      "content": "<p>Yes, I figured that was true once both my team mate and I separately had our scores get completely reversed on the test vs training data. </p>",
      "rawMarkdown": "Yes, I figured that was true once both my team mate and I separately had our scores get completely reversed on the test vs training data.",
      "votes": null
    },
    {
      "id": "753991",
      "postDate": "02/22/2020 22:49:26",
      "content": "<p>update: I was working on debugging my issues and found the following, \nctx: I'm using keras 2.3.1 TF 1.15 on aws GPU instnace with their pre configured DL image  </p>\n\n<p>1- 1st thing my model was using PIL for image load &amp; resize which is different from opencv (resizing interpolation strategy) I aligned that </p>\n\n<p>2- the keras &amp; TF versions are different between my training machine and kaggles kernel so I had to train a model on kaggle itself to see if that fixes things and it did enhance my score from ~0.60 -&gt; 0.5\n(the data set is a bit different in size of cropped faces 155 vs 224 on Kaggle) so I'll try to verify this </p>\n\n<p>at this point I'm feeling keras and TF are less portable but I'm doing more experiments to see what is happening</p>",
      "rawMarkdown": "update: I was working on debugging my issues and found the following, \nctx: I'm using keras 2.3.1 TF 1.15 on aws GPU instnace with their pre configured DL image  \n\n1- 1st thing my model was using PIL for image load &amp; resize which is different from opencv (resizing interpolation strategy) I aligned that \n\n2- the keras &amp; TF versions are different between my training machine and kaggles kernel so I had to train a model on kaggle itself to see if that fixes things and it did enhance my score from ~0.60 -&gt; 0.5\n(the data set is a bit different in size of cropped faces 155 vs 224 on Kaggle) so I'll try to verify this \n\nat this point I'm feeling keras and TF are less portable but I'm doing more experiments to see what is happening",
      "votes": null
    },
    {
      "id": "754070",
      "postDate": "02/23/2020 02:34:52",
      "content": "<p>For some reason, on kaggle, the model is easier to learn.</p>",
      "rawMarkdown": "For some reason, on kaggle, the model is easier to learn.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 749728,
      "author_name": "jamesphoward",
      "author_url": "",
      "post_date": "02/18/2020 21:56:41",
      "content": "<p>I don't think the training data is necessarily bad, I think the main difficulty is that there's no separate validation set with unique actors etc., so it is incredibly difficult to judge overfitting without posting to the leaderboard.</p>",
      "votes": null,
      "replies": [
        {
          "id": 749763,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "02/18/2020 22:26:39",
          "content": "<p>Yep, i agree. If we could isolate a good cv and make more educated sampling based on the actors, it would be better. If the host just provided the actors code in the training data and made sure there was no overlap, it would have saved us countless hours of grief. So, yes, training data is bad, but not for the reasons you mentioned. We all have 95% accuracy models that performed less than uniform 0.5 in our junk folder... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749880,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "02/19/2020 00:57:18",
          "content": "<pre><code>Thanks for a very concise description of the problem. I've never experienced this in a competition and it is very unsettling. Normally, one can do many simple, useful experiments on the home machine before submitting anything. It is particularly difficult when the submission process is so tedious and the submissions restricted.\n</code></pre>\n\n<p>Having said that, there are some positives. For example, I now give a fair amount of thought to what I want to try (and especially submit) instead of my usual shotgun approach.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 750863,
          "author_name": "meckdahl",
          "author_url": "",
          "post_date": "02/19/2020 18:53:59",
          "content": "<p>Of note, after a failed submission that I should have caught, and wasting a valuable submission.  I wrote a function to test the submission for easy errors before submitting.  I made it public yesterday if anyone wants a simple tool to prevent/limit failed submits.  It test the DF before writing to file.</p>\n\n<p><a href=\"https://www.kaggle.com/meckdahl/checksubmittable/\">https://www.kaggle.com/meckdahl/checksubmittable/</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 749792,
      "author_name": "olegtrott",
      "author_url": "",
      "post_date": "02/18/2020 22:45:33",
      "content": "<blockquote>\n  <p>and the LB data file structures are entirely different</p>\n</blockquote>\n\n<p>Different codecs or something?</p>\n\n<p>This is disturbing, because I sometimes come across MP4 formats (not in this training data) where even VLC throws errors (decoding failures, some hardware-software-codec combination that was supposed to work, but doesn't). Decoding can be problematic, especially where you won't have a chance to troubleshoot it.</p>\n\n<p><a href=\"/cristiancanton\">@cristiancanton</a>  could you comment on this?</p>\n\n<blockquote>\n  <p>Leader Board was vastly inconsistent with the provided data</p>\n</blockquote>\n\n<p>And the private test data will not be consistent with the public test data (or that's how I understood it).</p>",
      "votes": null,
      "replies": [
        {
          "id": 750859,
          "author_name": "meckdahl",
          "author_url": "",
          "post_date": "02/19/2020 18:49:29",
          "content": "<p>My best guess is the software that created the training data was entirely different and configured differently than the test data, from my testing comparisons.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 749800,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "02/18/2020 22:55:56",
      "content": "<p>In my opinion, there are really few mp4 that are not able decode. It is only a really small portion of it. It sure costed us some time but it is not that bad to say \"poorly designed\". And if you look at those top scorers, they actually have a pretty good score! \nI recommend you to find possible glitch in your inference kernel that might have problem when there's a bad mp4 file; or to improve CV(I am working on that).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 749865,
      "author_name": "basharallabadi",
      "author_url": "",
      "post_date": "02/19/2020 00:22:44",
      "content": "<p>I feel the same, I got training loss of 0.07 and cross validation loss of 0.14 but my inference model got 0.7 loss on LB :/  </p>",
      "votes": null,
      "replies": [
        {
          "id": 753991,
          "author_name": "basharallabadi",
          "author_url": "",
          "post_date": "02/22/2020 22:49:26",
          "content": "<p>update: I was working on debugging my issues and found the following, \nctx: I'm using keras 2.3.1 TF 1.15 on aws GPU instnace with their pre configured DL image  </p>\n\n<p>1- 1st thing my model was using PIL for image load &amp; resize which is different from opencv (resizing interpolation strategy) I aligned that </p>\n\n<p>2- the keras &amp; TF versions are different between my training machine and kaggles kernel so I had to train a model on kaggle itself to see if that fixes things and it did enhance my score from ~0.60 -&gt; 0.5\n(the data set is a bit different in size of cropped faces 155 vs 224 on Kaggle) so I'll try to verify this </p>\n\n<p>at this point I'm feeling keras and TF are less portable but I'm doing more experiments to see what is happening</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 754070,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "02/23/2020 02:34:52",
          "content": "<p>For some reason, on kaggle, the model is easier to learn.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 750415,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "02/19/2020 10:57:57",
      "content": "<p>Also note that the metric is logloss. Accuracy and logloss are pretty much unrelated. You can have 100% accuracy and score worse on logloss than a model that has 50% accuracy. </p>\n\n<p>(First case: predict 0.49 for all real videos and 0.5 for all fake videos = 100% accurate but logloss of ~0.68. Second case: predict 0.5 for all real videos and 1.0 for all fake videos = 50% accurate but logloss of ~0.35.)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 750868,
      "author_name": "meckdahl",
      "author_url": "",
      "post_date": "02/19/2020 18:56:12",
      "content": "<p>What might be valuable is some suggested secondary training datasets that better match the test dataset, if anyone has come across some.  That's on my to-do list.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 752209,
      "author_name": "skylord",
      "author_url": "",
      "post_date": "02/20/2020 19:42:26",
      "content": "<p>I agree with you, that the data is vry frustrating! I have spent days on improving a model only to see it fail  (or get a bad score) when I submit it. But that's representative of most <em>real world</em> datasets. And deepfakes are a social problem. These models should work in the wild. And as <a href=\"/unkownhihi\">@unkownhihi</a> points out the leaderboard has kagglers who have cracked it (to some degree!) </p>\n\n<p>After all, the sponsors won;t pay $1 mn for an easy problem! </p>",
      "votes": null,
      "replies": [
        {
          "id": 752324,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "02/20/2020 21:54:27",
          "content": "<p>The problem is difficult enough without having so much trouble getting proper cv. The selection of actors was the hosts decision and they could cluster it easily. Oh well, bygones</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 752687,
      "author_name": "alexis1024",
      "author_url": "",
      "post_date": "02/21/2020 09:58:40",
      "content": "<p>All I can say is you are not alone</p>",
      "votes": null,
      "replies": [
        {
          "id": 753017,
          "author_name": "meckdahl",
          "author_url": "",
          "post_date": "02/21/2020 16:06:35",
          "content": "<p>Yes, I figured that was true once both my team mate and I separately had our scores get completely reversed on the test vs training data. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "749716": "We are frustrated, and I'm sure in very good company.\n\nI feel like the sponsors of this contest fell well short of providing data to train on that accurately fit the test data.\n\nWe have looked into cutting edge, innovative ways of analyzing the data and have had impressive success, only to see the test data for Leader Board was vastly inconsistent with the provided data.\n\nAfter putting in well over 100 hours we have models that will do 95% accuracy on training data get results worse than just putting 0.5 for all.  I have developed an analysis technique of the structure of the underlying file that is even more accurate and the LB data file structures are entirely different, (not the leak).\n\nAm I wrong in feeling the training data should be indicative of the data tested against?\n\nI feel we have put time into a contest that was very poorly designed, and wasted much time due to this.\n\nMark",
    "749728": "I don't think the training data is necessarily bad, I think the main difficulty is that there's no separate validation set with unique actors etc., so it is incredibly difficult to judge overfitting without posting to the leaderboard.",
    "749763": "Yep, i agree. If we could isolate a good cv and make more educated sampling based on the actors, it would be better. If the host just provided the actors code in the training data and made sure there was no overlap, it would have saved us countless hours of grief. So, yes, training data is bad, but not for the reasons you mentioned. We all have 95% accuracy models that performed less than uniform 0.5 in our junk folder...",
    "749792": "&gt; and the LB data file structures are entirely different\n\nDifferent codecs or something?\n\nThis is disturbing, because I sometimes come across MP4 formats (not in this training data) where even VLC throws errors (decoding failures, some hardware-software-codec combination that was supposed to work, but doesn't). Decoding can be problematic, especially where you won't have a chance to troubleshoot it.\n\n@cristiancanton  could you comment on this?\n\n&gt; Leader Board was vastly inconsistent with the provided data\n\nAnd the private test data will not be consistent with the public test data (or that's how I understood it).",
    "749800": "In my opinion, there are really few mp4 that are not able decode. It is only a really small portion of it. It sure costed us some time but it is not that bad to say \"poorly designed\". And if you look at those top scorers, they actually have a pretty good score! \nI recommend you to find possible glitch in your inference kernel that might have problem when there's a bad mp4 file; or to improve CV(I am working on that).",
    "749865": "I feel the same, I got training loss of 0.07 and cross validation loss of 0.14 but my inference model got 0.7 loss on LB :/",
    "749880": "Thanks for a very concise description of the problem. I've never experienced this in a competition and it is very unsettling. Normally, one can do many simple, useful experiments on the home machine before submitting anything. It is particularly difficult when the submission process is so tedious and the submissions restricted.\n   Having said that, there are some positives. For example, I now give a fair amount of thought to what I want to try (and especially submit) instead of my usual shotgun approach.",
    "750415": "Also note that the metric is logloss. Accuracy and logloss are pretty much unrelated. You can have 100% accuracy and score worse on logloss than a model that has 50% accuracy. \n\n(First case: predict 0.49 for all real videos and 0.5 for all fake videos = 100% accurate but logloss of ~0.68. Second case: predict 0.5 for all real videos and 1.0 for all fake videos = 50% accurate but logloss of ~0.35.)",
    "750859": "My best guess is the software that created the training data was entirely different and configured differently than the test data, from my testing comparisons.",
    "750863": "Of note, after a failed submission that I should have caught, and wasting a valuable submission.  I wrote a function to test the submission for easy errors before submitting.  I made it public yesterday if anyone wants a simple tool to prevent/limit failed submits.  It test the DF before writing to file.\n\nhttps://www.kaggle.com/meckdahl/checksubmittable/",
    "750868": "What might be valuable is some suggested secondary training datasets that better match the test dataset, if anyone has come across some.  That's on my to-do list.",
    "752209": "I agree with you, that the data is vry frustrating! I have spent days on improving a model only to see it fail  (or get a bad score) when I submit it. But that's representative of most *real world* datasets. And deepfakes are a social problem. These models should work in the wild. And as @unkownhihi points out the leaderboard has kagglers who have cracked it (to some degree!) \n\nAfter all, the sponsors won;t pay $1 mn for an easy problem!",
    "752324": "The problem is difficult enough without having so much trouble getting proper cv. The selection of actors was the hosts decision and they could cluster it easily. Oh well, bygones",
    "752687": "All I can say is you are not alone",
    "753017": "Yes, I figured that was true once both my team mate and I separately had our scores get completely reversed on the test vs training data.",
    "753991": "update: I was working on debugging my issues and found the following, \nctx: I'm using keras 2.3.1 TF 1.15 on aws GPU instnace with their pre configured DL image  \n\n1- 1st thing my model was using PIL for image load &amp; resize which is different from opencv (resizing interpolation strategy) I aligned that \n\n2- the keras &amp; TF versions are different between my training machine and kaggles kernel so I had to train a model on kaggle itself to see if that fixes things and it did enhance my score from ~0.60 -&gt; 0.5\n(the data set is a bit different in size of cropped faces 155 vs 224 on Kaggle) so I'll try to verify this \n\nat this point I'm feeling keras and TF are less portable but I'm doing more experiments to see what is happening",
    "754070": "For some reason, on kaggle, the model is easier to learn."
  },
  "source": "meta"
}