{
  "id": 145982,
  "title": "113th place solution. Kaggle Team, whats up!?",
  "url": "/competitions/deepfake-detection-challenge/discussion/145982",
  "author_name": "Alex Shonenkov",
  "post_date": "2020-04-25T11:28:59.332000",
  "votes": 10,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi, everyone!</p>\n\n<p>I would like to understand what I have wrong in this competition. \nSo I made public my final kernel, that on public LB worked CORRECT! Also I made all datasets for this kernel public too.</p>\n\n<p>kernel:\n<a href=\"https://www.kaggle.com/shonenkov/final-kfold-inference-effb2\">https://www.kaggle.com/shonenkov/final-kfold-inference-effb2</a> </p>\n\n<p>models:\n<a href=\"https://www.kaggle.com/shonenkov/kfolddeepfakeeffb2-flip\">https://www.kaggle.com/shonenkov/kfolddeepfakeeffb2-flip</a>\n<a href=\"https://www.kaggle.com/shonenkov/kdold-deepfake-effb2\">https://www.kaggle.com/shonenkov/kdold-deepfake-effb2</a></p>\n\n<p>modificated face detecto:\n<a href=\"https://www.kaggle.com/shonenkov/face-detector\">https://www.kaggle.com/shonenkov/face-detector</a></p>\n\n<p>I don't have perfect solution, but it has 113th public place. </p>\n\n<p>What do you think about reasons for crushing my solution on private LB?</p>\n\n<p>Thank you in advance everyone for help to understand </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2F0d6d846e4b377e4d662c149498349e27%2Fpublic_place?generation=1587814817632556&amp;alt=media\" alt=\"\"></p>\n\n<p>[UPDATE]:</p>\n\n<p>So! After some discussions I have got some arguments. I do try/expect everywhere besides one moment! I don't try/expect for checking input signature and correctness input data. But on public data I have correct kernel! </p>\n\n<p>I think organizers must give data with input signature according to competition rules and public data</p>\n\n<p>What do you think about whose fault is it? </p>",
  "messages": [
    {
      "id": 820358,
      "postDate": "2020-04-25T11:28:59.333Z",
      "content": "<p>Hi, everyone!</p>\n\n<p>I would like to understand what I have wrong in this competition. \nSo I made public my final kernel, that on public LB worked CORRECT! Also I made all datasets for this kernel public too.</p>\n\n<p>kernel:\n<a href=\"https://www.kaggle.com/shonenkov/final-kfold-inference-effb2\">https://www.kaggle.com/shonenkov/final-kfold-inference-effb2</a> </p>\n\n<p>models:\n<a href=\"https://www.kaggle.com/shonenkov/kfolddeepfakeeffb2-flip\">https://www.kaggle.com/shonenkov/kfolddeepfakeeffb2-flip</a>\n<a href=\"https://www.kaggle.com/shonenkov/kdold-deepfake-effb2\">https://www.kaggle.com/shonenkov/kdold-deepfake-effb2</a></p>\n\n<p>modificated face detecto:\n<a href=\"https://www.kaggle.com/shonenkov/face-detector\">https://www.kaggle.com/shonenkov/face-detector</a></p>\n\n<p>I don't have perfect solution, but it has 113th public place. </p>\n\n<p>What do you think about reasons for crushing my solution on private LB?</p>\n\n<p>Thank you in advance everyone for help to understand </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2F0d6d846e4b377e4d662c149498349e27%2Fpublic_place?generation=1587814817632556&amp;alt=media\" alt=\"\"></p>\n\n<p>[UPDATE]:</p>\n\n<p>So! After some discussions I have got some arguments. I do try/expect everywhere besides one moment! I don't try/expect for checking input signature and correctness input data. But on public data I have correct kernel! </p>\n\n<p>I think organizers must give data with input signature according to competition rules and public data</p>\n\n<p>What do you think about whose fault is it? </p>",
      "rawMarkdown": "Hi, everyone!\n\nI would like to understand what I have wrong in this competition. \nSo I made public my final kernel, that on public LB worked CORRECT! Also I made all datasets for this kernel public too.\n\nkernel:\nhttps://www.kaggle.com/shonenkov/final-kfold-inference-effb2 \n\nmodels:\nhttps://www.kaggle.com/shonenkov/kfolddeepfakeeffb2-flip\nhttps://www.kaggle.com/shonenkov/kdold-deepfake-effb2\n\nmodificated face detecto:\nhttps://www.kaggle.com/shonenkov/face-detector\n\n\nI don't have perfect solution, but it has 113th public place. \n\nWhat do you think about reasons for crushing my solution on private LB?\n\nThank you in advance everyone for help to understand \n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2F0d6d846e4b377e4d662c149498349e27%2Fpublic_place?generation=1587814817632556&amp;alt=media)\n\n\n[UPDATE]:\n\nSo! After some discussions I have got some arguments. I do try/expect everywhere besides one moment! I don't try/expect for checking input signature and correctness input data. But on public data I have correct kernel! \n\nI think organizers must give data with input signature according to competition rules and public data\n\nWhat do you think about whose fault is it? \n",
      "votes": 9
    },
    {
      "id": 820374,
      "postDate": "2020-04-25T11:51:38.937Z",
      "content": "<p>It looks like you create a dataset of videos before iterating of them, and only if predicting a video fails do you set its prediction to 0.5.\nIt therefore appears that if a video fails before this point e.g. when building your <code>DatasetRetriever</code>that that video will end up without a prediction. This means it would be an invalid submission.</p>\n\n<p>Also, because you haven't try/except escaped this, I'm worried that an error when instantiating your <code>dataset = DatasetRetriever()</code> such as when using your <code>FaceDetector</code>could also be unhandled, but maybe that it escaped inside that class which we can't see as it's imported:</p>\n\n<p><code>\ndef process_dfs(df, num_workers=2):\n    def process_df(sub_df):\n        dataset = DatasetRetriever(sub_df)\n        result = deep_fake_predictor.predict(dataset)\n        return result\n</code></p>",
      "rawMarkdown": "It looks like you create a dataset of videos before iterating of them, and only if predicting a video fails do you set its prediction to 0.5.\nIt therefore appears that if a video fails before this point e.g. when building your `DatasetRetriever `that that video will end up without a prediction. This means it would be an invalid submission.\n\nAlso, because you haven't try/except escaped this, I'm worried that an error when instantiating your `dataset = DatasetRetriever()` such as when using your `FaceDetector `could also be unhandled, but maybe that it escaped inside that class which we can't see as it's imported:\n\n```\ndef process_dfs(df, num_workers=2):\n    def process_df(sub_df):\n        dataset = DatasetRetriever(sub_df)\n        result = deep_fake_predictor.predict(dataset)\n        return result\n```",
      "votes": 5,
      "replies": [
        {
          "id": 820383,
          "postDate": "2020-04-25T12:02:44.757Z",
          "content": "<p>Thank you for answer!</p>\n\n<p>In FaceDetector I used try except also:</p>\n\n<p>init.py\"&gt;https://www.kaggle.com/shonenkov/face-detector#<strong>init</strong>.py</p>\n\n<p>But I didn't  try/except on checking video signature. I expect video (according to rules of competition)</p>",
          "rawMarkdown": "Thank you for answer!\n\nIn FaceDetector I used try except also:\n\nhttps://www.kaggle.com/shonenkov/face-detector#__init__.py\n\nBut I didn't  try/except on checking video signature. I expect video (according to rules of competition)"
        }
      ]
    },
    {
      "id": 826472,
      "postDate": "2020-04-29T17:01:31.533Z",
      "content": "<p>Я не могу прокомментировать Ваш пост по сути, но хотел бы поддержать Вас... Я думаю, то что Вы делаете достаточно круто. Удачи!</p>",
      "rawMarkdown": "Я не могу прокомментировать Ваш пост по сути, но хотел бы поддержать Вас... Я думаю, то что Вы делаете достаточно круто. Удачи!",
      "votes": 1
    },
    {
      "id": 820365,
      "postDate": "2020-04-25T11:40:00.863Z",
      "content": "<p>It's common for people to do well on the public LB and get much worse scores on the private LB, usually due to overfitting.</p>",
      "rawMarkdown": "It's common for people to do well on the public LB and get much worse scores on the private LB, usually due to overfitting.",
      "votes": -8,
      "replies": [
        {
          "id": 820369,
          "postDate": "2020-04-25T11:44:13.277Z",
          "content": "<p>thank you for statistics! I have used cv kfold and checkpoints.</p>\n\n<p>Do you think It is overfitting?))</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2F51c8319fc288d0d8c2a105bdf4d25a6b%2F2020-04-25%2014-41-58.png?generation=1587814980045473&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "thank you for statistics! I have used cv kfold and checkpoints.\n\nDo you think It is overfitting?))\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2F51c8319fc288d0d8c2a105bdf4d25a6b%2F2020-04-25%2014-41-58.png?generation=1587814980045473&amp;alt=media)\n",
          "votes": 4
        }
      ]
    },
    {
      "id": 833931,
      "postDate": "2020-05-05T06:55:54.617Z",
      "content": "<p>Well, we were 67th position public, but ended up very bad like 0.5 something, with even try except loops and all, we had way too many models with perfect datas, there must had been a fault in the data pipeline, the models were excellent, we were even expecting a high silver or a gold but who knows what ended up that bad...</p>",
      "rawMarkdown": "Well, we were 67th position public, but ended up very bad like 0.5 something, with even try except loops and all, we had way too many models with perfect datas, there must had been a fault in the data pipeline, the models were excellent, we were even expecting a high silver or a gold but who knows what ended up that bad..."
    },
    {
      "id": 820410,
      "postDate": "2020-04-25T12:23:46.097Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 820374,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2020-04-25T11:51:38.937000",
      "content": "<p>It looks like you create a dataset of videos before iterating of them, and only if predicting a video fails do you set its prediction to 0.5.\nIt therefore appears that if a video fails before this point e.g. when building your <code>DatasetRetriever</code>that that video will end up without a prediction. This means it would be an invalid submission.</p>\n\n<p>Also, because you haven't try/except escaped this, I'm worried that an error when instantiating your <code>dataset = DatasetRetriever()</code> such as when using your <code>FaceDetector</code>could also be unhandled, but maybe that it escaped inside that class which we can't see as it's imported:</p>\n\n<p><code>\ndef process_dfs(df, num_workers=2):\n    def process_df(sub_df):\n        dataset = DatasetRetriever(sub_df)\n        result = deep_fake_predictor.predict(dataset)\n        return result\n</code></p>",
      "votes": 5,
      "replies": [
        {
          "id": 820383,
          "author_name": "Alex Shonenkov",
          "author_url": "",
          "post_date": "2020-04-25T12:02:44.757000",
          "content": "<p>Thank you for answer!</p>\n\n<p>In FaceDetector I used try except also:</p>\n\n<p>init.py\"&gt;https://www.kaggle.com/shonenkov/face-detector#<strong>init</strong>.py</p>\n\n<p>But I didn't  try/except on checking video signature. I expect video (according to rules of competition)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 826472,
      "author_name": " Igor Krasovskiy",
      "author_url": "",
      "post_date": "2020-04-29T17:01:31.533000",
      "content": "<p>Я не могу прокомментировать Ваш пост по сути, но хотел бы поддержать Вас... Я думаю, то что Вы делаете достаточно круто. Удачи!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 820365,
      "author_name": "Human Analog",
      "author_url": "",
      "post_date": "2020-04-25T11:40:00.863000",
      "content": "<p>It's common for people to do well on the public LB and get much worse scores on the private LB, usually due to overfitting.</p>",
      "votes": -8,
      "replies": [
        {
          "id": 820369,
          "author_name": "Alex Shonenkov",
          "author_url": "",
          "post_date": "2020-04-25T11:44:13.277000",
          "content": "<p>thank you for statistics! I have used cv kfold and checkpoints.</p>\n\n<p>Do you think It is overfitting?))</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2F51c8319fc288d0d8c2a105bdf4d25a6b%2F2020-04-25%2014-41-58.png?generation=1587814980045473&amp;alt=media\" alt=\"\"></p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 833931,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2020-05-05T06:55:54.617000",
      "content": "<p>Well, we were 67th position public, but ended up very bad like 0.5 something, with even try except loops and all, we had way too many models with perfect datas, there must had been a fault in the data pipeline, the models were excellent, we were even expecting a high silver or a gold but who knows what ended up that bad...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 820410,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-25T12:23:46.097000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "820358": "Hi, everyone!\n\nI would like to understand what I have wrong in this competition. \nSo I made public my final kernel, that on public LB worked CORRECT! Also I made all datasets for this kernel public too.\n\nkernel:\nhttps://www.kaggle.com/shonenkov/final-kfold-inference-effb2 \n\nmodels:\nhttps://www.kaggle.com/shonenkov/kfolddeepfakeeffb2-flip\nhttps://www.kaggle.com/shonenkov/kdold-deepfake-effb2\n\nmodificated face detecto:\nhttps://www.kaggle.com/shonenkov/face-detector\n\n\nI don't have perfect solution, but it has 113th public place. \n\nWhat do you think about reasons for crushing my solution on private LB?\n\nThank you in advance everyone for help to understand \n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2F0d6d846e4b377e4d662c149498349e27%2Fpublic_place?generation=1587814817632556&amp;alt=media)\n\n\n[UPDATE]:\n\nSo! After some discussions I have got some arguments. I do try/expect everywhere besides one moment! I don't try/expect for checking input signature and correctness input data. But on public data I have correct kernel! \n\nI think organizers must give data with input signature according to competition rules and public data\n\nWhat do you think about whose fault is it? \n",
    "820374": "It looks like you create a dataset of videos before iterating of them, and only if predicting a video fails do you set its prediction to 0.5.\nIt therefore appears that if a video fails before this point e.g. when building your `DatasetRetriever `that that video will end up without a prediction. This means it would be an invalid submission.\n\nAlso, because you haven't try/except escaped this, I'm worried that an error when instantiating your `dataset = DatasetRetriever()` such as when using your `FaceDetector `could also be unhandled, but maybe that it escaped inside that class which we can't see as it's imported:\n\n```\ndef process_dfs(df, num_workers=2):\n    def process_df(sub_df):\n        dataset = DatasetRetriever(sub_df)\n        result = deep_fake_predictor.predict(dataset)\n        return result\n```",
    "826472": "Я не могу прокомментировать Ваш пост по сути, но хотел бы поддержать Вас... Я думаю, то что Вы делаете достаточно круто. Удачи!",
    "820365": "It's common for people to do well on the public LB and get much worse scores on the private LB, usually due to overfitting.",
    "833931": "Well, we were 67th position public, but ended up very bad like 0.5 something, with even try except loops and all, we had way too many models with perfect datas, there must had been a fault in the data pipeline, the models were excellent, we were even expecting a high silver or a gold but who knows what ended up that bad...",
    "820410": ""
  }
}