{
  "id": 126691,
  "title": "Why you should NOT use the provided data chunks as valid/train split",
  "url": "/competitions/deepfake-detection-challenge/discussion/126691",
  "author_name": "Henrique Mendonça",
  "post_date": "2020-01-19T12:31:20.647000",
  "votes": 63,
  "comment_count": 24,
  "views": 0,
  "content": "<p>The host provided us with the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/data\">dataset split into smaller chunks</a> where the original and fake videos are always present within the same chunk. However, they don't guarantee that there's only 1 video per actor.</p>\n\n<p>In fact, many actors are present in several different videos and those videos may have split in different chunks/folders.\nThis kernel <a href=\"https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda\">https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda</a> tries to find the different occurrences of the same actors in different videos.</p>\n\n<p>Cluster 73 split in 9 data chunks:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2Ffd095a7fe546868cec22a73cf61576a2%2Ff01.png?generation=1579436371111386&amp;alt=media\" alt=\"\"></p>\n\n<p>Cluster 19 split in 11 data chunks:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F454955e04a9ae9885a9b8128f5822b85%2Ff02.png?generation=1579436389106706&amp;alt=media\" alt=\"\"></p>\n\n<p>There are many more examples in the <a href=\"https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda\">kernel</a>:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F8e2aab1b77a3912336a7c1f65592c09e%2Ff12.png?generation=1579510093495631&amp;alt=media\" alt=\"\">  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F9e0c2001b591000ba676d58dbb0c3e64%2Ff8.png?generation=1579510168289148&amp;alt=media\" alt=\"\">\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F811e40d1780162103b3c872af603fc06%2Ff7.png?generation=1579510209244082&amp;alt=media\" alt=\"\">  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F6a705862b2daffeb37ad82eb30086213%2Ff11.png?generation=1579510271861235&amp;alt=media\" alt=\"\"></p>\n\n<p>You can use the exported clusters, or a similar technique, to create your unbiased training and validation split.</p>",
  "messages": [
    {
      "id": 723037,
      "postDate": "2020-01-19T12:31:20.647Z",
      "content": "<p>The host provided us with the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/data\">dataset split into smaller chunks</a> where the original and fake videos are always present within the same chunk. However, they don't guarantee that there's only 1 video per actor.</p>\n\n<p>In fact, many actors are present in several different videos and those videos may have split in different chunks/folders.\nThis kernel <a href=\"https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda\">https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda</a> tries to find the different occurrences of the same actors in different videos.</p>\n\n<p>Cluster 73 split in 9 data chunks:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2Ffd095a7fe546868cec22a73cf61576a2%2Ff01.png?generation=1579436371111386&amp;alt=media\" alt=\"\"></p>\n\n<p>Cluster 19 split in 11 data chunks:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F454955e04a9ae9885a9b8128f5822b85%2Ff02.png?generation=1579436389106706&amp;alt=media\" alt=\"\"></p>\n\n<p>There are many more examples in the <a href=\"https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda\">kernel</a>:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F8e2aab1b77a3912336a7c1f65592c09e%2Ff12.png?generation=1579510093495631&amp;alt=media\" alt=\"\">  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F9e0c2001b591000ba676d58dbb0c3e64%2Ff8.png?generation=1579510168289148&amp;alt=media\" alt=\"\">\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F811e40d1780162103b3c872af603fc06%2Ff7.png?generation=1579510209244082&amp;alt=media\" alt=\"\">  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F6a705862b2daffeb37ad82eb30086213%2Ff11.png?generation=1579510271861235&amp;alt=media\" alt=\"\"></p>\n\n<p>You can use the exported clusters, or a similar technique, to create your unbiased training and validation split.</p>",
      "rawMarkdown": "The host provided us with the [dataset split into smaller chunks](https://www.kaggle.com/c/deepfake-detection-challenge/data) where the original and fake videos are always present within the same chunk. However, they don't guarantee that there's only 1 video per actor.\n\nIn fact, many actors are present in several different videos and those videos may have split in different chunks/folders.\nThis kernel https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda tries to find the different occurrences of the same actors in different videos.\n\nCluster 73 split in 9 data chunks:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2Ffd095a7fe546868cec22a73cf61576a2%2Ff01.png?generation=1579436371111386&amp;alt=media)\n\nCluster 19 split in 11 data chunks:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F454955e04a9ae9885a9b8128f5822b85%2Ff02.png?generation=1579436389106706&amp;alt=media)\n\nThere are many more examples in the [kernel](https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda):\n\n\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F8e2aab1b77a3912336a7c1f65592c09e%2Ff12.png?generation=1579510093495631&amp;alt=media)  ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F9e0c2001b591000ba676d58dbb0c3e64%2Ff8.png?generation=1579510168289148&amp;alt=media)\n  ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F811e40d1780162103b3c872af603fc06%2Ff7.png?generation=1579510209244082&amp;alt=media)  ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F6a705862b2daffeb37ad82eb30086213%2Ff11.png?generation=1579510271861235&amp;alt=media)\n\n \n\nYou can use the exported clusters, or a similar technique, to create your unbiased training and validation split.\n\n\n\n\n\n\n\n\n\n\n",
      "votes": 62
    },
    {
      "id": 729915,
      "postDate": "2020-01-26T20:09:27.253Z",
      "content": "<p>So I did some further analysis of the chunks and clusters. </p>\n\n<p>It seems every chunk has clusters that leak to other chunks, except chunk 0. It seems that chunk 0 only has that one actor. </p>",
      "rawMarkdown": "So I did some further analysis of the chunks and clusters. \n\nIt seems every chunk has clusters that leak to other chunks, except chunk 0. It seems that chunk 0 only has that one actor. \n",
      "votes": 5,
      "replies": [
        {
          "id": 730503,
          "postDate": "2020-01-27T15:04:57.823Z",
          "content": "<p>Good call <a href=\"/hooong\">@hooong</a>, chunk 0 seems like a good candidate. I'll update the kernel.\nBelow you can see, for each data chunk, how many clusters are spread across 2 or more chunks (including the chunk itself).</p>\n\n<p><code>\nchunk   n_clusters    nonunique_clusters\n0        1            [-1]\n40       4            [1, 368, 410, 415]\n9        4            [-1, 148, 343, 450]\n12       5            [-1, 1, 24, 165, 336]\n8        5            [1, 21, 23, 24, 27]\n3        6            [1, 10, 15, 75, 341, 343]\n</code></p>\n\n<p>Note that cluster -1 contains all videos where the face detection algorithm failed in the used dataset.</p>",
          "rawMarkdown": "Good call @hooong, chunk 0 seems like a good candidate. I'll update the kernel.\nBelow you can see, for each data chunk, how many clusters are spread across 2 or more chunks (including the chunk itself).\n\n```\nchunk \tn_clusters \t  nonunique_clusters\n0 \t \t 1 \t\t \t  [-1]\n40\t\t 4 \t\t\t  [1, 368, 410, 415]\n9 \t \t 4 \t\t \t  [-1, 148, 343, 450]\n12\t\t 5 \t\t\t  [-1, 1, 24, 165, 336]\n8 \t \t 5 \t\t\t  [1, 21, 23, 24, 27]\n3 \t \t 6 \t\t \t  [1, 10, 15, 75, 341, 343]\n```\n\nNote that cluster -1 contains all videos where the face detection algorithm failed in the used dataset.",
          "votes": 5
        },
        {
          "id": 730640,
          "postDate": "2020-01-27T17:59:20.677Z",
          "content": "<p><a href=\"https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda#Is-any-chunk-independent-from-the-others\">https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda#Is-any-chunk-independent-from-the-others</a>?</p>",
          "rawMarkdown": "https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda#Is-any-chunk-independent-from-the-others?",
          "votes": 2
        }
      ]
    },
    {
      "id": 723469,
      "postDate": "2020-01-20T04:51:54.073Z",
      "content": "<p>Great Job</p>",
      "rawMarkdown": "Great Job",
      "votes": 5
    },
    {
      "id": 752488,
      "postDate": "2020-02-21T04:22:24.060Z",
      "content": "<p>I think a simple actor_id field in the metadata would have been useful. It would have saved us this unnecessary clustering exercise. Very useful post though. Thanks!</p>",
      "rawMarkdown": "I think a simple actor_id field in the metadata would have been useful. It would have saved us this unnecessary clustering exercise. Very useful post though. Thanks!",
      "votes": 1
    },
    {
      "id": 723318,
      "postDate": "2020-01-19T21:03:22.773Z",
      "content": "<p>Good finding, Henrique! That explains why transfer learning doesn't work on dataset provided by the host. I had been training and validating on videos from these chunks: validation accuracy was increasing meanwhile score at test kept getting worse.</p>",
      "rawMarkdown": "Good finding, Henrique! That explains why transfer learning doesn't work on dataset provided by the host. I had been training and validating on videos from these chunks: validation accuracy was increasing meanwhile score at test kept getting worse.",
      "votes": 1
    },
    {
      "id": 723100,
      "postDate": "2020-01-19T14:07:15.087Z",
      "content": "<p>Hi Henrique, great job there. I downloaded your clusters on a previous version of your Kernel and it didn't included the chunks from 45 to 49, any reason for that?</p>",
      "rawMarkdown": "Hi Henrique, great job there. I downloaded your clusters on a previous version of your Kernel and it didn't included the chunks from 45 to 49, any reason for that?",
      "votes": 1,
      "replies": [
        {
          "id": 723125,
          "postDate": "2020-01-19T14:38:17.587Z",
          "content": "<p>Ciao <a href=\"/pedromb\">@pedromb</a> \nAs I was saying in another <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/126468#721738\">post</a>, the data comes from <a href=\"https://www.kaggle.com/unkownhihi/deepfake\">this dataset</a>, which was still being constantly updated by the author.\nThe latest version should have all folders:\n<code>Found 19154 real videos in 50 folders</code></p>",
          "rawMarkdown": "Ciao @pedromb \nAs I was saying in another [post](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/126468#721738), the data comes from [this dataset](https://www.kaggle.com/unkownhihi/deepfake), which was still being constantly updated by the author.\nThe latest version should have all folders:\n`Found 19154 real videos in 50 folders`",
          "votes": 2
        },
        {
          "id": 723130,
          "postDate": "2020-01-19T14:40:35.457Z",
          "content": "<p>Cool! Great job, I will try on my validation scheme and report if the results were better than using the splits by chunks :)</p>",
          "rawMarkdown": "Cool! Great job, I will try on my validation scheme and report if the results were better than using the splits by chunks :)",
          "votes": 1
        },
        {
          "id": 723429,
          "postDate": "2020-01-20T02:55:47.867Z",
          "content": "<p><a href=\"/pedromb\">@pedromb</a>  Did it work?</p>",
          "rawMarkdown": "@pedromb  Did it work?",
          "votes": 1
        },
        {
          "id": 723631,
          "postDate": "2020-01-20T08:56:13.013Z",
          "content": "<p>Unfortunately no, the results were worse then using the chunks as splits, maybe it needs some more work. </p>",
          "rawMarkdown": "Unfortunately no, the results were worse then using the chunks as splits, maybe it needs some more work. "
        },
        {
          "id": 723686,
          "postDate": "2020-01-20T10:37:06.190Z",
          "content": "<p>Thanks for the input <a href=\"/pedromb\">@pedromb</a> \nIf you better faces I'd strongly encourage you to run the clustering with that instead, as the used dataset has a lot of detection failures, and even about 1-2% of the videos have no faces at all (cluster -1).\nHaving said that, it's pretty costly to create clean experiments in deep learning just because there's a lot of randomness on each run.\nWould you mind sharing your experiment setup? I.e. what exactly got worst?</p>",
          "rawMarkdown": "Thanks for the input @pedromb \nIf you better faces I'd strongly encourage you to run the clustering with that instead, as the used dataset has a lot of detection failures, and even about 1-2% of the videos have no faces at all (cluster -1).\nHaving said that, it's pretty costly to create clean experiments in deep learning just because there's a lot of randomness on each run.\nWould you mind sharing your experiment setup? I.e. what exactly got worst?",
          "votes": 3
        },
        {
          "id": 723714,
          "postDate": "2020-01-20T11:07:39.377Z",
          "content": "<p>I kept the same setup I was running with chunks split.</p>\n\n<p>The splits were done using sklearn  GroupShuffleSplit, grouping by cluster.</p>\n\n<p>With the chunks split I got val: 0.27/ lb: 0.35, with the clusters split I got val: 0.19/ lb: 0.39. All the other variables stayed the same, on both case the validation split had ~2k real videos/~2k fake videos.</p>\n\n<p>The only difference is that the clusters I got from your notebook are missing around 400 videos, so in practice for the clusters experiment I had less training data, which can explain the difference in the LB, but the validation size was the same, I would expect to get a higher loss using the clusters split, which was not the case, so my conclusion is that there are some leak happening to validation. All the data processing and model was the same on both experiments.</p>\n\n<p>If I have more time I might try doing the clusters myself like you suggested, right now I'm gonna hit pause on the competition for the next couple of weeks as I have my final exams for my masters, but I might try it when I come back. </p>",
          "rawMarkdown": "I kept the same setup I was running with chunks split.\n\nThe splits were done using sklearn  GroupShuffleSplit, grouping by cluster.\n\nWith the chunks split I got val: 0.27/ lb: 0.35, with the clusters split I got val: 0.19/ lb: 0.39. All the other variables stayed the same, on both case the validation split had ~2k real videos/~2k fake videos.\n\nThe only difference is that the clusters I got from your notebook are missing around 400 videos, so in practice for the clusters experiment I had less training data, which can explain the difference in the LB, but the validation size was the same, I would expect to get a higher loss using the clusters split, which was not the case, so my conclusion is that there are some leak happening to validation. All the data processing and model was the same on both experiments.\n\nIf I have more time I might try doing the clusters myself like you suggested, right now I'm gonna hit pause on the competition for the next couple of weeks as I have my final exams for my masters, but I might try it when I come back. ",
          "votes": 8
        },
        {
          "id": 723752,
          "postDate": "2020-01-20T12:06:46.213Z",
          "content": "<p>That's really interesting, thanks for sharing.\nRandomly splitting the clusters might not be a good idea, as there are clusters with low quality images (non-faces).\nI am using only clusters within 1 chunk (and with a minimum number of videos) for validation, as there are still plenty of them. E.g.:\n<code>clusters[(clusters.n_videos &gt; 10) &amp; (clusters.n_chunks == 1)]</code></p>\n\n<p>There will still be some leak, but hopefully at least some actors will be unique to the validation.\nI'm also stuck with a lot of work, but will post my results as soon as I have some...</p>",
          "rawMarkdown": "That's really interesting, thanks for sharing.\nRandomly splitting the clusters might not be a good idea, as there are clusters with low quality images (non-faces).\nI am using only clusters within 1 chunk (and with a minimum number of videos) for validation, as there are still plenty of them. E.g.:\n`clusters[(clusters.n_videos &gt; 10) &amp; (clusters.n_chunks == 1)]`\n\nThere will still be some leak, but hopefully at least some actors will be unique to the validation.\nI'm also stuck with a lot of work, but will post my results as soon as I have some...",
          "votes": 5
        },
        {
          "id": 752180,
          "postDate": "2020-02-20T19:30:18.873Z",
          "content": "<p>Hey <a href=\"/hmendonca\">@hmendonca</a>, thanks for the discussion and the great notebook! Did you get a chance to look at this again?</p>",
          "rawMarkdown": "Hey @hmendonca, thanks for the discussion and the great notebook! Did you get a chance to look at this again?"
        }
      ]
    },
    {
      "id": 723295,
      "postDate": "2020-01-19T19:47:12.417Z",
      "content": "<p>Currently, I am assigning videos to clusters purely using the metadata which indicates the original video corresponding to a given fake. A given real video has a number of fakes assigned to it - defining one cluster. I am think that this is a pretty close approximation to the embedding solution, especially given that multiple face videos are only about 10% of the data. Is this reasoning incorrect?</p>",
      "rawMarkdown": "Currently, I am assigning videos to clusters purely using the metadata which indicates the original video corresponding to a given fake. A given real video has a number of fakes assigned to it - defining one cluster. I am think that this is a pretty close approximation to the embedding solution, especially given that multiple face videos are only about 10% of the data. Is this reasoning incorrect?",
      "replies": [
        {
          "id": 723733,
          "postDate": "2020-01-20T11:45:58.803Z",
          "content": "<p><a href=\"/petewills\">@petewills</a> as fair as I can see, all FAKE videos are in the same data chunks as their REAL originals.\nAlso in the clusters, I only cluster the REAL videos and then propagate the same cluster number to their fakes.</p>",
          "rawMarkdown": "@petewills as fair as I can see, all FAKE videos are in the same data chunks as their REAL originals.\nAlso in the clusters, I only cluster the REAL videos and then propagate the same cluster number to their fakes.",
          "votes": 4
        }
      ]
    },
    {
      "id": 778762,
      "postDate": "2020-03-18T17:54:58.020Z",
      "content": "<p>I went ahead and checked clusters visually, and it seems only a handful of actors is present across different chunks. I doubt it matters at all.</p>",
      "rawMarkdown": "I went ahead and checked clusters visually, and it seems only a handful of actors is present across different chunks. I doubt it matters at all."
    },
    {
      "id": 751976,
      "postDate": "2020-02-20T17:34:01.867Z",
      "content": "<p>Great</p>",
      "rawMarkdown": "Great"
    },
    {
      "id": 723300,
      "postDate": "2020-01-19T20:16:32.693Z",
      "rawMarkdown": "",
      "replies": [
        {
          "id": 723301,
          "postDate": "2020-01-19T20:18:52.143Z",
          "content": "<p>Nevermind; I just checked; appearently it also have common actors. </p>",
          "rawMarkdown": "Nevermind; I just checked; appearently it also have common actors. "
        }
      ]
    },
    {
      "id": 723058,
      "postDate": "2020-01-19T13:05:05.647Z",
      "content": "<p>good job</p>",
      "rawMarkdown": "good job"
    },
    {
      "id": 723323,
      "postDate": "2020-01-19T21:28:03.877Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 723893,
      "postDate": "2020-01-20T15:24:24.583Z",
      "content": "<p>Thanks for Sharing</p>",
      "rawMarkdown": "Thanks for Sharing\n"
    }
  ],
  "comments": [
    {
      "id": 729915,
      "author_name": "hongy",
      "author_url": "",
      "post_date": "2020-01-26T20:09:27.253000",
      "content": "<p>So I did some further analysis of the chunks and clusters. </p>\n\n<p>It seems every chunk has clusters that leak to other chunks, except chunk 0. It seems that chunk 0 only has that one actor. </p>",
      "votes": 5,
      "replies": [
        {
          "id": 730503,
          "author_name": "Henrique Mendonça",
          "author_url": "",
          "post_date": "2020-01-27T15:04:57.823000",
          "content": "<p>Good call <a href=\"/hooong\">@hooong</a>, chunk 0 seems like a good candidate. I'll update the kernel.\nBelow you can see, for each data chunk, how many clusters are spread across 2 or more chunks (including the chunk itself).</p>\n\n<p><code>\nchunk   n_clusters    nonunique_clusters\n0        1            [-1]\n40       4            [1, 368, 410, 415]\n9        4            [-1, 148, 343, 450]\n12       5            [-1, 1, 24, 165, 336]\n8        5            [1, 21, 23, 24, 27]\n3        6            [1, 10, 15, 75, 341, 343]\n</code></p>\n\n<p>Note that cluster -1 contains all videos where the face detection algorithm failed in the used dataset.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 730640,
          "author_name": "Henrique Mendonça",
          "author_url": "",
          "post_date": "2020-01-27T17:59:20.677000",
          "content": "<p><a href=\"https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda#Is-any-chunk-independent-from-the-others\">https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda#Is-any-chunk-independent-from-the-others</a>?</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 723469,
      "author_name": "amar jeet kushwaha",
      "author_url": "",
      "post_date": "2020-01-20T04:51:54.073000",
      "content": "<p>Great Job</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 752488,
      "author_name": "syllap",
      "author_url": "",
      "post_date": "2020-02-21T04:22:24.060000",
      "content": "<p>I think a simple actor_id field in the metadata would have been useful. It would have saved us this unnecessary clustering exercise. Very useful post though. Thanks!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 723318,
      "author_name": "Atsamaz Gatsoev",
      "author_url": "",
      "post_date": "2020-01-19T21:03:22.773000",
      "content": "<p>Good finding, Henrique! That explains why transfer learning doesn't work on dataset provided by the host. I had been training and validating on videos from these chunks: validation accuracy was increasing meanwhile score at test kept getting worse.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 723100,
      "author_name": "Pedro Bernardo",
      "author_url": "",
      "post_date": "2020-01-19T14:07:15.087000",
      "content": "<p>Hi Henrique, great job there. I downloaded your clusters on a previous version of your Kernel and it didn't included the chunks from 45 to 49, any reason for that?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 723125,
          "author_name": "Henrique Mendonça",
          "author_url": "",
          "post_date": "2020-01-19T14:38:17.587000",
          "content": "<p>Ciao <a href=\"/pedromb\">@pedromb</a> \nAs I was saying in another <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/126468#721738\">post</a>, the data comes from <a href=\"https://www.kaggle.com/unkownhihi/deepfake\">this dataset</a>, which was still being constantly updated by the author.\nThe latest version should have all folders:\n<code>Found 19154 real videos in 50 folders</code></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 723130,
          "author_name": "Pedro Bernardo",
          "author_url": "",
          "post_date": "2020-01-19T14:40:35.457000",
          "content": "<p>Cool! Great job, I will try on my validation scheme and report if the results were better than using the splits by chunks :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 723429,
          "author_name": "xxn-xx",
          "author_url": "",
          "post_date": "2020-01-20T02:55:47.867000",
          "content": "<p><a href=\"/pedromb\">@pedromb</a>  Did it work?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 723631,
          "author_name": "Pedro Bernardo",
          "author_url": "",
          "post_date": "2020-01-20T08:56:13.013000",
          "content": "<p>Unfortunately no, the results were worse then using the chunks as splits, maybe it needs some more work. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 723686,
          "author_name": "Henrique Mendonça",
          "author_url": "",
          "post_date": "2020-01-20T10:37:06.190000",
          "content": "<p>Thanks for the input <a href=\"/pedromb\">@pedromb</a> \nIf you better faces I'd strongly encourage you to run the clustering with that instead, as the used dataset has a lot of detection failures, and even about 1-2% of the videos have no faces at all (cluster -1).\nHaving said that, it's pretty costly to create clean experiments in deep learning just because there's a lot of randomness on each run.\nWould you mind sharing your experiment setup? I.e. what exactly got worst?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 723714,
          "author_name": "Pedro Bernardo",
          "author_url": "",
          "post_date": "2020-01-20T11:07:39.377000",
          "content": "<p>I kept the same setup I was running with chunks split.</p>\n\n<p>The splits were done using sklearn  GroupShuffleSplit, grouping by cluster.</p>\n\n<p>With the chunks split I got val: 0.27/ lb: 0.35, with the clusters split I got val: 0.19/ lb: 0.39. All the other variables stayed the same, on both case the validation split had ~2k real videos/~2k fake videos.</p>\n\n<p>The only difference is that the clusters I got from your notebook are missing around 400 videos, so in practice for the clusters experiment I had less training data, which can explain the difference in the LB, but the validation size was the same, I would expect to get a higher loss using the clusters split, which was not the case, so my conclusion is that there are some leak happening to validation. All the data processing and model was the same on both experiments.</p>\n\n<p>If I have more time I might try doing the clusters myself like you suggested, right now I'm gonna hit pause on the competition for the next couple of weeks as I have my final exams for my masters, but I might try it when I come back. </p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 723752,
          "author_name": "Henrique Mendonça",
          "author_url": "",
          "post_date": "2020-01-20T12:06:46.213000",
          "content": "<p>That's really interesting, thanks for sharing.\nRandomly splitting the clusters might not be a good idea, as there are clusters with low quality images (non-faces).\nI am using only clusters within 1 chunk (and with a minimum number of videos) for validation, as there are still plenty of them. E.g.:\n<code>clusters[(clusters.n_videos &gt; 10) &amp; (clusters.n_chunks == 1)]</code></p>\n\n<p>There will still be some leak, but hopefully at least some actors will be unique to the validation.\nI'm also stuck with a lot of work, but will post my results as soon as I have some...</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 752180,
          "author_name": "morg",
          "author_url": "",
          "post_date": "2020-02-20T19:30:18.873000",
          "content": "<p>Hey <a href=\"/hmendonca\">@hmendonca</a>, thanks for the discussion and the great notebook! Did you get a chance to look at this again?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 723295,
      "author_name": "pete",
      "author_url": "",
      "post_date": "2020-01-19T19:47:12.417000",
      "content": "<p>Currently, I am assigning videos to clusters purely using the metadata which indicates the original video corresponding to a given fake. A given real video has a number of fakes assigned to it - defining one cluster. I am think that this is a pretty close approximation to the embedding solution, especially given that multiple face videos are only about 10% of the data. Is this reasoning incorrect?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 723733,
          "author_name": "Henrique Mendonça",
          "author_url": "",
          "post_date": "2020-01-20T11:45:58.803000",
          "content": "<p><a href=\"/petewills\">@petewills</a> as fair as I can see, all FAKE videos are in the same data chunks as their REAL originals.\nAlso in the clusters, I only cluster the REAL videos and then propagate the same cluster number to their fakes.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 778762,
      "author_name": "Valeriy Mukhtarulin",
      "author_url": "",
      "post_date": "2020-03-18T17:54:58.020000",
      "content": "<p>I went ahead and checked clusters visually, and it seems only a handful of actors is present across different chunks. I doubt it matters at all.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 751976,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-20T17:34:01.867000",
      "content": "<p>Great</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 723300,
      "author_name": "Emre Bayram",
      "author_url": "",
      "post_date": "2020-01-19T20:16:32.693000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 723301,
          "author_name": "Emre Bayram",
          "author_url": "",
          "post_date": "2020-01-19T20:18:52.143000",
          "content": "<p>Nevermind; I just checked; appearently it also have common actors. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 723058,
      "author_name": "gaoxiaosu",
      "author_url": "",
      "post_date": "2020-01-19T13:05:05.647000",
      "content": "<p>good job</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 723323,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-19T21:28:03.877000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 723893,
      "author_name": "Ankit Saini",
      "author_url": "",
      "post_date": "2020-01-20T15:24:24.583000",
      "content": "<p>Thanks for Sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "723037": "The host provided us with the [dataset split into smaller chunks](https://www.kaggle.com/c/deepfake-detection-challenge/data) where the original and fake videos are always present within the same chunk. However, they don't guarantee that there's only 1 video per actor.\n\nIn fact, many actors are present in several different videos and those videos may have split in different chunks/folders.\nThis kernel https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda tries to find the different occurrences of the same actors in different videos.\n\nCluster 73 split in 9 data chunks:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2Ffd095a7fe546868cec22a73cf61576a2%2Ff01.png?generation=1579436371111386&amp;alt=media)\n\nCluster 19 split in 11 data chunks:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F454955e04a9ae9885a9b8128f5822b85%2Ff02.png?generation=1579436389106706&amp;alt=media)\n\nThere are many more examples in the [kernel](https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda):\n\n\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F8e2aab1b77a3912336a7c1f65592c09e%2Ff12.png?generation=1579510093495631&amp;alt=media)  ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F9e0c2001b591000ba676d58dbb0c3e64%2Ff8.png?generation=1579510168289148&amp;alt=media)\n  ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F811e40d1780162103b3c872af603fc06%2Ff7.png?generation=1579510209244082&amp;alt=media)  ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F6a705862b2daffeb37ad82eb30086213%2Ff11.png?generation=1579510271861235&amp;alt=media)\n\n \n\nYou can use the exported clusters, or a similar technique, to create your unbiased training and validation split.\n\n\n\n\n\n\n\n\n\n\n",
    "729915": "So I did some further analysis of the chunks and clusters. \n\nIt seems every chunk has clusters that leak to other chunks, except chunk 0. It seems that chunk 0 only has that one actor. \n",
    "723469": "Great Job",
    "752488": "I think a simple actor_id field in the metadata would have been useful. It would have saved us this unnecessary clustering exercise. Very useful post though. Thanks!",
    "723318": "Good finding, Henrique! That explains why transfer learning doesn't work on dataset provided by the host. I had been training and validating on videos from these chunks: validation accuracy was increasing meanwhile score at test kept getting worse.",
    "723100": "Hi Henrique, great job there. I downloaded your clusters on a previous version of your Kernel and it didn't included the chunks from 45 to 49, any reason for that?",
    "723295": "Currently, I am assigning videos to clusters purely using the metadata which indicates the original video corresponding to a given fake. A given real video has a number of fakes assigned to it - defining one cluster. I am think that this is a pretty close approximation to the embedding solution, especially given that multiple face videos are only about 10% of the data. Is this reasoning incorrect?",
    "778762": "I went ahead and checked clusters visually, and it seems only a handful of actors is present across different chunks. I doubt it matters at all.",
    "751976": "Great",
    "723300": "",
    "723058": "good job",
    "723323": "",
    "723893": "Thanks for Sharing\n"
  }
}