{
  "id": 140467,
  "title": "Splitting data into folds based on facially similar groupings",
  "url": "/competitions/deepfake-detection-challenge/discussion/140467",
  "author_name": "",
  "post_date": "2020-04-01T23:27:31.021386300Z",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Fellow Kagglers,</p>\n\n<p>first of all, thank you all for sharing amazing materials in discussions and kernels. For me it's a great privilege to learn from all of you! </p>\n\n<p>I'm curious if anyone was grouping videos based on facial similarity before splitting data into folds. If so, could you please share if it helped compared to chunk-based split. Thank you!</p>\n\n<p>I thought that feeding a classifier with true-fake image pairs containing the same face should reduce overfitting by complicating face memorization for the classifier. However, the facial grouping is unlikely to be perfect. Unfortunately I wasn't able to experiment with different data splitting strategies as I had started working on the competition too late (I arrived to the first reasonable draft for grouping authentic videos of the same person and matching authentic videos with facially similar fakes only on Friday before the deadline, and being short on time I just used it (with reduced recursion depth during matching for speed). \nI've shared technical details illustrating how I ran the facial-grouping split using data hosted by Kaggle in the <a href=\"https://www.kaggle.com/samusram/dfdc-organizing-data-into-folds-based-on-face\">notebook <em>DFDC: organizing data into folds based on face</em></a>, and my facial embeddings <a href=\"https://www.kaggle.com/samusram/dfdcfacenetembeddings\">here</a>.</p>\n\n<p><em>Some insights:</em>\n1. For good illumination and frontal faces the results seem to be not perfect but nice:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fd4293b39a99f2e9583107f003679aad6%2Fdovbqeieek.png?generation=1585782362358505&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F5866ab92fc52765fb3218761bdcd660a%2Fgiifpbniet.png?generation=1585782441238586&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F88daa3c1c5d292818e2fd4383b3b4dde%2Faewpytojhs.png?generation=1585781089355627&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F356a8989cda7e771e17c34bcc78a18cf%2Fabfnyenqdw.png?generation=1585782531638289&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F97889003220251c8a4efa79bbaeeeb5b%2Fjjxtumenxk.png?generation=1585782638358549&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fb8117f68ae5f305bb595dee4f4a115c3%2Fahesnzswur.png?generation=1585783086840703&amp;alt=media\" alt=\"\"></p>\n\n<ol>\n<li><p>Faces in profile are problematic\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F85ad33a6f605ca4bdbc822fc40ee3c79%2Feyhjqwikza.png?generation=1585783191825264&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fbf64c65eb01c3c10efdc1c2d8e6b0a15%2Fedldghszvx.png?generation=1585783275547876&amp;alt=media\" alt=\"\"></p></li>\n<li><p>Dark images are problematic </p></li>\n</ol>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F1e4e6fc58f630d08a2cfb5d2a5ff3454%2Fdark.png?generation=1585783418886137&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fa0d4d34a7e895d424d9f8e0de40c86b7%2Fdark%202.png?generation=1585783435880109&amp;alt=media\" alt=\"\"></p>\n\n<ol>\n<li>Glasses pose an additional challenge\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F98dec4dfbe724eb58d568fa35a6b748d%2Fglasses.png?generation=1585783478568394&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fd1392078eef14ae4f0d63bf9d478eb06%2Fhxmwbpwwlv%20glasses.png?generation=1585783492985607&amp;alt=media\" alt=\"\"></li>\n</ol>",
  "messages": [
    {
      "id": "794592",
      "postDate": "04/01/2020 23:27:31",
      "content": "<p>Fellow Kagglers,</p>\n\n<p>first of all, thank you all for sharing amazing materials in discussions and kernels. For me it's a great privilege to learn from all of you! </p>\n\n<p>I'm curious if anyone was grouping videos based on facial similarity before splitting data into folds. If so, could you please share if it helped compared to chunk-based split. Thank you!</p>\n\n<p>I thought that feeding a classifier with true-fake image pairs containing the same face should reduce overfitting by complicating face memorization for the classifier. However, the facial grouping is unlikely to be perfect. Unfortunately I wasn't able to experiment with different data splitting strategies as I had started working on the competition too late (I arrived to the first reasonable draft for grouping authentic videos of the same person and matching authentic videos with facially similar fakes only on Friday before the deadline, and being short on time I just used it (with reduced recursion depth during matching for speed). \nI've shared technical details illustrating how I ran the facial-grouping split using data hosted by Kaggle in the <a href=\"https://www.kaggle.com/samusram/dfdc-organizing-data-into-folds-based-on-face\">notebook <em>DFDC: organizing data into folds based on face</em></a>, and my facial embeddings <a href=\"https://www.kaggle.com/samusram/dfdcfacenetembeddings\">here</a>.</p>\n\n<p><em>Some insights:</em>\n1. For good illumination and frontal faces the results seem to be not perfect but nice:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fd4293b39a99f2e9583107f003679aad6%2Fdovbqeieek.png?generation=1585782362358505&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F5866ab92fc52765fb3218761bdcd660a%2Fgiifpbniet.png?generation=1585782441238586&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F88daa3c1c5d292818e2fd4383b3b4dde%2Faewpytojhs.png?generation=1585781089355627&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F356a8989cda7e771e17c34bcc78a18cf%2Fabfnyenqdw.png?generation=1585782531638289&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F97889003220251c8a4efa79bbaeeeb5b%2Fjjxtumenxk.png?generation=1585782638358549&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fb8117f68ae5f305bb595dee4f4a115c3%2Fahesnzswur.png?generation=1585783086840703&amp;alt=media\" alt=\"\"></p>\n\n<ol>\n<li><p>Faces in profile are problematic\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F85ad33a6f605ca4bdbc822fc40ee3c79%2Feyhjqwikza.png?generation=1585783191825264&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fbf64c65eb01c3c10efdc1c2d8e6b0a15%2Fedldghszvx.png?generation=1585783275547876&amp;alt=media\" alt=\"\"></p></li>\n<li><p>Dark images are problematic </p></li>\n</ol>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F1e4e6fc58f630d08a2cfb5d2a5ff3454%2Fdark.png?generation=1585783418886137&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fa0d4d34a7e895d424d9f8e0de40c86b7%2Fdark%202.png?generation=1585783435880109&amp;alt=media\" alt=\"\"></p>\n\n<ol>\n<li>Glasses pose an additional challenge\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F98dec4dfbe724eb58d568fa35a6b748d%2Fglasses.png?generation=1585783478568394&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fd1392078eef14ae4f0d63bf9d478eb06%2Fhxmwbpwwlv%20glasses.png?generation=1585783492985607&amp;alt=media\" alt=\"\"></li>\n</ol>",
      "rawMarkdown": "Fellow Kagglers,\n\nfirst of all, thank you all for sharing amazing materials in discussions and kernels. For me it's a great privilege to learn from all of you! \n\nI'm curious if anyone was grouping videos based on facial similarity before splitting data into folds. If so, could you please share if it helped compared to chunk-based split. Thank you!\n\nI thought that feeding a classifier with true-fake image pairs containing the same face should reduce overfitting by complicating face memorization for the classifier. However, the facial grouping is unlikely to be perfect. Unfortunately I wasn't able to experiment with different data splitting strategies as I had started working on the competition too late (I arrived to the first reasonable draft for grouping authentic videos of the same person and matching authentic videos with facially similar fakes only on Friday before the deadline, and being short on time I just used it (with reduced recursion depth during matching for speed). \nI've shared technical details illustrating how I ran the facial-grouping split using data hosted by Kaggle in the [notebook *DFDC: organizing data into folds based on face*](https://www.kaggle.com/samusram/dfdc-organizing-data-into-folds-based-on-face), and my facial embeddings [here](https://www.kaggle.com/samusram/dfdcfacenetembeddings).\n\n*Some insights:*\n1. For good illumination and frontal faces the results seem to be not perfect but nice:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fd4293b39a99f2e9583107f003679aad6%2Fdovbqeieek.png?generation=1585782362358505&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F5866ab92fc52765fb3218761bdcd660a%2Fgiifpbniet.png?generation=1585782441238586&amp;alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F88daa3c1c5d292818e2fd4383b3b4dde%2Faewpytojhs.png?generation=1585781089355627&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F356a8989cda7e771e17c34bcc78a18cf%2Fabfnyenqdw.png?generation=1585782531638289&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F97889003220251c8a4efa79bbaeeeb5b%2Fjjxtumenxk.png?generation=1585782638358549&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fb8117f68ae5f305bb595dee4f4a115c3%2Fahesnzswur.png?generation=1585783086840703&amp;alt=media)\n\n2. Faces in profile are problematic\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F85ad33a6f605ca4bdbc822fc40ee3c79%2Feyhjqwikza.png?generation=1585783191825264&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fbf64c65eb01c3c10efdc1c2d8e6b0a15%2Fedldghszvx.png?generation=1585783275547876&amp;alt=media)\n\n3. Dark images are problematic \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F1e4e6fc58f630d08a2cfb5d2a5ff3454%2Fdark.png?generation=1585783418886137&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fa0d4d34a7e895d424d9f8e0de40c86b7%2Fdark%202.png?generation=1585783435880109&amp;alt=media)\n\n4. Glasses pose an additional challenge\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F98dec4dfbe724eb58d568fa35a6b748d%2Fglasses.png?generation=1585783478568394&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fd1392078eef14ae4f0d63bf9d478eb06%2Fhxmwbpwwlv%20glasses.png?generation=1585783492985607&amp;alt=media)",
      "votes": null
    },
    {
      "id": "794756",
      "postDate": "04/02/2020 04:15:53",
      "content": "<p>Hi Raman, in my team, Henrique - <a href=\"/hmendonca\">@hmendonca</a>, has worked very hard on splitting data based on different actors; unfortunately, for us, it seems result in a minor improvement over simple random fold splitting. Perhaps Henrique could you please give some insights here ?</p>\n\n<p>BTW, you are amazing to have setup everything and get this score all by yourself within only several days !!</p>",
      "rawMarkdown": "Hi Raman, in my team, Henrique - @hmendonca, has worked very hard on splitting data based on different actors; unfortunately, for us, it seems result in a minor improvement over simple random fold splitting. Perhaps Henrique could you please give some insights here ?\n\nBTW, you are amazing to have setup everything and get this score all by yourself within only several days !!",
      "votes": null
    },
    {
      "id": "794764",
      "postDate": "04/02/2020 04:25:16",
      "content": "<p>Hi Jung, thank you for your reply! And thank you for your kind words, your ability to encourage with kind feedback is unique!</p>",
      "rawMarkdown": "Hi Jung, thank you for your reply! And thank you for your kind words, your ability to encourage with kind feedback is unique!",
      "votes": null
    },
    {
      "id": "794995",
      "postDate": "04/02/2020 09:27:16",
      "content": "<p><a href=\"/samusram\">@samusram</a> Nice work! It looks like you calculate the embeddings for all the fake videos as well. Is that right?</p>\n\n<p>Many people noticed that most actors were contained within a single folder, so splitting folders is almost as good as clustering the faces. However, clustering introduces a lot of noise (embedding failure) and the folders are quite safe. There are only a few prolific actors that appear in several folders, as you can see in the discussion and also in my public kernel. </p>\n\n<p>Therefore, we simply grouped the folders that had those common actors into 5 folds, i.e. keeping each main prolific actor within 1 single fold, when possible:</p>\n\n<p>```\nactor1 = [2, 7, 13, 24, 25, 27, 39, 47]  # folders\nactor2 = [6, 7, 10, 11, 17, 23, 37, 38]\nactor3 = [5, 16, 17, 18, 30, 35, 41]\nactor4 = [14, 26, 47]\nactor5 = [16, 17, 26, 28, 34]</p>\n\n<h2>proposed</h2>\n\n<p>actor_split = [\n    [0,  1,  2,  7, 13, 24, 25, 27, 39, 47],\n    [3,  6, 10, 11, 17, 23, 33, 37, 38],\n    [4,  5, 16, 18, 30, 32, 35, 41, 43],\n    [8, 12, 15, 20, 22, 31, 36, 40, 44],\n    [9, 14, 19, 21, 26, 28, 29, 34, 42],\n]\ntest = [45, 46, 48, 49]\n```</p>",
      "rawMarkdown": "samusram Nice work! It looks like you calculate the embeddings for all the fake videos as well. Is that right?\n\nMany people noticed that most actors were contained within a single folder, so splitting folders is almost as good as clustering the faces. However, clustering introduces a lot of noise (embedding failure) and the folders are quite safe. There are only a few prolific actors that appear in several folders, as you can see in the discussion and also in my public kernel. \n\nTherefore, we simply grouped the folders that had those common actors into 5 folds, i.e. keeping each main prolific actor within 1 single fold, when possible:\n\n```\nactor1 = [2, 7, 13, 24, 25, 27, 39, 47]  # folders\nactor2 = [6, 7, 10, 11, 17, 23, 37, 38]\nactor3 = [5, 16, 17, 18, 30, 35, 41]\nactor4 = [14, 26, 47]\nactor5 = [16, 17, 26, 28, 34]\n\n## proposed\nactor_split = [\n\t[0,  1,  2,  7, 13, 24, 25, 27, 39, 47],\n\t[3,  6, 10, 11, 17, 23, 33, 37, 38],\n\t[4,  5, 16, 18, 30, 32, 35, 41, 43],\n\t[8, 12, 15, 20, 22, 31, 36, 40, 44],\n\t[9, 14, 19, 21, 26, 28, 29, 34, 42],\n]\ntest = [45, 46, 48, 49]\n```",
      "votes": null
    },
    {
      "id": "795021",
      "postDate": "04/02/2020 09:58:14",
      "content": "<p><a href=\"/hmendonca\">@hmendonca</a> Thank you for the details! I should have thought of your elegant and more straightforward splitting strategy :)</p>\n\n<blockquote>\n  <p>It looks like you calculate the embeddings for all the fake videos as well. Is that right?</p>\n</blockquote>\n\n<p>Yes, Henrique. </p>\n\n<p>Based on a brief exploration it seems that unlike authentic videos, which usually have the same person being present in a single chunk except the corner cases you've kindly described, for the fake faces it doesn't hold. In the examples shared above I print the video chunk of the fake. You can see that fake videos with the same face are distributed over multiple chunks. It seems like the fake videos might be usually in the same folder as the original one. </p>",
      "rawMarkdown": "hmendonca Thank you for the details! I should have thought of your elegant and more straightforward splitting strategy :)\n\n&gt; It looks like you calculate the embeddings for all the fake videos as well. Is that right?\n\nYes, Henrique. \n\nBased on a brief exploration it seems that unlike authentic videos, which usually have the same person being present in a single chunk except the corner cases you've kindly described, for the fake faces it doesn't hold. In the examples shared above I print the video chunk of the fake. You can see that fake videos with the same face are distributed over multiple chunks. It seems like the fake videos might be usually in the same folder as the original one.",
      "votes": null
    },
    {
      "id": "795202",
      "postDate": "04/02/2020 14:02:36",
      "content": "<p>That's very interesting. Thanks.</p>\n\n<p>And yes, the fakes are always accompanied by their originals within the same folder, but no idea about the target actors (the inserted fake faces). Perhaps that was the leakage that we kept seeing, that caused the massive valid/LB gap!</p>",
      "rawMarkdown": "That's very interesting. Thanks.\n\nAnd yes, the fakes are always accompanied by their originals within the same folder, but no idea about the target actors (the inserted fake faces). Perhaps that was the leakage that we kept seeing, that caused the massive valid/LB gap!",
      "votes": null
    },
    {
      "id": "795227",
      "postDate": "04/02/2020 14:31:54",
      "content": "<p>Henrique, what was your validation/LB gap if you don't mind me asking? I had 0.05 - 0.08</p>",
      "rawMarkdown": "Henrique, what was your validation/LB gap if you don't mind me asking? I had 0.05 - 0.08",
      "votes": null
    },
    {
      "id": "796346",
      "postDate": "04/03/2020 14:06:59",
      "content": "<p>In my personal experiments, gaps can be as big as 0.1 or even 0.2 &gt;__&lt;</p>",
      "rawMarkdown": "In my personal experiments, gaps can be as big as 0.1 or even 0.2 &gt;__&lt;",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 794756,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "04/02/2020 04:15:53",
      "content": "<p>Hi Raman, in my team, Henrique - <a href=\"/hmendonca\">@hmendonca</a>, has worked very hard on splitting data based on different actors; unfortunately, for us, it seems result in a minor improvement over simple random fold splitting. Perhaps Henrique could you please give some insights here ?</p>\n\n<p>BTW, you are amazing to have setup everything and get this score all by yourself within only several days !!</p>",
      "votes": null,
      "replies": [
        {
          "id": 794764,
          "author_name": "samusram",
          "author_url": "",
          "post_date": "04/02/2020 04:25:16",
          "content": "<p>Hi Jung, thank you for your reply! And thank you for your kind words, your ability to encourage with kind feedback is unique!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 794995,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "04/02/2020 09:27:16",
          "content": "<p><a href=\"/samusram\">@samusram</a> Nice work! It looks like you calculate the embeddings for all the fake videos as well. Is that right?</p>\n\n<p>Many people noticed that most actors were contained within a single folder, so splitting folders is almost as good as clustering the faces. However, clustering introduces a lot of noise (embedding failure) and the folders are quite safe. There are only a few prolific actors that appear in several folders, as you can see in the discussion and also in my public kernel. </p>\n\n<p>Therefore, we simply grouped the folders that had those common actors into 5 folds, i.e. keeping each main prolific actor within 1 single fold, when possible:</p>\n\n<p>```\nactor1 = [2, 7, 13, 24, 25, 27, 39, 47]  # folders\nactor2 = [6, 7, 10, 11, 17, 23, 37, 38]\nactor3 = [5, 16, 17, 18, 30, 35, 41]\nactor4 = [14, 26, 47]\nactor5 = [16, 17, 26, 28, 34]</p>\n\n<h2>proposed</h2>\n\n<p>actor_split = [\n    [0,  1,  2,  7, 13, 24, 25, 27, 39, 47],\n    [3,  6, 10, 11, 17, 23, 33, 37, 38],\n    [4,  5, 16, 18, 30, 32, 35, 41, 43],\n    [8, 12, 15, 20, 22, 31, 36, 40, 44],\n    [9, 14, 19, 21, 26, 28, 29, 34, 42],\n]\ntest = [45, 46, 48, 49]\n```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 795021,
          "author_name": "samusram",
          "author_url": "",
          "post_date": "04/02/2020 09:58:14",
          "content": "<p><a href=\"/hmendonca\">@hmendonca</a> Thank you for the details! I should have thought of your elegant and more straightforward splitting strategy :)</p>\n\n<blockquote>\n  <p>It looks like you calculate the embeddings for all the fake videos as well. Is that right?</p>\n</blockquote>\n\n<p>Yes, Henrique. </p>\n\n<p>Based on a brief exploration it seems that unlike authentic videos, which usually have the same person being present in a single chunk except the corner cases you've kindly described, for the fake faces it doesn't hold. In the examples shared above I print the video chunk of the fake. You can see that fake videos with the same face are distributed over multiple chunks. It seems like the fake videos might be usually in the same folder as the original one. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 795202,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "04/02/2020 14:02:36",
          "content": "<p>That's very interesting. Thanks.</p>\n\n<p>And yes, the fakes are always accompanied by their originals within the same folder, but no idea about the target actors (the inserted fake faces). Perhaps that was the leakage that we kept seeing, that caused the massive valid/LB gap!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 795227,
          "author_name": "samusram",
          "author_url": "",
          "post_date": "04/02/2020 14:31:54",
          "content": "<p>Henrique, what was your validation/LB gap if you don't mind me asking? I had 0.05 - 0.08</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 796346,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "04/03/2020 14:06:59",
          "content": "<p>In my personal experiments, gaps can be as big as 0.1 or even 0.2 &gt;__&lt;</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "794592": "Fellow Kagglers,\n\nfirst of all, thank you all for sharing amazing materials in discussions and kernels. For me it's a great privilege to learn from all of you! \n\nI'm curious if anyone was grouping videos based on facial similarity before splitting data into folds. If so, could you please share if it helped compared to chunk-based split. Thank you!\n\nI thought that feeding a classifier with true-fake image pairs containing the same face should reduce overfitting by complicating face memorization for the classifier. However, the facial grouping is unlikely to be perfect. Unfortunately I wasn't able to experiment with different data splitting strategies as I had started working on the competition too late (I arrived to the first reasonable draft for grouping authentic videos of the same person and matching authentic videos with facially similar fakes only on Friday before the deadline, and being short on time I just used it (with reduced recursion depth during matching for speed). \nI've shared technical details illustrating how I ran the facial-grouping split using data hosted by Kaggle in the [notebook *DFDC: organizing data into folds based on face*](https://www.kaggle.com/samusram/dfdc-organizing-data-into-folds-based-on-face), and my facial embeddings [here](https://www.kaggle.com/samusram/dfdcfacenetembeddings).\n\n*Some insights:*\n1. For good illumination and frontal faces the results seem to be not perfect but nice:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fd4293b39a99f2e9583107f003679aad6%2Fdovbqeieek.png?generation=1585782362358505&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F5866ab92fc52765fb3218761bdcd660a%2Fgiifpbniet.png?generation=1585782441238586&amp;alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F88daa3c1c5d292818e2fd4383b3b4dde%2Faewpytojhs.png?generation=1585781089355627&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F356a8989cda7e771e17c34bcc78a18cf%2Fabfnyenqdw.png?generation=1585782531638289&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F97889003220251c8a4efa79bbaeeeb5b%2Fjjxtumenxk.png?generation=1585782638358549&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fb8117f68ae5f305bb595dee4f4a115c3%2Fahesnzswur.png?generation=1585783086840703&amp;alt=media)\n\n2. Faces in profile are problematic\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F85ad33a6f605ca4bdbc822fc40ee3c79%2Feyhjqwikza.png?generation=1585783191825264&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fbf64c65eb01c3c10efdc1c2d8e6b0a15%2Fedldghszvx.png?generation=1585783275547876&amp;alt=media)\n\n3. Dark images are problematic \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F1e4e6fc58f630d08a2cfb5d2a5ff3454%2Fdark.png?generation=1585783418886137&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fa0d4d34a7e895d424d9f8e0de40c86b7%2Fdark%202.png?generation=1585783435880109&amp;alt=media)\n\n4. Glasses pose an additional challenge\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2F98dec4dfbe724eb58d568fa35a6b748d%2Fglasses.png?generation=1585783478568394&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F493138%2Fd1392078eef14ae4f0d63bf9d478eb06%2Fhxmwbpwwlv%20glasses.png?generation=1585783492985607&amp;alt=media)",
    "794756": "Hi Raman, in my team, Henrique - @hmendonca, has worked very hard on splitting data based on different actors; unfortunately, for us, it seems result in a minor improvement over simple random fold splitting. Perhaps Henrique could you please give some insights here ?\n\nBTW, you are amazing to have setup everything and get this score all by yourself within only several days !!",
    "794764": "Hi Jung, thank you for your reply! And thank you for your kind words, your ability to encourage with kind feedback is unique!",
    "794995": "samusram Nice work! It looks like you calculate the embeddings for all the fake videos as well. Is that right?\n\nMany people noticed that most actors were contained within a single folder, so splitting folders is almost as good as clustering the faces. However, clustering introduces a lot of noise (embedding failure) and the folders are quite safe. There are only a few prolific actors that appear in several folders, as you can see in the discussion and also in my public kernel. \n\nTherefore, we simply grouped the folders that had those common actors into 5 folds, i.e. keeping each main prolific actor within 1 single fold, when possible:\n\n```\nactor1 = [2, 7, 13, 24, 25, 27, 39, 47]  # folders\nactor2 = [6, 7, 10, 11, 17, 23, 37, 38]\nactor3 = [5, 16, 17, 18, 30, 35, 41]\nactor4 = [14, 26, 47]\nactor5 = [16, 17, 26, 28, 34]\n\n## proposed\nactor_split = [\n\t[0,  1,  2,  7, 13, 24, 25, 27, 39, 47],\n\t[3,  6, 10, 11, 17, 23, 33, 37, 38],\n\t[4,  5, 16, 18, 30, 32, 35, 41, 43],\n\t[8, 12, 15, 20, 22, 31, 36, 40, 44],\n\t[9, 14, 19, 21, 26, 28, 29, 34, 42],\n]\ntest = [45, 46, 48, 49]\n```",
    "795021": "hmendonca Thank you for the details! I should have thought of your elegant and more straightforward splitting strategy :)\n\n&gt; It looks like you calculate the embeddings for all the fake videos as well. Is that right?\n\nYes, Henrique. \n\nBased on a brief exploration it seems that unlike authentic videos, which usually have the same person being present in a single chunk except the corner cases you've kindly described, for the fake faces it doesn't hold. In the examples shared above I print the video chunk of the fake. You can see that fake videos with the same face are distributed over multiple chunks. It seems like the fake videos might be usually in the same folder as the original one.",
    "795202": "That's very interesting. Thanks.\n\nAnd yes, the fakes are always accompanied by their originals within the same folder, but no idea about the target actors (the inserted fake faces). Perhaps that was the leakage that we kept seeing, that caused the massive valid/LB gap!",
    "795227": "Henrique, what was your validation/LB gap if you don't mind me asking? I had 0.05 - 0.08",
    "796346": "In my personal experiments, gaps can be as big as 0.1 or even 0.2 &gt;__&lt;"
  },
  "source": "meta"
}