{
  "id": 132437,
  "title": "Clustering-based Validation Scheme",
  "url": "/competitions/deepfake-detection-challenge/discussion/132437",
  "author_name": "",
  "post_date": "2020-02-26T00:01:05.024581Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Anyone recently successful with clustering-based validation set generation? This notebook is exciting: <a href=\"https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda\">https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda</a>, to start with, but if someone can confirm then it would be really helpful, before seriously putting some time on it!</p>\n\n<p>P.S. I'm currently using folder-based approach and getting OK results. But, as the competition is getting closer to an end, I don't want to miss out on any opportunity to try different approaches :D</p>",
  "messages": [
    {
      "id": "756673",
      "postDate": "02/26/2020 00:01:05",
      "content": "<p>Anyone recently successful with clustering-based validation set generation? This notebook is exciting: <a href=\"https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda\">https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda</a>, to start with, but if someone can confirm then it would be really helpful, before seriously putting some time on it!</p>\n\n<p>P.S. I'm currently using folder-based approach and getting OK results. But, as the competition is getting closer to an end, I don't want to miss out on any opportunity to try different approaches :D</p>",
      "rawMarkdown": "Anyone recently successful with clustering-based validation set generation? This notebook is exciting: https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda, to start with, but if someone can confirm then it would be really helpful, before seriously putting some time on it!\n\nP.S. I'm currently using folder-based approach and getting OK results. But, as the competition is getting closer to an end, I don't want to miss out on any opportunity to try different approaches :D",
      "votes": null
    },
    {
      "id": "756678",
      "postDate": "02/26/2020 00:06:01",
      "content": "<p>Good question. I've been following the comments on this and I don't see any real evidence that it's been put to good use. I'm wondering now about using the 0 chunk for cv.</p>",
      "rawMarkdown": "Good question. I've been following the comments on this and I don't see any real evidence that it's been put to good use. I'm wondering now about using the 0 chunk for cv.",
      "votes": null
    },
    {
      "id": "756701",
      "postDate": "02/26/2020 01:13:54",
      "content": "<p>If we visualize the clusters by folder id, may be we can get some insights.</p>",
      "rawMarkdown": "If we visualize the clusters by folder id, may be we can get some insights.",
      "votes": null
    },
    {
      "id": "756884",
      "postDate": "02/26/2020 07:01:43",
      "content": "<p>Haven't tried clustering but I have another thought. Instead of training and inference across all face poses  have you tried just limiting the approach to certain poses i.e. front pose only? In my mind that would reduce the variation and  bias that can exist across different poses. It does however assume that videos in the LB have frames with front poses (which I assume would be the case to make a deep fake).</p>",
      "rawMarkdown": "Haven't tried clustering but I have another thought. Instead of training and inference across all face poses  have you tried just limiting the approach to certain poses i.e. front pose only? In my mind that would reduce the variation and  bias that can exist across different poses. It does however assume that videos in the LB have frames with front poses (which I assume would be the case to make a deep fake).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 756678,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "02/26/2020 00:06:01",
      "content": "<p>Good question. I've been following the comments on this and I don't see any real evidence that it's been put to good use. I'm wondering now about using the 0 chunk for cv.</p>",
      "votes": null,
      "replies": [
        {
          "id": 756701,
          "author_name": "debanga",
          "author_url": "",
          "post_date": "02/26/2020 01:13:54",
          "content": "<p>If we visualize the clusters by folder id, may be we can get some insights.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 756884,
      "author_name": "maralski",
      "author_url": "",
      "post_date": "02/26/2020 07:01:43",
      "content": "<p>Haven't tried clustering but I have another thought. Instead of training and inference across all face poses  have you tried just limiting the approach to certain poses i.e. front pose only? In my mind that would reduce the variation and  bias that can exist across different poses. It does however assume that videos in the LB have frames with front poses (which I assume would be the case to make a deep fake).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "756673": "Anyone recently successful with clustering-based validation set generation? This notebook is exciting: https://www.kaggle.com/hmendonca/proper-clustering-with-facenet-embeddings-eda, to start with, but if someone can confirm then it would be really helpful, before seriously putting some time on it!\n\nP.S. I'm currently using folder-based approach and getting OK results. But, as the competition is getting closer to an end, I don't want to miss out on any opportunity to try different approaches :D",
    "756678": "Good question. I've been following the comments on this and I don't see any real evidence that it's been put to good use. I'm wondering now about using the 0 chunk for cv.",
    "756701": "If we visualize the clusters by folder id, may be we can get some insights.",
    "756884": "Haven't tried clustering but I have another thought. Instead of training and inference across all face poses  have you tried just limiting the approach to certain poses i.e. front pose only? In my mind that would reduce the variation and  bias that can exist across different poses. It does however assume that videos in the LB have frames with front poses (which I assume would be the case to make a deep fake)."
  },
  "source": "meta"
}