{
  "id": 129832,
  "title": "All actors",
  "url": "/competitions/deepfake-detection-challenge/discussion/129832",
  "author_name": "nosound",
  "post_date": "2020-02-11T00:27:22.729000",
  "votes": 66,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Somehow I finished up selecting a single image per actor, which I want to share with you (only real videos). Overall, my clustering effort resulted in 437 different people in the train videos, all of them in one picture below. </p>\n\n<p>As I grew increasingly bored by that task I started to select more funny pictures, instead of beautiful ones. Some of them are hilarious :)</p>\n\n<p>Because the clustering procedure is not perfect, my estimation is that there are about 10-20 pairs which are actually the same person, and about 5 actors who are not appearing in this family shot.</p>\n\n<p>Using this opportunity I want to say thanks to this beautiful bunch of people for the important work they did.</p>\n\n<p><img src=\"https://i.imgur.com/X77RVwJ.jpg\" alt=\"actors\"></p>",
  "messages": [
    {
      "id": 741793,
      "postDate": "2020-02-11T00:27:22.730Z",
      "content": "<p>Somehow I finished up selecting a single image per actor, which I want to share with you (only real videos). Overall, my clustering effort resulted in 437 different people in the train videos, all of them in one picture below. </p>\n\n<p>As I grew increasingly bored by that task I started to select more funny pictures, instead of beautiful ones. Some of them are hilarious :)</p>\n\n<p>Because the clustering procedure is not perfect, my estimation is that there are about 10-20 pairs which are actually the same person, and about 5 actors who are not appearing in this family shot.</p>\n\n<p>Using this opportunity I want to say thanks to this beautiful bunch of people for the important work they did.</p>\n\n<p><img src=\"https://i.imgur.com/X77RVwJ.jpg\" alt=\"actors\"></p>",
      "rawMarkdown": "Somehow I finished up selecting a single image per actor, which I want to share with you (only real videos). Overall, my clustering effort resulted in 437 different people in the train videos, all of them in one picture below. \n\nAs I grew increasingly bored by that task I started to select more funny pictures, instead of beautiful ones. Some of them are hilarious :)\n\nBecause the clustering procedure is not perfect, my estimation is that there are about 10-20 pairs which are actually the same person, and about 5 actors who are not appearing in this family shot.\n\nUsing this opportunity I want to say thanks to this beautiful bunch of people for the important work they did.\n\n![actors](https://i.imgur.com/X77RVwJ.jpg)",
      "votes": 66
    },
    {
      "id": 741842,
      "postDate": "2020-02-11T01:00:06.753Z",
      "content": "<p>Did your research support the rumour that using a single directory is a good way to make a CV?</p>",
      "rawMarkdown": "Did your research support the rumour that using a single directory is a good way to make a CV?",
      "votes": 5,
      "replies": [
        {
          "id": 742472,
          "postDate": "2020-02-11T10:07:36.937Z",
          "content": "<p>Yes and no, if I remember correctly about 70% of the actors have a single folder home, but the rest are distributed across 2-5 folders. That is why I started this clustering,- I was not satisfied by splitting by folders.</p>",
          "rawMarkdown": "Yes and no, if I remember correctly about 70% of the actors have a single folder home, but the rest are distributed across 2-5 folders. That is why I started this clustering,- I was not satisfied by splitting by folders.",
          "votes": 4
        }
      ]
    },
    {
      "id": 742916,
      "postDate": "2020-02-11T16:34:16.560Z",
      "content": "<p>Amazing job <a href=\"/zaharch\">@zaharch</a> !! 👍 \nDid you try UMAP?</p>",
      "rawMarkdown": "Amazing job @zaharch !! 👍 \nDid you try UMAP?",
      "votes": 3,
      "replies": [
        {
          "id": 742941,
          "postDate": "2020-02-11T16:49:59.773Z",
          "content": "<p>Thank you! No, I didn't, is it good? But part of my clustering procedure is t-SNE which fulfills a similar role.</p>",
          "rawMarkdown": "Thank you! No, I didn't, is it good? But part of my clustering procedure is t-SNE which fulfills a similar role.",
          "votes": 1
        }
      ]
    },
    {
      "id": 742411,
      "postDate": "2020-02-11T09:31:15.400Z",
      "content": "<p>Have you tried doing a CV split based on the actors? Also, any chance you have a csv file with a corresponding actor to each original video :) ? Good job btw, have you done any manual work to get this result or was all machine learning based?</p>",
      "rawMarkdown": "Have you tried doing a CV split based on the actors? Also, any chance you have a csv file with a corresponding actor to each original video :) ? Good job btw, have you done any manual work to get this result or was all machine learning based?",
      "votes": 1,
      "replies": [
        {
          "id": 742464,
          "postDate": "2020-02-11T10:03:08.017Z",
          "content": "<p>The clustering I did for the train/valid splitting, yes. Not sure yet how helpful it is. I reviewed all clusters manually, separating some and merging others, but I did it fast so there are some un-merged clusters still present.</p>\n\n<p>Regarding sharing the csv, I considered it but decided against, as it is a significant amount of work and I am participating competitively.</p>",
          "rawMarkdown": "The clustering I did for the train/valid splitting, yes. Not sure yet how helpful it is. I reviewed all clusters manually, separating some and merging others, but I did it fast so there are some un-merged clusters still present.\n\nRegarding sharing the csv, I considered it but decided against, as it is a significant amount of work and I am participating competitively.",
          "votes": 4
        },
        {
          "id": 742534,
          "postDate": "2020-02-11T11:08:03.400Z",
          "content": "<p>Fair enough, I was doing something similar, but every step that I improved the clusters the CV result still didn't track well enough the LB score. In fact it didn't have any advantage over splitting by folders,  so I just gave up eventually.</p>\n\n<p>Having a good validation set was extremely important for me as I was saving the weights for the model under the best CV score, so I spent quite some time doing that but the results were not helpful, let me know if you have better luck then me ;)</p>",
          "rawMarkdown": "Fair enough, I was doing something similar, but every step that I improved the clusters the CV result still didn't track well enough the LB score. In fact it didn't have any advantage over splitting by folders,  so I just gave up eventually.\n\nHaving a good validation set was extremely important for me as I was saving the weights for the model under the best CV score, so I spent quite some time doing that but the results were not helpful, let me know if you have better luck then me ;)",
          "votes": 3
        }
      ]
    },
    {
      "id": 742044,
      "postDate": "2020-02-11T03:43:05.607Z",
      "content": "<p>You are always amazing, <a href=\"/zaharch\">@zaharch</a> </p>",
      "rawMarkdown": "You are always amazing, @zaharch ",
      "votes": 1,
      "replies": [
        {
          "id": 742454,
          "postDate": "2020-02-11T09:56:07.753Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!",
          "votes": 1
        }
      ]
    },
    {
      "id": 759678,
      "postDate": "2020-02-29T10:27:50.383Z",
      "content": "<p>Great work! <a href=\"/zaharch\">@zaharch</a> Any chance you have the mapping clustered actor -&gt; videos as a metadata file?</p>",
      "rawMarkdown": "Great work! @zaharch Any chance you have the mapping clustered actor -&gt; videos as a metadata file?",
      "replies": [
        {
          "id": 759697,
          "postDate": "2020-02-29T11:01:05.287Z",
          "content": "<p>Thank you! Regarding the sharing, please see my another comment in this thread.</p>",
          "rawMarkdown": "Thank you! Regarding the sharing, please see my another comment in this thread.",
          "votes": 2
        },
        {
          "id": 759704,
          "postDate": "2020-02-29T11:07:23.960Z",
          "content": "<p>Just read your other comment. Totally fair indeed. Will try to make my own as well. Best of luck! </p>",
          "rawMarkdown": "Just read your other comment. Totally fair indeed. Will try to make my own as well. Best of luck! ",
          "votes": 1
        },
        {
          "id": 760047,
          "postDate": "2020-02-29T19:17:07.503Z",
          "content": "<p>Sure, but please note that value from it is probably very limited. For example, I don't think you need to invest your time in it if you are above LB 0.3.</p>",
          "rawMarkdown": "Sure, but please note that value from it is probably very limited. For example, I don't think you need to invest your time in it if you are above LB 0.3.",
          "votes": 1
        },
        {
          "id": 760067,
          "postDate": "2020-02-29T20:01:22.647Z",
          "content": "<p>I am not there yet (just getting started), thanks for the tip! :) </p>",
          "rawMarkdown": "I am not there yet (just getting started), thanks for the tip! :) "
        }
      ]
    },
    {
      "id": 748024,
      "postDate": "2020-02-17T05:48:34.417Z",
      "content": "<p>Thanks for sharing this analysis. Do you know if the same actors will be present in the test dataset or the same ones?</p>",
      "rawMarkdown": "Thanks for sharing this analysis. Do you know if the same actors will be present in the test dataset or the same ones?"
    },
    {
      "id": 742056,
      "postDate": "2020-02-11T03:56:17.273Z",
      "content": "<p>I wonder what that black screen supposed to mean🤔 😜 </p>",
      "rawMarkdown": "I wonder what that black screen supposed to mean🤔 😜 ",
      "replies": [
        {
          "id": 742152,
          "postDate": "2020-02-11T05:19:42.020Z",
          "content": "<p>Great work btw.</p>",
          "rawMarkdown": "Great work btw."
        },
        {
          "id": 742453,
          "postDate": "2020-02-11T09:55:50.837Z",
          "content": "<p>It is two different people in a (very) dark rooms. If you brighten the pictures you will see them. </p>",
          "rawMarkdown": "It is two different people in a (very) dark rooms. If you brighten the pictures you will see them. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 743673,
      "postDate": "2020-02-12T07:21:04.567Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 741797,
      "postDate": "2020-02-11T00:30:04.767Z",
      "content": "<p>Thanks for sharing!!!</p>",
      "rawMarkdown": "Thanks for sharing!!!",
      "votes": 1
    },
    {
      "id": 743690,
      "postDate": "2020-02-12T07:40:28.867Z",
      "content": "<p>WOW, you are AMAZING!! Thank you.</p>",
      "rawMarkdown": "WOW, you are AMAZING!! Thank you."
    }
  ],
  "comments": [
    {
      "id": 741842,
      "author_name": "pete",
      "author_url": "",
      "post_date": "2020-02-11T01:00:06.753000",
      "content": "<p>Did your research support the rumour that using a single directory is a good way to make a CV?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 742472,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-02-11T10:07:36.937000",
          "content": "<p>Yes and no, if I remember correctly about 70% of the actors have a single folder home, but the rest are distributed across 2-5 folders. That is why I started this clustering,- I was not satisfied by splitting by folders.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 742916,
      "author_name": "Henrique Mendonça",
      "author_url": "",
      "post_date": "2020-02-11T16:34:16.560000",
      "content": "<p>Amazing job <a href=\"/zaharch\">@zaharch</a> !! 👍 \nDid you try UMAP?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 742941,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-02-11T16:49:59.773000",
          "content": "<p>Thank you! No, I didn't, is it good? But part of my clustering procedure is t-SNE which fulfills a similar role.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 742411,
      "author_name": "Pedro Bernardo",
      "author_url": "",
      "post_date": "2020-02-11T09:31:15.400000",
      "content": "<p>Have you tried doing a CV split based on the actors? Also, any chance you have a csv file with a corresponding actor to each original video :) ? Good job btw, have you done any manual work to get this result or was all machine learning based?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 742464,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-02-11T10:03:08.017000",
          "content": "<p>The clustering I did for the train/valid splitting, yes. Not sure yet how helpful it is. I reviewed all clusters manually, separating some and merging others, but I did it fast so there are some un-merged clusters still present.</p>\n\n<p>Regarding sharing the csv, I considered it but decided against, as it is a significant amount of work and I am participating competitively.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 742534,
          "author_name": "Pedro Bernardo",
          "author_url": "",
          "post_date": "2020-02-11T11:08:03.400000",
          "content": "<p>Fair enough, I was doing something similar, but every step that I improved the clusters the CV result still didn't track well enough the LB score. In fact it didn't have any advantage over splitting by folders,  so I just gave up eventually.</p>\n\n<p>Having a good validation set was extremely important for me as I was saving the weights for the model under the best CV score, so I spent quite some time doing that but the results were not helpful, let me know if you have better luck then me ;)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 742044,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2020-02-11T03:43:05.607000",
      "content": "<p>You are always amazing, <a href=\"/zaharch\">@zaharch</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 742454,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-02-11T09:56:07.753000",
          "content": "<p>Thank you!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 759678,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-02-29T10:27:50.383000",
      "content": "<p>Great work! <a href=\"/zaharch\">@zaharch</a> Any chance you have the mapping clustered actor -&gt; videos as a metadata file?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 759697,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-02-29T11:01:05.287000",
          "content": "<p>Thank you! Regarding the sharing, please see my another comment in this thread.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 759704,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2020-02-29T11:07:23.960000",
          "content": "<p>Just read your other comment. Totally fair indeed. Will try to make my own as well. Best of luck! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 760047,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-02-29T19:17:07.503000",
          "content": "<p>Sure, but please note that value from it is probably very limited. For example, I don't think you need to invest your time in it if you are above LB 0.3.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 760067,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2020-02-29T20:01:22.647000",
          "content": "<p>I am not there yet (just getting started), thanks for the tip! :) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 748024,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-02-17T05:48:34.417000",
      "content": "<p>Thanks for sharing this analysis. Do you know if the same actors will be present in the test dataset or the same ones?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 742056,
      "author_name": "Shangqiu Li",
      "author_url": "",
      "post_date": "2020-02-11T03:56:17.273000",
      "content": "<p>I wonder what that black screen supposed to mean🤔 😜 </p>",
      "votes": 0,
      "replies": [
        {
          "id": 742152,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-11T05:19:42.020000",
          "content": "<p>Great work btw.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 742453,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-02-11T09:55:50.837000",
          "content": "<p>It is two different people in a (very) dark rooms. If you brighten the pictures you will see them. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 743673,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-12T07:21:04.567000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 741797,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2020-02-11T00:30:04.767000",
      "content": "<p>Thanks for sharing!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 743690,
      "author_name": "Zungmann",
      "author_url": "",
      "post_date": "2020-02-12T07:40:28.867000",
      "content": "<p>WOW, you are AMAZING!! Thank you.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "741793": "Somehow I finished up selecting a single image per actor, which I want to share with you (only real videos). Overall, my clustering effort resulted in 437 different people in the train videos, all of them in one picture below. \n\nAs I grew increasingly bored by that task I started to select more funny pictures, instead of beautiful ones. Some of them are hilarious :)\n\nBecause the clustering procedure is not perfect, my estimation is that there are about 10-20 pairs which are actually the same person, and about 5 actors who are not appearing in this family shot.\n\nUsing this opportunity I want to say thanks to this beautiful bunch of people for the important work they did.\n\n![actors](https://i.imgur.com/X77RVwJ.jpg)",
    "741842": "Did your research support the rumour that using a single directory is a good way to make a CV?",
    "742916": "Amazing job @zaharch !! 👍 \nDid you try UMAP?",
    "742411": "Have you tried doing a CV split based on the actors? Also, any chance you have a csv file with a corresponding actor to each original video :) ? Good job btw, have you done any manual work to get this result or was all machine learning based?",
    "742044": "You are always amazing, @zaharch ",
    "759678": "Great work! @zaharch Any chance you have the mapping clustered actor -&gt; videos as a metadata file?",
    "748024": "Thanks for sharing this analysis. Do you know if the same actors will be present in the test dataset or the same ones?",
    "742056": "I wonder what that black screen supposed to mean🤔 😜 ",
    "743673": "",
    "741797": "Thanks for sharing!!!",
    "743690": "WOW, you are AMAZING!! Thank you."
  }
}