{
  "id": 130155,
  "title": "wondering what program you use to go over the images? Linux",
  "url": "/competitions/deepfake-detection-challenge/discussion/130155",
  "author_name": "",
  "post_date": "2020-02-12T13:10:18.992550300Z",
  "votes": null,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I have over a million faces using linux, but I haven't found a program to clean the data... I am stuck here, the code I used to get the faces it has many false positive and other faces I don't think they will be good for training.... what program you guys used to go over all images?</p>",
  "messages": [
    {
      "id": "743992",
      "postDate": "02/12/2020 13:10:18",
      "content": "<p>I have over a million faces using linux, but I haven't found a program to clean the data... I am stuck here, the code I used to get the faces it has many false positive and other faces I don't think they will be good for training.... what program you guys used to go over all images?</p>",
      "rawMarkdown": "I have over a million faces using linux, but I haven't found a program to clean the data... I am stuck here, the code I used to get the faces it has many false positive and other faces I don't think they will be good for training.... what program you guys used to go over all images?",
      "votes": null
    },
    {
      "id": "744025",
      "postDate": "02/12/2020 13:40:30",
      "content": "<p>I use my Linux box over ssh and it doesn't run a GUI, so instead I created some Python code that outputs a bunch of HTML files. Then I use twisted as a webserver and view the HTML files via a web browser. Works quite well and you can do all kinds of interesting things, such as making a page for all face crops that are relatively large but have a low score (most likely false positives), etc.</p>",
      "rawMarkdown": "I use my Linux box over ssh and it doesn't run a GUI, so instead I created some Python code that outputs a bunch of HTML files. Then I use twisted as a webserver and view the HTML files via a web browser. Works quite well and you can do all kinds of interesting things, such as making a page for all face crops that are relatively large but have a low score (most likely false positives), etc.",
      "votes": null
    },
    {
      "id": "744090",
      "postDate": "02/12/2020 14:50:16",
      "content": "<p>Hi there! thanks for the response! good idea! but I will have to see how it works... but you gave me an idea, use windows in my other machine connect to my linux box and use a program to parse imges. Any other ideas are welcome.</p>",
      "rawMarkdown": "Hi there! thanks for the response! good idea! but I will have to see how it works... but you gave me an idea, use windows in my other machine connect to my linux box and use a program to parse imges. Any other ideas are welcome.",
      "votes": null
    },
    {
      "id": "745140",
      "postDate": "02/13/2020 14:44:07",
      "content": "<p>I found KPhotoAlbun, loads all the 1M photos no problem takes hours to load but works.</p>",
      "rawMarkdown": "I found KPhotoAlbun, loads all the 1M photos no problem takes hours to load but works.",
      "votes": null
    },
    {
      "id": "745414",
      "postDate": "02/13/2020 19:49:12",
      "content": "<p>Any FS / OS will struggle past 10K files within a single directory ... try to create sub directories (e.g /train/part1/videoID/frameID.jpg )</p>",
      "rawMarkdown": "Any FS / OS will struggle past 10K files within a single directory ... try to create sub directories (e.g /train/part1/videoID/frameID.jpg )",
      "votes": null
    },
    {
      "id": "745436",
      "postDate": "02/13/2020 20:10:12",
      "content": "<blockquote>\n  <p>Any FS / OS will struggle past 10K files within a single directory</p>\n</blockquote>\n\n<p>That may have been true for older file systems, but I'd be surprised if this was still an issue on a modern OS.</p>",
      "rawMarkdown": "&gt; Any FS / OS will struggle past 10K files within a single directory\n\nThat may have been true for older file systems, but I'd be surprised if this was still an issue on a modern OS.",
      "votes": null
    },
    {
      "id": "745462",
      "postDate": "02/13/2020 20:48:33",
      "content": "<p><a href=\"/ma7moud\">@ma7moud</a>, Good point. But, I have currently 900K faces in one train directory (although I use only a subset due to data balancing) from which my data loader loads during training. I haven't really experimented if that has reduced my read/write performance.</p>",
      "rawMarkdown": "ma7moud, Good point. But, I have currently 900K faces in one train directory (although I use only a subset due to data balancing) from which my data loader loads during training. I haven't really experimented if that has reduced my read/write performance.",
      "votes": null
    },
    {
      "id": "745962",
      "postDate": "02/14/2020 12:35:41",
      "content": "<p>I see this behaviour on Ubuntu 18.04 , windows 10, macos catalina and even on cloud virtualized storage like GCS buckets / Google Drive  ... there are actually several issues open on repos for Google Drive and GCS connector apis complaining about this (like I/O errors and what not) ... The problem is rather not in the code base of the OS itself but rather enforced timeouts for UI for each of those methods to get back results so it times out and throws an IO error of some sort (and those are even present for native OS tools ... like file explorers and what not) ... I'm totally guessing here but these problem might stem from UI rendering frameworks forcing thread timeouts ... but like <a href=\"/humananalog\">@humananalog</a> mentioned ... no UI no Problem ... I use ssh then run python with glob / os.walk and it works fine for larger directories.</p>\n\n<p>TBH I'm considering switching from raw files to <em>insert your framework of choice name here</em> records ... because loading multiple files has its own overhead (plus decoding of file format).</p>",
      "rawMarkdown": "I see this behaviour on Ubuntu 18.04 , windows 10, macos catalina and even on cloud virtualized storage like GCS buckets / Google Drive  ... there are actually several issues open on repos for Google Drive and GCS connector apis complaining about this (like I/O errors and what not) ... The problem is rather not in the code base of the OS itself but rather enforced timeouts for UI for each of those methods to get back results so it times out and throws an IO error of some sort (and those are even present for native OS tools ... like file explorers and what not) ... I'm totally guessing here but these problem might stem from UI rendering frameworks forcing thread timeouts ... but like @humananalog mentioned ... no UI no Problem ... I use ssh then run python with glob / os.walk and it works fine for larger directories.\n\nTBH I'm considering switching from raw files to *insert your framework of choice name here* records ... because loading multiple files has its own overhead (plus decoding of file format).",
      "votes": null
    },
    {
      "id": "746377",
      "postDate": "02/14/2020 23:07:37",
      "content": "<p>I was able to load 1Million faces with KPhotoAlbum no problem, it just takes couple of hours to load all images.</p>",
      "rawMarkdown": "I was able to load 1Million faces with KPhotoAlbum no problem, it just takes couple of hours to load all images.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 744025,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "02/12/2020 13:40:30",
      "content": "<p>I use my Linux box over ssh and it doesn't run a GUI, so instead I created some Python code that outputs a bunch of HTML files. Then I use twisted as a webserver and view the HTML files via a web browser. Works quite well and you can do all kinds of interesting things, such as making a page for all face crops that are relatively large but have a low score (most likely false positives), etc.</p>",
      "votes": null,
      "replies": [
        {
          "id": 744090,
          "author_name": "oscarrangel",
          "author_url": "",
          "post_date": "02/12/2020 14:50:16",
          "content": "<p>Hi there! thanks for the response! good idea! but I will have to see how it works... but you gave me an idea, use windows in my other machine connect to my linux box and use a program to parse imges. Any other ideas are welcome.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 745140,
      "author_name": "oscarrangel",
      "author_url": "",
      "post_date": "02/13/2020 14:44:07",
      "content": "<p>I found KPhotoAlbun, loads all the 1M photos no problem takes hours to load but works.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 745414,
      "author_name": "ma7moud",
      "author_url": "",
      "post_date": "02/13/2020 19:49:12",
      "content": "<p>Any FS / OS will struggle past 10K files within a single directory ... try to create sub directories (e.g /train/part1/videoID/frameID.jpg )</p>",
      "votes": null,
      "replies": [
        {
          "id": 745436,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "02/13/2020 20:10:12",
          "content": "<blockquote>\n  <p>Any FS / OS will struggle past 10K files within a single directory</p>\n</blockquote>\n\n<p>That may have been true for older file systems, but I'd be surprised if this was still an issue on a modern OS.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745462,
          "author_name": "debanga",
          "author_url": "",
          "post_date": "02/13/2020 20:48:33",
          "content": "<p><a href=\"/ma7moud\">@ma7moud</a>, Good point. But, I have currently 900K faces in one train directory (although I use only a subset due to data balancing) from which my data loader loads during training. I haven't really experimented if that has reduced my read/write performance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745962,
          "author_name": "ma7moud",
          "author_url": "",
          "post_date": "02/14/2020 12:35:41",
          "content": "<p>I see this behaviour on Ubuntu 18.04 , windows 10, macos catalina and even on cloud virtualized storage like GCS buckets / Google Drive  ... there are actually several issues open on repos for Google Drive and GCS connector apis complaining about this (like I/O errors and what not) ... The problem is rather not in the code base of the OS itself but rather enforced timeouts for UI for each of those methods to get back results so it times out and throws an IO error of some sort (and those are even present for native OS tools ... like file explorers and what not) ... I'm totally guessing here but these problem might stem from UI rendering frameworks forcing thread timeouts ... but like <a href=\"/humananalog\">@humananalog</a> mentioned ... no UI no Problem ... I use ssh then run python with glob / os.walk and it works fine for larger directories.</p>\n\n<p>TBH I'm considering switching from raw files to <em>insert your framework of choice name here</em> records ... because loading multiple files has its own overhead (plus decoding of file format).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746377,
          "author_name": "oscarrangel",
          "author_url": "",
          "post_date": "02/14/2020 23:07:37",
          "content": "<p>I was able to load 1Million faces with KPhotoAlbum no problem, it just takes couple of hours to load all images.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "743992": "I have over a million faces using linux, but I haven't found a program to clean the data... I am stuck here, the code I used to get the faces it has many false positive and other faces I don't think they will be good for training.... what program you guys used to go over all images?",
    "744025": "I use my Linux box over ssh and it doesn't run a GUI, so instead I created some Python code that outputs a bunch of HTML files. Then I use twisted as a webserver and view the HTML files via a web browser. Works quite well and you can do all kinds of interesting things, such as making a page for all face crops that are relatively large but have a low score (most likely false positives), etc.",
    "744090": "Hi there! thanks for the response! good idea! but I will have to see how it works... but you gave me an idea, use windows in my other machine connect to my linux box and use a program to parse imges. Any other ideas are welcome.",
    "745140": "I found KPhotoAlbun, loads all the 1M photos no problem takes hours to load but works.",
    "745414": "Any FS / OS will struggle past 10K files within a single directory ... try to create sub directories (e.g /train/part1/videoID/frameID.jpg )",
    "745436": "&gt; Any FS / OS will struggle past 10K files within a single directory\n\nThat may have been true for older file systems, but I'd be surprised if this was still an issue on a modern OS.",
    "745462": "ma7moud, Good point. But, I have currently 900K faces in one train directory (although I use only a subset due to data balancing) from which my data loader loads during training. I haven't really experimented if that has reduced my read/write performance.",
    "745962": "I see this behaviour on Ubuntu 18.04 , windows 10, macos catalina and even on cloud virtualized storage like GCS buckets / Google Drive  ... there are actually several issues open on repos for Google Drive and GCS connector apis complaining about this (like I/O errors and what not) ... The problem is rather not in the code base of the OS itself but rather enforced timeouts for UI for each of those methods to get back results so it times out and throws an IO error of some sort (and those are even present for native OS tools ... like file explorers and what not) ... I'm totally guessing here but these problem might stem from UI rendering frameworks forcing thread timeouts ... but like @humananalog mentioned ... no UI no Problem ... I use ssh then run python with glob / os.walk and it works fine for larger directories.\n\nTBH I'm considering switching from raw files to *insert your framework of choice name here* records ... because loading multiple files has its own overhead (plus decoding of file format).",
    "746377": "I was able to load 1Million faces with KPhotoAlbum no problem, it just takes couple of hours to load all images."
  },
  "source": "meta"
}