{
  "id": 121355,
  "title": "Do we have more options aside cloud computing?",
  "url": "/competitions/deepfake-detection-challenge/discussion/121355",
  "author_name": "",
  "post_date": "2019-12-12T16:33:58.561608900Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>People is talking about the huge hardware requirements needed in order to get some decent results in this competition. The only option seems to be cloud computing but, what if I have a decent PC at home? Of course I won't be able to train super large models (mainly because of RAM and VRAM limitations) but maybe I can get some decent results (top 3-5%) using \"simpler\" pipelines.</p>\n\n<p>For me a decent PC is something with &gt;= intel i7 or equivalent, &gt;= RTX 2080 SUPER and 32-64 GB of RAM.</p>",
  "messages": [
    {
      "id": "693659",
      "postDate": "12/12/2019 16:33:58",
      "content": "<p>People is talking about the huge hardware requirements needed in order to get some decent results in this competition. The only option seems to be cloud computing but, what if I have a decent PC at home? Of course I won't be able to train super large models (mainly because of RAM and VRAM limitations) but maybe I can get some decent results (top 3-5%) using \"simpler\" pipelines.</p>\n\n<p>For me a decent PC is something with &gt;= intel i7 or equivalent, &gt;= RTX 2080 SUPER and 32-64 GB of RAM.</p>",
      "rawMarkdown": "People is talking about the huge hardware requirements needed in order to get some decent results in this competition. The only option seems to be cloud computing but, what if I have a decent PC at home? Of course I won't be able to train super large models (mainly because of RAM and VRAM limitations) but maybe I can get some decent results (top 3-5%) using \"simpler\" pipelines.\n\nFor me a decent PC is something with &gt;= intel i7 or equivalent, &gt;= RTX 2080 SUPER and 32-64 GB of RAM.",
      "votes": null
    },
    {
      "id": "693794",
      "postDate": "12/12/2019 19:46:02",
      "content": "<p>I trained loads of models on single GPUs even in clusters.</p>\n\n<p>Training on your PC at home will result in the PC being pretty slow to handle during the training. Considering the amount of data and I/O needed this will be at least hours at a time. (Mine basically froze due to 100% of the GPU and a lot of CPU being in use.)</p>\n\n<p>You'll also want to have good cooling and possibly a PSU or enough space for model snapshots. In addition make sure your power supply is sufficient for the continued load. I for one have a rather small PSU that will start to buckle at maximum GPU load at over extended times.</p>\n\n<p>But of course in theory, if your system is set up well and you have time where you don't need it, your home PC is fine.</p>\n\n<p>An idea is also to subsample the real data, in my meta-data analysis I had a quick look and it seems real and fake labels are a 20/80 split. Essentially you could fathom just not using part of the 80% to balance your problem and reduce data load.</p>",
      "rawMarkdown": "I trained loads of models on single GPUs even in clusters.\n\nTraining on your PC at home will result in the PC being pretty slow to handle during the training. Considering the amount of data and I/O needed this will be at least hours at a time. (Mine basically froze due to 100% of the GPU and a lot of CPU being in use.)\n\nYou'll also want to have good cooling and possibly a PSU or enough space for model snapshots. In addition make sure your power supply is sufficient for the continued load. I for one have a rather small PSU that will start to buckle at maximum GPU load at over extended times.\n\nBut of course in theory, if your system is set up well and you have time where you don't need it, your home PC is fine.\n\nAn idea is also to subsample the real data, in my meta-data analysis I had a quick look and it seems real and fake labels are a 20/80 split. Essentially you could fathom just not using part of the 80% to balance your problem and reduce data load.",
      "votes": null
    },
    {
      "id": "693843",
      "postDate": "12/12/2019 20:56:44",
      "content": "<p>You don't need to train on all the data or have a lot of compute to compete in this competition. Have a look at this <a href=\"https://arxiv.org/pdf/1911.00686v2.pdf\">paper</a> they are able to train a good model for Deep Fake detection with few samples and a simple pipeline.</p>\n\n<p>Also, the time limits on kernel running time rule out any model that takes more than 9 hours for training and inference. </p>",
      "rawMarkdown": "You don't need to train on all the data or have a lot of compute to compete in this competition. Have a look at this [paper](https://arxiv.org/pdf/1911.00686v2.pdf) they are able to train a good model for Deep Fake detection with few samples and a simple pipeline.\n\nAlso, the time limits on kernel running time rule out any model that takes more than 9 hours for training and inference.",
      "votes": null
    },
    {
      "id": "694912",
      "postDate": "12/14/2019 10:45:03",
      "content": "<p>Interesting paper. As you say, in this competition is important to keep a good level of efficiency in terms of model and data size and time.</p>\n\n<p>P.D: Are you sure about that 9 hours limit? I thought the limit was only for inference given that you can train the model wherever you want.</p>",
      "rawMarkdown": "Interesting paper. As you say, in this competition is important to keep a good level of efficiency in terms of model and data size and time.\n\nP.D: Are you sure about that 9 hours limit? I thought the limit was only for inference given that you can train the model wherever you want.",
      "votes": null
    },
    {
      "id": "694916",
      "postDate": "12/14/2019 10:49:43",
      "content": "<p>I think the key for using a PC of these characteristics in this competition is to handle memory (RAM, VRAM and disk) in a really efficient way. Expect a nice electricity bill too :)</p>",
      "rawMarkdown": "I think the key for using a PC of these characteristics in this competition is to handle memory (RAM, VRAM and disk) in a really efficient way. Expect a nice electricity bill too :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 693794,
      "author_name": "jesperdramsch",
      "author_url": "",
      "post_date": "12/12/2019 19:46:02",
      "content": "<p>I trained loads of models on single GPUs even in clusters.</p>\n\n<p>Training on your PC at home will result in the PC being pretty slow to handle during the training. Considering the amount of data and I/O needed this will be at least hours at a time. (Mine basically froze due to 100% of the GPU and a lot of CPU being in use.)</p>\n\n<p>You'll also want to have good cooling and possibly a PSU or enough space for model snapshots. In addition make sure your power supply is sufficient for the continued load. I for one have a rather small PSU that will start to buckle at maximum GPU load at over extended times.</p>\n\n<p>But of course in theory, if your system is set up well and you have time where you don't need it, your home PC is fine.</p>\n\n<p>An idea is also to subsample the real data, in my meta-data analysis I had a quick look and it seems real and fake labels are a 20/80 split. Essentially you could fathom just not using part of the 80% to balance your problem and reduce data load.</p>",
      "votes": null,
      "replies": [
        {
          "id": 694916,
          "author_name": "kewontong",
          "author_url": "",
          "post_date": "12/14/2019 10:49:43",
          "content": "<p>I think the key for using a PC of these characteristics in this competition is to handle memory (RAM, VRAM and disk) in a really efficient way. Expect a nice electricity bill too :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 693843,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "12/12/2019 20:56:44",
      "content": "<p>You don't need to train on all the data or have a lot of compute to compete in this competition. Have a look at this <a href=\"https://arxiv.org/pdf/1911.00686v2.pdf\">paper</a> they are able to train a good model for Deep Fake detection with few samples and a simple pipeline.</p>\n\n<p>Also, the time limits on kernel running time rule out any model that takes more than 9 hours for training and inference. </p>",
      "votes": null,
      "replies": [
        {
          "id": 694912,
          "author_name": "kewontong",
          "author_url": "",
          "post_date": "12/14/2019 10:45:03",
          "content": "<p>Interesting paper. As you say, in this competition is important to keep a good level of efficiency in terms of model and data size and time.</p>\n\n<p>P.D: Are you sure about that 9 hours limit? I thought the limit was only for inference given that you can train the model wherever you want.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "693659": "People is talking about the huge hardware requirements needed in order to get some decent results in this competition. The only option seems to be cloud computing but, what if I have a decent PC at home? Of course I won't be able to train super large models (mainly because of RAM and VRAM limitations) but maybe I can get some decent results (top 3-5%) using \"simpler\" pipelines.\n\nFor me a decent PC is something with &gt;= intel i7 or equivalent, &gt;= RTX 2080 SUPER and 32-64 GB of RAM.",
    "693794": "I trained loads of models on single GPUs even in clusters.\n\nTraining on your PC at home will result in the PC being pretty slow to handle during the training. Considering the amount of data and I/O needed this will be at least hours at a time. (Mine basically froze due to 100% of the GPU and a lot of CPU being in use.)\n\nYou'll also want to have good cooling and possibly a PSU or enough space for model snapshots. In addition make sure your power supply is sufficient for the continued load. I for one have a rather small PSU that will start to buckle at maximum GPU load at over extended times.\n\nBut of course in theory, if your system is set up well and you have time where you don't need it, your home PC is fine.\n\nAn idea is also to subsample the real data, in my meta-data analysis I had a quick look and it seems real and fake labels are a 20/80 split. Essentially you could fathom just not using part of the 80% to balance your problem and reduce data load.",
    "693843": "You don't need to train on all the data or have a lot of compute to compete in this competition. Have a look at this [paper](https://arxiv.org/pdf/1911.00686v2.pdf) they are able to train a good model for Deep Fake detection with few samples and a simple pipeline.\n\nAlso, the time limits on kernel running time rule out any model that takes more than 9 hours for training and inference.",
    "694912": "Interesting paper. As you say, in this competition is important to keep a good level of efficiency in terms of model and data size and time.\n\nP.D: Are you sure about that 9 hours limit? I thought the limit was only for inference given that you can train the model wherever you want.",
    "694916": "I think the key for using a PC of these characteristics in this competition is to handle memory (RAM, VRAM and disk) in a really efficient way. Expect a nice electricity bill too :)"
  },
  "source": "meta"
}