{
  "id": 130637,
  "title": "What infra are you using?",
  "url": "/competitions/deepfake-detection-challenge/discussion/130637",
  "author_name": "",
  "post_date": "2020-02-15T13:05:47.391725400Z",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>As mentioned in the competition instructions, it is clear that training should be done outside Kaggle kernels. \nSince I am joining this competition, I was wondering which infrastructure are you using to train your models? Are you using a cloud instance or local setup? If so, could you share some details. </p>\n\n<p>Much appreciated. :)</p>",
  "messages": [
    {
      "id": "746722",
      "postDate": "02/15/2020 13:05:47",
      "content": "<p>As mentioned in the competition instructions, it is clear that training should be done outside Kaggle kernels. \nSince I am joining this competition, I was wondering which infrastructure are you using to train your models? Are you using a cloud instance or local setup? If so, could you share some details. </p>\n\n<p>Much appreciated. :)</p>",
      "rawMarkdown": "As mentioned in the competition instructions, it is clear that training should be done outside Kaggle kernels. \nSince I am joining this competition, I was wondering which infrastructure are you using to train your models? Are you using a cloud instance or local setup? If so, could you share some details. \n\nMuch appreciated. :)",
      "votes": null
    },
    {
      "id": "746762",
      "postDate": "02/15/2020 14:05:39",
      "content": "<p>I'm using a local setup with a single RTX Titan</p>",
      "rawMarkdown": "I'm using a local setup with a single RTX Titan",
      "votes": null
    },
    {
      "id": "746765",
      "postDate": "02/15/2020 14:10:25",
      "content": "<p>Awesome setup, thanks for sharing. Do you know how a single RTX Titan compares to a single 1080Ti in performance? \nI could do a Google search but I am being lazy. ;)</p>",
      "rawMarkdown": "Awesome setup, thanks for sharing. Do you know how a single RTX Titan compares to a single 1080Ti in performance? \nI could do a Google search but I am being lazy. ;)",
      "votes": null
    },
    {
      "id": "746766",
      "postDate": "02/15/2020 14:16:09",
      "content": "<p>The RTX Titan is similar in performance to a 2080 Ti, so I would guess around double the performance, maybe a bit less?</p>\n\n<p>However, the advantage of the Titan isn't so much in speed but the 24Gb of RAM, so you can have larger batch sizes. However, in my experience in this competition, using larger/more complicated models doesn't help, they just overfit. I don't think my score would be any worse if I only had access to a 1080 Ti, I'd just have to be more patient :)</p>",
      "rawMarkdown": "The RTX Titan is similar in performance to a 2080 Ti, so I would guess around double the performance, maybe a bit less?\n\nHowever, the advantage of the Titan isn't so much in speed but the 24Gb of RAM, so you can have larger batch sizes. However, in my experience in this competition, using larger/more complicated models doesn't help, they just overfit. I don't think my score would be any worse if I only had access to a 1080 Ti, I'd just have to be more patient :)",
      "votes": null
    },
    {
      "id": "746799",
      "postDate": "02/15/2020 15:02:52",
      "content": "<p>I use a 1080 Ti. My current training set is 1M images and training a model takes about 10 - 15 hours so I do it overnight. (As <a href=\"/jamesphoward\">@jamesphoward</a> says, larger models just overfit, and the models I'm using are probably too large.)</p>",
      "rawMarkdown": "I use a 1080 Ti. My current training set is 1M images and training a model takes about 10 - 15 hours so I do it overnight. (As @jamesphoward says, larger models just overfit, and the models I'm using are probably too large.)",
      "votes": null
    },
    {
      "id": "746850",
      "postDate": "02/15/2020 16:14:50",
      "content": "<p>That's good news for me. Thanks for sharing ! </p>",
      "rawMarkdown": "That's good news for me. Thanks for sharing !",
      "votes": null
    },
    {
      "id": "746853",
      "postDate": "02/15/2020 16:16:21",
      "content": "<p>I use a 1080 ti, I have 1.6 million images to train so, I am training in very small chunks... (400k on validation set xD)</p>",
      "rawMarkdown": "I use a 1080 ti, I have 1.6 million images to train so, I am training in very small chunks... (400k on validation set xD)",
      "votes": null
    },
    {
      "id": "746958",
      "postDate": "02/15/2020 19:29:26",
      "content": "<p>I was using AWS Sagemaker first, mainly it looked like the easiest / most productive way. In the end, I only used it to do the pre-processing to extract the faces from videos. The whole then uploaded to S3. Since Sagemaker had problems with TF 2, I (was forced to) switch to plain EC2 Instances and install everything manually. The side effect benefit is also lower pricing (through spot instances). What I miss though is really the comfort of kaggle. Simply \"committing\" your work and see the nice result the day after is simply really cool. Here I have to use logging in training callbacks and such.</p>\n\n<p>Another feature which I think would be handy is AWS Lambda for mass parallel preprocessing (but it's problematic due to the image size constraint (no TF). </p>\n\n<p>I think the GCP ecosystem is a bit easier to deal wih. AWS is quite ...hardcore.</p>",
      "rawMarkdown": "I was using AWS Sagemaker first, mainly it looked like the easiest / most productive way. In the end, I only used it to do the pre-processing to extract the faces from videos. The whole then uploaded to S3. Since Sagemaker had problems with TF 2, I (was forced to) switch to plain EC2 Instances and install everything manually. The side effect benefit is also lower pricing (through spot instances). What I miss though is really the comfort of kaggle. Simply \"committing\" your work and see the nice result the day after is simply really cool. Here I have to use logging in training callbacks and such.\n\nAnother feature which I think would be handy is AWS Lambda for mass parallel preprocessing (but it's problematic due to the image size constraint (no TF). \n\nI think the GCP ecosystem is a bit easier to deal wih. AWS is quite ...hardcore.",
      "votes": null
    },
    {
      "id": "746992",
      "postDate": "02/15/2020 20:31:40",
      "content": "<p>I got my current score by preprocessing on a cloud instance(cropping face) and training/inference on kaggle kernels.\n<a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/129770\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/129770</a></p>",
      "rawMarkdown": "I got my current score by preprocessing on a cloud instance(cropping face) and training/inference on kaggle kernels.\nhttps://www.kaggle.com/c/deepfake-detection-challenge/discussion/129770",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 746762,
      "author_name": "jamesphoward",
      "author_url": "",
      "post_date": "02/15/2020 14:05:39",
      "content": "<p>I'm using a local setup with a single RTX Titan</p>",
      "votes": null,
      "replies": [
        {
          "id": 746765,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "02/15/2020 14:10:25",
          "content": "<p>Awesome setup, thanks for sharing. Do you know how a single RTX Titan compares to a single 1080Ti in performance? \nI could do a Google search but I am being lazy. ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746766,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "02/15/2020 14:16:09",
          "content": "<p>The RTX Titan is similar in performance to a 2080 Ti, so I would guess around double the performance, maybe a bit less?</p>\n\n<p>However, the advantage of the Titan isn't so much in speed but the 24Gb of RAM, so you can have larger batch sizes. However, in my experience in this competition, using larger/more complicated models doesn't help, they just overfit. I don't think my score would be any worse if I only had access to a 1080 Ti, I'd just have to be more patient :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 746799,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "02/15/2020 15:02:52",
      "content": "<p>I use a 1080 Ti. My current training set is 1M images and training a model takes about 10 - 15 hours so I do it overnight. (As <a href=\"/jamesphoward\">@jamesphoward</a> says, larger models just overfit, and the models I'm using are probably too large.)</p>",
      "votes": null,
      "replies": [
        {
          "id": 746850,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "02/15/2020 16:14:50",
          "content": "<p>That's good news for me. Thanks for sharing ! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 746853,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "02/15/2020 16:16:21",
      "content": "<p>I use a 1080 ti, I have 1.6 million images to train so, I am training in very small chunks... (400k on validation set xD)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 746958,
      "author_name": "dagnelies",
      "author_url": "",
      "post_date": "02/15/2020 19:29:26",
      "content": "<p>I was using AWS Sagemaker first, mainly it looked like the easiest / most productive way. In the end, I only used it to do the pre-processing to extract the faces from videos. The whole then uploaded to S3. Since Sagemaker had problems with TF 2, I (was forced to) switch to plain EC2 Instances and install everything manually. The side effect benefit is also lower pricing (through spot instances). What I miss though is really the comfort of kaggle. Simply \"committing\" your work and see the nice result the day after is simply really cool. Here I have to use logging in training callbacks and such.</p>\n\n<p>Another feature which I think would be handy is AWS Lambda for mass parallel preprocessing (but it's problematic due to the image size constraint (no TF). </p>\n\n<p>I think the GCP ecosystem is a bit easier to deal wih. AWS is quite ...hardcore.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 746992,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "02/15/2020 20:31:40",
      "content": "<p>I got my current score by preprocessing on a cloud instance(cropping face) and training/inference on kaggle kernels.\n<a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/129770\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/129770</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "746722": "As mentioned in the competition instructions, it is clear that training should be done outside Kaggle kernels. \nSince I am joining this competition, I was wondering which infrastructure are you using to train your models? Are you using a cloud instance or local setup? If so, could you share some details. \n\nMuch appreciated. :)",
    "746762": "I'm using a local setup with a single RTX Titan",
    "746765": "Awesome setup, thanks for sharing. Do you know how a single RTX Titan compares to a single 1080Ti in performance? \nI could do a Google search but I am being lazy. ;)",
    "746766": "The RTX Titan is similar in performance to a 2080 Ti, so I would guess around double the performance, maybe a bit less?\n\nHowever, the advantage of the Titan isn't so much in speed but the 24Gb of RAM, so you can have larger batch sizes. However, in my experience in this competition, using larger/more complicated models doesn't help, they just overfit. I don't think my score would be any worse if I only had access to a 1080 Ti, I'd just have to be more patient :)",
    "746799": "I use a 1080 Ti. My current training set is 1M images and training a model takes about 10 - 15 hours so I do it overnight. (As @jamesphoward says, larger models just overfit, and the models I'm using are probably too large.)",
    "746850": "That's good news for me. Thanks for sharing !",
    "746853": "I use a 1080 ti, I have 1.6 million images to train so, I am training in very small chunks... (400k on validation set xD)",
    "746958": "I was using AWS Sagemaker first, mainly it looked like the easiest / most productive way. In the end, I only used it to do the pre-processing to extract the faces from videos. The whole then uploaded to S3. Since Sagemaker had problems with TF 2, I (was forced to) switch to plain EC2 Instances and install everything manually. The side effect benefit is also lower pricing (through spot instances). What I miss though is really the comfort of kaggle. Simply \"committing\" your work and see the nice result the day after is simply really cool. Here I have to use logging in training callbacks and such.\n\nAnother feature which I think would be handy is AWS Lambda for mass parallel preprocessing (but it's problematic due to the image size constraint (no TF). \n\nI think the GCP ecosystem is a bit easier to deal wih. AWS is quite ...hardcore.",
    "746992": "I got my current score by preprocessing on a cloud instance(cropping face) and training/inference on kaggle kernels.\nhttps://www.kaggle.com/c/deepfake-detection-challenge/discussion/129770"
  },
  "source": "meta"
}