{
  "id": 132082,
  "title": "Let's talk about GPU Spec & Training Time",
  "url": "/competitions/deepfake-detection-challenge/discussion/132082",
  "author_name": "gw song",
  "post_date": "2020-02-24T00:27:09.370000",
  "votes": 4,
  "comment_count": 20,
  "views": 0,
  "content": "<p>I want to share hardware spec and training time on this board.</p>\n\n<p>Without the info of the model size (and also batch size), \ncomputation time only would not tell much.</p>\n\n<p>But I am inspired by the article \"Best single model\" in this board. </p>\n\n<p>So, Let me start</p>\n\n<ul>\n<li><p>GPU Spec \n. Tesla P100-16GB RAM (Azure)</p></li>\n<li><p>Training Time\n. 4Days / 10epochs\n  ( After 10 epochs, it start to overfit to my validation data and private set of the LB </p></li>\n<li><p>Data Size\n. 500K Frame</p></li>\n</ul>\n\n<p>Oh, and my best single model scored 0.403 on LB</p>",
  "messages": [
    {
      "id": 754929,
      "postDate": "2020-02-24T08:22:14.750Z",
      "content": "<p>For score of 0.38* Single model\nGPU Spec - kaggle gpu\nTraining Time - No more than 4 hours,\nData Size -  ~35k videos\nCredits to train - <a href=\"/unkownhihi\">@unkownhihi</a></p>",
      "rawMarkdown": "For score of 0.38* Single model\nGPU Spec - kaggle gpu\nTraining Time - No more than 4 hours,\nData Size -  ~35k videos\nCredits to train - @unkownhihi",
      "votes": 5,
      "replies": [
        {
          "id": 755220,
          "postDate": "2020-02-24T15:08:02.137Z",
          "content": "<p>Wow, thank you so much!</p>\n\n<p>I'm really surprised to your comment that you just use kaggle GPU and train only 4-hours.\nIt's really push me try to find a way to train wisely.</p>\n\n<p>Did you train from scratch or pre-trained one?</p>\n\n<p>And, how many frames did you use for your training?\n\"~35K videos\" means \"35K x frames per video\"?</p>",
          "rawMarkdown": "Wow, thank you so much!\n\nI'm really surprised to your comment that you just use kaggle GPU and train only 4-hours.\nIt's really push me try to find a way to train wisely.\n\nDid you train from scratch or pre-trained one?\n\nAnd, how many frames did you use for your training?\n\"~35K videos\" means \"35K x frames per video\"?\n\n"
        },
        {
          "id": 755348,
          "postDate": "2020-02-24T17:29:05.093Z",
          "content": "<p>10 frames per video</p>\n\n<p>It was a model with imagenet weights, no layers freezed.</p>",
          "rawMarkdown": "10 frames per video\n\nIt was a model with imagenet weights, no layers freezed."
        },
        {
          "id": 755381,
          "postDate": "2020-02-24T18:15:58.970Z",
          "content": "<p>You're using this dataset? <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128954\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128954</a> It has 35k videos</p>",
          "rawMarkdown": "You're using this dataset? https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128954 It has 35k videos"
        },
        {
          "id": 755402,
          "postDate": "2020-02-24T18:41:40.660Z",
          "content": "<p>No, we created our unique dataset.</p>",
          "rawMarkdown": "No, we created our unique dataset."
        },
        {
          "id": 755759,
          "postDate": "2020-02-25T05:23:01.407Z",
          "content": "<p>Thank you for sharing.\nI start from scratch bc it give me a better result than imgnet-pretrained at the first time.\nBut I might reconsider it</p>",
          "rawMarkdown": "Thank you for sharing.\nI start from scratch bc it give me a better result than imgnet-pretrained at the first time.\nBut I might reconsider it"
        }
      ]
    },
    {
      "id": 755366,
      "postDate": "2020-02-24T17:53:09.337Z",
      "content": "<p>Score: 0.34 Single model\nGPU: 1080 Ti\nTraining time: 5 min/epoch\nData size: 2 x 11K videos train set, 2 x 5K videos validation</p>",
      "rawMarkdown": "Score: 0.34 Single model\nGPU: 1080 Ti\nTraining time: 5 min/epoch\nData size: 2 x 11K videos train set, 2 x 5K videos validation",
      "votes": 4,
      "replies": [
        {
          "id": 755768,
          "postDate": "2020-02-25T05:31:17.340Z",
          "content": "<p>Wow, you guys really amazing to train such a short time and get superb result. :)</p>\n\n<p>I should try with smaller model.</p>\n\n<p>Thank you!</p>",
          "rawMarkdown": "Wow, you guys really amazing to train such a short time and get superb result. :)\n\nI should try with smaller model.\n\nThank you!"
        },
        {
          "id": 756033,
          "postDate": "2020-02-25T11:04:52.513Z",
          "content": "<p><a href=\"/ngcferreira\">@ngcferreira</a> do you want to merge?, We have tons of models to ensemble, I also have a 1080 ti, why 11k videos?, you can do much better if you go higher.</p>",
          "rawMarkdown": "@ngcferreira do you want to merge?, We have tons of models to ensemble, I also have a 1080 ti, why 11k videos?, you can do much better if you go higher."
        },
        {
          "id": 756483,
          "postDate": "2020-02-25T18:53:21.133Z",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> Sorry I made a mistake, I only counted the real videos. I actually divided the training data in 3 sets, a training, a validation and a test set. My training set has 11k real videos, and all the respective fake ones, the validation one has 5K real videos + the fakes, and the test set has 2K real videos + fakes. In each epoch, I choose all real training videos, and randomly choose the same amount of fake videos. I do the save with the validation set. I use the test set see the log loss of the trained model.</p>\n\n<p>I didn't have much time to dedicate to this competition, so I still want to try a few more things before I decide if I want to merge with others. I will contact you later on, before the 3rd of March.</p>",
          "rawMarkdown": "@harshitsheoran Sorry I made a mistake, I only counted the real videos. I actually divided the training data in 3 sets, a training, a validation and a test set. My training set has 11k real videos, and all the respective fake ones, the validation one has 5K real videos + the fakes, and the test set has 2K real videos + fakes. In each epoch, I choose all real training videos, and randomly choose the same amount of fake videos. I do the save with the validation set. I use the test set see the log loss of the trained model.\n\nI didn't have much time to dedicate to this competition, so I still want to try a few more things before I decide if I want to merge with others. I will contact you later on, before the 3rd of March."
        },
        {
          "id": 756499,
          "postDate": "2020-02-25T19:11:02.183Z",
          "content": "<p>Please do it within 7 days, <a href=\"/ngcferreira\">@ngcferreira</a> Thank you for your reply, till the time you get that score with a single model, by any number of videos, it does not bother me.</p>",
          "rawMarkdown": "Please do it within 7 days, @ngcferreira Thank you for your reply, till the time you get that score with a single model, by any number of videos, it does not bother me."
        }
      ]
    },
    {
      "id": 754707,
      "postDate": "2020-02-24T00:27:09.370Z",
      "content": "<p>I want to share hardware spec and training time on this board.</p>\n\n<p>Without the info of the model size (and also batch size), \ncomputation time only would not tell much.</p>\n\n<p>But I am inspired by the article \"Best single model\" in this board. </p>\n\n<p>So, Let me start</p>\n\n<ul>\n<li><p>GPU Spec \n. Tesla P100-16GB RAM (Azure)</p></li>\n<li><p>Training Time\n. 4Days / 10epochs\n  ( After 10 epochs, it start to overfit to my validation data and private set of the LB </p></li>\n<li><p>Data Size\n. 500K Frame</p></li>\n</ul>\n\n<p>Oh, and my best single model scored 0.403 on LB</p>",
      "rawMarkdown": "I want to share hardware spec and training time on this board.\n\nWithout the info of the model size (and also batch size), \ncomputation time only would not tell much.\n\nBut I am inspired by the article \"Best single model\" in this board. \n\nSo, Let me start\n\n\n  - GPU Spec \n    . Tesla P100-16GB RAM (Azure)\n\n  - Training Time\n    . 4Days / 10epochs\n      ( After 10 epochs, it start to overfit to my validation data and private set of the LB \n\n  - Data Size\n   . 500K Frame\n\n\nOh, and my best single model scored 0.403 on LB",
      "votes": 4
    },
    {
      "id": 755434,
      "postDate": "2020-02-24T19:18:09.393Z",
      "content": "<p>Score: 0.47 single model \nTPU : v2-8 \nTraining time: 30 seconds per epoch for 30 epochs so around 20 mins with tpu setup time \nBatchsize: 512 \nimages/epoch : 512 * 64 = 32768\nUsing 90/10 train val split on all videos </p>\n\n<p>I think its more about the dataset bc you can get really good results with a super simple backbone if your data is good. </p>",
      "rawMarkdown": "Score: 0.47 single model \nTPU : v2-8 \nTraining time: 30 seconds per epoch for 30 epochs so around 20 mins with tpu setup time \nBatchsize: 512 \nimages/epoch : 512 * 64 = 32768\nUsing 90/10 train val split on all videos \n\nI think its more about the dataset bc you can get really good results with a super simple backbone if your data is good. ",
      "votes": 1
    },
    {
      "id": 755762,
      "postDate": "2020-02-25T05:24:44.013Z",
      "content": "<p>Score: Submission CSV Not Found (but i hope i can be first!)\nGPU: RTX 2070\nTraining time: 6 min/ epoch\nData size: 17Kvideo &amp; Voice data\nbut it takes submission fail for 6days... TT\ni think it has to many hard code.. that is why it makes.. error.. TT</p>",
      "rawMarkdown": "Score: Submission CSV Not Found (but i hope i can be first!)\nGPU: RTX 2070\nTraining time: 6 min/ epoch\nData size: 17Kvideo &amp; Voice data\nbut it takes submission fail for 6days... TT\ni think it has to many hard code.. that is why it makes.. error.. TT",
      "replies": [
        {
          "id": 755773,
          "postDate": "2020-02-25T05:36:00.790Z",
          "content": "<p>Did you create so many output files for by-product of your algorithm processing?</p>",
          "rawMarkdown": "Did you create so many output files for by-product of your algorithm processing?"
        },
        {
          "id": 755776,
          "postDate": "2020-02-25T05:40:53.510Z",
          "content": "<p>yes output version is 39th... but until now i didn't success to submit T.T</p>",
          "rawMarkdown": "yes output version is 39th... but until now i didn't success to submit T.T"
        },
        {
          "id": 755819,
          "postDate": "2020-02-25T06:29:19.293Z",
          "content": "<p>No, I mean create many by-product files for one-time submission.\nIf you create too many file on your processing, that error can occur.</p>",
          "rawMarkdown": "No, I mean create many by-product files for one-time submission.\nIf you create too many file on your processing, that error can occur."
        },
        {
          "id": 755823,
          "postDate": "2020-02-25T06:32:04.053Z",
          "content": "<p>yes i make many by product files.... T.T in my one-time submission..</p>",
          "rawMarkdown": "yes i make many by product files.... T.T in my one-time submission.."
        },
        {
          "id": 756061,
          "postDate": "2020-02-25T11:29:25.630Z",
          "content": "<p>I'm not sure but same error message can occur when it encounter out of resouces.\nHow about to reduce RAM/GPU/GPU-MEMORY/Storage usage.</p>\n\n<p>I hope your problem to be solved.</p>",
          "rawMarkdown": "I'm not sure but same error message can occur when it encounter out of resouces.\nHow about to reduce RAM/GPU/GPU-MEMORY/Storage usage.\n\nI hope your problem to be solved."
        }
      ]
    },
    {
      "id": 755179,
      "postDate": "2020-02-24T14:22:52.437Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 755177,
      "postDate": "2020-02-24T14:18:48.147Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 754929,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2020-02-24T08:22:14.750000",
      "content": "<p>For score of 0.38* Single model\nGPU Spec - kaggle gpu\nTraining Time - No more than 4 hours,\nData Size -  ~35k videos\nCredits to train - <a href=\"/unkownhihi\">@unkownhihi</a></p>",
      "votes": 5,
      "replies": [
        {
          "id": 755220,
          "author_name": "gw song",
          "author_url": "",
          "post_date": "2020-02-24T15:08:02.137000",
          "content": "<p>Wow, thank you so much!</p>\n\n<p>I'm really surprised to your comment that you just use kaggle GPU and train only 4-hours.\nIt's really push me try to find a way to train wisely.</p>\n\n<p>Did you train from scratch or pre-trained one?</p>\n\n<p>And, how many frames did you use for your training?\n\"~35K videos\" means \"35K x frames per video\"?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755348,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-02-24T17:29:05.093000",
          "content": "<p>10 frames per video</p>\n\n<p>It was a model with imagenet weights, no layers freezed.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755381,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-02-24T18:15:58.970000",
          "content": "<p>You're using this dataset? <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128954\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128954</a> It has 35k videos</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755402,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-02-24T18:41:40.660000",
          "content": "<p>No, we created our unique dataset.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755759,
          "author_name": "gw song",
          "author_url": "",
          "post_date": "2020-02-25T05:23:01.407000",
          "content": "<p>Thank you for sharing.\nI start from scratch bc it give me a better result than imgnet-pretrained at the first time.\nBut I might reconsider it</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 755366,
      "author_name": "Nuno Ferreira",
      "author_url": "",
      "post_date": "2020-02-24T17:53:09.337000",
      "content": "<p>Score: 0.34 Single model\nGPU: 1080 Ti\nTraining time: 5 min/epoch\nData size: 2 x 11K videos train set, 2 x 5K videos validation</p>",
      "votes": 4,
      "replies": [
        {
          "id": 755768,
          "author_name": "gw song",
          "author_url": "",
          "post_date": "2020-02-25T05:31:17.340000",
          "content": "<p>Wow, you guys really amazing to train such a short time and get superb result. :)</p>\n\n<p>I should try with smaller model.</p>\n\n<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756033,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-02-25T11:04:52.513000",
          "content": "<p><a href=\"/ngcferreira\">@ngcferreira</a> do you want to merge?, We have tons of models to ensemble, I also have a 1080 ti, why 11k videos?, you can do much better if you go higher.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756483,
          "author_name": "Nuno Ferreira",
          "author_url": "",
          "post_date": "2020-02-25T18:53:21.133000",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> Sorry I made a mistake, I only counted the real videos. I actually divided the training data in 3 sets, a training, a validation and a test set. My training set has 11k real videos, and all the respective fake ones, the validation one has 5K real videos + the fakes, and the test set has 2K real videos + fakes. In each epoch, I choose all real training videos, and randomly choose the same amount of fake videos. I do the save with the validation set. I use the test set see the log loss of the trained model.</p>\n\n<p>I didn't have much time to dedicate to this competition, so I still want to try a few more things before I decide if I want to merge with others. I will contact you later on, before the 3rd of March.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756499,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-02-25T19:11:02.183000",
          "content": "<p>Please do it within 7 days, <a href=\"/ngcferreira\">@ngcferreira</a> Thank you for your reply, till the time you get that score with a single model, by any number of videos, it does not bother me.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 755434,
      "author_name": "hongy",
      "author_url": "",
      "post_date": "2020-02-24T19:18:09.393000",
      "content": "<p>Score: 0.47 single model \nTPU : v2-8 \nTraining time: 30 seconds per epoch for 30 epochs so around 20 mins with tpu setup time \nBatchsize: 512 \nimages/epoch : 512 * 64 = 32768\nUsing 90/10 train val split on all videos </p>\n\n<p>I think its more about the dataset bc you can get really good results with a super simple backbone if your data is good. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 755762,
      "author_name": "Choo",
      "author_url": "",
      "post_date": "2020-02-25T05:24:44.013000",
      "content": "<p>Score: Submission CSV Not Found (but i hope i can be first!)\nGPU: RTX 2070\nTraining time: 6 min/ epoch\nData size: 17Kvideo &amp; Voice data\nbut it takes submission fail for 6days... TT\ni think it has to many hard code.. that is why it makes.. error.. TT</p>",
      "votes": 0,
      "replies": [
        {
          "id": 755773,
          "author_name": "gw song",
          "author_url": "",
          "post_date": "2020-02-25T05:36:00.790000",
          "content": "<p>Did you create so many output files for by-product of your algorithm processing?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755776,
          "author_name": "Choo",
          "author_url": "",
          "post_date": "2020-02-25T05:40:53.510000",
          "content": "<p>yes output version is 39th... but until now i didn't success to submit T.T</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755819,
          "author_name": "gw song",
          "author_url": "",
          "post_date": "2020-02-25T06:29:19.293000",
          "content": "<p>No, I mean create many by-product files for one-time submission.\nIf you create too many file on your processing, that error can occur.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755823,
          "author_name": "Choo",
          "author_url": "",
          "post_date": "2020-02-25T06:32:04.053000",
          "content": "<p>yes i make many by product files.... T.T in my one-time submission..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756061,
          "author_name": "gw song",
          "author_url": "",
          "post_date": "2020-02-25T11:29:25.630000",
          "content": "<p>I'm not sure but same error message can occur when it encounter out of resouces.\nHow about to reduce RAM/GPU/GPU-MEMORY/Storage usage.</p>\n\n<p>I hope your problem to be solved.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 755179,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-24T14:22:52.437000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 755177,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-24T14:18:48.147000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "754929": "For score of 0.38* Single model\nGPU Spec - kaggle gpu\nTraining Time - No more than 4 hours,\nData Size -  ~35k videos\nCredits to train - @unkownhihi",
    "755366": "Score: 0.34 Single model\nGPU: 1080 Ti\nTraining time: 5 min/epoch\nData size: 2 x 11K videos train set, 2 x 5K videos validation",
    "754707": "I want to share hardware spec and training time on this board.\n\nWithout the info of the model size (and also batch size), \ncomputation time only would not tell much.\n\nBut I am inspired by the article \"Best single model\" in this board. \n\nSo, Let me start\n\n\n  - GPU Spec \n    . Tesla P100-16GB RAM (Azure)\n\n  - Training Time\n    . 4Days / 10epochs\n      ( After 10 epochs, it start to overfit to my validation data and private set of the LB \n\n  - Data Size\n   . 500K Frame\n\n\nOh, and my best single model scored 0.403 on LB",
    "755434": "Score: 0.47 single model \nTPU : v2-8 \nTraining time: 30 seconds per epoch for 30 epochs so around 20 mins with tpu setup time \nBatchsize: 512 \nimages/epoch : 512 * 64 = 32768\nUsing 90/10 train val split on all videos \n\nI think its more about the dataset bc you can get really good results with a super simple backbone if your data is good. ",
    "755762": "Score: Submission CSV Not Found (but i hope i can be first!)\nGPU: RTX 2070\nTraining time: 6 min/ epoch\nData size: 17Kvideo &amp; Voice data\nbut it takes submission fail for 6days... TT\ni think it has to many hard code.. that is why it makes.. error.. TT",
    "755179": "",
    "755177": ""
  }
}