{
  "id": 130438,
  "title": "Beginner Queries",
  "url": "/competitions/deepfake-detection-challenge/discussion/130438",
  "author_name": "",
  "post_date": "2020-02-14T05:48:06.830272800Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi, \nI understand that we have to upload model and make submission. I am not sure about two things:\n1) Do we have to build whole preprocessing pipeline in Kaggle notebook of those test data and then predict on the preprocessed data using the model we trained locally? As preprocessing is already taking a huge amount of time (capturing 100 frames per video, resizing and converting to numpy, normalizing) so preprocessing + prediction might exceed 9 hours. And if I upload the trained model, what if it exceeds 1 GB (Inception +  LSTM layers)? Is there any other way?\n2) Can we download published test data, make predictions locally (e.g. on our laptop), create csv submission, upload that csv to Kaggle and see result in public leaderboard?</p>",
  "messages": [
    {
      "id": "745721",
      "postDate": "02/14/2020 05:48:06",
      "content": "<p>Hi, \nI understand that we have to upload model and make submission. I am not sure about two things:\n1) Do we have to build whole preprocessing pipeline in Kaggle notebook of those test data and then predict on the preprocessed data using the model we trained locally? As preprocessing is already taking a huge amount of time (capturing 100 frames per video, resizing and converting to numpy, normalizing) so preprocessing + prediction might exceed 9 hours. And if I upload the trained model, what if it exceeds 1 GB (Inception +  LSTM layers)? Is there any other way?\n2) Can we download published test data, make predictions locally (e.g. on our laptop), create csv submission, upload that csv to Kaggle and see result in public leaderboard?</p>",
      "rawMarkdown": "Hi, \nI understand that we have to upload model and make submission. I am not sure about two things:\n1) Do we have to build whole preprocessing pipeline in Kaggle notebook of those test data and then predict on the preprocessed data using the model we trained locally? As preprocessing is already taking a huge amount of time (capturing 100 frames per video, resizing and converting to numpy, normalizing) so preprocessing + prediction might exceed 9 hours. And if I upload the trained model, what if it exceeds 1 GB (Inception +  LSTM layers)? Is there any other way?\n2) Can we download published test data, make predictions locally (e.g. on our laptop), create csv submission, upload that csv to Kaggle and see result in public leaderboard?",
      "votes": null
    },
    {
      "id": "745825",
      "postDate": "02/14/2020 08:50:35",
      "content": "<p>Answers to questions:\n1) Yes. You will have to make the entire pipeline of reading the videos and then predicting the class by using your model.\n2) The final predictions are to be made in a private dataset which you can't see. So you'll have to commit your notebook and submit that for the purposes of running it on the private dataset.</p>",
      "rawMarkdown": "Answers to questions:\n1) Yes. You will have to make the entire pipeline of reading the videos and then predicting the class by using your model.\n2) The final predictions are to be made in a private dataset which you can't see. So you'll have to commit your notebook and submit that for the purposes of running it on the private dataset.",
      "votes": null
    },
    {
      "id": "746429",
      "postDate": "02/15/2020 02:02:11",
      "content": "<p>It is not really possible to exceed 1 GB unless you're doing model stacking. FYI, my best scoring model weights is around 80MB. Unless you're stacking &gt;12 models, you will be fine. You do not need to worry about that.\nIf your model weights is more than 500MB, there's some serious problem going on. Choose your backbone model wisely.</p>",
      "rawMarkdown": "It is not really possible to exceed 1 GB unless you're doing model stacking. FYI, my best scoring model weights is around 80MB. Unless you're stacking &gt;12 models, you will be fine. You do not need to worry about that.\nIf your model weights is more than 500MB, there's some serious problem going on. Choose your backbone model wisely.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 745825,
      "author_name": "akashnandi",
      "author_url": "",
      "post_date": "02/14/2020 08:50:35",
      "content": "<p>Answers to questions:\n1) Yes. You will have to make the entire pipeline of reading the videos and then predicting the class by using your model.\n2) The final predictions are to be made in a private dataset which you can't see. So you'll have to commit your notebook and submit that for the purposes of running it on the private dataset.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 746429,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "02/15/2020 02:02:11",
      "content": "<p>It is not really possible to exceed 1 GB unless you're doing model stacking. FYI, my best scoring model weights is around 80MB. Unless you're stacking &gt;12 models, you will be fine. You do not need to worry about that.\nIf your model weights is more than 500MB, there's some serious problem going on. Choose your backbone model wisely.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "745721": "Hi, \nI understand that we have to upload model and make submission. I am not sure about two things:\n1) Do we have to build whole preprocessing pipeline in Kaggle notebook of those test data and then predict on the preprocessed data using the model we trained locally? As preprocessing is already taking a huge amount of time (capturing 100 frames per video, resizing and converting to numpy, normalizing) so preprocessing + prediction might exceed 9 hours. And if I upload the trained model, what if it exceeds 1 GB (Inception +  LSTM layers)? Is there any other way?\n2) Can we download published test data, make predictions locally (e.g. on our laptop), create csv submission, upload that csv to Kaggle and see result in public leaderboard?",
    "745825": "Answers to questions:\n1) Yes. You will have to make the entire pipeline of reading the videos and then predicting the class by using your model.\n2) The final predictions are to be made in a private dataset which you can't see. So you'll have to commit your notebook and submit that for the purposes of running it on the private dataset.",
    "746429": "It is not really possible to exceed 1 GB unless you're doing model stacking. FYI, my best scoring model weights is around 80MB. Unless you're stacking &gt;12 models, you will be fine. You do not need to worry about that.\nIf your model weights is more than 500MB, there's some serious problem going on. Choose your backbone model wisely."
  },
  "source": "meta"
}