{
  "id": 127503,
  "title": "Can we PreProcess full data",
  "url": "/competitions/deepfake-detection-challenge/discussion/127503",
  "author_name": "",
  "post_date": "2020-01-24T10:18:52.839280600Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Is it possible to pre process the kaggle full data set using kaggle kernel .Please point me to some link.</p>",
  "messages": [
    {
      "id": "728001",
      "postDate": "01/24/2020 10:18:52",
      "content": "<p>Is it possible to pre process the kaggle full data set using kaggle kernel .Please point me to some link.</p>",
      "rawMarkdown": "Is it possible to pre process the kaggle full data set using kaggle kernel .Please point me to some link.",
      "votes": null
    },
    {
      "id": "728645",
      "postDate": "01/25/2020 02:31:06",
      "content": "<p>No. It is not possible. The kaggle kernel max storage space is less than 10GB.</p>",
      "rawMarkdown": "No. It is not possible. The kaggle kernel max storage space is less than 10GB.",
      "votes": null
    },
    {
      "id": "729107",
      "postDate": "01/25/2020 19:10:27",
      "content": "<p>You have to do it in chunks. For example preprocess 100 -&gt; then run it through the model/predict -&gt; then delete the 100 preprocessed -&gt; preprocess the next 100 -&gt; run it through the model -&gt; and so on</p>",
      "rawMarkdown": "You have to do it in chunks. For example preprocess 100 -&gt; then run it through the model/predict -&gt; then delete the 100 preprocessed -&gt; preprocess the next 100 -&gt; run it through the model -&gt; and so on",
      "votes": null
    },
    {
      "id": "729293",
      "postDate": "01/26/2020 01:26:43",
      "content": "<p>Not possible. You can't even download the chunks.</p>",
      "rawMarkdown": "Not possible. You can't even download the chunks.",
      "votes": null
    },
    {
      "id": "729315",
      "postDate": "01/26/2020 02:23:18",
      "content": "<p>Edit\nAccording to <a href=\"https://www.kaggle.com/docs/kernels#technical-specifications\">technical specification</a>, you have 16G temp storage and 16G RAM. So maybe it's possible to do it with some trick. I thought about that till I got AWS credits.</p>",
      "rawMarkdown": "Edit\nAccording to [technical specification](https://www.kaggle.com/docs/kernels#technical-specifications), you have 16G temp storage and 16G RAM. So maybe it's possible to do it with some trick. I thought about that till I got AWS credits.",
      "votes": null
    },
    {
      "id": "729584",
      "postDate": "01/26/2020 12:07:38",
      "content": "<p>Still, you won't be able to unzip the 10 GB file ...or at least not in one go. :(</p>\n\n<p>But you could indeed download it and process it by selective extraction ...it's perhaps even not a bad idea ...this AWS stuff is so cumbersome and time consuming.</p>",
      "rawMarkdown": "Still, you won't be able to unzip the 10 GB file ...or at least not in one go. :(\n\nBut you could indeed download it and process it by selective extraction ...it's perhaps even not a bad idea ...this AWS stuff is so cumbersome and time consuming.",
      "votes": null
    },
    {
      "id": "730562",
      "postDate": "01/27/2020 16:18:42",
      "content": "<p>Oh I see he's talking about the full training data -- I thought just test set.. I do the 4000 image test set like this.</p>",
      "rawMarkdown": "Oh I see he's talking about the full training data -- I thought just test set.. I do the 4000 image test set like this.",
      "votes": null
    },
    {
      "id": "731495",
      "postDate": "01/28/2020 17:42:20",
      "content": "<p><a href=\"https://www.kaggle.com/sciarrilli/dfdc-f150\">https://www.kaggle.com/sciarrilli/dfdc-f150</a></p>",
      "rawMarkdown": "https://www.kaggle.com/sciarrilli/dfdc-f150",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 728645,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "01/25/2020 02:31:06",
      "content": "<p>No. It is not possible. The kaggle kernel max storage space is less than 10GB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 729315,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "01/26/2020 02:23:18",
          "content": "<p>Edit\nAccording to <a href=\"https://www.kaggle.com/docs/kernels#technical-specifications\">technical specification</a>, you have 16G temp storage and 16G RAM. So maybe it's possible to do it with some trick. I thought about that till I got AWS credits.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 729584,
          "author_name": "dagnelies",
          "author_url": "",
          "post_date": "01/26/2020 12:07:38",
          "content": "<p>Still, you won't be able to unzip the 10 GB file ...or at least not in one go. :(</p>\n\n<p>But you could indeed download it and process it by selective extraction ...it's perhaps even not a bad idea ...this AWS stuff is so cumbersome and time consuming.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 729107,
      "author_name": "sethkitchen",
      "author_url": "",
      "post_date": "01/25/2020 19:10:27",
      "content": "<p>You have to do it in chunks. For example preprocess 100 -&gt; then run it through the model/predict -&gt; then delete the 100 preprocessed -&gt; preprocess the next 100 -&gt; run it through the model -&gt; and so on</p>",
      "votes": null,
      "replies": [
        {
          "id": 729293,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "01/26/2020 01:26:43",
          "content": "<p>Not possible. You can't even download the chunks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 730562,
          "author_name": "sethkitchen",
          "author_url": "",
          "post_date": "01/27/2020 16:18:42",
          "content": "<p>Oh I see he's talking about the full training data -- I thought just test set.. I do the 4000 image test set like this.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 731495,
      "author_name": "sciarrilli",
      "author_url": "",
      "post_date": "01/28/2020 17:42:20",
      "content": "<p><a href=\"https://www.kaggle.com/sciarrilli/dfdc-f150\">https://www.kaggle.com/sciarrilli/dfdc-f150</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "728001": "Is it possible to pre process the kaggle full data set using kaggle kernel .Please point me to some link.",
    "728645": "No. It is not possible. The kaggle kernel max storage space is less than 10GB.",
    "729107": "You have to do it in chunks. For example preprocess 100 -&gt; then run it through the model/predict -&gt; then delete the 100 preprocessed -&gt; preprocess the next 100 -&gt; run it through the model -&gt; and so on",
    "729293": "Not possible. You can't even download the chunks.",
    "729315": "Edit\nAccording to [technical specification](https://www.kaggle.com/docs/kernels#technical-specifications), you have 16G temp storage and 16G RAM. So maybe it's possible to do it with some trick. I thought about that till I got AWS credits.",
    "729584": "Still, you won't be able to unzip the 10 GB file ...or at least not in one go. :(\n\nBut you could indeed download it and process it by selective extraction ...it's perhaps even not a bad idea ...this AWS stuff is so cumbersome and time consuming.",
    "730562": "Oh I see he's talking about the full training data -- I thought just test set.. I do the 4000 image test set like this.",
    "731495": "https://www.kaggle.com/sciarrilli/dfdc-f150"
  },
  "source": "meta"
}