{
  "id": 123801,
  "title": "How to efficiently read and resize paraquet files?",
  "url": "/competitions/bengaliai-cv19/discussion/123801",
  "author_name": "",
  "post_date": "2019-12-30T13:16:09.469274200Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi,\nHow are members reading train and test paraquet files into csv and still able to work on kaggle kernel?\nPersonally, this is first time, I am using paraquet and kernel is crashing when loading. Though, I am able to read the file and train the model in separate kernels, but still size comes around 13 Gb for image resize 46x46. Now upon increasing the image size, which ideally one would like to improve the model accuracy, the single kernel with GPU too crashes. I am stil to try the CPU kernel though. Any inputs are welcomed.</p>",
  "messages": [
    {
      "id": "706494",
      "postDate": "12/30/2019 13:16:09",
      "content": "<p>Hi,\nHow are members reading train and test paraquet files into csv and still able to work on kaggle kernel?\nPersonally, this is first time, I am using paraquet and kernel is crashing when loading. Though, I am able to read the file and train the model in separate kernels, but still size comes around 13 Gb for image resize 46x46. Now upon increasing the image size, which ideally one would like to improve the model accuracy, the single kernel with GPU too crashes. I am stil to try the CPU kernel though. Any inputs are welcomed.</p>",
      "rawMarkdown": "Hi,\nHow are members reading train and test paraquet files into csv and still able to work on kaggle kernel?\nPersonally, this is first time, I am using paraquet and kernel is crashing when loading. Though, I am able to read the file and train the model in separate kernels, but still size comes around 13 Gb for image resize 46x46. Now upon increasing the image size, which ideally one would like to improve the model accuracy, the single kernel with GPU too crashes. I am stil to try the CPU kernel though. Any inputs are welcomed.",
      "votes": null
    },
    {
      "id": "706501",
      "postDate": "12/30/2019 13:28:14",
      "content": "<p>my suggestion for submission kernel:</p>\n\n<ol>\n<li>in net model</li>\n</ol>\n\n<p>```\nclass Net(nn.Module): \n    def <strong>init</strong>(self, num_class=(168,11,7)):\n        super(Net, self).<strong>init</strong>() \n        e = ResNet34()\n        ... etc ... </p>\n\n<pre><code>def forward(self, x):\n    batch_size,C,H,W = x.shape\n    x = F.interpolate(x,size=(148,256), mode='bilinear',align_corners=False) # resize, crop here!!!!\n\n    ... etc ...\n</code></pre>\n\n<p>```</p>\n\n<ol>\n<li>in submission code:</li>\n</ol>\n\n<p>```</p>\n\n<pre><code>    df  = pd.read_parquet(DATA_DIR+'/train_image_data_%d.parquet'%i, engine='pyarrow') \n    ... etc ...\n\n\n    for b in range(0,len(df),batch_size):\n\n        B = min(len(df),b+batch_size)-b\n        image = df.iloc[b:b+B, range(1,32332+1)].values\n        image_id = df.iloc[b:b+B, 0].values\n\n        input = torch.from_numpy(image).float().cuda()\n        logit  = net(input)\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "my suggestion for submission kernel:\n\n1. in net model\n\n```\nclass Net(nn.Module): \n    def __init__(self, num_class=(168,11,7)):\n        super(Net, self).__init__() \n        e = ResNet34()\n        ... etc ... \n \n    def forward(self, x):\n        batch_size,C,H,W = x.shape\n        x = F.interpolate(x,size=(148,256), mode='bilinear',align_corners=False) # resize, crop here!!!!\n \n        ... etc ...\n\n\n```\n\n2. in submission code:\n\n```\n\n        df  = pd.read_parquet(DATA_DIR+'/train_image_data_%d.parquet'%i, engine='pyarrow') \n        ... etc ...\n\n \n        for b in range(0,len(df),batch_size):\n\n            B = min(len(df),b+batch_size)-b\n            image = df.iloc[b:b+B, range(1,32332+1)].values\n            image_id = df.iloc[b:b+B, 0].values\n\n            input = torch.from_numpy(image).float().cuda()\n            logit  = net(input)\n\n```",
      "votes": null
    },
    {
      "id": "706519",
      "postDate": "12/30/2019 14:10:08",
      "content": "<p>In submission kernel you may drop each  big dataframe after processing: <code>df.drop(df.index.values,inplace=True)</code>  This operation substantially cleans memory.</p>",
      "rawMarkdown": "In submission kernel you may drop each  big dataframe after processing: ` df.drop(df.index.values,inplace=True)`  This operation substantially cleans memory.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 706501,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/30/2019 13:28:14",
      "content": "<p>my suggestion for submission kernel:</p>\n\n<ol>\n<li>in net model</li>\n</ol>\n\n<p>```\nclass Net(nn.Module): \n    def <strong>init</strong>(self, num_class=(168,11,7)):\n        super(Net, self).<strong>init</strong>() \n        e = ResNet34()\n        ... etc ... </p>\n\n<pre><code>def forward(self, x):\n    batch_size,C,H,W = x.shape\n    x = F.interpolate(x,size=(148,256), mode='bilinear',align_corners=False) # resize, crop here!!!!\n\n    ... etc ...\n</code></pre>\n\n<p>```</p>\n\n<ol>\n<li>in submission code:</li>\n</ol>\n\n<p>```</p>\n\n<pre><code>    df  = pd.read_parquet(DATA_DIR+'/train_image_data_%d.parquet'%i, engine='pyarrow') \n    ... etc ...\n\n\n    for b in range(0,len(df),batch_size):\n\n        B = min(len(df),b+batch_size)-b\n        image = df.iloc[b:b+B, range(1,32332+1)].values\n        image_id = df.iloc[b:b+B, 0].values\n\n        input = torch.from_numpy(image).float().cuda()\n        logit  = net(input)\n</code></pre>\n\n<p>```</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 706519,
      "author_name": "andreyzotov",
      "author_url": "",
      "post_date": "12/30/2019 14:10:08",
      "content": "<p>In submission kernel you may drop each  big dataframe after processing: <code>df.drop(df.index.values,inplace=True)</code>  This operation substantially cleans memory.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "706494": "Hi,\nHow are members reading train and test paraquet files into csv and still able to work on kaggle kernel?\nPersonally, this is first time, I am using paraquet and kernel is crashing when loading. Though, I am able to read the file and train the model in separate kernels, but still size comes around 13 Gb for image resize 46x46. Now upon increasing the image size, which ideally one would like to improve the model accuracy, the single kernel with GPU too crashes. I am stil to try the CPU kernel though. Any inputs are welcomed.",
    "706501": "my suggestion for submission kernel:\n\n1. in net model\n\n```\nclass Net(nn.Module): \n    def __init__(self, num_class=(168,11,7)):\n        super(Net, self).__init__() \n        e = ResNet34()\n        ... etc ... \n \n    def forward(self, x):\n        batch_size,C,H,W = x.shape\n        x = F.interpolate(x,size=(148,256), mode='bilinear',align_corners=False) # resize, crop here!!!!\n \n        ... etc ...\n\n\n```\n\n2. in submission code:\n\n```\n\n        df  = pd.read_parquet(DATA_DIR+'/train_image_data_%d.parquet'%i, engine='pyarrow') \n        ... etc ...\n\n \n        for b in range(0,len(df),batch_size):\n\n            B = min(len(df),b+batch_size)-b\n            image = df.iloc[b:b+B, range(1,32332+1)].values\n            image_id = df.iloc[b:b+B, 0].values\n\n            input = torch.from_numpy(image).float().cuda()\n            logit  = net(input)\n\n```",
    "706519": "In submission kernel you may drop each  big dataframe after processing: ` df.drop(df.index.values,inplace=True)`  This operation substantially cleans memory."
  },
  "source": "meta"
}