{
  "id": 122778,
  "title": "Using fastai to load parquet file",
  "url": "/competitions/bengaliai-cv19/discussion/122778",
  "author_name": "Anoop Kumar yadav",
  "post_date": "2019-12-22T20:06:00.859000",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Is there any documentation about handling parquet file</p>",
  "messages": [
    {
      "id": 705292,
      "postDate": "2019-12-28T18:07:22.083Z",
      "content": "<p>Most users in this competition handle parquets files using the Pandas library: <a href=\"https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read_parquet.html\">https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read_parquet.html</a> that hence returns a dataframe to use for your fastai implementation.</p>",
      "rawMarkdown": "Most users in this competition handle parquets files using the Pandas library: https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read_parquet.html that hence returns a dataframe to use for your fastai implementation.",
      "votes": 1,
      "replies": [
        {
          "id": 705955,
          "postDate": "2019-12-29T18:01:34.360Z",
          "content": "<p>I can read parquet from pandas but fastai takes images and image path but how to give image data from pandas. I know there is a function in fastai in which I can give dataframe as input but that only take Image name but here I have Image data in dataframe how to handle that.</p>",
          "rawMarkdown": "I can read parquet from pandas but fastai takes images and image path but how to give image data from pandas. I know there is a function in fastai in which I can give dataframe as input but that only take Image name but here I have Image data in dataframe how to handle that.",
          "votes": 1
        },
        {
          "id": 706725,
          "postDate": "2019-12-30T18:27:12.543Z",
          "content": "<p>I think I may see what you are talking about. I personally do not know about fastai but I encountered your issue when using the keras <code>flow_from_dataframe(..)</code> method where you have to provide a dataframe with the path to images as one of the columns...</p>\n\n<p>What I have done in my project is just to <strong>save as .png (or any format) the images</strong> in the train parquets on my personal computer (or you could create a private Kaggle Dataset if you only use Kaggle) then perform training locally ! \nAfter this, I <strong>save my model and import it in a Kaggle notebook used only for inference</strong>.\nIf I understand Fast.ai right, it will be helpful to find the correct model/model architecture for the problem then train it ? This process could be done on your computer and you could then just save this model in the optimal form then load it using Pytorch in your inference kernel.</p>\n\n<p>I hope I was helpful even if I didnt quite answer your question !</p>\n\n<p>I wish you good luck for the rest of the competition !</p>",
          "rawMarkdown": "I think I may see what you are talking about. I personally do not know about fastai but I encountered your issue when using the keras `flow_from_dataframe(..)` method where you have to provide a dataframe with the path to images as one of the columns...\n\nWhat I have done in my project is just to **save as .png (or any format) the images** in the train parquets on my personal computer (or you could create a private Kaggle Dataset if you only use Kaggle) then perform training locally ! \nAfter this, I **save my model and import it in a Kaggle notebook used only for inference**.\nIf I understand Fast.ai right, it will be helpful to find the correct model/model architecture for the problem then train it ? This process could be done on your computer and you could then just save this model in the optimal form then load it using Pytorch in your inference kernel.\n\nI hope I was helpful even if I didnt quite answer your question !\n\nI wish you good luck for the rest of the competition !"
        }
      ]
    },
    {
      "id": 700912,
      "postDate": "2019-12-22T20:06:00.860Z",
      "content": "<p>Is there any documentation about handling parquet file</p>",
      "rawMarkdown": "Is there any documentation about handling parquet file",
      "votes": 1
    },
    {
      "id": 708933,
      "postDate": "2020-01-02T20:58:06.850Z",
      "content": "<p>In another competition I first converted pandas dataframe into arrays and reshaped them to images, and then used customized ImageList to put the arrays into a databunch. You can try to adjust my code snippets for this particular competition:</p>\n\n<p>```\nclass ArrayImageList(ImageList):\n    @classmethod\n    def from_array(cls, array):\n        return cls(list(enumerate(array)))</p>\n\n<pre><code>def label_from_array(self, array, label_cls=None, **kwargs):\n    indices,_ = zip(*self.items)\n    labels = array[list(indices)]\n    return self._label_from_list(labels, label_cls=label_cls, **kwargs)\n\ndef get(self, i):\n    n = self.items[i]\n    if len(n.shape) == 1: n = n[1]\n    n = torch.tensor(n).float()/255\n    return Image(n)\n</code></pre>\n\n<p>```</p>\n\n<p><code>\nx_train = np.reshape(df_train.values[:,1:],(-1,1,sz,sz))\ny_train = df_train.values[:,0]\n</code></p>\n\n<p><code>\ndata = (ArrayImageList.from_array(x_train)\n        .split_by_idx(idx_val)\n        .label_from_array(y_train)\n        .transform(get_transforms(do_flip=False), size=sz, padding_mode='zeros')\n        .databunch(bs=bs).normalize(stats))\n</code></p>\n\n<p>Though, in my opinion for this competition parquets is not a right way of doing things. I'm not sure if you would have enough RAM if decide to run your code at kaggle to load all 4 files and produce a databunch on GPU nodes. Also, while overhead of reading images is avoided, you add another overhead of image preprocessing, which must be done because the input data is a kind of low quality, and you need to do at least cropping at each image generation.</p>",
      "rawMarkdown": "In another competition I first converted pandas dataframe into arrays and reshaped them to images, and then used customized ImageList to put the arrays into a databunch. You can try to adjust my code snippets for this particular competition:\n\n```\nclass ArrayImageList(ImageList):\n    @classmethod\n    def from_array(cls, array):\n        return cls(list(enumerate(array)))\n    \n    def label_from_array(self, array, label_cls=None, **kwargs):\n        indices,_ = zip(*self.items)\n        labels = array[list(indices)]\n        return self._label_from_list(labels, label_cls=label_cls, **kwargs)\n    \n    def get(self, i):\n        n = self.items[i]\n        if len(n.shape) == 1: n = n[1]\n        n = torch.tensor(n).float()/255\n        return Image(n)\n```\n\n```\nx_train = np.reshape(df_train.values[:,1:],(-1,1,sz,sz))\ny_train = df_train.values[:,0]\n```\n\n```\ndata = (ArrayImageList.from_array(x_train)\n        .split_by_idx(idx_val)\n        .label_from_array(y_train)\n        .transform(get_transforms(do_flip=False), size=sz, padding_mode='zeros')\n        .databunch(bs=bs).normalize(stats))\n```\n\nThough, in my opinion for this competition parquets is not a right way of doing things. I'm not sure if you would have enough RAM if decide to run your code at kaggle to load all 4 files and produce a databunch on GPU nodes. Also, while overhead of reading images is avoided, you add another overhead of image preprocessing, which must be done because the input data is a kind of low quality, and you need to do at least cropping at each image generation.",
      "votes": 2
    },
    {
      "id": 705947,
      "postDate": "2019-12-29T17:50:34.783Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 705292,
      "author_name": "Thomas Di Martino",
      "author_url": "",
      "post_date": "2019-12-28T18:07:22.083000",
      "content": "<p>Most users in this competition handle parquets files using the Pandas library: <a href=\"https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read_parquet.html\">https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read_parquet.html</a> that hence returns a dataframe to use for your fastai implementation.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 705955,
          "author_name": "Anoop Kumar yadav",
          "author_url": "",
          "post_date": "2019-12-29T18:01:34.360000",
          "content": "<p>I can read parquet from pandas but fastai takes images and image path but how to give image data from pandas. I know there is a function in fastai in which I can give dataframe as input but that only take Image name but here I have Image data in dataframe how to handle that.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 706725,
          "author_name": "Thomas Di Martino",
          "author_url": "",
          "post_date": "2019-12-30T18:27:12.543000",
          "content": "<p>I think I may see what you are talking about. I personally do not know about fastai but I encountered your issue when using the keras <code>flow_from_dataframe(..)</code> method where you have to provide a dataframe with the path to images as one of the columns...</p>\n\n<p>What I have done in my project is just to <strong>save as .png (or any format) the images</strong> in the train parquets on my personal computer (or you could create a private Kaggle Dataset if you only use Kaggle) then perform training locally ! \nAfter this, I <strong>save my model and import it in a Kaggle notebook used only for inference</strong>.\nIf I understand Fast.ai right, it will be helpful to find the correct model/model architecture for the problem then train it ? This process could be done on your computer and you could then just save this model in the optimal form then load it using Pytorch in your inference kernel.</p>\n\n<p>I hope I was helpful even if I didnt quite answer your question !</p>\n\n<p>I wish you good luck for the rest of the competition !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 708933,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2020-01-02T20:58:06.850000",
      "content": "<p>In another competition I first converted pandas dataframe into arrays and reshaped them to images, and then used customized ImageList to put the arrays into a databunch. You can try to adjust my code snippets for this particular competition:</p>\n\n<p>```\nclass ArrayImageList(ImageList):\n    @classmethod\n    def from_array(cls, array):\n        return cls(list(enumerate(array)))</p>\n\n<pre><code>def label_from_array(self, array, label_cls=None, **kwargs):\n    indices,_ = zip(*self.items)\n    labels = array[list(indices)]\n    return self._label_from_list(labels, label_cls=label_cls, **kwargs)\n\ndef get(self, i):\n    n = self.items[i]\n    if len(n.shape) == 1: n = n[1]\n    n = torch.tensor(n).float()/255\n    return Image(n)\n</code></pre>\n\n<p>```</p>\n\n<p><code>\nx_train = np.reshape(df_train.values[:,1:],(-1,1,sz,sz))\ny_train = df_train.values[:,0]\n</code></p>\n\n<p><code>\ndata = (ArrayImageList.from_array(x_train)\n        .split_by_idx(idx_val)\n        .label_from_array(y_train)\n        .transform(get_transforms(do_flip=False), size=sz, padding_mode='zeros')\n        .databunch(bs=bs).normalize(stats))\n</code></p>\n\n<p>Though, in my opinion for this competition parquets is not a right way of doing things. I'm not sure if you would have enough RAM if decide to run your code at kaggle to load all 4 files and produce a databunch on GPU nodes. Also, while overhead of reading images is avoided, you add another overhead of image preprocessing, which must be done because the input data is a kind of low quality, and you need to do at least cropping at each image generation.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 705947,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-29T17:50:34.783000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "705292": "Most users in this competition handle parquets files using the Pandas library: https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read_parquet.html that hence returns a dataframe to use for your fastai implementation.",
    "700912": "Is there any documentation about handling parquet file",
    "708933": "In another competition I first converted pandas dataframe into arrays and reshaped them to images, and then used customized ImageList to put the arrays into a databunch. You can try to adjust my code snippets for this particular competition:\n\n```\nclass ArrayImageList(ImageList):\n    @classmethod\n    def from_array(cls, array):\n        return cls(list(enumerate(array)))\n    \n    def label_from_array(self, array, label_cls=None, **kwargs):\n        indices,_ = zip(*self.items)\n        labels = array[list(indices)]\n        return self._label_from_list(labels, label_cls=label_cls, **kwargs)\n    \n    def get(self, i):\n        n = self.items[i]\n        if len(n.shape) == 1: n = n[1]\n        n = torch.tensor(n).float()/255\n        return Image(n)\n```\n\n```\nx_train = np.reshape(df_train.values[:,1:],(-1,1,sz,sz))\ny_train = df_train.values[:,0]\n```\n\n```\ndata = (ArrayImageList.from_array(x_train)\n        .split_by_idx(idx_val)\n        .label_from_array(y_train)\n        .transform(get_transforms(do_flip=False), size=sz, padding_mode='zeros')\n        .databunch(bs=bs).normalize(stats))\n```\n\nThough, in my opinion for this competition parquets is not a right way of doing things. I'm not sure if you would have enough RAM if decide to run your code at kaggle to load all 4 files and produce a databunch on GPU nodes. Also, while overhead of reading images is avoided, you add another overhead of image preprocessing, which must be done because the input data is a kind of low quality, and you need to do at least cropping at each image generation.",
    "705947": ""
  }
}