{
  "id": 163478,
  "title": "HELP : Keyerror while enumerating over DataLoader outputs",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/163478",
  "author_name": "",
  "post_date": "2020-07-02T07:55:25.943544700Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>```\ntrain_dl = DataLoader(train_ds, batch_size, shuffle=True, num_workers=4, pin_memory=True)</p>\n\n<p>val_dl = DataLoader(val_ds, batch_size*2, shuffle = False,num_workers=4, pin_memory=True)\n```\nI have created the above train &amp; validation dataloaders and I am getting following error while enumerating over those mentioned dataloaders</p>\n\n<p>```\nKeyError: Caught KeyError in DataLoader worker process 0.\nOriginal Traceback (most recent call last):\n  File \"/opt/conda/lib/python3.7/site-packages/pandas/core/indexes/base.py\", line 2646, in get_loc\n    return self._engine.get_loc(key)\n  File \"pandas/_libs/index.pyx\", line 111, in pandas._libs.index.IndexEngine.get_loc\n  File \"pandas/_libs/index.pyx\", line 138, in pandas._libs.index.IndexEngine.get_loc\n  File \"pandas/_libs/hashtable_class_helper.pxi\", line 998, in pandas._libs.hashtable.Int64HashTable.get_item\n  File \"pandas/_libs/hashtable_class_helper.pxi\", line 1005, in pandas._libs.hashtable.Int64HashTable.get_item\nKeyError: 1405</p>\n\n<p>During handling of the above exception, another exception occurred:</p>\n\n<p>Traceback (most recent call last):\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/worker.py\", line 178, in _worker_loop\n    data = fetcher.fetch(index)\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/fetch.py\", line 44, in fetch\n    data = [self.dataset[idx] for idx in possibly_batched_index]\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/fetch.py\", line 44, in \n    data = [self.dataset[idx] for idx in possibly_batched_index]\n  File \"\", line 11, in <strong>getitem</strong>\n    row = self.df.loc[idx]\n  File \"/opt/conda/lib/python3.7/site-packages/pandas/core/indexing.py\", line 1768, in <strong>getitem</strong>\n    return self._getitem_axis(maybe_callable, axis=axis)\n  File \"/opt/conda/lib/python3.7/site-packages/pandas/core/indexing.py\", line 1965, in _getitem_axis\n    return self._get_label(key, axis=axis)\n  File \"/opt/conda/lib/python3.7/site-packages/pandas/core/indexing.py\", line 625, in _get_label\n    return self.obj._xs(label, axis=axis)\n  File \"/opt/conda/lib/python3.7/site-packages/pandas/core/generic.py\", line 3537, in xs\n    loc = self.index.get_loc(key)\n  File \"/opt/conda/lib/python3.7/site-packages/pandas/core/indexes/base.py\", line 2648, in get_loc\n    return self._engine.get_loc(self._maybe_cast_indexer(key))\n  File \"pandas/_libs/index.pyx\", line 111, in pandas._libs.index.IndexEngine.get_loc\n  File \"pandas/_libs/index.pyx\", line 138, in pandas._libs.index.IndexEngine.get_loc\n  File \"pandas/_libs/hashtable_class_helper.pxi\", line 998, in pandas._libs.hashtable.Int64HashTable.get_item\n  File \"pandas/_libs/hashtable_class_helper.pxi\", line 1005, in pandas._libs.hashtable.Int64HashTable.get_item\nKeyError: 1405</p>\n\n<p>```\nPlease let me know on how to resolve this</p>",
  "messages": [
    {
      "id": "912085",
      "postDate": "07/02/2020 07:55:25",
      "content": "<p>```\ntrain_dl = DataLoader(train_ds, batch_size, shuffle=True, num_workers=4, pin_memory=True)</p>\n\n<p>val_dl = DataLoader(val_ds, batch_size*2, shuffle = False,num_workers=4, pin_memory=True)\n```\nI have created the above train &amp; validation dataloaders and I am getting following error while enumerating over those mentioned dataloaders</p>\n\n<p>```\nKeyError: Caught KeyError in DataLoader worker process 0.\nOriginal Traceback (most recent call last):\n  File \"/opt/conda/lib/python3.7/site-packages/pandas/core/indexes/base.py\", line 2646, in get_loc\n    return self._engine.get_loc(key)\n  File \"pandas/_libs/index.pyx\", line 111, in pandas._libs.index.IndexEngine.get_loc\n  File \"pandas/_libs/index.pyx\", line 138, in pandas._libs.index.IndexEngine.get_loc\n  File \"pandas/_libs/hashtable_class_helper.pxi\", line 998, in pandas._libs.hashtable.Int64HashTable.get_item\n  File \"pandas/_libs/hashtable_class_helper.pxi\", line 1005, in pandas._libs.hashtable.Int64HashTable.get_item\nKeyError: 1405</p>\n\n<p>During handling of the above exception, another exception occurred:</p>\n\n<p>Traceback (most recent call last):\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/worker.py\", line 178, in _worker_loop\n    data = fetcher.fetch(index)\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/fetch.py\", line 44, in fetch\n    data = [self.dataset[idx] for idx in possibly_batched_index]\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/fetch.py\", line 44, in \n    data = [self.dataset[idx] for idx in possibly_batched_index]\n  File \"\", line 11, in <strong>getitem</strong>\n    row = self.df.loc[idx]\n  File \"/opt/conda/lib/python3.7/site-packages/pandas/core/indexing.py\", line 1768, in <strong>getitem</strong>\n    return self._getitem_axis(maybe_callable, axis=axis)\n  File \"/opt/conda/lib/python3.7/site-packages/pandas/core/indexing.py\", line 1965, in _getitem_axis\n    return self._get_label(key, axis=axis)\n  File \"/opt/conda/lib/python3.7/site-packages/pandas/core/indexing.py\", line 625, in _get_label\n    return self.obj._xs(label, axis=axis)\n  File \"/opt/conda/lib/python3.7/site-packages/pandas/core/generic.py\", line 3537, in xs\n    loc = self.index.get_loc(key)\n  File \"/opt/conda/lib/python3.7/site-packages/pandas/core/indexes/base.py\", line 2648, in get_loc\n    return self._engine.get_loc(self._maybe_cast_indexer(key))\n  File \"pandas/_libs/index.pyx\", line 111, in pandas._libs.index.IndexEngine.get_loc\n  File \"pandas/_libs/index.pyx\", line 138, in pandas._libs.index.IndexEngine.get_loc\n  File \"pandas/_libs/hashtable_class_helper.pxi\", line 998, in pandas._libs.hashtable.Int64HashTable.get_item\n  File \"pandas/_libs/hashtable_class_helper.pxi\", line 1005, in pandas._libs.hashtable.Int64HashTable.get_item\nKeyError: 1405</p>\n\n<p>```\nPlease let me know on how to resolve this</p>",
      "rawMarkdown": "```\ntrain_dl = DataLoader(train_ds, batch_size, shuffle=True, num_workers=4, pin_memory=True)\n\nval_dl = DataLoader(val_ds, batch_size*2, shuffle = False,num_workers=4, pin_memory=True)\n```\nI have created the above train &amp; validation dataloaders and I am getting following error while enumerating over those mentioned dataloaders\n\n```\nKeyError: Caught KeyError in DataLoader worker process 0.\nOriginal Traceback (most recent call last):\n  File \"/opt/conda/lib/python3.7/site-packages/pandas/core/indexes/base.py\", line 2646, in get_loc\n    return self._engine.get_loc(key)\n  File \"pandas/_libs/index.pyx\", line 111, in pandas._libs.index.IndexEngine.get_loc\n  File \"pandas/_libs/index.pyx\", line 138, in pandas._libs.index.IndexEngine.get_loc\n  File \"pandas/_libs/hashtable_class_helper.pxi\", line 998, in pandas._libs.hashtable.Int64HashTable.get_item\n  File \"pandas/_libs/hashtable_class_helper.pxi\", line 1005, in pandas._libs.hashtable.Int64HashTable.get_item\nKeyError: 1405\n\nDuring handling of the above exception, another exception occurred:\n\nTraceback (most recent call last):\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/worker.py\", line 178, in _worker_loop\n    data = fetcher.fetch(index)\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/fetch.py\", line 44, in fetch\n    data = [self.dataset[idx] for idx in possibly_batched_index]\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/fetch.py\", line 44, in",
      "votes": null
    },
    {
      "id": "912164",
      "postDate": "07/02/2020 09:16:07",
      "content": "<p>Error is due to \"train_ds\"(Dataset). Something might be wrong in your CustomDataset function</p>",
      "rawMarkdown": "Error is due to \"train_ds\"(Dataset). Something might be wrong in your CustomDataset function",
      "votes": null
    },
    {
      "id": "912409",
      "postDate": "07/02/2020 13:10:08",
      "content": "<p>This is my custom dataset function</p>\n\n<p>```\nclass MelanomaClassificationDataset(Dataset):\n    def <strong>init</strong>(self, df, root_dir, transform=None):\n        self.df = df\n        self.transform = transform\n        self.root_dir = root_dir</p>\n\n<pre><code>def __len__(self):\n    return len(self.df)    \n\ndef __getitem__(self, idx):\n    row = self.df.loc[idx]\n    img_id, img_label = row['image_name'], row['target']\n    img_fname = self.root_dir + \"/\" + str(img_id) + \".png\"\n    img = Image.open(img_fname)\n    if self.transform:\n        img = self.transform(img)\n    return img, img_label\n</code></pre>\n\n<p><code>\n</code>\ndata_df = data_df.sample(frac=1).reset_index(drop = True)\nX = data_df.iloc[:,:-1]\ny = data_df.iloc[:,-1]</p>\n\n<p>print(\"X_columns :\", X.columns.tolist())\nprint(\"y_column :\", y.name)</p>\n\n<p>X_train, X_test, y_train, y_test = train_test_split(X, y,\n                                                    stratify=y, \n                                                    test_size=0.30 ,random_state = 42)</p>\n\n<p>train_df = X_train.copy()\ntrain_df['target'] = y_train</p>\n\n<p>val_df = X_test.copy()\nval_df['target'] = y_test</p>\n\n<p>train_ds = MelanomaClassificationDataset(train_df, TRAIN_RESIZED, transform=train_tfms)\nval_ds = MelanomaClassificationDataset(val_df, TRAIN_RESIZED, transform=valid_tfms)\n```</p>\n\n<p>Can you please tell me the mistake in the above code... I am unable to find it</p>\n\n<p>Thanks in advance</p>",
      "rawMarkdown": "This is my custom dataset function\n\n```\nclass MelanomaClassificationDataset(Dataset):\n    def __init__(self, df, root_dir, transform=None):\n        self.df = df\n        self.transform = transform\n        self.root_dir = root_dir\n        \n    def __len__(self):\n        return len(self.df)    \n    \n    def __getitem__(self, idx):\n        row = self.df.loc[idx]\n        img_id, img_label = row['image_name'], row['target']\n        img_fname = self.root_dir + \"/\" + str(img_id) + \".png\"\n        img = Image.open(img_fname)\n        if self.transform:\n            img = self.transform(img)\n        return img, img_label\n```\n```\ndata_df = data_df.sample(frac=1).reset_index(drop = True)\nX = data_df.iloc[:,:-1]\ny = data_df.iloc[:,-1]\n\nprint(\"X_columns :\", X.columns.tolist())\nprint(\"y_column :\", y.name)\n\nX_train, X_test, y_train, y_test = train_test_split(X, y,\n                                                    stratify=y, \n                                                    test_size=0.30 ,random_state = 42)\n\ntrain_df = X_train.copy()\ntrain_df['target'] = y_train\n\nval_df = X_test.copy()\nval_df['target'] = y_test\n\ntrain_ds = MelanomaClassificationDataset(train_df, TRAIN_RESIZED, transform=train_tfms)\nval_ds = MelanomaClassificationDataset(val_df, TRAIN_RESIZED, transform=valid_tfms)\n```\n\nCan you please tell me the mistake in the above code... I am unable to find it\n\nThanks in advance",
      "votes": null
    },
    {
      "id": "917062",
      "postDate": "07/06/2020 07:57:50",
      "content": "<p>I do able to resolve this issue, by resetting the index of train &amp; validation dataframes post train &amp; test split operation</p>\n\n<p>```\ndata_df = data_df.sample(frac=1).reset_index(drop = True)\nX = data_df.iloc[:,:-1]\ny = data_df.iloc[:,-1]</p>\n\n<p>print(\"X_columns :\", X.columns.tolist())\nprint(\"y_column :\", y.name)</p>\n\n<p>X_train, X_test, y_train, y_test = train_test_split(X, y,\n                                                    stratify=y, \n                                                    test_size=0.30 ,random_state = 42)</p>\n\n<p>train_df = X_train.copy()\ntrain_df['target'] = y_train\ntrain_df = train_df.reset_index()</p>\n\n<p>val_df = X_test.copy()\nval_df['target'] = y_test\nval_df = val_df.reset_index()\n```</p>",
      "rawMarkdown": "I do able to resolve this issue, by resetting the index of train &amp; validation dataframes post train &amp; test split operation\n\n```\ndata_df = data_df.sample(frac=1).reset_index(drop = True)\nX = data_df.iloc[:,:-1]\ny = data_df.iloc[:,-1]\n\nprint(\"X_columns :\", X.columns.tolist())\nprint(\"y_column :\", y.name)\n\nX_train, X_test, y_train, y_test = train_test_split(X, y,\n                                                    stratify=y, \n                                                    test_size=0.30 ,random_state = 42)\n\ntrain_df = X_train.copy()\ntrain_df['target'] = y_train\ntrain_df = train_df.reset_index()\n\nval_df = X_test.copy()\nval_df['target'] = y_test\nval_df = val_df.reset_index()\n```",
      "votes": null
    },
    {
      "id": "984952",
      "postDate": "08/25/2020 11:43:37",
      "content": "<p>Thank you for this fix.</p>",
      "rawMarkdown": "Thank you for this fix.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 912164,
      "author_name": "ishalgarg",
      "author_url": "",
      "post_date": "07/02/2020 09:16:07",
      "content": "<p>Error is due to \"train_ds\"(Dataset). Something might be wrong in your CustomDataset function</p>",
      "votes": null,
      "replies": [
        {
          "id": 912409,
          "author_name": "vineel369",
          "author_url": "",
          "post_date": "07/02/2020 13:10:08",
          "content": "<p>This is my custom dataset function</p>\n\n<p>```\nclass MelanomaClassificationDataset(Dataset):\n    def <strong>init</strong>(self, df, root_dir, transform=None):\n        self.df = df\n        self.transform = transform\n        self.root_dir = root_dir</p>\n\n<pre><code>def __len__(self):\n    return len(self.df)    \n\ndef __getitem__(self, idx):\n    row = self.df.loc[idx]\n    img_id, img_label = row['image_name'], row['target']\n    img_fname = self.root_dir + \"/\" + str(img_id) + \".png\"\n    img = Image.open(img_fname)\n    if self.transform:\n        img = self.transform(img)\n    return img, img_label\n</code></pre>\n\n<p><code>\n</code>\ndata_df = data_df.sample(frac=1).reset_index(drop = True)\nX = data_df.iloc[:,:-1]\ny = data_df.iloc[:,-1]</p>\n\n<p>print(\"X_columns :\", X.columns.tolist())\nprint(\"y_column :\", y.name)</p>\n\n<p>X_train, X_test, y_train, y_test = train_test_split(X, y,\n                                                    stratify=y, \n                                                    test_size=0.30 ,random_state = 42)</p>\n\n<p>train_df = X_train.copy()\ntrain_df['target'] = y_train</p>\n\n<p>val_df = X_test.copy()\nval_df['target'] = y_test</p>\n\n<p>train_ds = MelanomaClassificationDataset(train_df, TRAIN_RESIZED, transform=train_tfms)\nval_ds = MelanomaClassificationDataset(val_df, TRAIN_RESIZED, transform=valid_tfms)\n```</p>\n\n<p>Can you please tell me the mistake in the above code... I am unable to find it</p>\n\n<p>Thanks in advance</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 917062,
      "author_name": "vineel369",
      "author_url": "",
      "post_date": "07/06/2020 07:57:50",
      "content": "<p>I do able to resolve this issue, by resetting the index of train &amp; validation dataframes post train &amp; test split operation</p>\n\n<p>```\ndata_df = data_df.sample(frac=1).reset_index(drop = True)\nX = data_df.iloc[:,:-1]\ny = data_df.iloc[:,-1]</p>\n\n<p>print(\"X_columns :\", X.columns.tolist())\nprint(\"y_column :\", y.name)</p>\n\n<p>X_train, X_test, y_train, y_test = train_test_split(X, y,\n                                                    stratify=y, \n                                                    test_size=0.30 ,random_state = 42)</p>\n\n<p>train_df = X_train.copy()\ntrain_df['target'] = y_train\ntrain_df = train_df.reset_index()</p>\n\n<p>val_df = X_test.copy()\nval_df['target'] = y_test\nval_df = val_df.reset_index()\n```</p>",
      "votes": null,
      "replies": [
        {
          "id": 984952,
          "author_name": "uadhikari",
          "author_url": "",
          "post_date": "08/25/2020 11:43:37",
          "content": "<p>Thank you for this fix.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "912085": "```\ntrain_dl = DataLoader(train_ds, batch_size, shuffle=True, num_workers=4, pin_memory=True)\n\nval_dl = DataLoader(val_ds, batch_size*2, shuffle = False,num_workers=4, pin_memory=True)\n```\nI have created the above train &amp; validation dataloaders and I am getting following error while enumerating over those mentioned dataloaders\n\n```\nKeyError: Caught KeyError in DataLoader worker process 0.\nOriginal Traceback (most recent call last):\n  File \"/opt/conda/lib/python3.7/site-packages/pandas/core/indexes/base.py\", line 2646, in get_loc\n    return self._engine.get_loc(key)\n  File \"pandas/_libs/index.pyx\", line 111, in pandas._libs.index.IndexEngine.get_loc\n  File \"pandas/_libs/index.pyx\", line 138, in pandas._libs.index.IndexEngine.get_loc\n  File \"pandas/_libs/hashtable_class_helper.pxi\", line 998, in pandas._libs.hashtable.Int64HashTable.get_item\n  File \"pandas/_libs/hashtable_class_helper.pxi\", line 1005, in pandas._libs.hashtable.Int64HashTable.get_item\nKeyError: 1405\n\nDuring handling of the above exception, another exception occurred:\n\nTraceback (most recent call last):\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/worker.py\", line 178, in _worker_loop\n    data = fetcher.fetch(index)\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/fetch.py\", line 44, in fetch\n    data = [self.dataset[idx] for idx in possibly_batched_index]\n  File \"/opt/conda/lib/python3.7/site-packages/torch/utils/data/_utils/fetch.py\", line 44, in",
    "912164": "Error is due to \"train_ds\"(Dataset). Something might be wrong in your CustomDataset function",
    "912409": "This is my custom dataset function\n\n```\nclass MelanomaClassificationDataset(Dataset):\n    def __init__(self, df, root_dir, transform=None):\n        self.df = df\n        self.transform = transform\n        self.root_dir = root_dir\n        \n    def __len__(self):\n        return len(self.df)    \n    \n    def __getitem__(self, idx):\n        row = self.df.loc[idx]\n        img_id, img_label = row['image_name'], row['target']\n        img_fname = self.root_dir + \"/\" + str(img_id) + \".png\"\n        img = Image.open(img_fname)\n        if self.transform:\n            img = self.transform(img)\n        return img, img_label\n```\n```\ndata_df = data_df.sample(frac=1).reset_index(drop = True)\nX = data_df.iloc[:,:-1]\ny = data_df.iloc[:,-1]\n\nprint(\"X_columns :\", X.columns.tolist())\nprint(\"y_column :\", y.name)\n\nX_train, X_test, y_train, y_test = train_test_split(X, y,\n                                                    stratify=y, \n                                                    test_size=0.30 ,random_state = 42)\n\ntrain_df = X_train.copy()\ntrain_df['target'] = y_train\n\nval_df = X_test.copy()\nval_df['target'] = y_test\n\ntrain_ds = MelanomaClassificationDataset(train_df, TRAIN_RESIZED, transform=train_tfms)\nval_ds = MelanomaClassificationDataset(val_df, TRAIN_RESIZED, transform=valid_tfms)\n```\n\nCan you please tell me the mistake in the above code... I am unable to find it\n\nThanks in advance",
    "917062": "I do able to resolve this issue, by resetting the index of train &amp; validation dataframes post train &amp; test split operation\n\n```\ndata_df = data_df.sample(frac=1).reset_index(drop = True)\nX = data_df.iloc[:,:-1]\ny = data_df.iloc[:,-1]\n\nprint(\"X_columns :\", X.columns.tolist())\nprint(\"y_column :\", y.name)\n\nX_train, X_test, y_train, y_test = train_test_split(X, y,\n                                                    stratify=y, \n                                                    test_size=0.30 ,random_state = 42)\n\ntrain_df = X_train.copy()\ntrain_df['target'] = y_train\ntrain_df = train_df.reset_index()\n\nval_df = X_test.copy()\nval_df['target'] = y_test\nval_df = val_df.reset_index()\n```",
    "984952": "Thank you for this fix."
  },
  "source": "meta"
}