{
  "id": 266816,
  "title": "TPU error Please help ",
  "url": "/competitions/landmark-recognition-2021/discussion/266816",
  "author_name": "",
  "post_date": "2021-08-20T15:25:43.196060900Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I am building the pipeline to train model on tpu but end up getting this error please help.</p>\n<p>Exception in device=TPU:0: /pytorch/xla/third_party/tensorflow/bazel-tensorflow/tensorflow/compiler/xla/xla_client/debug_macros.h:27 : Check failed: status.status() == ::tensorflow::Status::OK() (Invalid argument: Automatic shape inference not supported: bf16[64,81313] and bf16[64,81313,1] vs. OK)<br>\n*** Begin stack trace ***<br>\n    tensorflow::CurrentStackTrace()<br>\n    xla::Shape const* ConsumeValue(stream_executor::port::StatusOr&amp;&amp;)<br>\n    torch_xla::XlaHelpers::ShapeOfXlaOp(xla::XlaOp)</p>\n<pre><code>torch_xla::BuildNllLoss(xla::XlaOp, xla::XlaOp, xla::XlaOp, int, torch_xla::ReductionMode)\n\ntorch_xla::ir::ops::InferOutputShape(absl::lts_2020_02_25::Span&lt;xla::Shape const&gt;, std::function&lt;xla::XlaOp (absl::lts_2020_02_25::Span&lt;xla::XlaOp const&gt;)&gt; const&amp;)\n\ntorch_xla::ir::Node::GetOpShape(std::function&lt;xla::Shape ()&gt; const&amp;) const\ntorch_xla::ir::Node::Node(torch_xla::ir::OpKind, absl::lts_2020_02_25::Span&lt;torch_xla::ir::Value const&gt;, std::function&lt;xla::Shape ()&gt; const&amp;, unsigned long, absl::lts_2020_02_25::uint128)\ntorch_xla::ir::ops::NllLoss::NllLoss(torch_xla::ir::Value const&amp;, torch_xla::ir::Value const&amp;, absl::lts_2020_02_25::optional&lt;torch_xla::ir::Value&gt; const&amp;, torch_xla::ReductionMode, int)\ntorch_xla::XLATensor::nll_loss(torch_xla::XLATensor const&amp;, torch_xla::XLATensor const&amp;, torch_xla::XLATensor const&amp;, long, int)\ntorch_xla::AtenXlaType::nll_loss_forward(at::Tensor const&amp;, at::Tensor const&amp;, c10::optional&lt;at::Tensor&gt; const&amp;, long, long)\nc10::impl::wrap_kernel_functor_unboxed_&lt;c10::impl::detail::WrapFunctionIntoRuntimeFunctor_&lt;std::tuple&lt;at::Tensor, at::Tensor&gt; (*)(at::Tensor const&amp;, at::Tensor const&amp;, c10::optional&lt;at::Tensor&gt; const&amp;, long, long), std::tuple&lt;at::Tensor, at::Tensor&gt;, c10::guts::typelist::typelist&lt;at::Tensor const&amp;, at::Tensor const&amp;, c10::optional&lt;at::Tensor&gt; const&amp;, long, long&gt; &gt;, std::tuple&lt;at::Tensor, at::Tensor&gt; (at::Tensor const&amp;, at::Tensor const&amp;, c10::optional&lt;at::Tensor&gt; const&amp;, long, long)&gt;::call(c10::OperatorKernel*, at::Tensor const&amp;, at::Tensor const&amp;, c10::optional&lt;at::Tensor&gt; const&amp;, long, long)\n</code></pre>",
  "messages": [
    {
      "id": "1483336",
      "postDate": "08/20/2021 15:25:43",
      "content": "<p>I am building the pipeline to train model on tpu but end up getting this error please help.</p>\n<p>Exception in device=TPU:0: /pytorch/xla/third_party/tensorflow/bazel-tensorflow/tensorflow/compiler/xla/xla_client/debug_macros.h:27 : Check failed: status.status() == ::tensorflow::Status::OK() (Invalid argument: Automatic shape inference not supported: bf16[64,81313] and bf16[64,81313,1] vs. OK)<br>\n*** Begin stack trace ***<br>\n    tensorflow::CurrentStackTrace()<br>\n    xla::Shape const* ConsumeValue(stream_executor::port::StatusOr&amp;&amp;)<br>\n    torch_xla::XlaHelpers::ShapeOfXlaOp(xla::XlaOp)</p>\n<pre><code>torch_xla::BuildNllLoss(xla::XlaOp, xla::XlaOp, xla::XlaOp, int, torch_xla::ReductionMode)\n\ntorch_xla::ir::ops::InferOutputShape(absl::lts_2020_02_25::Span&lt;xla::Shape const&gt;, std::function&lt;xla::XlaOp (absl::lts_2020_02_25::Span&lt;xla::XlaOp const&gt;)&gt; const&amp;)\n\ntorch_xla::ir::Node::GetOpShape(std::function&lt;xla::Shape ()&gt; const&amp;) const\ntorch_xla::ir::Node::Node(torch_xla::ir::OpKind, absl::lts_2020_02_25::Span&lt;torch_xla::ir::Value const&gt;, std::function&lt;xla::Shape ()&gt; const&amp;, unsigned long, absl::lts_2020_02_25::uint128)\ntorch_xla::ir::ops::NllLoss::NllLoss(torch_xla::ir::Value const&amp;, torch_xla::ir::Value const&amp;, absl::lts_2020_02_25::optional&lt;torch_xla::ir::Value&gt; const&amp;, torch_xla::ReductionMode, int)\ntorch_xla::XLATensor::nll_loss(torch_xla::XLATensor const&amp;, torch_xla::XLATensor const&amp;, torch_xla::XLATensor const&amp;, long, int)\ntorch_xla::AtenXlaType::nll_loss_forward(at::Tensor const&amp;, at::Tensor const&amp;, c10::optional&lt;at::Tensor&gt; const&amp;, long, long)\nc10::impl::wrap_kernel_functor_unboxed_&lt;c10::impl::detail::WrapFunctionIntoRuntimeFunctor_&lt;std::tuple&lt;at::Tensor, at::Tensor&gt; (*)(at::Tensor const&amp;, at::Tensor const&amp;, c10::optional&lt;at::Tensor&gt; const&amp;, long, long), std::tuple&lt;at::Tensor, at::Tensor&gt;, c10::guts::typelist::typelist&lt;at::Tensor const&amp;, at::Tensor const&amp;, c10::optional&lt;at::Tensor&gt; const&amp;, long, long&gt; &gt;, std::tuple&lt;at::Tensor, at::Tensor&gt; (at::Tensor const&amp;, at::Tensor const&amp;, c10::optional&lt;at::Tensor&gt; const&amp;, long, long)&gt;::call(c10::OperatorKernel*, at::Tensor const&amp;, at::Tensor const&amp;, c10::optional&lt;at::Tensor&gt; const&amp;, long, long)\n</code></pre>",
      "rawMarkdown": "I am building the pipeline to train model on tpu but end up getting this error please help.\n\n\nException in device=TPU:0: /pytorch/xla/third_party/tensorflow/bazel-tensorflow/tensorflow/compiler/xla/xla_client/debug_macros.h:27 : Check failed: status.status() == ::tensorflow::Status::OK() (Invalid argument: Automatic shape inference not supported: bf16[64,81313] and bf16[64,81313,1] vs. OK)\n*** Begin stack trace ***\n\ttensorflow::CurrentStackTrace()\n\txla::Shape const* ConsumeValue<xla::Shape const*>(stream_executor::port::StatusOr<xla::Shape const*>&&)\n\ttorch_xla::XlaHelpers::ShapeOfXlaOp(xla::XlaOp)\n\t\n\ttorch_xla::BuildNllLoss(xla::XlaOp, xla::XlaOp, xla::XlaOp, int, torch_xla::ReductionMode)\n\t\n\ttorch_xla::ir::ops::InferOutputShape(absl::lts_2020_02_25::Span<xla::Shape const>, std::function<xla::XlaOp (absl::lts_2020_02_25::Span<xla::XlaOp const>)> const&)\n\t\n\ttorch_xla::ir::Node::GetOpShape(std::function<xla::Shape ()> const&) const\n\ttorch_xla::ir::Node::Node(torch_xla::ir::OpKind, absl::lts_2020_02_25::Span<torch_xla::ir::Value const>, std::function<xla::Shape ()> const&, unsigned long, absl::lts_2020_02_25::uint128)\n\ttorch_xla::ir::ops::NllLoss::NllLoss(torch_xla::ir::Value const&, torch_xla::ir::Value const&, absl::lts_2020_02_25::optional<torch_xla::ir::Value> const&, torch_xla::ReductionMode, int)\n\ttorch_xla::XLATensor::nll_loss(torch_xla::XLATensor const&, torch_xla::XLATensor const&, torch_xla::XLATensor const&, long, int)\n\ttorch_xla::AtenXlaType::nll_loss_forward(at::Tensor const&, at::Tensor const&, c10::optional<at::Tensor> const&, long, long)\n\tc10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoRuntimeFunctor_<std::tuple<at::Tensor, at::Tensor> (*)(at::Tensor const&, at::Tensor const&, c10::optional<at::Tensor> const&, long, long), std::tuple<at::Tensor, at::Tensor>, c10::guts::typelist::typelist<at::Tensor const&, at::Tensor const&, c10::optional<at::Tensor> const&, long, long> >, std::tuple<at::Tensor, at::Tensor> (at::Tensor const&, at::Tensor const&, c10::optional<at::Tensor> const&, long, long)>::call(c10::OperatorKernel*, at::Tensor const&, at::Tensor const&, c10::optional<at::Tensor> const&, long, long)",
      "votes": null
    },
    {
      "id": "1484686",
      "postDate": "08/21/2021 13:58:24",
      "content": "<p>The most eye-catching part of the error for me is</p>\n<pre><code>Automatic shape inference not supported\n</code></pre>\n<p>TPU's generally require explicit shapes. Without seeing your code, I would suggest looking at parts of your code where images are read and call a reshape on it to explicitly state the shape. I had roughly the same error once and fixed it this way.<br>\nThis would look roughly like this:</p>\n<pre><code>image = tf.io.decode_jpeg(file_path) # You know this image is 256x256, but the TPU doesn't!\nimage = tf.reshape(image, [256, 256, 3]) # Now the TPU knows it is 256x256!\n</code></pre>",
      "rawMarkdown": "The most eye-catching part of the error for me is\n\n```\nAutomatic shape inference not supported\n```\n\nTPU's generally require explicit shapes. Without seeing your code, I would suggest looking at parts of your code where images are read and call a reshape on it to explicitly state the shape. I had roughly the same error once and fixed it this way.\nThis would look roughly like this:\n\n```\nimage = tf.io.decode_jpeg(file_path) # You know this image is 256x256, but the TPU doesn't!\nimage = tf.reshape(image, [256, 256, 3]) # Now the TPU knows it is 256x256!\n```",
      "votes": null
    },
    {
      "id": "1484853",
      "postDate": "08/21/2021 15:57:04",
      "content": "<p>class Dataset:<br>\n    def <strong>init</strong>(self, df):<br>\n        self.df = df<br>\n        self.transforms = transforms.Compose(transforms.Resize((256,256)))</p>\n<pre><code>def __len__(self):\n    return len(self.df)\n\ndef __getitem__(self, index):\n    label = self.df.iloc[index].landmark_id\n    filename = self.df.iloc[index].path\n    image = Image.open(filename).convert('L')\n    image = cv2.resize(np.array(image), dsize=(256,256))     \n    data_tensor = torch.tensor(image).float().unsqueeze(0)\n    #data_tensor = data_tensor.permute(2, 0, 1)\n\n    return (data_tensor, torch.tensor(label))\n</code></pre>\n<p>This is my Dataset class. Please have a look.</p>",
      "rawMarkdown": "class Dataset:\n    def __init__(self, df):\n        self.df = df\n        self.transforms = transforms.Compose(transforms.Resize((256,256)))\n        \n    def __len__(self):\n        return len(self.df)\n\n    def __getitem__(self, index):\n        label = self.df.iloc[index].landmark_id\n        filename = self.df.iloc[index].path\n        image = Image.open(filename).convert('L')\n        image = cv2.resize(np.array(image), dsize=(256,256))     \n        data_tensor = torch.tensor(image).float().unsqueeze(0)\n        #data_tensor = data_tensor.permute(2, 0, 1)\n        \n        return (data_tensor, torch.tensor(label))\n\n\nThis is my Dataset class. Please have a look.",
      "votes": null
    },
    {
      "id": "1484866",
      "postDate": "08/21/2021 16:14:08",
      "content": "<p>I have little experience with Pytorch, but you can try the following:</p>\n<pre><code>def __getitem__(self, index):\n    label = self.df.iloc[index].landmark_id\n    filename = self.df.iloc[index].path\n    image = Image.open(filename).convert('L')\n    image = cv2.resize(np.array(image), dsize=(256,256))     \n    data_tensor = torch.tensor(image).float()\n    # Explicitly state the shape so the compiler can inference it\n    data_tensor = torch.reshape(data_tensor, (256,256,3)).unsqueeze(0)\n    #data_tensor = data_tensor.permute(2, 0, 1)\n\n    return (data_tensor, torch.tensor(label))\n</code></pre>\n<p>However, looking more closely at your code, I think the error is in your label. Your output layer has 81313 neurons, meaning the maximum label value is 81312 (zero based). The <code>landmark_id</code> is however not continuous, resulting in label values far larger than 81312. You should create continuous label values as follows.</p>\n<pre><code>train['label'] = train['landmark_id'].astype('category').cat.codes\n</code></pre>\n<p>Another quick way to check if that's the error is to hardcode to label to 0, thus replacing <code>label = self.df.iloc[index].landmark_id</code> with <code>label = 0</code></p>\n<p>What also surprises me is the combination of Pytorch and Tensorflow, you seem to use both frameworks. This is possible, but its asking for weird errors in practice.</p>\n<p>Hope this helps!</p>\n<p>P.S. I am working on a training notebook as well, planning on publishing it somewhere next week.</p>",
      "rawMarkdown": "I have little experience with Pytorch, but you can try the following:\n\n```\ndef __getitem__(self, index):\n    label = self.df.iloc[index].landmark_id\n    filename = self.df.iloc[index].path\n    image = Image.open(filename).convert('L')\n    image = cv2.resize(np.array(image), dsize=(256,256))     \n    data_tensor = torch.tensor(image).float()\n    # Explicitly state the shape so the compiler can inference it\n    data_tensor = torch.reshape(data_tensor, (256,256,3)).unsqueeze(0)\n    #data_tensor = data_tensor.permute(2, 0, 1)\n\n    return (data_tensor, torch.tensor(label))\n```\n\nHowever, looking more closely at your code, I think the error is in your label. Your output layer has 81313 neurons, meaning the maximum label value is 81312 (zero based). The `landmark_id` is however not continuous, resulting in label values far larger than 81312. You should create continuous label values as follows.\n\n```\ntrain['label'] = train['landmark_id'].astype('category').cat.codes\n```\n\nAnother quick way to check if that's the error is to hardcode to label to 0, thus replacing `label = self.df.iloc[index].landmark_id` with `label = 0`\n\nWhat also surprises me is the combination of Pytorch and Tensorflow, you seem to use both frameworks. This is possible, but its asking for weird errors in practice.\n\nHope this helps!\n\nP.S. I am working on a training notebook as well, planning on publishing it somewhere next week.",
      "votes": null
    },
    {
      "id": "1485101",
      "postDate": "08/21/2021 19:30:11",
      "content": "<p>One last thought, try to print the labels in your dataset using the following:</p>\n<pre><code>images, labels = next(iter(your_dataset))\nprint(labels[0])\n</code></pre>\n<p>It looks like your labels are wrapped as the error message states <code>bf16[64,81313] and bf16[64,81313,1]</code>. Thus your labels might be [42] instead of just 42, unsqueezing the label in the dataloader should do the trick if this is the case.</p>",
      "rawMarkdown": "One last thought, try to print the labels in your dataset using the following:\n\n```\nimages, labels = next(iter(your_dataset))\nprint(labels[0])\n```\n\nIt looks like your labels are wrapped as the error message states `bf16[64,81313] and bf16[64,81313,1]`. Thus your labels might be [42] instead of just 42, unsqueezing the label in the dataloader should do the trick if this is the case.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1484686,
      "author_name": "markwijkhuizen",
      "author_url": "",
      "post_date": "08/21/2021 13:58:24",
      "content": "<p>The most eye-catching part of the error for me is</p>\n<pre><code>Automatic shape inference not supported\n</code></pre>\n<p>TPU's generally require explicit shapes. Without seeing your code, I would suggest looking at parts of your code where images are read and call a reshape on it to explicitly state the shape. I had roughly the same error once and fixed it this way.<br>\nThis would look roughly like this:</p>\n<pre><code>image = tf.io.decode_jpeg(file_path) # You know this image is 256x256, but the TPU doesn't!\nimage = tf.reshape(image, [256, 256, 3]) # Now the TPU knows it is 256x256!\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1484853,
          "author_name": "riyajm",
          "author_url": "",
          "post_date": "08/21/2021 15:57:04",
          "content": "<p>class Dataset:<br>\n    def <strong>init</strong>(self, df):<br>\n        self.df = df<br>\n        self.transforms = transforms.Compose(transforms.Resize((256,256)))</p>\n<pre><code>def __len__(self):\n    return len(self.df)\n\ndef __getitem__(self, index):\n    label = self.df.iloc[index].landmark_id\n    filename = self.df.iloc[index].path\n    image = Image.open(filename).convert('L')\n    image = cv2.resize(np.array(image), dsize=(256,256))     \n    data_tensor = torch.tensor(image).float().unsqueeze(0)\n    #data_tensor = data_tensor.permute(2, 0, 1)\n\n    return (data_tensor, torch.tensor(label))\n</code></pre>\n<p>This is my Dataset class. Please have a look.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1484866,
          "author_name": "markwijkhuizen",
          "author_url": "",
          "post_date": "08/21/2021 16:14:08",
          "content": "<p>I have little experience with Pytorch, but you can try the following:</p>\n<pre><code>def __getitem__(self, index):\n    label = self.df.iloc[index].landmark_id\n    filename = self.df.iloc[index].path\n    image = Image.open(filename).convert('L')\n    image = cv2.resize(np.array(image), dsize=(256,256))     \n    data_tensor = torch.tensor(image).float()\n    # Explicitly state the shape so the compiler can inference it\n    data_tensor = torch.reshape(data_tensor, (256,256,3)).unsqueeze(0)\n    #data_tensor = data_tensor.permute(2, 0, 1)\n\n    return (data_tensor, torch.tensor(label))\n</code></pre>\n<p>However, looking more closely at your code, I think the error is in your label. Your output layer has 81313 neurons, meaning the maximum label value is 81312 (zero based). The <code>landmark_id</code> is however not continuous, resulting in label values far larger than 81312. You should create continuous label values as follows.</p>\n<pre><code>train['label'] = train['landmark_id'].astype('category').cat.codes\n</code></pre>\n<p>Another quick way to check if that's the error is to hardcode to label to 0, thus replacing <code>label = self.df.iloc[index].landmark_id</code> with <code>label = 0</code></p>\n<p>What also surprises me is the combination of Pytorch and Tensorflow, you seem to use both frameworks. This is possible, but its asking for weird errors in practice.</p>\n<p>Hope this helps!</p>\n<p>P.S. I am working on a training notebook as well, planning on publishing it somewhere next week.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1485101,
          "author_name": "markwijkhuizen",
          "author_url": "",
          "post_date": "08/21/2021 19:30:11",
          "content": "<p>One last thought, try to print the labels in your dataset using the following:</p>\n<pre><code>images, labels = next(iter(your_dataset))\nprint(labels[0])\n</code></pre>\n<p>It looks like your labels are wrapped as the error message states <code>bf16[64,81313] and bf16[64,81313,1]</code>. Thus your labels might be [42] instead of just 42, unsqueezing the label in the dataloader should do the trick if this is the case.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1483336": "I am building the pipeline to train model on tpu but end up getting this error please help.\n\n\nException in device=TPU:0: /pytorch/xla/third_party/tensorflow/bazel-tensorflow/tensorflow/compiler/xla/xla_client/debug_macros.h:27 : Check failed: status.status() == ::tensorflow::Status::OK() (Invalid argument: Automatic shape inference not supported: bf16[64,81313] and bf16[64,81313,1] vs. OK)\n*** Begin stack trace ***\n\ttensorflow::CurrentStackTrace()\n\txla::Shape const* ConsumeValue<xla::Shape const*>(stream_executor::port::StatusOr<xla::Shape const*>&&)\n\ttorch_xla::XlaHelpers::ShapeOfXlaOp(xla::XlaOp)\n\t\n\ttorch_xla::BuildNllLoss(xla::XlaOp, xla::XlaOp, xla::XlaOp, int, torch_xla::ReductionMode)\n\t\n\ttorch_xla::ir::ops::InferOutputShape(absl::lts_2020_02_25::Span<xla::Shape const>, std::function<xla::XlaOp (absl::lts_2020_02_25::Span<xla::XlaOp const>)> const&)\n\t\n\ttorch_xla::ir::Node::GetOpShape(std::function<xla::Shape ()> const&) const\n\ttorch_xla::ir::Node::Node(torch_xla::ir::OpKind, absl::lts_2020_02_25::Span<torch_xla::ir::Value const>, std::function<xla::Shape ()> const&, unsigned long, absl::lts_2020_02_25::uint128)\n\ttorch_xla::ir::ops::NllLoss::NllLoss(torch_xla::ir::Value const&, torch_xla::ir::Value const&, absl::lts_2020_02_25::optional<torch_xla::ir::Value> const&, torch_xla::ReductionMode, int)\n\ttorch_xla::XLATensor::nll_loss(torch_xla::XLATensor const&, torch_xla::XLATensor const&, torch_xla::XLATensor const&, long, int)\n\ttorch_xla::AtenXlaType::nll_loss_forward(at::Tensor const&, at::Tensor const&, c10::optional<at::Tensor> const&, long, long)\n\tc10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoRuntimeFunctor_<std::tuple<at::Tensor, at::Tensor> (*)(at::Tensor const&, at::Tensor const&, c10::optional<at::Tensor> const&, long, long), std::tuple<at::Tensor, at::Tensor>, c10::guts::typelist::typelist<at::Tensor const&, at::Tensor const&, c10::optional<at::Tensor> const&, long, long> >, std::tuple<at::Tensor, at::Tensor> (at::Tensor const&, at::Tensor const&, c10::optional<at::Tensor> const&, long, long)>::call(c10::OperatorKernel*, at::Tensor const&, at::Tensor const&, c10::optional<at::Tensor> const&, long, long)",
    "1484686": "The most eye-catching part of the error for me is\n\n```\nAutomatic shape inference not supported\n```\n\nTPU's generally require explicit shapes. Without seeing your code, I would suggest looking at parts of your code where images are read and call a reshape on it to explicitly state the shape. I had roughly the same error once and fixed it this way.\nThis would look roughly like this:\n\n```\nimage = tf.io.decode_jpeg(file_path) # You know this image is 256x256, but the TPU doesn't!\nimage = tf.reshape(image, [256, 256, 3]) # Now the TPU knows it is 256x256!\n```",
    "1484853": "class Dataset:\n    def __init__(self, df):\n        self.df = df\n        self.transforms = transforms.Compose(transforms.Resize((256,256)))\n        \n    def __len__(self):\n        return len(self.df)\n\n    def __getitem__(self, index):\n        label = self.df.iloc[index].landmark_id\n        filename = self.df.iloc[index].path\n        image = Image.open(filename).convert('L')\n        image = cv2.resize(np.array(image), dsize=(256,256))     \n        data_tensor = torch.tensor(image).float().unsqueeze(0)\n        #data_tensor = data_tensor.permute(2, 0, 1)\n        \n        return (data_tensor, torch.tensor(label))\n\n\nThis is my Dataset class. Please have a look.",
    "1484866": "I have little experience with Pytorch, but you can try the following:\n\n```\ndef __getitem__(self, index):\n    label = self.df.iloc[index].landmark_id\n    filename = self.df.iloc[index].path\n    image = Image.open(filename).convert('L')\n    image = cv2.resize(np.array(image), dsize=(256,256))     \n    data_tensor = torch.tensor(image).float()\n    # Explicitly state the shape so the compiler can inference it\n    data_tensor = torch.reshape(data_tensor, (256,256,3)).unsqueeze(0)\n    #data_tensor = data_tensor.permute(2, 0, 1)\n\n    return (data_tensor, torch.tensor(label))\n```\n\nHowever, looking more closely at your code, I think the error is in your label. Your output layer has 81313 neurons, meaning the maximum label value is 81312 (zero based). The `landmark_id` is however not continuous, resulting in label values far larger than 81312. You should create continuous label values as follows.\n\n```\ntrain['label'] = train['landmark_id'].astype('category').cat.codes\n```\n\nAnother quick way to check if that's the error is to hardcode to label to 0, thus replacing `label = self.df.iloc[index].landmark_id` with `label = 0`\n\nWhat also surprises me is the combination of Pytorch and Tensorflow, you seem to use both frameworks. This is possible, but its asking for weird errors in practice.\n\nHope this helps!\n\nP.S. I am working on a training notebook as well, planning on publishing it somewhere next week.",
    "1485101": "One last thought, try to print the labels in your dataset using the following:\n\n```\nimages, labels = next(iter(your_dataset))\nprint(labels[0])\n```\n\nIt looks like your labels are wrapped as the error message states `bf16[64,81313] and bf16[64,81313,1]`. Thus your labels might be [42] instead of just 42, unsqueezing the label in the dataloader should do the trick if this is the case."
  },
  "source": "meta"
}