{
  "id": 402684,
  "title": "TPU VM Tensorflow 2.12.0 problem with ConvNeX",
  "url": "/competitions/birdclef-2023/discussion/402684",
  "author_name": "Andrij",
  "post_date": "2023-04-19T11:49:55.125000",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi to all! <br>\nI want to try ConvNeX. <br>\nFor this I installed tf 2.12.0 which includes ConvNeX models. But now the training time of one epoch is 50+ hours. Does anyone know what the problem is? <br>\nMaybe someone was able to run the ConvNeX on TPU VM in a different way. <br>\nHere is my code notebook to understand the problem:<a href=\"https://www.kaggle.com/code/aikhmelnytskyy/birdclef23-trainconvnext/notebook\" target=\"_blank\">https://www.kaggle.com/code/aikhmelnytskyy/birdclef23-trainconvnext/notebook</a>. <br>\nThanks for the answer!</p>",
  "messages": [
    {
      "id": 2230301,
      "postDate": "2023-04-22T07:59:33.180Z",
      "content": "<p><a href=\"https://www.kaggle.com/aikhmelnytskyy\" target=\"_blank\">@aikhmelnytskyy</a> <br>\nI've just tried ConvNeXt on TPU-VM on an image dataset and it works just fine; <a href=\"https://drive.google.com/file/d/14ii3dhjPGZXJHHOZrEw4Tlw9P07S31SX/view?usp=sharing\" target=\"_blank\">gist</a>. (this gist will be removed in future). Tips, don't install tf 2.12 manually, it comes by default now. </p>",
      "rawMarkdown": "@aikhmelnytskyy \nI've just tried ConvNeXt on TPU-VM on an image dataset and it works just fine; [gist](https://drive.google.com/file/d/14ii3dhjPGZXJHHOZrEw4Tlw9P07S31SX/view?usp=sharing). (this gist will be removed in future). Tips, don't install tf 2.12 manually, it comes by default now. ",
      "votes": 1
    },
    {
      "id": 2227118,
      "postDate": "2023-04-19T14:13:13.327Z",
      "content": "<p>trains fine for me if I do it via <a href=\"https://github.com/leondgarse/keras_cv_attention_models\" target=\"_blank\">https://github.com/leondgarse/keras_cv_attention_models</a></p>",
      "rawMarkdown": "trains fine for me if I do it via https://github.com/leondgarse/keras_cv_attention_models",
      "votes": 1,
      "replies": [
        {
          "id": 2227128,
          "postDate": "2023-04-19T14:19:55.567Z",
          "content": "<p>Thank you. I tried Keras_cv_attention_models but I get an error: 2023-04-19 14:43:04.927404: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:903] model_pruner failed: INVALID_ARGUMENT: Graph does not contain terminal node Adam/Adam/AssignAddVariableOp.<br>\n2023-04-19 14:43:07.408151: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:903] model_pruner failed: INVALID_ARGUMENT: Graph does not contain terminal node Adam/Adam/AssignAddVariableOp.<br>\n2023-04-19 14:43:32.480805: E tensorflow/core/tpu/kernels/tpu_compilation_cache_external.cc:113] Reshape's input dynamic dimension is decomposed into multiple output dynamic dimensions, but the constraint is ambiguous and XLA can't infer the output dimension %reshape.11975 = f32[32,32,96,512]{3,2,1,0} reshape(f32[98304,512]{1,0} %transpose.11967), metadata={op_type=\"Reshape\" op_name=\"model/convnext_base/stack1_block1_up_dense/Tensordot\"}. <br>\n2023-04-19 14:43:32.481013: F tensorflow/core/tpu/kernels/tpu_program_group.cc:86] Check failed: xla_tpu_programs.size() &gt; 0 (0 vs. 0)<br>\n<a href=\"https://symbolize.stripped_domain/r/?trace=7ff24442ace1,7ff24442ad5f,7ff2053827a7,7ff20b6e149d,7ff20b74e769,7ff20b74f259,7ff20b745c96,7ff20b747f5c,7ff2256691be,7ff22560be4f,7ff2255f6921,7ff20b6b5444,7ff20b6b3116,7ff225d7728e,7ff2443d7ea6&amp;map=947919304d8b3d847ddb4dfdea237b00f15d2086:7ff202077000-7ff217973140,886380cfe14c768b793f7ee12153f640442e39c4:7ff224ac3000-7ff22673a6c0\" target=\"_blank\">https://symbolize.stripped_domain/r/?trace=7ff24442ace1,7ff24442ad5f,7ff2053827a7,7ff20b6e149d,7ff20b74e769,7ff20b74f259,7ff20b745c96,7ff20b747f5c,7ff2256691be,7ff22560be4f,7ff2255f6921,7ff20b6b5444,7ff20b6b3116,7ff225d7728e,7ff2443d7ea6&amp;map=947919304d8b3d847ddb4dfdea237b00f15d2086:7ff202077000-7ff217973140,886380cfe14c768b793f7ee12153f640442e39c4:7ff224ac3000-7ff22673a6c0</a> <br>\n*** SIGABRT received by PID 14 (TID 1029) on cpu 9 from PID 14; stack trace: ***<br>\nPC: @     0x7ff24442ace1  (unknown)  raise<br>\n    @     0x7ff201410ed3        992  (unknown)<br>\n    @     0x7ff24442ad60       3904  (unknown)<br>\n    @     0x7ff2053827a8        896  tensorflow::tpu::TpuProgramGroup::Initialize()<br>\n    @     0x7ff20b6e149e       1696  tensorflow::tpu::TpuCompilationCacheExternal::InitializeEntry()<br>\n    @     0x7ff20b74e76a       1072  tensorflow::tpu::TpuCompilationCacheInterface::CompileIfKeyAbsentHelper()<br>\n    @     0x7ff20b74f25a        128  tensorflow::tpu::TpuCompilationCacheInterface::CompileIfKeyAbsent()<br>\n    @     0x7ff20b745c97       1312  tensorflow::tpu::TpuCompileOpKernelCommon::ComputeInternal()<br>\n    @     0x7ff20b747f5d        624  tensorflow::tpu::TpuCompileOpKernelCommon::Compute()<br>\n    @     0x7ff2256691bf        496  tensorflow::ThreadPoolDevice::Compute()<br>\n    @     0x7ff22560be50       2496  tensorflow::(anonymous namespace)::ExecutorState&lt;&gt;::Process()<br>\n    @     0x7ff2255f6922         48  std::_Function_handler&lt;&gt;::_M_invoke()<br>\n    @     0x7ff20b6b5445        160  Eigen::ThreadPoolTempl&lt;&gt;::WorkerLoop()<br>\n    @     0x7ff20b6b3117         64  std::_Function_handler&lt;&gt;::_M_invoke()<br>\n    @     0x7ff225d7728f         96  tensorflow::(anonymous namespace)::PThread::ThreadFn()<br>\n    @     0x7ff2443d7ea7  (unknown)  start_thread<br>\n<a href=\"https://symbolize.stripped_domain/r/?trace=7ff24442ace1,7ff201410ed2,7ff24442ad5f,7ff2053827a7,7ff20b6e149d,7ff20b74e769,7ff20b74f259,7ff20b745c96,7ff20b747f5c,7ff2256691be,7ff22560be4f,7ff2255f6921,7ff20b6b5444,7ff20b6b3116,7ff225d7728e,7ff2443d7ea6&amp;map=947919304d8b3d847ddb4dfdea237b00f15d2086:7ff202077000-7ff217973140,886380cfe14c768b793f7ee12153f640442e39c4:7ff224ac3000-7ff22673a6c0,4707200934d6baba849e70d87aad4e2c:7ff1ed106000-7ff201788eb0\" target=\"_blank\">https://symbolize.stripped_domain/r/?trace=7ff24442ace1,7ff201410ed2,7ff24442ad5f,7ff2053827a7,7ff20b6e149d,7ff20b74e769,7ff20b74f259,7ff20b745c96,7ff20b747f5c,7ff2256691be,7ff22560be4f,7ff2255f6921,7ff20b6b5444,7ff20b6b3116,7ff225d7728e,7ff2443d7ea6&amp;map=947919304d8b3d847ddb4dfdea237b00f15d2086:7ff202077000-7ff217973140,886380cfe14c768b793f7ee12153f640442e39c4:7ff224ac3000-7ff22673a6c0,4707200934d6baba849e70d87aad4e2c:7ff1ed106000-7ff201788eb0</a> <br>\nE0419 14:43:32.847925    1029 coredump_hook.cc:365] RAW: Remote crash data gathering hook invoked.<br>\nE0419 14:43:32.848074    1029 client.cc:222] RAW: Coroner client retries enabled (b/136286901), will retry for up to 30 sec.<br>\nE0419 14:43:32.848112    1029 coredump_hook.cc:473] RAW: Sending fingerprint to remote end.<br>\nE0419 14:43:32.848181    1029 coredump_socket.cc:124] RAW: Stat failed errno=2 on socket /var/google/services/logmanagerd/remote_coredump.socket<br>\nE0419 14:43:32.848211    1029 coredump_hook.cc:477] RAW: Cannot send fingerprint to Coroner: [NOT_FOUND] Missing crash reporting socket. Is the listener running?<br>\nE0419 14:43:32.848236    1029 coredump_hook.cc:547] RAW: Dumping core locally.</p>",
          "rawMarkdown": "Thank you. I tried Keras_cv_attention_models but I get an error: 2023-04-19 14:43:04.927404: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:903] model_pruner failed: INVALID_ARGUMENT: Graph does not contain terminal node Adam/Adam/AssignAddVariableOp.\n2023-04-19 14:43:07.408151: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:903] model_pruner failed: INVALID_ARGUMENT: Graph does not contain terminal node Adam/Adam/AssignAddVariableOp.\n2023-04-19 14:43:32.480805: E tensorflow/core/tpu/kernels/tpu_compilation_cache_external.cc:113] Reshape's input dynamic dimension is decomposed into multiple output dynamic dimensions, but the constraint is ambiguous and XLA can't infer the output dimension %reshape.11975 = f32[32,32,96,512]{3,2,1,0} reshape(f32[98304,512]{1,0} %transpose.11967), metadata={op_type=\"Reshape\" op_name=\"model/convnext_base/stack1_block1_up_dense/Tensordot\"}. \n2023-04-19 14:43:32.481013: F tensorflow/core/tpu/kernels/tpu_program_group.cc:86] Check failed: xla_tpu_programs.size() > 0 (0 vs. 0)\nhttps://symbolize.stripped_domain/r/?trace=7ff24442ace1,7ff24442ad5f,7ff2053827a7,7ff20b6e149d,7ff20b74e769,7ff20b74f259,7ff20b745c96,7ff20b747f5c,7ff2256691be,7ff22560be4f,7ff2255f6921,7ff20b6b5444,7ff20b6b3116,7ff225d7728e,7ff2443d7ea6&map=947919304d8b3d847ddb4dfdea237b00f15d2086:7ff202077000-7ff217973140,886380cfe14c768b793f7ee12153f640442e39c4:7ff224ac3000-7ff22673a6c0 \n*** SIGABRT received by PID 14 (TID 1029) on cpu 9 from PID 14; stack trace: ***\nPC: @     0x7ff24442ace1  (unknown)  raise\n    @     0x7ff201410ed3        992  (unknown)\n    @     0x7ff24442ad60       3904  (unknown)\n    @     0x7ff2053827a8        896  tensorflow::tpu::TpuProgramGroup::Initialize()\n    @     0x7ff20b6e149e       1696  tensorflow::tpu::TpuCompilationCacheExternal::InitializeEntry()\n    @     0x7ff20b74e76a       1072  tensorflow::tpu::TpuCompilationCacheInterface::CompileIfKeyAbsentHelper()\n    @     0x7ff20b74f25a        128  tensorflow::tpu::TpuCompilationCacheInterface::CompileIfKeyAbsent()\n    @     0x7ff20b745c97       1312  tensorflow::tpu::TpuCompileOpKernelCommon::ComputeInternal()\n    @     0x7ff20b747f5d        624  tensorflow::tpu::TpuCompileOpKernelCommon::Compute()\n    @     0x7ff2256691bf        496  tensorflow::ThreadPoolDevice::Compute()\n    @     0x7ff22560be50       2496  tensorflow::(anonymous namespace)::ExecutorState<>::Process()\n    @     0x7ff2255f6922         48  std::_Function_handler<>::_M_invoke()\n    @     0x7ff20b6b5445        160  Eigen::ThreadPoolTempl<>::WorkerLoop()\n    @     0x7ff20b6b3117         64  std::_Function_handler<>::_M_invoke()\n    @     0x7ff225d7728f         96  tensorflow::(anonymous namespace)::PThread::ThreadFn()\n    @     0x7ff2443d7ea7  (unknown)  start_thread\nhttps://symbolize.stripped_domain/r/?trace=7ff24442ace1,7ff201410ed2,7ff24442ad5f,7ff2053827a7,7ff20b6e149d,7ff20b74e769,7ff20b74f259,7ff20b745c96,7ff20b747f5c,7ff2256691be,7ff22560be4f,7ff2255f6921,7ff20b6b5444,7ff20b6b3116,7ff225d7728e,7ff2443d7ea6&map=947919304d8b3d847ddb4dfdea237b00f15d2086:7ff202077000-7ff217973140,886380cfe14c768b793f7ee12153f640442e39c4:7ff224ac3000-7ff22673a6c0,4707200934d6baba849e70d87aad4e2c:7ff1ed106000-7ff201788eb0 \nE0419 14:43:32.847925    1029 coredump_hook.cc:365] RAW: Remote crash data gathering hook invoked.\nE0419 14:43:32.848074    1029 client.cc:222] RAW: Coroner client retries enabled (b/136286901), will retry for up to 30 sec.\nE0419 14:43:32.848112    1029 coredump_hook.cc:473] RAW: Sending fingerprint to remote end.\nE0419 14:43:32.848181    1029 coredump_socket.cc:124] RAW: Stat failed errno=2 on socket /var/google/services/logmanagerd/remote_coredump.socket\nE0419 14:43:32.848211    1029 coredump_hook.cc:477] RAW: Cannot send fingerprint to Coroner: [NOT_FOUND] Missing crash reporting socket. Is the listener running?\nE0419 14:43:32.848236    1029 coredump_hook.cc:547] RAW: Dumping core locally.",
          "replies": [
            {
              "id": 2228154,
              "postDate": "2023-04-20T10:08:18.073Z",
              "content": "<p><a href=\"https://www.kaggle.com/aikhmelnytskyy\" target=\"_blank\">@aikhmelnytskyy</a> looks very familiar 😃 but I don't remember how I sorted it out, sorry. Maybe try a different way to define a backbone…</p>",
              "rawMarkdown": "@aikhmelnytskyy looks very familiar 😃 but I don't remember how I sorted it out, sorry. Maybe try a different way to define a backbone...",
              "votes": 1
            },
            {
              "id": 2234037,
              "postDate": "2023-04-24T19:47:53.667Z",
              "content": "<p>lower batch size.</p>",
              "rawMarkdown": "lower batch size."
            }
          ]
        }
      ]
    },
    {
      "id": 2226935,
      "postDate": "2023-04-19T11:49:55.127Z",
      "content": "<p>Hi to all! <br>\nI want to try ConvNeX. <br>\nFor this I installed tf 2.12.0 which includes ConvNeX models. But now the training time of one epoch is 50+ hours. Does anyone know what the problem is? <br>\nMaybe someone was able to run the ConvNeX on TPU VM in a different way. <br>\nHere is my code notebook to understand the problem:<a href=\"https://www.kaggle.com/code/aikhmelnytskyy/birdclef23-trainconvnext/notebook\" target=\"_blank\">https://www.kaggle.com/code/aikhmelnytskyy/birdclef23-trainconvnext/notebook</a>. <br>\nThanks for the answer!</p>",
      "rawMarkdown": "Hi to all! \nI want to try ConvNeX. \nFor this I installed tf 2.12.0 which includes ConvNeX models. But now the training time of one epoch is 50+ hours. Does anyone know what the problem is? \nMaybe someone was able to run the ConvNeX on TPU VM in a different way. \nHere is my code notebook to understand the problem:https://www.kaggle.com/code/aikhmelnytskyy/birdclef23-trainconvnext/notebook. \nThanks for the answer!",
      "votes": 1
    },
    {
      "id": 2229574,
      "postDate": "2023-04-21T13:29:58.287Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2230301,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2023-04-22T07:59:33.180000",
      "content": "<p><a href=\"https://www.kaggle.com/aikhmelnytskyy\" target=\"_blank\">@aikhmelnytskyy</a> <br>\nI've just tried ConvNeXt on TPU-VM on an image dataset and it works just fine; <a href=\"https://drive.google.com/file/d/14ii3dhjPGZXJHHOZrEw4Tlw9P07S31SX/view?usp=sharing\" target=\"_blank\">gist</a>. (this gist will be removed in future). Tips, don't install tf 2.12 manually, it comes by default now. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2227118,
      "author_name": "Aindriú",
      "author_url": "",
      "post_date": "2023-04-19T14:13:13.327000",
      "content": "<p>trains fine for me if I do it via <a href=\"https://github.com/leondgarse/keras_cv_attention_models\" target=\"_blank\">https://github.com/leondgarse/keras_cv_attention_models</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2227128,
          "author_name": "Andrij",
          "author_url": "",
          "post_date": "2023-04-19T14:19:55.567000",
          "content": "<p>Thank you. I tried Keras_cv_attention_models but I get an error: 2023-04-19 14:43:04.927404: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:903] model_pruner failed: INVALID_ARGUMENT: Graph does not contain terminal node Adam/Adam/AssignAddVariableOp.<br>\n2023-04-19 14:43:07.408151: E tensorflow/core/grappler/optimizers/meta_optimizer.cc:903] model_pruner failed: INVALID_ARGUMENT: Graph does not contain terminal node Adam/Adam/AssignAddVariableOp.<br>\n2023-04-19 14:43:32.480805: E tensorflow/core/tpu/kernels/tpu_compilation_cache_external.cc:113] Reshape's input dynamic dimension is decomposed into multiple output dynamic dimensions, but the constraint is ambiguous and XLA can't infer the output dimension %reshape.11975 = f32[32,32,96,512]{3,2,1,0} reshape(f32[98304,512]{1,0} %transpose.11967), metadata={op_type=\"Reshape\" op_name=\"model/convnext_base/stack1_block1_up_dense/Tensordot\"}. <br>\n2023-04-19 14:43:32.481013: F tensorflow/core/tpu/kernels/tpu_program_group.cc:86] Check failed: xla_tpu_programs.size() &gt; 0 (0 vs. 0)<br>\n<a href=\"https://symbolize.stripped_domain/r/?trace=7ff24442ace1,7ff24442ad5f,7ff2053827a7,7ff20b6e149d,7ff20b74e769,7ff20b74f259,7ff20b745c96,7ff20b747f5c,7ff2256691be,7ff22560be4f,7ff2255f6921,7ff20b6b5444,7ff20b6b3116,7ff225d7728e,7ff2443d7ea6&amp;map=947919304d8b3d847ddb4dfdea237b00f15d2086:7ff202077000-7ff217973140,886380cfe14c768b793f7ee12153f640442e39c4:7ff224ac3000-7ff22673a6c0\" target=\"_blank\">https://symbolize.stripped_domain/r/?trace=7ff24442ace1,7ff24442ad5f,7ff2053827a7,7ff20b6e149d,7ff20b74e769,7ff20b74f259,7ff20b745c96,7ff20b747f5c,7ff2256691be,7ff22560be4f,7ff2255f6921,7ff20b6b5444,7ff20b6b3116,7ff225d7728e,7ff2443d7ea6&amp;map=947919304d8b3d847ddb4dfdea237b00f15d2086:7ff202077000-7ff217973140,886380cfe14c768b793f7ee12153f640442e39c4:7ff224ac3000-7ff22673a6c0</a> <br>\n*** SIGABRT received by PID 14 (TID 1029) on cpu 9 from PID 14; stack trace: ***<br>\nPC: @     0x7ff24442ace1  (unknown)  raise<br>\n    @     0x7ff201410ed3        992  (unknown)<br>\n    @     0x7ff24442ad60       3904  (unknown)<br>\n    @     0x7ff2053827a8        896  tensorflow::tpu::TpuProgramGroup::Initialize()<br>\n    @     0x7ff20b6e149e       1696  tensorflow::tpu::TpuCompilationCacheExternal::InitializeEntry()<br>\n    @     0x7ff20b74e76a       1072  tensorflow::tpu::TpuCompilationCacheInterface::CompileIfKeyAbsentHelper()<br>\n    @     0x7ff20b74f25a        128  tensorflow::tpu::TpuCompilationCacheInterface::CompileIfKeyAbsent()<br>\n    @     0x7ff20b745c97       1312  tensorflow::tpu::TpuCompileOpKernelCommon::ComputeInternal()<br>\n    @     0x7ff20b747f5d        624  tensorflow::tpu::TpuCompileOpKernelCommon::Compute()<br>\n    @     0x7ff2256691bf        496  tensorflow::ThreadPoolDevice::Compute()<br>\n    @     0x7ff22560be50       2496  tensorflow::(anonymous namespace)::ExecutorState&lt;&gt;::Process()<br>\n    @     0x7ff2255f6922         48  std::_Function_handler&lt;&gt;::_M_invoke()<br>\n    @     0x7ff20b6b5445        160  Eigen::ThreadPoolTempl&lt;&gt;::WorkerLoop()<br>\n    @     0x7ff20b6b3117         64  std::_Function_handler&lt;&gt;::_M_invoke()<br>\n    @     0x7ff225d7728f         96  tensorflow::(anonymous namespace)::PThread::ThreadFn()<br>\n    @     0x7ff2443d7ea7  (unknown)  start_thread<br>\n<a href=\"https://symbolize.stripped_domain/r/?trace=7ff24442ace1,7ff201410ed2,7ff24442ad5f,7ff2053827a7,7ff20b6e149d,7ff20b74e769,7ff20b74f259,7ff20b745c96,7ff20b747f5c,7ff2256691be,7ff22560be4f,7ff2255f6921,7ff20b6b5444,7ff20b6b3116,7ff225d7728e,7ff2443d7ea6&amp;map=947919304d8b3d847ddb4dfdea237b00f15d2086:7ff202077000-7ff217973140,886380cfe14c768b793f7ee12153f640442e39c4:7ff224ac3000-7ff22673a6c0,4707200934d6baba849e70d87aad4e2c:7ff1ed106000-7ff201788eb0\" target=\"_blank\">https://symbolize.stripped_domain/r/?trace=7ff24442ace1,7ff201410ed2,7ff24442ad5f,7ff2053827a7,7ff20b6e149d,7ff20b74e769,7ff20b74f259,7ff20b745c96,7ff20b747f5c,7ff2256691be,7ff22560be4f,7ff2255f6921,7ff20b6b5444,7ff20b6b3116,7ff225d7728e,7ff2443d7ea6&amp;map=947919304d8b3d847ddb4dfdea237b00f15d2086:7ff202077000-7ff217973140,886380cfe14c768b793f7ee12153f640442e39c4:7ff224ac3000-7ff22673a6c0,4707200934d6baba849e70d87aad4e2c:7ff1ed106000-7ff201788eb0</a> <br>\nE0419 14:43:32.847925    1029 coredump_hook.cc:365] RAW: Remote crash data gathering hook invoked.<br>\nE0419 14:43:32.848074    1029 client.cc:222] RAW: Coroner client retries enabled (b/136286901), will retry for up to 30 sec.<br>\nE0419 14:43:32.848112    1029 coredump_hook.cc:473] RAW: Sending fingerprint to remote end.<br>\nE0419 14:43:32.848181    1029 coredump_socket.cc:124] RAW: Stat failed errno=2 on socket /var/google/services/logmanagerd/remote_coredump.socket<br>\nE0419 14:43:32.848211    1029 coredump_hook.cc:477] RAW: Cannot send fingerprint to Coroner: [NOT_FOUND] Missing crash reporting socket. Is the listener running?<br>\nE0419 14:43:32.848236    1029 coredump_hook.cc:547] RAW: Dumping core locally.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2228154,
              "author_name": "Aindriú",
              "author_url": "",
              "post_date": "2023-04-20T10:08:18.073000",
              "content": "<p><a href=\"https://www.kaggle.com/aikhmelnytskyy\" target=\"_blank\">@aikhmelnytskyy</a> looks very familiar 😃 but I don't remember how I sorted it out, sorry. Maybe try a different way to define a backbone…</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2234037,
              "author_name": "nymfree",
              "author_url": "",
              "post_date": "2023-04-24T19:47:53.667000",
              "content": "<p>lower batch size.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2229574,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-21T13:29:58.287000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2230301": "@aikhmelnytskyy \nI've just tried ConvNeXt on TPU-VM on an image dataset and it works just fine; [gist](https://drive.google.com/file/d/14ii3dhjPGZXJHHOZrEw4Tlw9P07S31SX/view?usp=sharing). (this gist will be removed in future). Tips, don't install tf 2.12 manually, it comes by default now. ",
    "2227118": "trains fine for me if I do it via https://github.com/leondgarse/keras_cv_attention_models",
    "2226935": "Hi to all! \nI want to try ConvNeX. \nFor this I installed tf 2.12.0 which includes ConvNeX models. But now the training time of one epoch is 50+ hours. Does anyone know what the problem is? \nMaybe someone was able to run the ConvNeX on TPU VM in a different way. \nHere is my code notebook to understand the problem:https://www.kaggle.com/code/aikhmelnytskyy/birdclef23-trainconvnext/notebook. \nThanks for the answer!",
    "2229574": ""
  }
}