{
  "id": 585236,
  "title": "CUDA Speed Up for vel-to-seis and Some Findings",
  "url": "/competitions/waveform-inversion/discussion/585236",
  "author_name": "lhwcv",
  "post_date": "2025-06-19T01:48:52.745000",
  "votes": 14,
  "comment_count": 17,
  "views": 0,
  "content": "<p>see:   <a href=\"https://github.com/lhwcv/cuda_vel_forward/tree/main\" target=\"_blank\">https://github.com/lhwcv/cuda_vel_forward/tree/main</a></p>\n<ul>\n<li>(500, 1, 70, 70) --&gt; (500, 5, 1000, 70): about 1min on RTX 5090</li>\n</ul>\n<p>Findings:</p>\n<p>MAE is highly correlated with the reconstruction error, achieving a Pearson correlation coefficient of 0.76 on the validation set. This insight can guide training—for instance, by jointly training a Reward Model alongside the main network.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F33616902d6bf69863f8757493e135c3c%2F1.jpg?generation=1750297814708888&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 3227415,
      "postDate": "2025-06-19T01:48:52.747Z",
      "content": "<p>see:   <a href=\"https://github.com/lhwcv/cuda_vel_forward/tree/main\" target=\"_blank\">https://github.com/lhwcv/cuda_vel_forward/tree/main</a></p>\n<ul>\n<li>(500, 1, 70, 70) --&gt; (500, 5, 1000, 70): about 1min on RTX 5090</li>\n</ul>\n<p>Findings:</p>\n<p>MAE is highly correlated with the reconstruction error, achieving a Pearson correlation coefficient of 0.76 on the validation set. This insight can guide training—for instance, by jointly training a Reward Model alongside the main network.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F33616902d6bf69863f8757493e135c3c%2F1.jpg?generation=1750297814708888&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "see:   https://github.com/lhwcv/cuda_vel_forward/tree/main\n\n-  (500, 1, 70, 70) --> (500, 5, 1000, 70): about 1min on RTX 5090\n\nFindings:\n\nMAE is highly correlated with the reconstruction error, achieving a Pearson correlation coefficient of 0.76 on the validation set. This insight can guide training—for instance, by jointly training a Reward Model alongside the main network.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F33616902d6bf69863f8757493e135c3c%2F1.jpg?generation=1750297814708888&alt=media)\n\n",
      "votes": 14
    },
    {
      "id": 3228229,
      "postDate": "2025-06-20T00:10:08.823Z",
      "content": "<p>i find a torch implementation (that is differentiable) here:<br>\n<a href=\"https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_f.py\" target=\"_blank\">https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_f.py</a></p>\n<p>I only check a few examples and the forwarding model error (compared to ground truth seismic file) is low.  error is exact zero for float32</p>\n<p>please check more samples if you want to use.</p>\n<hr>\n<p>I belive this is the code used by openFWI team<br>\n<a href=\"https://github.com/lanl/OpenFWI/issues/10:\" target=\"_blank\">https://github.com/lanl/OpenFWI/issues/10:</a><br>\n\". However, you can generate the FWI-F by utilizing the OpenFWI velocity model and performing seismic forward modeling with a random frequency range of 5 to 25 Hz.\"</p>",
      "rawMarkdown": "i find a torch implementation (that is differentiable) here:\nhttps://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_f.py\n\nI only check a few examples and the forwarding model error (compared to ground truth seismic file) is low.  error is exact zero for float32\n\n\nplease check more samples if you want to use.\n\n\n\n---\nI belive this is the code used by openFWI team\nhttps://github.com/lanl/OpenFWI/issues/10:\n\". However, you can generate the FWI-F by utilizing the OpenFWI velocity model and performing seismic forward modeling with a random frequency range of 5 to 25 Hz.\"\n",
      "votes": 3,
      "replies": [
        {
          "id": 3228292,
          "postDate": "2025-06-20T02:51:20.433Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> how test set have different distribution for last source? i.e CureFault_B</p>",
          "rawMarkdown": "@hengck23 how test set have different distribution for last source? i.e CureFault_B",
          "replies": [
            {
              "id": 3228295,
              "postDate": "2025-06-20T02:55:17.970Z",
              "content": "<p>there are many ways, and one way is validation and test error.</p>",
              "rawMarkdown": "there are many ways, and one way is validation and test error.\n",
              "votes": 1
            }
          ]
        },
        {
          "id": 3228297,
          "postDate": "2025-06-20T02:57:37.870Z",
          "rawMarkdown": "",
          "isDeleted": true,
          "replies": [
            {
              "id": 3228304,
              "postDate": "2025-06-20T03:00:52.443Z",
              "content": "<p>you should rewrite the code for parallel process. the aythor prcess one slice at a time, but torch is tensor lib and we should process whole array.</p>\n<p>lapaican operator may be implement as conv, maybe</p>\n<p>-- </p>\n<p>further all dimension is fixed(including the dt time loop). u can consult chatgpt on fast torch code or customised cuda kernel. i estimate 20 to 30x speedup possible</p>",
              "rawMarkdown": "you should rewrite the code for parallel process. the aythor prcess one slice at a time, but torch is tensor lib and we should process whole array.\n\nlapaican operator may be implement as conv, maybe\n\n-- \n\nfurther all dimension is fixed(including the dt time loop). u can consult chatgpt on fast torch code or customised cuda kernel. i estimate 20 to 30x speedup possible"
            }
          ]
        },
        {
          "id": 3228306,
          "postDate": "2025-06-20T03:02:07.673Z",
          "content": "<p>Hello, may I ask if you have compared their speeds? I used the poster's code and it takes nearly 5 minutes to complete one val. My batchsize is 16<br>\n22%|█████████▌ | 14/63 [00:37&lt;5, 2.63s/it], cpu: h20</p>",
          "rawMarkdown": "Hello, may I ask if you have compared their speeds? I used the poster's code and it takes nearly 5 minutes to complete one val. My batchsize is 16\n22%|█████████▌ | 14/63 [00:37<5, 2.63s/it], cpu: h20"
        }
      ]
    },
    {
      "id": 3229432,
      "postDate": "2025-06-21T13:36:29.940Z",
      "content": "<p>Thanks for the nice code. Unfortunately, I am getting much larger deviations on the training set than in the original \"Improved Vel to Seis\" code:<br>\nMean errors  of ~0.032123 as compared to ~0.000009 as it was in the original notebook. </p>\n<p>Is there a room for improvement?</p>",
      "rawMarkdown": "Thanks for the nice code. Unfortunately, I am getting much larger deviations on the training set than in the original \"Improved Vel to Seis\" code:\nMean errors  of ~0.032123 as compared to ~0.000009 as it was in the original notebook. \n\nIs there a room for improvement?",
      "votes": 2,
      "replies": [
        {
          "id": 3230536,
          "postDate": "2025-06-23T07:01:18.407Z",
          "content": "<p>Check <a href=\"https://www.kaggle.com/code/gguillard/torch-vel-to-seis/\" target=\"_blank\">this notebook</a>.</p>",
          "rawMarkdown": "Check [this notebook](https://www.kaggle.com/code/gguillard/torch-vel-to-seis/).",
          "votes": 1
        }
      ]
    },
    {
      "id": 3227549,
      "postDate": "2025-06-19T04:55:50.137Z",
      "content": "<p>maybe a side effect of this is:<br>\nsubmit =max( reconstruct error( model1), reconstruct error( model2), …reconstruct error( ensemble), …)</p>\n<p>maybe better then ensembling</p>",
      "rawMarkdown": "maybe a side effect of this is:\nsubmit =max( reconstruct error( model1), reconstruct error( model2), ...reconstruct error( ensemble), ...)\n\nmaybe better then ensembling",
      "votes": 2
    },
    {
      "id": 3227424,
      "postDate": "2025-06-19T02:13:20.457Z",
      "content": "<p>or use reconstruction error to select pseudo label ?</p>",
      "rawMarkdown": "or use reconstruction error to select pseudo label ?",
      "votes": 2,
      "replies": [
        {
          "id": 3227425,
          "postDate": "2025-06-19T02:14:14.530Z",
          "content": "<p>yes choose those for example err &lt; 0.01</p>",
          "rawMarkdown": "yes choose those for example err < 0.01",
          "votes": 2,
          "replies": [
            {
              "id": 3227512,
              "postDate": "2025-06-19T04:19:56.730Z",
              "content": "<p>assume given a test sample xt. i make prediction v1. i rconstruct x1</p>\n<p>now v1 is not the same as ground truth vt, there is some error in my prediction</p>\n<p>i add (x1,v1) to my train set. next time my model will NOT predict v1 given xt. it ONLY predict v1 if given x1. say the new model now predict v2.</p>\n<p>hopefully the model will predict v3,v4 …. until converge to vt.</p>\n<p>in my method no need really to select based on err. this is my guess only, currently doing experiment</p>\n<hr>\n<p>note that this is NOT  pseudo label .</p>\n<p>we are lower \"probability\" of p(v wrong | x test), and \"hoping\" this will increase p(v correct | x test)</p>\n<p>whether it works or not would depends of distance(xt,x1) is same as distance(v correct,v1) or not.</p>\n<p>i.e. assume  distance(xt,x1) = very small, then even given train sample(x1,v1), the model still predict model(xt) = v1</p>\n<p>ideally, we want large distance(v correct,v1) to give large distance(xt,x1) </p>",
              "rawMarkdown": "assume given a test sample xt. i make prediction v1. i rconstruct x1\n\nnow v1 is not the same as ground truth vt, there is some error in my prediction\n\ni add (x1,v1) to my train set. next time my model will NOT predict v1 given xt. it ONLY predict v1 if given x1. say the new model now predict v2.\n\nhopefully the model will predict v3,v4 .... until converge to vt.\n\nin my method no need really to select based on err. this is my guess only, currently doing experiment\n\n---\n\nnote that this is NOT  pseudo label .\n\nwe are lower \"probability\" of p(v wrong | x test), and \"hoping\" this will increase p(v correct | x test)\n\nwhether it works or not would depends of distance(xt,x1) is same as distance(v correct,v1) or not.\n\ni.e. assume  distance(xt,x1) = very small, then even given train sample(x1,v1), the model still predict model(xt) = v1\n\nideally, we want large distance(v correct,v1) to give large distance(xt,x1) ",
              "votes": 1
            },
            {
              "id": 3227544,
              "postDate": "2025-06-19T04:48:55.177Z",
              "content": "<p>analogy:</p>\n<hr>\n<p>ME: \"how many legs the cat has?\"<br>\nMODEL : \"8\"<br>\nME: \"spider has 8 legs\"</p>\n<hr>\n<p>ME: \"how many legs the cat has?\"<br>\nMODEL : \"2\" #he will not say 8 again<br>\nME: \"bird has 2 legs\"</p>\n<hr>\n<p>ME: \"how many legs the cat has?\"<br>\nMODEL : \"…\"  #not 8, not 2</p>\n<hr>\n<p>there is a catch: this only happens if the model knows cat is not spider and cat is not bird. if model thinks cat is bird, it would not work</p>",
              "rawMarkdown": "analogy:\n\n---\n\nME: \"how many legs the cat has?\"\nMODEL : \"8\"\nME: \"spider has 8 legs\"\n\n---\n\nME: \"how many legs the cat has?\"\nMODEL : \"2\" #he will not say 8 again\nME: \"bird has 2 legs\"\n\n---\n\n\nME: \"how many legs the cat has?\"\nMODEL : \"...\"  #not 8, not 2\n\n---\n\nthere is a catch: this only happens if the model knows cat is not spider and cat is not bird. if model thinks cat is bird, it would not work\n",
              "votes": 1
            },
            {
              "id": 3227557,
              "postDate": "2025-06-19T05:02:38.780Z",
              "content": "<p>Yes good idea, let’s try</p>",
              "rawMarkdown": "Yes good idea, let’s try"
            }
          ]
        }
      ]
    },
    {
      "id": 3227966,
      "postDate": "2025-06-19T13:55:17.123Z",
      "content": "<p>`# B, 1, 70, 70<br>\nlbl = np.random.rand(1, 1, 70, 70)<br>\nlbl = torch.from_numpy(lbl).float()<br>\na = Vel_Forward()</p>\n<h1>B, 5, 1000,70</h1>\n<p>gen_arr = a(lbl).cpu().numpy()<br>\nprint(gen_arr.shape)#(1, 5, 1001, 70)<br>\n`<br>\nHello. May I ask why it doesn't match the shape you expected</p>",
      "rawMarkdown": "`# B, 1, 70, 70\nlbl = np.random.rand(1, 1, 70, 70)\nlbl = torch.from_numpy(lbl).float()\na = Vel_Forward()\n# B, 5, 1000,70\ngen_arr = a(lbl).cpu().numpy()\nprint(gen_arr.shape)#(1, 5, 1001, 70)\n`\nHello. May I ask why it doesn't match the shape you expected",
      "replies": [
        {
          "id": 3227968,
          "postDate": "2025-06-19T13:58:36.977Z",
          "content": "<p>Modify nt to 1000 instead of  1001, and comment without [:,1:,:,:]</p>",
          "rawMarkdown": "Modify nt to 1000 instead of  1001, and comment without [:,1:,:,:]",
          "replies": [
            {
              "id": 3228303,
              "postDate": "2025-06-20T03:00:49.217Z",
              "content": "<p>Hello, do you think my speed is normal? One val takes nearly 5 minutes and my batchsize is 16<br>\n22%|█████████▌ | 14/63 [00:37&lt;5, 2.63s/it], cpu: h20</p>",
              "rawMarkdown": "Hello, do you think my speed is normal? One val takes nearly 5 minutes and my batchsize is 16\n22%|█████████▌ | 14/63 [00:37<5, 2.63s/it], cpu: h20"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3228229,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-06-20T00:10:08.823000",
      "content": "<p>i find a torch implementation (that is differentiable) here:<br>\n<a href=\"https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_f.py\" target=\"_blank\">https://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_f.py</a></p>\n<p>I only check a few examples and the forwarding model error (compared to ground truth seismic file) is low.  error is exact zero for float32</p>\n<p>please check more samples if you want to use.</p>\n<hr>\n<p>I belive this is the code used by openFWI team<br>\n<a href=\"https://github.com/lanl/OpenFWI/issues/10:\" target=\"_blank\">https://github.com/lanl/OpenFWI/issues/10:</a><br>\n\". However, you can generate the FWI-F by utilizing the OpenFWI velocity model and performing seismic forward modeling with a random frequency range of 5 to 25 Hz.\"</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3228292,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2025-06-20T02:51:20.433000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> how test set have different distribution for last source? i.e CureFault_B</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3228295,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-06-20T02:55:17.970000",
              "content": "<p>there are many ways, and one way is validation and test error.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3228297,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-06-20T02:57:37.870000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 3228304,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-06-20T03:00:52.443000",
              "content": "<p>you should rewrite the code for parallel process. the aythor prcess one slice at a time, but torch is tensor lib and we should process whole array.</p>\n<p>lapaican operator may be implement as conv, maybe</p>\n<p>-- </p>\n<p>further all dimension is fixed(including the dt time loop). u can consult chatgpt on fast torch code or customised cuda kernel. i estimate 20 to 30x speedup possible</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3228306,
          "author_name": "water joe",
          "author_url": "",
          "post_date": "2025-06-20T03:02:07.673000",
          "content": "<p>Hello, may I ask if you have compared their speeds? I used the poster's code and it takes nearly 5 minutes to complete one val. My batchsize is 16<br>\n22%|█████████▌ | 14/63 [00:37&lt;5, 2.63s/it], cpu: h20</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3229432,
      "author_name": "Sviatoslav Bilokin",
      "author_url": "",
      "post_date": "2025-06-21T13:36:29.940000",
      "content": "<p>Thanks for the nice code. Unfortunately, I am getting much larger deviations on the training set than in the original \"Improved Vel to Seis\" code:<br>\nMean errors  of ~0.032123 as compared to ~0.000009 as it was in the original notebook. </p>\n<p>Is there a room for improvement?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3230536,
          "author_name": "gguillard",
          "author_url": "",
          "post_date": "2025-06-23T07:01:18.407000",
          "content": "<p>Check <a href=\"https://www.kaggle.com/code/gguillard/torch-vel-to-seis/\" target=\"_blank\">this notebook</a>.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3227549,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-06-19T04:55:50.137000",
      "content": "<p>maybe a side effect of this is:<br>\nsubmit =max( reconstruct error( model1), reconstruct error( model2), …reconstruct error( ensemble), …)</p>\n<p>maybe better then ensembling</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3227424,
      "author_name": "atom1231",
      "author_url": "",
      "post_date": "2025-06-19T02:13:20.457000",
      "content": "<p>or use reconstruction error to select pseudo label ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3227425,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2025-06-19T02:14:14.530000",
          "content": "<p>yes choose those for example err &lt; 0.01</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3227512,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-06-19T04:19:56.730000",
              "content": "<p>assume given a test sample xt. i make prediction v1. i rconstruct x1</p>\n<p>now v1 is not the same as ground truth vt, there is some error in my prediction</p>\n<p>i add (x1,v1) to my train set. next time my model will NOT predict v1 given xt. it ONLY predict v1 if given x1. say the new model now predict v2.</p>\n<p>hopefully the model will predict v3,v4 …. until converge to vt.</p>\n<p>in my method no need really to select based on err. this is my guess only, currently doing experiment</p>\n<hr>\n<p>note that this is NOT  pseudo label .</p>\n<p>we are lower \"probability\" of p(v wrong | x test), and \"hoping\" this will increase p(v correct | x test)</p>\n<p>whether it works or not would depends of distance(xt,x1) is same as distance(v correct,v1) or not.</p>\n<p>i.e. assume  distance(xt,x1) = very small, then even given train sample(x1,v1), the model still predict model(xt) = v1</p>\n<p>ideally, we want large distance(v correct,v1) to give large distance(xt,x1) </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3227544,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-06-19T04:48:55.177000",
              "content": "<p>analogy:</p>\n<hr>\n<p>ME: \"how many legs the cat has?\"<br>\nMODEL : \"8\"<br>\nME: \"spider has 8 legs\"</p>\n<hr>\n<p>ME: \"how many legs the cat has?\"<br>\nMODEL : \"2\" #he will not say 8 again<br>\nME: \"bird has 2 legs\"</p>\n<hr>\n<p>ME: \"how many legs the cat has?\"<br>\nMODEL : \"…\"  #not 8, not 2</p>\n<hr>\n<p>there is a catch: this only happens if the model knows cat is not spider and cat is not bird. if model thinks cat is bird, it would not work</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3227557,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2025-06-19T05:02:38.780000",
              "content": "<p>Yes good idea, let’s try</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3227966,
      "author_name": "water joe",
      "author_url": "",
      "post_date": "2025-06-19T13:55:17.123000",
      "content": "<p>`# B, 1, 70, 70<br>\nlbl = np.random.rand(1, 1, 70, 70)<br>\nlbl = torch.from_numpy(lbl).float()<br>\na = Vel_Forward()</p>\n<h1>B, 5, 1000,70</h1>\n<p>gen_arr = a(lbl).cpu().numpy()<br>\nprint(gen_arr.shape)#(1, 5, 1001, 70)<br>\n`<br>\nHello. May I ask why it doesn't match the shape you expected</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3227968,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2025-06-19T13:58:36.977000",
          "content": "<p>Modify nt to 1000 instead of  1001, and comment without [:,1:,:,:]</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3228303,
              "author_name": "water joe",
              "author_url": "",
              "post_date": "2025-06-20T03:00:49.217000",
              "content": "<p>Hello, do you think my speed is normal? One val takes nearly 5 minutes and my batchsize is 16<br>\n22%|█████████▌ | 14/63 [00:37&lt;5, 2.63s/it], cpu: h20</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3227415": "see:   https://github.com/lhwcv/cuda_vel_forward/tree/main\n\n-  (500, 1, 70, 70) --> (500, 5, 1000, 70): about 1min on RTX 5090\n\nFindings:\n\nMAE is highly correlated with the reconstruction error, achieving a Pearson correlation coefficient of 0.76 on the validation set. This insight can guide training—for instance, by jointly training a Reward Model alongside the main network.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F33616902d6bf69863f8757493e135c3c%2F1.jpg?generation=1750297814708888&alt=media)\n\n",
    "3228229": "i find a torch implementation (that is differentiable) here:\nhttps://github.com/lu-group/fourier-deeponet-fwi/blob/main/data/cfa/data_gen_f.py\n\nI only check a few examples and the forwarding model error (compared to ground truth seismic file) is low.  error is exact zero for float32\n\n\nplease check more samples if you want to use.\n\n\n\n---\nI belive this is the code used by openFWI team\nhttps://github.com/lanl/OpenFWI/issues/10:\n\". However, you can generate the FWI-F by utilizing the OpenFWI velocity model and performing seismic forward modeling with a random frequency range of 5 to 25 Hz.\"\n",
    "3229432": "Thanks for the nice code. Unfortunately, I am getting much larger deviations on the training set than in the original \"Improved Vel to Seis\" code:\nMean errors  of ~0.032123 as compared to ~0.000009 as it was in the original notebook. \n\nIs there a room for improvement?",
    "3227549": "maybe a side effect of this is:\nsubmit =max( reconstruct error( model1), reconstruct error( model2), ...reconstruct error( ensemble), ...)\n\nmaybe better then ensembling",
    "3227424": "or use reconstruction error to select pseudo label ?",
    "3227966": "`# B, 1, 70, 70\nlbl = np.random.rand(1, 1, 70, 70)\nlbl = torch.from_numpy(lbl).float()\na = Vel_Forward()\n# B, 5, 1000,70\ngen_arr = a(lbl).cpu().numpy()\nprint(gen_arr.shape)#(1, 5, 1001, 70)\n`\nHello. May I ask why it doesn't match the shape you expected"
  }
}