{
  "id": 587529,
  "title": "14th place solution",
  "url": "/competitions/waveform-inversion/writeups/ruby-14th-place-solution",
  "author_name": "",
  "post_date": "2025-07-02T11:55:25.730Z",
  "votes": 24,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Thanks Kaggle and organizer for hosting this competition, this is the most interesting competition that I ever entered. And also great thanks to people who sharing ideas and codes, especially:<br>\n<a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> for modeling ideas<br>\n<a href=\"https://www.kaggle.com/jaewook704\" target=\"_blank\">@jaewook704</a> <a href=\"https://www.kaggle.com/manatoyo\" target=\"_blank\">@manatoyo</a> for forward simulation code<br>\n<a href=\"https://www.kaggle.com/bguberfain\" target=\"_blank\">@bguberfain</a> for starter notebook</p>\n<p><strong>Data Processing</strong></p>\n<ol>\n<li>sign(x)*log(1+|x|)</li>\n<li>reorder each receiver based on CMP(central middle point) wrt source along receiver dimension to make input better aligned with target. To keep receiver dimension 70, odd and even indices are split into two channels. (5,1000,70)-&gt;(10,1000,70) <br>\n(-18MAE when ~160MAE  on 7k samples)</li>\n<li>add (x,y) coordinate embedding as extra channels since FWI task is spatial variant</li>\n</ol>\n<p><strong>Model</strong></p>\n<ol>\n<li>I stack multiple U-net to mimic common iterative methods in physics / math. Here each U-net works as a single iteration step with a global receptive field. Scaling depth by stacking more U-nets works better than just scaling width or increase number of conv layers.</li>\n<li>Based on my experiment on 40k samples, larger model always lead to better result. The largest model I can train : stack 5 U-nets with depth 4 channel dim 128. Not sure how much further gain is possible by continue scaling up.</li>\n<li>I tried Convnext (using Bartley's model or replace 1st U-net in my model) without success and I don’t have enough computation resources to train CAFormer to fully converge so I end up without using any pretrained backbone.</li>\n<li>Other details: few step stride conv to downsample input data; intermediate conv layer(from Bartely work); down by avg_pool up by bilinear; batch norm (consistent better than other type of normalization once converged ); skip connect</li>\n</ol>\n<p><strong>Loss</strong></p>\n<ol>\n<li>MAE with smaller weight to deeper positions (1~1/4 linearly) since I suspect it be more noisy (84-&gt;81 MAE on 10k)</li>\n</ol>\n<p><strong>Augmentation</strong></p>\n<ol>\n<li>Symmetric augmentation (-10 MAE when ~110 on 10k samples)<ul>\n<li>source is placed at [0, 17, 34, 52, 69], 34 after flip will be 35 this causes conflict in feature meaning. So I insert one more channel to represent source at 35 and fill by zero. Feature meaning is then self consistent before and after flip, though model still need to be trained to learn to handle this.</li>\n<li>used as TTA: 15.28-&gt;14.71</li></ul></li>\n<li>Velocity map scaling(-10 MAE when ~100 on 10k samples)<ul>\n<li>Scale target velocity map by *(1+alpha) then compress seis data along time axis to /(1+alpha) </li>\n<li>After play with forward simulation code I find scale velocity also affect sesi data maginitue.I can’t track this in closed form so I use an empirical rule: *(1+alpha)^0.26. Based on validation this only has minor effects if any.</li>\n<li>Here we are actually scaling time, this causes the augmented dataset with different source freq, so this doesn’t generate data with the same distribution as OpenFWI but luckily it still helps.</li></ul></li>\n</ol>\n<p><strong>Reconstruction Error Optimization</strong></p>\n<ol>\n<li>notation: G - true inverse model, F - true forward model, M - our NN inverse model</li>\n<li>here I minimize ||F(M(x_k)) - x ||^2 wrt x_k start from x rather than change model M</li>\n<li>assume M is already a good estimation of G in terms of local change:  F(M(x+dx))-F(M(x))~=F(G(x+dx))-F(G(x))=dx then we can use simple iteration rule to reduce error without gradient of F:<br>\nx_k=x_k-lambda*(F(M(x_k)) - x ) <br>\nthat is we can manipulate x directly with change in y since f here close to identity map</li>\n<li>When used together with TTA: M(x):=(M(x)+Flip(M(Flip(x))))/2</li>\n<li>performance:(each iter costs around 1.5h for whole test)<br>\n17.17-&gt;14.45 (lambda0.85, 1 iters) <br>\n17.17-&gt;13.49 (lambda0.7, 3 iters) <br>\n14.71-&gt;12.34 (lambda 0.6, 5 iters) for model already finetuned with data generate by such iterative optimization</li>\n</ol>\n<p><strong>Training stages(speed performance reported on single 4080s)</strong></p>\n<ol>\n<li>100 epoch on OpenFWI, 12 days <ul>\n<li>AdamW, weight decay 1e-4, batch size 28, init lr 28/64*3e-4, 80% init lr+ 20% cos lr, EMA in last epoch(only minor diff, droped in later stages)</li>\n<li>predict test set and run forward simulation to generate extra data, only symmetric TTA used in this stage</li>\n<li>val/LB: 21.0/22.4(sym TTA)</li></ul></li>\n<li>6+20 epoch on OpenFWI + *4 copy of stage 1 generated data, 6 days <ul>\n<li>after first 6 epoch my computer restarted so I continue training from it</li>\n<li>only cos lr is used without init constant phase, init lr 28/64*1.5e-4</li>\n<li>val/LB: 17.17/???  (sym TTA)<br>\nval/LB: 14.45/14.9 (sym TTA+lambda0.85, 1 iter)<br>\nval/LB: 13.49/???  (sym TTA+lambda0.7, 3 iters)</li>\n<li>predict test and run forward simulation to generate extra data (sym TTA+lambda0.85, 2 iters)</li></ul></li>\n<li>10 epoch on OpenFWI + *10 copy of stage 2 generated data, 3 days <ul>\n<li>only cos lr is used without init constant phase, init lr 28/64*1.5e-4</li>\n<li>val/LB: 14.71/???  (sym TTA)<br>\nval/LB: 12.34/12.7 (sym TTA+lambda0.6, 5 iters)</li></ul></li>\n</ol>",
  "messages": [
    {
      "id": "3237944",
      "postDate": "07/01/2025 12:16:13",
      "content": "<p>Thanks Kaggle and organizer for hosting this competition, this is the most interesting competition that I ever entered. And also great thanks to people who sharing ideas and codes, especially:<br>\n<a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> for modeling ideas<br>\n<a href=\"https://www.kaggle.com/jaewook704\" target=\"_blank\">@jaewook704</a> <a href=\"https://www.kaggle.com/manatoyo\" target=\"_blank\">@manatoyo</a> for forward simulation code<br>\n<a href=\"https://www.kaggle.com/bguberfain\" target=\"_blank\">@bguberfain</a> for starter notebook</p>\n<p><strong>Data Processing</strong></p>\n<ol>\n<li>sign(x)*log(1+|x|)</li>\n<li>reorder each receiver based on CMP(central middle point) wrt source along receiver dimension to make input better aligned with target. To keep receiver dimension 70, odd and even indices are split into two channels. (5,1000,70)-&gt;(10,1000,70) <br>\n(-18MAE when ~160MAE  on 7k samples)</li>\n<li>add (x,y) coordinate embedding as extra channels since FWI task is spatial variant</li>\n</ol>\n<p><strong>Model</strong></p>\n<ol>\n<li>I stack multiple U-net to mimic common iterative methods in physics / math. Here each U-net works as a single iteration step with a global receptive field. Scaling depth by stacking more U-nets works better than just scaling width or increase number of conv layers.</li>\n<li>Based on my experiment on 40k samples, larger model always lead to better result. The largest model I can train : stack 5 U-nets with depth 4 channel dim 128. Not sure how much further gain is possible by continue scaling up.</li>\n<li>I tried Convnext (using Bartley's model or replace 1st U-net in my model) without success and I don’t have enough computation resources to train CAFormer to fully converge so I end up without using any pretrained backbone.</li>\n<li>Other details: few step stride conv to downsample input data; intermediate conv layer(from Bartely work); down by avg_pool up by bilinear; batch norm (consistent better than other type of normalization once converged ); skip connect</li>\n</ol>\n<p><strong>Loss</strong></p>\n<ol>\n<li>MAE with smaller weight to deeper positions (1~1/4 linearly) since I suspect it be more noisy (84-&gt;81 MAE on 10k)</li>\n</ol>\n<p><strong>Augmentation</strong></p>\n<ol>\n<li>Symmetric augmentation (-10 MAE when ~110 on 10k samples)<ul>\n<li>source is placed at [0, 17, 34, 52, 69], 34 after flip will be 35 this causes conflict in feature meaning. So I insert one more channel to represent source at 35 and fill by zero. Feature meaning is then self consistent before and after flip, though model still need to be trained to learn to handle this.</li>\n<li>used as TTA: 15.28-&gt;14.71</li></ul></li>\n<li>Velocity map scaling(-10 MAE when ~100 on 10k samples)<ul>\n<li>Scale target velocity map by *(1+alpha) then compress seis data along time axis to /(1+alpha) </li>\n<li>After play with forward simulation code I find scale velocity also affect sesi data maginitue.I can’t track this in closed form so I use an empirical rule: *(1+alpha)^0.26. Based on validation this only has minor effects if any.</li>\n<li>Here we are actually scaling time, this causes the augmented dataset with different source freq, so this doesn’t generate data with the same distribution as OpenFWI but luckily it still helps.</li></ul></li>\n</ol>\n<p><strong>Reconstruction Error Optimization</strong></p>\n<ol>\n<li>notation: G - true inverse model, F - true forward model, M - our NN inverse model</li>\n<li>here I minimize ||F(M(x_k)) - x ||^2 wrt x_k start from x rather than change model M</li>\n<li>assume M is already a good estimation of G in terms of local change:  F(M(x+dx))-F(M(x))~=F(G(x+dx))-F(G(x))=dx then we can use simple iteration rule to reduce error without gradient of F:<br>\nx_k=x_k-lambda*(F(M(x_k)) - x ) <br>\nthat is we can manipulate x directly with change in y since f here close to identity map</li>\n<li>When used together with TTA: M(x):=(M(x)+Flip(M(Flip(x))))/2</li>\n<li>performance:(each iter costs around 1.5h for whole test)<br>\n17.17-&gt;14.45 (lambda0.85, 1 iters) <br>\n17.17-&gt;13.49 (lambda0.7, 3 iters) <br>\n14.71-&gt;12.34 (lambda 0.6, 5 iters) for model already finetuned with data generate by such iterative optimization</li>\n</ol>\n<p><strong>Training stages(speed performance reported on single 4080s)</strong></p>\n<ol>\n<li>100 epoch on OpenFWI, 12 days <ul>\n<li>AdamW, weight decay 1e-4, batch size 28, init lr 28/64*3e-4, 80% init lr+ 20% cos lr, EMA in last epoch(only minor diff, droped in later stages)</li>\n<li>predict test set and run forward simulation to generate extra data, only symmetric TTA used in this stage</li>\n<li>val/LB: 21.0/22.4(sym TTA)</li></ul></li>\n<li>6+20 epoch on OpenFWI + *4 copy of stage 1 generated data, 6 days <ul>\n<li>after first 6 epoch my computer restarted so I continue training from it</li>\n<li>only cos lr is used without init constant phase, init lr 28/64*1.5e-4</li>\n<li>val/LB: 17.17/???  (sym TTA)<br>\nval/LB: 14.45/14.9 (sym TTA+lambda0.85, 1 iter)<br>\nval/LB: 13.49/???  (sym TTA+lambda0.7, 3 iters)</li>\n<li>predict test and run forward simulation to generate extra data (sym TTA+lambda0.85, 2 iters)</li></ul></li>\n<li>10 epoch on OpenFWI + *10 copy of stage 2 generated data, 3 days <ul>\n<li>only cos lr is used without init constant phase, init lr 28/64*1.5e-4</li>\n<li>val/LB: 14.71/???  (sym TTA)<br>\nval/LB: 12.34/12.7 (sym TTA+lambda0.6, 5 iters)</li></ul></li>\n</ol>",
      "rawMarkdown": "Thanks Kaggle and organizer for hosting this competition, this is the most interesting competition that I ever entered. And also great thanks to people who sharing ideas and codes, especially:\n@brendanartley for modeling ideas\n@jaewook704 @manatoyo for forward simulation code\n@bguberfain for starter notebook\n\n**Data Processing**\n1. sign(x)*log(1+|x|)\n2. reorder each receiver based on CMP(central middle point) wrt source along receiver dimension to make input better aligned with target. To keep receiver dimension 70, odd and even indices are split into two channels. (5,1000,70)->(10,1000,70) \n(-18MAE when ~160MAE  on 7k samples)\n3. add (x,y) coordinate embedding as extra channels since FWI task is spatial variant\n\n**Model**\n1. I stack multiple U-net to mimic common iterative methods in physics / math. Here each U-net works as a single iteration step with a global receptive field. Scaling depth by stacking more U-nets works better than just scaling width or increase number of conv layers.\n2. Based on my experiment on 40k samples, larger model always lead to better result. The largest model I can train : stack 5 U-nets with depth 4 channel dim 128. Not sure how much further gain is possible by continue scaling up.\n3. I tried Convnext (using Bartley's model or replace 1st U-net in my model) without success and I don’t have enough computation resources to train CAFormer to fully converge so I end up without using any pretrained backbone.\n4. Other details: few step stride conv to downsample input data; intermediate conv layer(from Bartely work); down by avg_pool up by bilinear; batch norm (consistent better than other type of normalization once converged ); skip connect\n\n**Loss**\n1. MAE with smaller weight to deeper positions (1~1/4 linearly) since I suspect it be more noisy (84->81 MAE on 10k)\n\n**Augmentation**\n1. Symmetric augmentation (-10 MAE when ~110 on 10k samples)\n  - source is placed at [0, 17, 34, 52, 69], 34 after flip will be 35 this causes conflict in feature meaning. So I insert one more channel to represent source at 35 and fill by zero. Feature meaning is then self consistent before and after flip, though model still need to be trained to learn to handle this.\n  - used as TTA: 15.28->14.71\n2. Velocity map scaling(-10 MAE when ~100 on 10k samples)\n  - Scale target velocity map by *(1+alpha) then compress seis data along time axis to /(1+alpha) \n  - After play with forward simulation code I find scale velocity also affect sesi data maginitue.I can’t track this in closed form so I use an empirical rule: *(1+alpha)^0.26. Based on validation this only has minor effects if any.\n  - Here we are actually scaling time, this causes the augmented dataset with different source freq, so this doesn’t generate data with the same distribution as OpenFWI but luckily it still helps.\n\n**Reconstruction Error Optimization**\n1. notation: G - true inverse model, F - true forward model, M - our NN inverse model\n2. here I minimize ||F(M(x_k)) - x ||^2 wrt x_k start from x rather than change model M\n3. assume M is already a good estimation of G in terms of local change:  F(M(x+dx))-F(M(x))~=F(G(x+dx))-F(G(x))=dx then we can use simple iteration rule to reduce error without gradient of F:\nx_k=x_k-lambda*(F(M(x_k)) - x ) \nthat is we can manipulate x directly with change in y since f here close to identity map\n4. When used together with TTA: M(x):=(M(x)+Flip(M(Flip(x))))/2\n5. performance:(each iter costs around 1.5h for whole test)\n17.17->14.45 (lambda0.85, 1 iters) \n17.17->13.49 (lambda0.7, 3 iters) \n14.71->12.34 (lambda 0.6, 5 iters) for model already finetuned with data generate by such iterative optimization\n\n**Training stages(speed performance reported on single 4080s)**\n1. 100 epoch on OpenFWI, 12 days \n  - AdamW, weight decay 1e-4, batch size 28, init lr 28/64*3e-4, 80% init lr+ 20% cos lr, EMA in last epoch(only minor diff, droped in later stages)\n  - predict test set and run forward simulation to generate extra data, only symmetric TTA used in this stage\n  - val/LB: 21.0/22.4(sym TTA)\n2. 6+20 epoch on OpenFWI + *4 copy of stage 1 generated data, 6 days \n  - after first 6 epoch my computer restarted so I continue training from it\n  - only cos lr is used without init constant phase, init lr 28/64*1.5e-4\n  - val/LB: 17.17/???  (sym TTA)\n     val/LB: 14.45/14.9 (sym TTA+lambda0.85, 1 iter)\n     val/LB: 13.49/???  (sym TTA+lambda0.7, 3 iters)\n  - predict test and run forward simulation to generate extra data (sym TTA+lambda0.85, 2 iters)\n3. 10 epoch on OpenFWI + *10 copy of stage 2 generated data, 3 days \n  - only cos lr is used without init constant phase, init lr 28/64*1.5e-4\n  - val/LB: 14.71/???  (sym TTA)\n   val/LB: 12.34/12.7 (sym TTA+lambda0.6, 5 iters)",
      "votes": null
    },
    {
      "id": "3238077",
      "postDate": "07/01/2025 14:20:13",
      "content": "<p>Thank you for sharing, this is great stuff!</p>",
      "rawMarkdown": "Thank you for sharing, this is great stuff!",
      "votes": null
    },
    {
      "id": "3238605",
      "postDate": "07/02/2025 02:48:48",
      "content": "<p>Reconstruction Error Optimization part is really smart, <br>\ni tried selecting best prediciton from x1, x2,…,xn, np.median(x1,x2,…,xn) according to their reconstruction error, but just worse than simple np.median(x1,x2,…,xn), so didn't go further~</p>",
      "rawMarkdown": "Reconstruction Error Optimization part is really smart, \ni tried selecting best prediciton from x1, x2,...,xn, np.median(x1,x2,...,xn) according to their reconstruction error, but just worse than simple np.median(x1,x2,...,xn), so didn't go further~",
      "votes": null
    },
    {
      "id": "3238615",
      "postDate": "07/02/2025 03:04:05",
      "content": "<p>Thanks for your comment, wish you would reach 1st with one submission next time!</p>",
      "rawMarkdown": "Thanks for your comment, wish you would reach 1st with one submission next time!",
      "votes": null
    },
    {
      "id": "3238954",
      "postDate": "07/02/2025 11:02:42",
      "content": "<p>Genius!! I enjoyed dead heat with you in the last 24 hours.  I used 8 × H100 GPUs while you used mathematics.</p>",
      "rawMarkdown": "Genius!! I enjoyed dead heat with you in the last 24 hours.  I used 8 × H100 GPUs while you used mathematics.",
      "votes": null
    },
    {
      "id": "3239016",
      "postDate": "07/02/2025 11:51:25",
      "content": "<p>Thanks, it is my honor to receive such an evaluation from you！</p>",
      "rawMarkdown": "Thanks, it is my honor to receive such an evaluation from you！",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3238077,
      "author_name": "taylorsamarel",
      "author_url": "",
      "post_date": "07/01/2025 14:20:13",
      "content": "<p>Thank you for sharing, this is great stuff!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3238605,
      "author_name": "hydantess",
      "author_url": "",
      "post_date": "07/02/2025 02:48:48",
      "content": "<p>Reconstruction Error Optimization part is really smart, <br>\ni tried selecting best prediciton from x1, x2,…,xn, np.median(x1,x2,…,xn) according to their reconstruction error, but just worse than simple np.median(x1,x2,…,xn), so didn't go further~</p>",
      "votes": null,
      "replies": [
        {
          "id": 3238615,
          "author_name": "w5833946",
          "author_url": "",
          "post_date": "07/02/2025 03:04:05",
          "content": "<p>Thanks for your comment, wish you would reach 1st with one submission next time!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3238954,
      "author_name": "junkoda",
      "author_url": "",
      "post_date": "07/02/2025 11:02:42",
      "content": "<p>Genius!! I enjoyed dead heat with you in the last 24 hours.  I used 8 × H100 GPUs while you used mathematics.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3239016,
          "author_name": "w5833946",
          "author_url": "",
          "post_date": "07/02/2025 11:51:25",
          "content": "<p>Thanks, it is my honor to receive such an evaluation from you！</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3237944": "Thanks Kaggle and organizer for hosting this competition, this is the most interesting competition that I ever entered. And also great thanks to people who sharing ideas and codes, especially:\n@brendanartley for modeling ideas\n@jaewook704 @manatoyo for forward simulation code\n@bguberfain for starter notebook\n\n**Data Processing**\n1. sign(x)*log(1+|x|)\n2. reorder each receiver based on CMP(central middle point) wrt source along receiver dimension to make input better aligned with target. To keep receiver dimension 70, odd and even indices are split into two channels. (5,1000,70)->(10,1000,70) \n(-18MAE when ~160MAE  on 7k samples)\n3. add (x,y) coordinate embedding as extra channels since FWI task is spatial variant\n\n**Model**\n1. I stack multiple U-net to mimic common iterative methods in physics / math. Here each U-net works as a single iteration step with a global receptive field. Scaling depth by stacking more U-nets works better than just scaling width or increase number of conv layers.\n2. Based on my experiment on 40k samples, larger model always lead to better result. The largest model I can train : stack 5 U-nets with depth 4 channel dim 128. Not sure how much further gain is possible by continue scaling up.\n3. I tried Convnext (using Bartley's model or replace 1st U-net in my model) without success and I don’t have enough computation resources to train CAFormer to fully converge so I end up without using any pretrained backbone.\n4. Other details: few step stride conv to downsample input data; intermediate conv layer(from Bartely work); down by avg_pool up by bilinear; batch norm (consistent better than other type of normalization once converged ); skip connect\n\n**Loss**\n1. MAE with smaller weight to deeper positions (1~1/4 linearly) since I suspect it be more noisy (84->81 MAE on 10k)\n\n**Augmentation**\n1. Symmetric augmentation (-10 MAE when ~110 on 10k samples)\n  - source is placed at [0, 17, 34, 52, 69], 34 after flip will be 35 this causes conflict in feature meaning. So I insert one more channel to represent source at 35 and fill by zero. Feature meaning is then self consistent before and after flip, though model still need to be trained to learn to handle this.\n  - used as TTA: 15.28->14.71\n2. Velocity map scaling(-10 MAE when ~100 on 10k samples)\n  - Scale target velocity map by *(1+alpha) then compress seis data along time axis to /(1+alpha) \n  - After play with forward simulation code I find scale velocity also affect sesi data maginitue.I can’t track this in closed form so I use an empirical rule: *(1+alpha)^0.26. Based on validation this only has minor effects if any.\n  - Here we are actually scaling time, this causes the augmented dataset with different source freq, so this doesn’t generate data with the same distribution as OpenFWI but luckily it still helps.\n\n**Reconstruction Error Optimization**\n1. notation: G - true inverse model, F - true forward model, M - our NN inverse model\n2. here I minimize ||F(M(x_k)) - x ||^2 wrt x_k start from x rather than change model M\n3. assume M is already a good estimation of G in terms of local change:  F(M(x+dx))-F(M(x))~=F(G(x+dx))-F(G(x))=dx then we can use simple iteration rule to reduce error without gradient of F:\nx_k=x_k-lambda*(F(M(x_k)) - x ) \nthat is we can manipulate x directly with change in y since f here close to identity map\n4. When used together with TTA: M(x):=(M(x)+Flip(M(Flip(x))))/2\n5. performance:(each iter costs around 1.5h for whole test)\n17.17->14.45 (lambda0.85, 1 iters) \n17.17->13.49 (lambda0.7, 3 iters) \n14.71->12.34 (lambda 0.6, 5 iters) for model already finetuned with data generate by such iterative optimization\n\n**Training stages(speed performance reported on single 4080s)**\n1. 100 epoch on OpenFWI, 12 days \n  - AdamW, weight decay 1e-4, batch size 28, init lr 28/64*3e-4, 80% init lr+ 20% cos lr, EMA in last epoch(only minor diff, droped in later stages)\n  - predict test set and run forward simulation to generate extra data, only symmetric TTA used in this stage\n  - val/LB: 21.0/22.4(sym TTA)\n2. 6+20 epoch on OpenFWI + *4 copy of stage 1 generated data, 6 days \n  - after first 6 epoch my computer restarted so I continue training from it\n  - only cos lr is used without init constant phase, init lr 28/64*1.5e-4\n  - val/LB: 17.17/???  (sym TTA)\n     val/LB: 14.45/14.9 (sym TTA+lambda0.85, 1 iter)\n     val/LB: 13.49/???  (sym TTA+lambda0.7, 3 iters)\n  - predict test and run forward simulation to generate extra data (sym TTA+lambda0.85, 2 iters)\n3. 10 epoch on OpenFWI + *10 copy of stage 2 generated data, 3 days \n  - only cos lr is used without init constant phase, init lr 28/64*1.5e-4\n  - val/LB: 14.71/???  (sym TTA)\n   val/LB: 12.34/12.7 (sym TTA+lambda0.6, 5 iters)",
    "3238077": "Thank you for sharing, this is great stuff!",
    "3238605": "Reconstruction Error Optimization part is really smart, \ni tried selecting best prediciton from x1, x2,...,xn, np.median(x1,x2,...,xn) according to their reconstruction error, but just worse than simple np.median(x1,x2,...,xn), so didn't go further~",
    "3238615": "Thanks for your comment, wish you would reach 1st with one submission next time!",
    "3238954": "Genius!! I enjoyed dead heat with you in the last 24 hours.  I used 8 × H100 GPUs while you used mathematics.",
    "3239016": "Thanks, it is my honor to receive such an evaluation from you！"
  },
  "source": "meta"
}