{
  "id": 92440,
  "title": "MAE sucks",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/92440",
  "author_name": "Scirpus",
  "post_date": "2019-05-16T14:18:18.207000",
  "votes": 29,
  "comment_count": 29,
  "views": 0,
  "content": "<p>I found that using sqrt(time_to_failure) and using MSE gives much better performance on MAE once I square the predictions using Genetic Programming - anyone else tried this using more traditional approaches?\nblind:  1.8747610611086327\nvisible:  1.8005036417515352</p>\n\n<p>Edit: I ran it again as GP is a stochastic process to make sure it wasn't a fluke\nblindII:  1.877302349827226\nvisII:  1.8013807996259996\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/532255/13228/mae.png\" alt=\"My MAE\"></p>",
  "messages": [
    {
      "id": 532255,
      "postDate": "2019-05-16T14:18:18.207Z",
      "content": "<p>I found that using sqrt(time_to_failure) and using MSE gives much better performance on MAE once I square the predictions using Genetic Programming - anyone else tried this using more traditional approaches?\nblind:  1.8747610611086327\nvisible:  1.8005036417515352</p>\n\n<p>Edit: I ran it again as GP is a stochastic process to make sure it wasn't a fluke\nblindII:  1.877302349827226\nvisII:  1.8013807996259996\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/532255/13228/mae.png\" alt=\"My MAE\"></p>",
      "rawMarkdown": "I found that using sqrt(time_to_failure) and using MSE gives much better performance on MAE once I square the predictions using Genetic Programming - anyone else tried this using more traditional approaches?\nblind:  1.8747610611086327\nvisible:  1.8005036417515352\n\nEdit: I ran it again as GP is a stochastic process to make sure it wasn't a fluke\nblindII:  1.877302349827226\nvisII:  1.8013807996259996\n![My MAE](https://storage.googleapis.com/kaggle-forum-message-attachments/532255/13228/mae.png)",
      "votes": 28
    },
    {
      "id": 532282,
      "postDate": "2019-05-16T15:10:08.110Z",
      "content": "<p>I have - using lgb to predict on sqrt(y) scale and then squaring the predictions gives truly horrible results.</p>\n\n<p>Edit: it was horrible in local cv, but really nice on LB. I am genuinely confused atm.</p>",
      "rawMarkdown": "I have - using lgb to predict on sqrt(y) scale and then squaring the predictions gives truly horrible results.\n\nEdit: it was horrible in local cv, but really nice on LB. I am genuinely confused atm.",
      "votes": 7,
      "replies": [
        {
          "id": 532361,
          "postDate": "2019-05-16T18:08:15.227Z",
          "content": "<p>Wow you have zoomed up the LB quickly ;)  Looks like your confused state was only temporary - congratulations ;)</p>",
          "rawMarkdown": "Wow you have zoomed up the LB quickly ;)  Looks like your confused state was only temporary - congratulations ;)",
          "votes": 1
        },
        {
          "id": 532374,
          "postDate": "2019-05-16T18:56:06.327Z",
          "content": "<p>Thanks :-) Let's see if it lasts. The confusion persists though: same data, same cv (fixed seed), same sqrt/square transform and two models: LGB 1.305 cv -&gt; 1.398 lb, XGB 1.486 cv -&gt; 1.361 lb...</p>",
          "rawMarkdown": "Thanks :-) Let's see if it lasts. The confusion persists though: same data, same cv (fixed seed), same sqrt/square transform and two models: LGB 1.305 cv -&gt; 1.398 lb, XGB 1.486 cv -&gt; 1.361 lb...",
          "votes": 2
        },
        {
          "id": 532380,
          "postDate": "2019-05-16T19:32:56.163Z",
          "content": "<p>try blending both together Konrad and let us know LB results</p>",
          "rawMarkdown": "try blending both together Konrad and let us know LB results"
        },
        {
          "id": 532384,
          "postDate": "2019-05-16T19:41:22.497Z",
          "content": "<p>What a pleasant surprise</p>",
          "rawMarkdown": "What a pleasant surprise"
        },
        {
          "id": 532386,
          "postDate": "2019-05-16T19:41:32.243Z",
          "content": "<p><a href=\"/konradb\">@konradb</a>, congrats these are really good numbers. If you do not mind me asking how many features are you using? If not its okay, you can share after the competition.</p>",
          "rawMarkdown": "@konradb, congrats these are really good numbers. If you do not mind me asking how many features are you using? If not its okay, you can share after the competition."
        },
        {
          "id": 532393,
          "postDate": "2019-05-16T20:37:37.383Z",
          "content": "<p>@Corey: I will - first thing tomorrow, once I have new subs :-)</p>",
          "rawMarkdown": "@Corey: I will - first thing tomorrow, once I have new subs :-)"
        },
        {
          "id": 532394,
          "postDate": "2019-05-16T20:38:30.630Z",
          "content": "<p><a href=\"/sheriytm\">@sheriytm</a> 500ish - i took several of my own, raided the kernels for inspiration and then started trimming.</p>",
          "rawMarkdown": "@sheriytm 500ish - i took several of my own, raided the kernels for inspiration and then started trimming.",
          "votes": 1
        },
        {
          "id": 532396,
          "postDate": "2019-05-16T21:06:53.100Z",
          "content": "<p>Thanks <a href=\"/konradb\">@konradb</a>, that means I am on the right path. For my best model so far, I have trimmed down to 150 and on my way up again with a different set of features.</p>",
          "rawMarkdown": "Thanks @konradb, that means I am on the right path. For my best model so far, I have trimmed down to 150 and on my way up again with a different set of features."
        },
        {
          "id": 532412,
          "postDate": "2019-05-16T22:07:00.947Z",
          "content": "<p>Good luck :-)</p>",
          "rawMarkdown": "Good luck :-)"
        },
        {
          "id": 532597,
          "postDate": "2019-05-17T10:54:48.163Z",
          "content": "<p>I've run about 1000 models (with different lgb paramteters) and ploted the MAE and MSE (attached in this post). The correlation of MAE and MSE  looks very different depending on the fold I'm using to evaluate.</p>\n\n<p>(I took the folds spliting the time series into 3 continuos sections)</p>",
          "rawMarkdown": "I've run about 1000 models (with different lgb paramteters) and ploted the MAE and MSE (attached in this post). The correlation of MAE and MSE  looks very different depending on the fold I'm using to evaluate.\n\n(I took the folds spliting the time series into 3 continuos sections)",
          "votes": 3
        },
        {
          "id": 532684,
          "postDate": "2019-05-17T14:28:20.840Z",
          "content": "<p>it's not surprise. \"look on your data' :)</p>",
          "rawMarkdown": "it's not surprise. \"look on your data' :)"
        }
      ]
    },
    {
      "id": 532262,
      "postDate": "2019-05-16T14:27:39.150Z",
      "content": "<p>this may be because MAE has a second derivative of 0, but MSE has a nonzero second derivative; so MSE has some concept of momentum when it is training. i also find it bizarre that different objectives can work better</p>",
      "rawMarkdown": "this may be because MAE has a second derivative of 0, but MSE has a nonzero second derivative; so MSE has some concept of momentum when it is training. i also find it bizarre that different objectives can work better",
      "votes": 6,
      "replies": [
        {
          "id": 532373,
          "postDate": "2019-05-16T18:55:23.727Z",
          "content": "<p>If i remember correctly, minimizing KL-divergence boils down to minimizing mse for regression problems, could play a role, too.</p>",
          "rawMarkdown": "If i remember correctly, minimizing KL-divergence boils down to minimizing mse for regression problems, could play a role, too."
        }
      ]
    },
    {
      "id": 533046,
      "postDate": "2019-05-18T10:39:21.633Z",
      "content": "<p>What is difference between blind and visible MAE..?</p>",
      "rawMarkdown": "What is difference between blind and visible MAE..?",
      "votes": 3
    },
    {
      "id": 532328,
      "postDate": "2019-05-16T16:43:46.083Z",
      "content": "<p>The function in my automated parameter tuning kernel includes both RMSE and MAE as possible evaluation metrics. I've found it selects RMSE more often than not. </p>",
      "rawMarkdown": "The function in my automated parameter tuning kernel includes both RMSE and MAE as possible evaluation metrics. I've found it selects RMSE more often than not. ",
      "votes": 3,
      "replies": [
        {
          "id": 532523,
          "postDate": "2019-05-17T07:03:15.200Z",
          "content": "<p>Are you transforming the target to sqrt(target) then training then squaring the predictions?</p>",
          "rawMarkdown": "Are you transforming the target to sqrt(target) then training then squaring the predictions?",
          "votes": 1
        },
        {
          "id": 532669,
          "postDate": "2019-05-17T13:54:43.350Z",
          "content": "<p>No, I just found it curious that when given the option, an automated CV-minimising function would usually decline to use MAE. I think it's a strange choice of metric if we're being tested on earthquake samples that are harder to predict - I'd assume they would be longer sections, which as we've found are the hardest samples to model accurately. MSE/RMSE would give a far better indication of our success.</p>",
          "rawMarkdown": "No, I just found it curious that when given the option, an automated CV-minimising function would usually decline to use MAE. I think it's a strange choice of metric if we're being tested on earthquake samples that are harder to predict - I'd assume they would be longer sections, which as we've found are the hardest samples to model accurately. MSE/RMSE would give a far better indication of our success.",
          "votes": 1
        },
        {
          "id": 532675,
          "postDate": "2019-05-17T14:08:51.403Z",
          "content": "<p>Are you discussing MAE as a metric or as an objective function?  Not the same necessarily.</p>",
          "rawMarkdown": "Are you discussing MAE as a metric or as an objective function?  Not the same necessarily.",
          "votes": 4
        },
        {
          "id": 532738,
          "postDate": "2019-05-17T16:02:04.010Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> I was talking about evaluation metric and should have made the distinction. MAE has never given me the best results as an objective.</p>",
          "rawMarkdown": "@cpmpml I was talking about evaluation metric and should have made the distinction. MAE has never given me the best results as an objective.",
          "votes": 1
        },
        {
          "id": 532783,
          "postDate": "2019-05-17T17:33:34.123Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 532484,
      "postDate": "2019-05-17T04:36:05.237Z",
      "content": "<p>I am using RMSE and also MAE, it is more beneficial when we blended both model so you also try because according to your graph of MAE LB score is affected more in under 100 iterations. </p>",
      "rawMarkdown": "I am using RMSE and also MAE, it is more beneficial when we blended both model so you also try because according to your graph of MAE LB score is affected more in under 100 iterations. ",
      "votes": 1,
      "replies": [
        {
          "id": 532522,
          "postDate": "2019-05-17T07:03:09.030Z",
          "content": "<p>Are you transforming the target to sqrt(target) then training then squaring the predictions?</p>",
          "rawMarkdown": "Are you transforming the target to sqrt(target) then training then squaring the predictions?",
          "votes": 1
        }
      ]
    },
    {
      "id": 534588,
      "postDate": "2019-05-21T13:58:45.970Z",
      "content": "<p>I've experienced the same phenomenon with LightGBM. Using a simple model with CV of 2.0104 and a LB of 1.462, I get a CV of <strong>3.1949</strong> and a LB of 1.418.</p>\n\n<p>As discussed, I predicted the sqrt(y) with LightGBM and squared the results at the end. Any idea of where this huge gap could come from? I was trusting my CV till I tried this...</p>",
      "rawMarkdown": "I've experienced the same phenomenon with LightGBM. Using a simple model with CV of 2.0104 and a LB of 1.462, I get a CV of **3.1949** and a LB of 1.418.\n\nAs discussed, I predicted the sqrt(y) with LightGBM and squared the results at the end. Any idea of where this huge gap could come from? I was trusting my CV till I tried this...",
      "votes": 2,
      "replies": [
        {
          "id": 534593,
          "postDate": "2019-05-21T14:11:37.957Z",
          "content": "<p>Definitely trust your cv - the thing I was trying to show is that 13% of ~2600 samples is so tiny for a regression problem that one good/bad prediction will make a significant change in your Public LB MAE score.  I am not even sure that a consistent CV score is possible with loads of features so I am just going to submit both my best CV score and my best Public Score and hope for the best!</p>\n\n<p>The only thing I can guarantee is that there will be loads of comments after the competition ends of the flavour &gt; I would have gotten a gold medal if I chose this submission.</p>",
          "rawMarkdown": "Definitely trust your cv - the thing I was trying to show is that 13% of ~2600 samples is so tiny for a regression problem that one good/bad prediction will make a significant change in your Public LB MAE score.  I am not even sure that a consistent CV score is possible with loads of features so I am just going to submit both my best CV score and my best Public Score and hope for the best!\n\nThe only thing I can guarantee is that there will be loads of comments after the competition ends of the flavour &gt; I would have gotten a gold medal if I chose this submission.",
          "votes": 3
        },
        {
          "id": 534642,
          "postDate": "2019-05-21T15:44:38.123Z",
          "content": "<p>That makes sense. Thanks for the advice.</p>",
          "rawMarkdown": "That makes sense. Thanks for the advice.",
          "votes": 2
        }
      ]
    },
    {
      "id": 532260,
      "postDate": "2019-05-16T14:26:07.650Z",
      "content": "<p>Interesting, looking at target transform was next on my todolist.</p>",
      "rawMarkdown": "Interesting, looking at target transform was next on my todolist.",
      "votes": 1,
      "replies": [
        {
          "id": 533158,
          "postDate": "2019-05-18T14:55:01.093Z",
          "content": "<p>looks like it worked ;)</p>",
          "rawMarkdown": "looks like it worked ;)"
        },
        {
          "id": 533281,
          "postDate": "2019-05-18T21:41:13.497Z",
          "content": "<p>Nope</p>",
          "rawMarkdown": "Nope",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 532282,
      "author_name": "Konrad Banachewicz",
      "author_url": "",
      "post_date": "2019-05-16T15:10:08.110000",
      "content": "<p>I have - using lgb to predict on sqrt(y) scale and then squaring the predictions gives truly horrible results.</p>\n\n<p>Edit: it was horrible in local cv, but really nice on LB. I am genuinely confused atm.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 532361,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-05-16T18:08:15.227000",
          "content": "<p>Wow you have zoomed up the LB quickly ;)  Looks like your confused state was only temporary - congratulations ;)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 532374,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2019-05-16T18:56:06.327000",
          "content": "<p>Thanks :-) Let's see if it lasts. The confusion persists though: same data, same cv (fixed seed), same sqrt/square transform and two models: LGB 1.305 cv -&gt; 1.398 lb, XGB 1.486 cv -&gt; 1.361 lb...</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 532380,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2019-05-16T19:32:56.163000",
          "content": "<p>try blending both together Konrad and let us know LB results</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 532384,
          "author_name": "Ilham Firdausi Putra",
          "author_url": "",
          "post_date": "2019-05-16T19:41:22.497000",
          "content": "<p>What a pleasant surprise</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 532386,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-05-16T19:41:32.243000",
          "content": "<p><a href=\"/konradb\">@konradb</a>, congrats these are really good numbers. If you do not mind me asking how many features are you using? If not its okay, you can share after the competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 532393,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2019-05-16T20:37:37.383000",
          "content": "<p>@Corey: I will - first thing tomorrow, once I have new subs :-)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 532394,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2019-05-16T20:38:30.630000",
          "content": "<p><a href=\"/sheriytm\">@sheriytm</a> 500ish - i took several of my own, raided the kernels for inspiration and then started trimming.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 532396,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-05-16T21:06:53.100000",
          "content": "<p>Thanks <a href=\"/konradb\">@konradb</a>, that means I am on the right path. For my best model so far, I have trimmed down to 150 and on my way up again with a different set of features.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 532412,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2019-05-16T22:07:00.947000",
          "content": "<p>Good luck :-)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 532597,
          "author_name": "Matias Thayer",
          "author_url": "",
          "post_date": "2019-05-17T10:54:48.163000",
          "content": "<p>I've run about 1000 models (with different lgb paramteters) and ploted the MAE and MSE (attached in this post). The correlation of MAE and MSE  looks very different depending on the fold I'm using to evaluate.</p>\n\n<p>(I took the folds spliting the time series into 3 continuos sections)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 532684,
          "author_name": "DevJpyO",
          "author_url": "",
          "post_date": "2019-05-17T14:28:20.840000",
          "content": "<p>it's not surprise. \"look on your data' :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 532262,
      "author_name": "CoreyJamesLevinson",
      "author_url": "",
      "post_date": "2019-05-16T14:27:39.150000",
      "content": "<p>this may be because MAE has a second derivative of 0, but MSE has a nonzero second derivative; so MSE has some concept of momentum when it is training. i also find it bizarre that different objectives can work better</p>",
      "votes": 6,
      "replies": [
        {
          "id": 532373,
          "author_name": "Sven Hinderer",
          "author_url": "",
          "post_date": "2019-05-16T18:55:23.727000",
          "content": "<p>If i remember correctly, minimizing KL-divergence boils down to minimizing mse for regression problems, could play a role, too.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 533046,
      "author_name": "Reverse Flash",
      "author_url": "",
      "post_date": "2019-05-18T10:39:21.633000",
      "content": "<p>What is difference between blind and visible MAE..?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 532328,
      "author_name": "RNA",
      "author_url": "",
      "post_date": "2019-05-16T16:43:46.083000",
      "content": "<p>The function in my automated parameter tuning kernel includes both RMSE and MAE as possible evaluation metrics. I've found it selects RMSE more often than not. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 532523,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-05-17T07:03:15.200000",
          "content": "<p>Are you transforming the target to sqrt(target) then training then squaring the predictions?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 532669,
          "author_name": "RNA",
          "author_url": "",
          "post_date": "2019-05-17T13:54:43.350000",
          "content": "<p>No, I just found it curious that when given the option, an automated CV-minimising function would usually decline to use MAE. I think it's a strange choice of metric if we're being tested on earthquake samples that are harder to predict - I'd assume they would be longer sections, which as we've found are the hardest samples to model accurately. MSE/RMSE would give a far better indication of our success.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 532675,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-17T14:08:51.403000",
          "content": "<p>Are you discussing MAE as a metric or as an objective function?  Not the same necessarily.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 532738,
          "author_name": "RNA",
          "author_url": "",
          "post_date": "2019-05-17T16:02:04.010000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> I was talking about evaluation metric and should have made the distinction. MAE has never given me the best results as an objective.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 532783,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-17T17:33:34.123000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 532484,
      "author_name": "Himanshu Soni",
      "author_url": "",
      "post_date": "2019-05-17T04:36:05.237000",
      "content": "<p>I am using RMSE and also MAE, it is more beneficial when we blended both model so you also try because according to your graph of MAE LB score is affected more in under 100 iterations. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 532522,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-05-17T07:03:09.030000",
          "content": "<p>Are you transforming the target to sqrt(target) then training then squaring the predictions?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 534588,
      "author_name": "Ricard Delgado",
      "author_url": "",
      "post_date": "2019-05-21T13:58:45.970000",
      "content": "<p>I've experienced the same phenomenon with LightGBM. Using a simple model with CV of 2.0104 and a LB of 1.462, I get a CV of <strong>3.1949</strong> and a LB of 1.418.</p>\n\n<p>As discussed, I predicted the sqrt(y) with LightGBM and squared the results at the end. Any idea of where this huge gap could come from? I was trusting my CV till I tried this...</p>",
      "votes": 2,
      "replies": [
        {
          "id": 534593,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-05-21T14:11:37.957000",
          "content": "<p>Definitely trust your cv - the thing I was trying to show is that 13% of ~2600 samples is so tiny for a regression problem that one good/bad prediction will make a significant change in your Public LB MAE score.  I am not even sure that a consistent CV score is possible with loads of features so I am just going to submit both my best CV score and my best Public Score and hope for the best!</p>\n\n<p>The only thing I can guarantee is that there will be loads of comments after the competition ends of the flavour &gt; I would have gotten a gold medal if I chose this submission.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 534642,
          "author_name": "Ricard Delgado",
          "author_url": "",
          "post_date": "2019-05-21T15:44:38.123000",
          "content": "<p>That makes sense. Thanks for the advice.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 532260,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-05-16T14:26:07.650000",
      "content": "<p>Interesting, looking at target transform was next on my todolist.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 533158,
          "author_name": "Massoud Hosseinali",
          "author_url": "",
          "post_date": "2019-05-18T14:55:01.093000",
          "content": "<p>looks like it worked ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 533281,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-18T21:41:13.497000",
          "content": "<p>Nope</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "532255": "I found that using sqrt(time_to_failure) and using MSE gives much better performance on MAE once I square the predictions using Genetic Programming - anyone else tried this using more traditional approaches?\nblind:  1.8747610611086327\nvisible:  1.8005036417515352\n\nEdit: I ran it again as GP is a stochastic process to make sure it wasn't a fluke\nblindII:  1.877302349827226\nvisII:  1.8013807996259996\n![My MAE](https://storage.googleapis.com/kaggle-forum-message-attachments/532255/13228/mae.png)",
    "532282": "I have - using lgb to predict on sqrt(y) scale and then squaring the predictions gives truly horrible results.\n\nEdit: it was horrible in local cv, but really nice on LB. I am genuinely confused atm.",
    "532262": "this may be because MAE has a second derivative of 0, but MSE has a nonzero second derivative; so MSE has some concept of momentum when it is training. i also find it bizarre that different objectives can work better",
    "533046": "What is difference between blind and visible MAE..?",
    "532328": "The function in my automated parameter tuning kernel includes both RMSE and MAE as possible evaluation metrics. I've found it selects RMSE more often than not. ",
    "532484": "I am using RMSE and also MAE, it is more beneficial when we blended both model so you also try because according to your graph of MAE LB score is affected more in under 100 iterations. ",
    "534588": "I've experienced the same phenomenon with LightGBM. Using a simple model with CV of 2.0104 and a LB of 1.462, I get a CV of **3.1949** and a LB of 1.418.\n\nAs discussed, I predicted the sqrt(y) with LightGBM and squared the results at the end. Any idea of where this huge gap could come from? I was trusting my CV till I tried this...",
    "532260": "Interesting, looking at target transform was next on my todolist."
  }
}